Skill-based accounting test for hiring: work-sample scenarios, rubric, and benchmarks
An accounting mis-hire is one of the quietest expensive mistakes a small business can make: the books look fine right up until they don't. Resumes can't protect you, because "owned month-end close" appears on the resume of people who can and can't actually do it. This page gives you a work-sample screen that settles the question: nine realistic scenarios with answer guidance, a 0-4 rubric, weighted scoring to 100, role benchmarks, and an integrity gate. If you're hiring an accountant now, the accountant hiring guide covers pay, sourcing, and interviews for the role. Hiring for the CFO seat instead? Our CFO hiring guide covers what that role requires on top of the technical baseline this screen checks.
What to screen for
A strong accounting hire produces accurate, supportable, timely output under real constraints: incomplete data, deadlines, and pressure. Five domains cover it:
- Financial accounting and reporting judgment: accrual thinking, revenue and expense recognition, capitalize versus expense, adjusting entries
- Close execution and reconciliation: bank and GL reconciliations with discrepancies, root causes, clean tie-outs
- Controls and auditability: preventive versus detective controls, approvals, segregation of duties
- Data and spreadsheet fluency: error checking, summarization, variance analysis
- Professional judgment: integrity, escalation, clear communication under pressure
Recommended structure, 75 to 120 minutes total: a work-sample case (35 to 45 minutes, 40 points), a data reasoning task (20 to 30 minutes, 25 points), a controls scenario (15 to 20 minutes, 20 points), and judgment prompts (10 to 15 minutes, 15 points).
The test: 9 scenarios with answer guidance
1. Month-end close: adjusting entries
Scenario: Closing December books, the candidate finds: a $24,000 vendor invoice dated December 20 for Q1 services (January through March), posted entirely to expense in December; and an $18,000 customer payment received December 28 for services delivered in January, recorded as December revenue.
Task: Propose the adjusting entries and explain the financial statement impact.
Strong answer: debit prepaid expenses and credit expense for $24,000 (then amortize monthly over Q1), and debit revenue and credit unearned revenue for $18,000. December expenses drop $24,000 and revenue drops $18,000, both moving to the periods they belong to. The explanation should be clean enough for a non-accountant owner to follow.
2. Bank reconciliation with discrepancies
Scenario: Bank statement balance $112,450; GL cash balance $105,730. Identified items: outstanding checks $9,600, deposits in transit $14,500, an unrecorded $120 bank fee, and a $2,440 customer wire received but not posted to the GL.
Task: Work the reconciliation, book the required entries, and state what remains unexplained.
Strong answer: adjusts the bank side (add deposits in transit, subtract outstanding checks) and the book side (record the fee and the wire), then explicitly identifies that a difference remains and says what they'd investigate next rather than forcing it to tie. That honesty under a deadline is precisely the trait you're screening for; a candidate who plugs the gap or declares it reconciled is your red flag.
3. Revenue recognition under sales pressure
Scenario: A customer signs a $60,000 contract for a 12-month subscription starting February 1 and pays in full on January 15. The sales team asks the candidate to "help revenue this quarter."
Task: Describe recognition from January 15 through March 31 and how they'd communicate the rationale.
Strong answer: nothing recognized in January (the cash is a contract liability), $5,000 per month starting February, so $10,000 by March 31. Equally important: a respectful, firm explanation to sales about why, and no hedging about whether the rules bend.
4. Capitalize versus expense
Scenario: $8,000 of computer equipment, $6,500 of annual software licenses, and $3,000 of separately invoiced implementation services.
Task: Classify each, explain why, and outline the follow-on accounting.
Strong answer: capitalize the equipment (subject to your capitalization threshold) and depreciate over its useful life; expense the annual licenses over the subscription period; treatment of implementation costs depends on policy and the nature of the service, and a strong candidate says so and asks about your threshold rather than asserting one universal rule.
5. Variance analysis with limited information
Scenario: Operating expenses up 18% month over month. The summary shows headcount up 3, contractor spend up $22,000, travel up $9,000, software up $6,000.
Task: Write 5 to 8 bullets identifying likely drivers, the data to request next, and how to present it to a non-finance leader.
Strong answer: quantifies each driver against the total change, distinguishes one-time from recurring, names the follow-up questions (are the contractors project-based, was the travel an event), and ends with a plain-language summary an owner could read in a minute.
6. Error-finding in journal entries
Scenario: Three material entries to review. 1) Debit accounts receivable $15,000, credit cash $15,000. 2) Debit depreciation expense $4,000, credit accumulated depreciation $4,000. 3) Debit cash $2,500, credit unearned revenue $2,500.
Task: Identify which are wrong or suspicious and what to verify.
Strong answer: Entry 1 is the suspect: debiting AR against cash isn't a normal transaction shape (a billing credits revenue; a customer payment debits cash and credits AR), so the candidate should question what it's trying to record. Entries 2 and 3 are standard, with a verification note on the depreciation schedule and the deferral terms. You're testing whether transaction shapes look wrong to them on sight.
7. Internal controls in a lean team
Scenario: In accounts payable, one person can create vendors, enter invoices, and approve payments up to $10,000. The CFO says, "We're lean; it's fine."
Task: Identify the risks, categorize recommended controls as preventive or detective, and propose a remediation plan realistic for a small team.
Strong answer: names the segregation-of-duties risk and the fraud pattern it enables (fictitious vendors), then proposes layered, right-sized controls: someone else approves new vendors, a second set of eyes on payment runs above a threshold, and a monthly exception report. Generic "add more controls" answers score 2; small-team-realistic designs score 4.
8. Audit request triage
Scenario: An auditor requests support for 25 samples with a 48-hour deadline, mid-close, manager out.
Task: A step-by-step plan for the next 2 hours.
Strong answer: triage the list by what's fast versus slow to pull, communicate early with the auditor about sequencing and any at-risk items, protect the close-critical tasks, and leave a documented trail. You're watching for prioritization and proactive communication, not heroics.
9. Ethics under pressure
Scenario: A senior leader asks the candidate to "move" an expense into next month to hit targets: "We'll reverse it later."
Task: How do they respond, and what do they do if the pressure continues?
Strong answer: a clear no with a professional explanation of why, an offer of legitimate alternatives if any exist, and a named escalation path if pressure continues. Anything that entertains the adjustment "just this once" gates the whole assessment. This is the question protecting you from the expensive quiet mistake.
Scoring
Score each prompt 0 to 4 against these anchors, then weight: work-sample case (scenarios 1 to 4) 40 points, data reasoning (5) 25 points, controls (7) 20 points, professional judgment (8 and 9) 15 points. Scenario 6 works well as an unscored calibration item or bonus.
- 0, not demonstrated: misunderstands the task; material errors; can't explain reasoning
- 1, inconsistent: partial approach, multiple errors, weak assumptions
- 2, competent: mostly correct, minor errors, reasonable explanation
- 3, strong: correct and complete, clear logic, anticipates follow-up issues
- 4, advanced: accuracy plus judgment, strong documentation habits, business-aligned communication
Gates, defined in advance and applied consistently: if the ethics scenario scores below 2, require a follow-up conversation and additional review before proceeding, whatever the total. For GL and close roles, if scenarios 1 and 2 average below 2, require more evidence before advancing.
If two people can score: score work-sample sections independently after a 15-minute calibration, and reconcile any totals more than 10 points apart with documented rationale.
Benchmarks and how to read results
- Below 50, foundational gaps: inconsistent accrual logic or reconciliation discipline. Higher error and rework risk without close supervision. Not a close-ownership hire.
- 50 to 69, emerging: handles routine tasks with guidance; misses edge cases and messy data. Workable for entry AP and AR roles with strong review; plan weekly reconciliation reviews and a template-driven close checklist into onboarding.
- 70 to 84, role-ready: competent judgment, accurate mechanics, sound prioritization. Fits staff and many senior contexts with ownership of reconciliations and parts of close.
- 85 to 100, advanced: reliable under ambiguity, communicates clearly, anticipates control and audit needs. Fits senior accountant and close-owner roles; add a role-specific technical memo exercise for reporting-heavy positions.
Suggested minimums by role, as starting points to refine with your own outcome data: entry AP/AR 55 to 65, staff accountant 70 or better, senior accountant or close owner 80 or better.
Use the domain pattern, not just the total, to design onboarding: a strong scorer with a weak controls answer gets controls shadowing in month one; a strong analyst with slower mechanics gets the close checklist first.
Fairness, accessibility, and documentation
Standardize the prompts, time boxes, allowed tools, and scoring anchors for every candidate at the same stage. Offer reasonable accommodations, keep materials screen-reader friendly, and monitor pass rates by group where legally appropriate, using selection-rate ratios like the four-fifths rule as a first screen. If disparities appear, check whether content over-measures irrelevant skills and whether scoring drifted between raters. Keep an audit trail: the role-to-skill blueprint, test versions, rubrics, calibration notes, and decision rationale. Assessment results are structured inputs alongside interviews and references; no single score should ever be the whole decision.
Run this screen automatically with Truffle
The reason most owners don't run work-sample screens is logistics: sending files, chasing submissions, scoring nine scenarios per candidate by hand. Truffle is a candidate screening platform that combines talent assessments with resume screening and one-way video interviews, so this screen runs as part of the same pipeline that reads the resumes. Load your scenarios into an assessment, use a one-way interview for the judgment questions so you can hear how a candidate handles the ethics scenario in their own words, and AI scores every response against the criteria you set, reasoning shown. You review a ranked shortlist with the evidence attached instead of a folder of spreadsheets. You still make every call.
Plans start at $49 a month. The 7-day free trial includes 30 credits and no credit card. Start free trial.