Evaluating talent: a practical framework and scorecard for hiring
Why gut-feel evaluation fails (and what to build instead)
When you don't have a recruiter, candidate evaluation defaults to gut feel: read the resume, have a chat, see who you liked. It feels fine right up until the likable candidate misses every deadline, and on a small team you feel that miss for months.
Evaluation breaks in one of three places:
- You haven't defined "good" in job-relevant terms, so every interview measures charm.
- You collect weak signals: unstructured chats, generic questions, resume keywords.
- You score inconsistently: no written criteria, no anchors, notes that say "seemed sharp."
This page gives you the fix: a model for what to measure, a scorecard you can copy, 9 scenarios to pressure-test your process, and the fairness checks that keep it defensible. It works for hiring and for internal decisions like promotions.
What to measure: the SBPV model
S, skills: can they do the work? Demonstrated ability in job-relevant tasks. Evidence: work samples, skills tests, portfolio review, simulations.
B, behaviors: how do they work? Observable behaviors like prioritization, follow-through, and how they handle friction. Evidence: structured behavioral interviews, reference checks.
P, potential: how much can they grow into? Capacity for more complexity and ambiguity; learning agility. Evidence: stretch scenarios, problem-solving exercises, trajectory.
V, values: do they strengthen the team? Alignment with your defined non-negotiables (ethics, customer care, standards). Evidence: a structured values interview with anchored examples. This is values alignment against criteria you wrote down, not "would I get a beer with them." Unstructured "fit" is where bias hides.
Skills and behaviors are current-state. Potential is future-state. Values are guardrails.
How to measure: the methods menu
- Structured interviews: behaviors, values, role judgment.
- Work samples and simulations: skills and decision quality. The highest-signal method for most roles.
- Skills tests: discrete skills, kept job-relevant.
- Reference checks (structured): past behaviors in context.
- Case studies: problem framing, but watch for coached candidates.
The vocabulary that makes this auditable: job-relevant criteria from a quick job analysis, behaviorally anchored rating scales (written descriptions of what a 1 vs a 5 looks like), calibration (aligning raters), and adverse impact monitoring. If you can explain what you measured, how, and why it's job-relevant, your process scales and survives scrutiny.
The scorecard (copy and use)
Role: ____ Level: ____
Must-meet criteria (gates):
- ____ (skills or behaviors). Weight: __
- ____ (skills or behaviors). Weight: __
Differentiators:
- ____ (skills, behaviors, or potential). Weight: __
- ____. Weight: __
Values (non-negotiables, defined in advance):
- ____. Weight: __
Rating scale: 1 below, 2 mixed, 3 meets, 4 strong, 5 exceptional.
Evidence requirement: every rating needs 2-3 bullet points of observable evidence: things the candidate said, did, or produced. "I felt good about her" is not evidence. "Walked through a reconciliation error she caught and how she fixed the process" is.
Decision rules
- No hire if any non-negotiable value lands below "meets."
- Hire only if all must-meet criteria hit the bar and at least one high-leverage criterion shows a strong signal.
- Any rating above "meets" requires documented evidence, not an impression.
The 20-minute debrief agenda
- Silent review (2 min): everyone re-reads their own scorecard notes.
- Criterion readout (10 min): one criterion at a time, evidence only.
- Risk check (3 min): any must-meet gaps? any values concerns?
- Decision rule (3 min): hire or not yet, with the rationale captured.
- Next steps (2 min): feedback plan, data gaps, process improvements.
Even if "everyone" is you and one trusted teammate, run the agenda. Evidence-first readouts are what stop the loudest voice (including your own) from deciding by default.
Pressure-test your process: 9 scenarios
Use these to audit how you'd actually handle the situations that break hiring processes. For each, a strong answer states what you'd measure (SBPV), how, how you'd score it, and what decision rule applies.
Scenario 1: define a success profile
You're hiring a customer success manager. The brief so far: "someone proactive who can manage accounts and reduce churn."
Task: write a one-page success profile with 6-8 criteria maximum, each labeled must-have or trainable, with what strong evidence looks like.
Strong answers separate retention skill from stakeholder behaviors, use measurable outcomes (renewal influence, onboarding completion), and define baseline vs standout.
Scenario 2: choose methods under volume
You're hiring account executives from a large pile with 10 days to fill and at most 2.5 hours of interview time per candidate.
Task: design the minimum viable stage plan with what each stage measures.
Watch-outs: over-indexing on unstructured "chat" screens, and case studies that advantage coached candidates.
Scenario 3: build an anchored scorecard
Task: for a backend engineer, write 5-6 criteria and full anchors for one of them, like system design.
Anchor quality example: a 1 proposes components but misses key constraints; a 3 delivers a coherent architecture that handles core constraints with basic trade-offs; a 5 anticipates failure modes, capacity, security, and rollout.
Scenario 4: performance vs readiness (promotion)
Two internal candidates for team lead. A: highest output, sometimes abrasive, resists feedback. B: solid output, coaches others informally, handles ambiguity well.
Task: define what evidence you'd require before promoting either.
The classic failure: promoting the strongest individual contributor without validating any leadership behavior.
Scenario 5: a split panel
Feedback is split: "great communicator, hire," "weak on the core skill, no hire," and "seems smart, unsure."
Task: write the calibration agenda and decision rule that resolves this without politics.
Strong answers go criterion by criterion on evidence, separate interviewer drift from candidate signal, and decide explicitly between "collect more data" and "decide now."
Scenario 6: adverse impact appears
Your funnel shows a drop-off for one demographic group at the take-home stage.
Task: list 5 actions for the next 30 days that reduce the impact while preserving job-relevance.
High-signal actions: audit instructions, time expectations, and rubric clarity; add structured scoring with blind review where feasible; offer accommodations and alternative time windows; monitor pass-through rates by stage.
Scenario 7: rejected candidate asks for feedback
Task: write a feedback note aligned to the scorecard that is specific, respectful, and low-risk.
Good notes reference job-related criteria and observed evidence, avoid subjective judgments like "not a fit," and don't debate.
Scenario 8: internal mobility
Employees want to move into an operations analyst role.
Task: propose a lightweight skills evaluation, how results map to learning plans, and how you'll stop managers from quietly blocking moves.
Scenario 9: prove it's working
Task: define the numbers you'd track to show structured evaluation is paying off.
Useful ones: 90-day retention, ramp time, percentage of hires meeting expectations at 180 days, candidate drop-off by stage, and rating variance between interviewers.
Where does your process stand?
Score yourself honestly against the scenarios:
- Instinct-driven. Decisions ride on resumes and unstructured chats. Highest mis-hire risk and bias exposure. Next 30 days: write a success profile for one role, replace "overall impression" with criterion scoring, and standardize a short debrief.
- Structured starter. Common questions and a basic rubric exist, but anchors are thin and scoring varies by interviewer. Next: write anchors for your top 3 criteria, take evidence-based notes, and add must-meet decision rules.
- Reliable evaluator. Success profiles, job-relevant methods, consistent scoring, disciplined debriefs. Next: check score variance between raters quarterly, track how hires perform against their scores, and monitor adverse impact.
- Evaluation architect. The full loop runs across hiring and internal moves, with documentation and a feedback loop from outcomes back into criteria. Keep a reusable library of scorecards, question banks, and work samples per role family.
Process benchmarks worth adopting at any size: 5-8 criteria per scorecard (more dilutes reliability), defined questions and rubrics for most interviews, and small panels of prepared interviewers rather than long loops.
Compliance and fairness notes
- Derive criteria from the job. A one-page job analysis (essential tasks, top behaviors, minimum standards) is your job-relatedness documentation.
- Administer consistently: same questions, same methods, same rubric for every candidate in a role.
- Require evidence behind ratings. It reduces bias and it's your paper trail if a decision is challenged.
- Monitor pass-through rates by stage and group; the 4/5ths rule is a reasonable first flag.
- Provide accommodations, keep candidate data minimal, and define retention periods.
- Structured evaluation reduces inconsistency. Don't claim more than that, even to yourself: it's a discipline, not a guarantee.
Run this framework automatically with Truffle
Here's the honest catch: everything above is a system, and systems need an operator. When you're the owner and the operator and it's 9pm, the scorecard is the first thing that gets skipped.
Truffle is a candidate screening platform that combines resume screening, one-way video interviews, and talent assessments, which means your criteria get applied to every applicant whether it's candidate 4 or candidate 84. You define what matters in the role. Truffle screens the whole pile against it, ranks candidates with the evidence shown, and surfaces 30-second Candidate Shorts so your review time goes to the people worth meeting. AI does the recruiter's homework. The scorecard, the debrief, and the decision stay yours.
The 7-day free trial includes 30 credits, no card required. Start free trial