Field Notes
Interviewing & screening practices Aug 2026 9 min read

Why a great interview doesn't mean a good hire

A 2017 meta-analysis of 2,515 interviews found charm moves the interview score and barely moves the job performance that follows it. Nerves do the same thing in reverse. Here's the part of the interview that structure alone can't fix, and what actually closes the gap.

A calm, confident candidate and a visibly nervous candidate answering the same interview question, illustrating how delivery style can outweigh job-relevant skill
AI summary
  • A 2017 meta-analysis of 2,515 interviews found candidates who used self-promotion and charm tactics scored meaningfully higher in the interview room (corrected r = .24) but not once they were on the job (r = .18, not significant).
  • Interview anxiety pulls interview ratings down by a similar margin (r = -.19 across two decades of research), while a 2019 study found it has close to no relationship with how the same people actually performed once hired.
  • A tighter question set fixes what gets asked. It doesn't remove the room. Closing that gap takes a step that isn't scored on how someone comes across live, which is what a talent assessment is built to add.

A candidate who compliments your team, name-drops something from your website, and tells you exactly what you want to hear scores better in an interview. That part isn’t a guess. A 2017 meta-analysis of 2,515 interviews found self-promotion and similar charm tactics move the interview rating by a real, measurable amount, a corrected .24. Track the same people into the job and that number drops to .18, and it stops being statistically significant. (Peck & Levashina, 2017, Frontiers in Psychology)

Something is showing up in the interview room that has almost nothing to do with the work.

That’s this post’s actual subject. Not whether interviews lie, and not whether the questions were good enough. Two ordinary traits, being easy to like and being calm while getting judged, quietly steer the interview score away from the thing you’re actually trying to measure, and they do it whether the interview is a rambling chat or a tightly scripted one. If you screen candidates yourself with no one to catch it when it happens, you’ve probably hired the charm and passed on the nerves at least once. You just never found out, because the nervous one never got the offer to prove you wrong.

The charm that gets someone hired doesn’t show up in their work

Picture two candidates for a front-desk role that pulled forty applications this week. Both have the same two years of experience. One sits forward, laughs at your jokes, and tells you your clinic feels like exactly the kind of place they’ve wanted to work. The other answers the same questions correctly but flatly, hands folded, eyes on the desk more than on you.

You will probably prefer the first one. Most people do.

What the numbers actually show

Interview researchers call what that candidate is doing impression management, a catalog of tactics that includes self-promotion, ingratiation, and telling an interviewer what they want to hear. A meta-analysis by Peck and Levashina pulled together 20 studies and 2,515 participants and found those self-focused tactics correlate with interview ratings at a corrected .24, a real and significant lift.

Then they checked a smaller set of studies, 730 people across six of them, that followed the same candidates into the job and compared the same tactics against their actual performance ratings. That number came in at .18 and wasn’t statistically significant. The charm that won the room didn’t reliably show up in the work six months later.

None of this makes the flat, quiet candidate a better hire by default. It means the thing that separated the two in your mind, in the moment, was mostly a performance skill with its own weak and inconsistent relationship to the skill you’re actually paying for.

The nervous candidate you passed on might have outperformed everyone you hired instead

Run the same interview back with the anxiety turned up instead of down. A 2018 meta-analysis in the Canadian Journal of Behavioural Science pulled together decades of interview research and found a moderate negative relationship, r = -.19, between how anxious a candidate seemed and how the interviewer scored them. Shaky hands, a cracking voice, a pause before answering: all of it pulls the interview score down, independent of what the candidate actually said.

Here’s the part that should bother you more. A 2019 study in the International Journal of Selection and Assessment followed real applicants for Resident Assistant positions past the interview and into the job, checking their measured anxiety against supervisor ratings of their actual work. The relationship was close to zero. The nerves that cost them points in the room told the interviewer almost nothing about how they’d do the job.

That’s a genuinely strange result to sit with. The exact trait an interviewer is reading, penalizing, and using to justify passing on someone doesn’t predict the thing the interview exists to predict. You’re not screening out weaker candidates when you screen out nervous ones. You’re screening out people who happen to find a stranger’s evaluation stressful, which describes a lot of qualified people applying for a role that matters to them.

The hire you never get to check

The frustrating shape of this problem is that you rarely see the miss. The charming candidate who underperforms gets hired, and you watch it happen over the next few months. The nervous candidate who would have outperformed them never gets the job, so there’s no month five where you find out you were wrong. One error is visible and correctable. The other is invisible by design.

A tighter question set fixes what gets asked, not the room

If you’ve read our breakdown of structured versus unstructured interviews, you already know the standard fix for most interview problems: same questions, same order, same rubric, scored before the debrief. It’s real, and it’s worth doing regardless of anything in this post.

It’s also not the fix for this specific problem, and it’s worth being honest about why. Structure standardizes what gets asked and how it gets scored on paper. It says nothing about how the answer gets delivered. A nervous candidate can get the exact same five questions as everyone else and still deliver every answer with a shaking voice and a long pause before starting. A charming candidate can hit the same rubric with a smile and a compliment stitched into the middle of it. Neither study behind these numbers isolated unstructured interviews. Both looked at interview ratings broadly, structured ones included, because charm and nerves are properties of delivery, not properties of the question.

Structuring the questions is still the right move. It just answers a different question than the one this post is asking. It tells you the same thing got asked of everyone. It doesn’t tell you the room stopped mattering.

What a screening step without a live audience actually catches

The honest fix isn’t a better interviewer or a longer rubric. It’s a screening step that never puts the candidate in a room with someone watching them in real time, because that’s the exact condition under which charm and nerves do their damage. Truffle is an AI screening platform that combines resume screening, one-way video interviews, and talent assessments, and the assessment layer is built for exactly this gap.

Say you’re hiring for that same front-desk role. Instead of asking a candidate to describe how they’d handle an overbooked morning with two upset customers and a phone that won’t stop ringing, you put the actual scenario in front of them and let them work through it alone, on their own time, with nobody watching. That’s a Situational Judgment Test, and it scores the candidate’s approach against how your best people already handle the same situation, not against a generic answer key and not against how composed they looked while answering.

That doesn’t erase bias in hiring. It closes off one specific channel. There’s no interviewer in the room to be charmed, and no audience for nerves to spike in front of. The score comes from what someone would actually do, not from how steady their voice was while saying it.

The same logic applies to a Personality assessment built on validated Big Five research, which reads temperament and work style instead of live delivery, and to an Environment Fit assessment, which checks whether what a candidate wants from a role matches what the role actually is. None of the three replaces your judgment. AI surfaces what each candidate’s answers show against the criteria you set. You’re still the one deciding who gets the offer, with one more signal that wasn’t shaped by who happened to be more comfortable in your office that afternoon.

Where this actually changes your process

You don’t need to run every candidate through an assessment to get the benefit.

The two ends of your pile

The mismatch this post describes matters most at the two ends of your pile: the polished candidate you’re tempted to fast-track past everyone else, and the qualified-on-paper candidate whose interview left you cold.

For the first, a short scenario assessment is a cheap check before you commit. If the approach behind the charm holds up against how your team actually works, good, you’ve confirmed something real. If it doesn’t, you’ve caught it before the offer instead of during month three. For the second, the same assessment gives a nervous candidate a shot at showing you the thing the interview never let them show, because nobody’s watching them show it.

Let the interview do what it’s good at

Keep the interview for what it’s actually good at. Whether you’d want this person on your team, how they communicate under normal circumstances, whether the resume and the one-way interview answers line up with a real person. Just stop asking it to also measure something it was never built to isolate.

Gut-feel hiring runs into a version of this same wall once a team gets past four or five people: the read that used to work stops being checked against anything real. Adding a scenario assessment to your screening process is the same move in miniature, giving your read on a candidate’s charm or nerves something to check itself against before the offer goes out, not after.

The interview answers a different question than the one you’re asking it

None of this means interviews are broken or that you should stop running them. It means the interview has always been better at telling you who’s comfortable being evaluated than who’s good at the job, and the two happen to correlate just often enough that most owners never notice the gap until a hire goes sideways for no obvious reason.

The fix isn’t a sharper read on people. It’s adding a signal that was never in the room to begin with, scored against what your own best people actually do, not against who told the best story about it.

Ready to see it on your next role? Truffle’s talent assessments sit in the same view as your resumes and one-way interviews, so a candidate’s approach to a real scenario shows up next to how they interviewed, not as a separate test you have to reconcile yourself. See plans and pricing, or start with a 7-day free trial, 30 credits, no card required.

Frequently asked questions about interview bias in hiring

Can a nervous candidate still be a great hire?

Often, yes. A 2019 study following Resident Assistant candidates into the job found interview anxiety had close to no relationship with their actual performance once hired, even though the same anxiety pulled their interview scores down. Nerves in the room are a weak signal for how someone will do the work.

Does being likeable in an interview mean someone will be a good employee?

Not reliably. A 2017 meta-analysis found self-promotion and charm tactics raised interview scores by a real, statistically significant amount, but the same tactics showed no significant relationship with the person’s job performance ratings afterward. Likeability predicts more likeability, not more skill.

Doesn’t a structured interview already fix this?

It fixes a different problem. Structure standardizes the questions and the scoring rubric, which genuinely improves an interview’s ability to predict performance. It doesn’t change how an answer sounds when someone delivers it nervously or persuasively, because that’s a property of the room, not the question set.

How does a talent assessment avoid the same bias?

By removing the audience. A situational judgment assessment presents a real scenario and scores the candidate’s approach against how your team already handles it, completed alone rather than in front of an interviewer. There’s no one to charm and no one to be anxious in front of, so the score reflects the approach instead of the delivery.

End of dispatch

Founder, Truffle

Sean began his career in leadership at Best Buy Canada before scaling SimpleTexting from $1MM to $40MM ARR. As COO at Sinch, he led 750+ people and $300MM ARR. A marathoner and sun-chaser, he thrives on big challenges.

More from Field Notes

Truffle is an AI screening platform built for the AI age

Start free trial

7 days · 30 credits · no card required

Start typing to search 300+ pages on hiretruffle.com.