Field Notes
Candidate screening software Jul 2026 9 min read

What is a situational judgment test? What the score actually means

Situational judgment tests explained: what they measure, what the real validity research says, and why whose answer key you're using matters more than the score.

What is a situational judgment test? What the score actually means
AI summary
  • A situational judgment test (SJT) puts a candidate in a realistic work scenario and asks them to rank possible responses. Meta-analyses put its criterion validity around .26 to .34, a real number, but one that depends heavily on how the test was built and scored.
  • Research on coaching and faking shows scores can shift by roughly half a standard deviation from coaching alone, and close to a full standard deviation from faking, which is exactly what happens when a test uses one generic, reusable answer key.
  • The question worth asking before you trust any SJT score isn't whether SJTs are validated. It's whose answer key you're actually scoring against: a portable one anyone can prep for, or one built around how your own team handles the job.

Ask anyone selling a situational judgment test whether it’s scientifically validated, and you’ll get a yes, usually with a number attached. The research behind that number is real. What almost nobody selling you the test explains is which test the number is actually about.

A situational judgment test, or SJT, puts a candidate in a realistic work scenario, a coworker is behind on a deadline, a customer is furious, a shift just lost two people, and asks them to rank a set of possible responses from best to worst. It’s a multiple-choice test for judgment instead of knowledge.

The pitch is simple: SJTs are validated, so using one gives you real signal on how someone will handle the job before you ever meet them. That part is true, for a specific kind of test, built and scored a specific way. It’s not automatically true for every test wearing the SJT label, and the gap between the two is bigger than most hiring guides let on.

How a situational judgment test works

Every SJT has the same two pieces: a scenario and a set of responses. The scenario describes a situation the candidate is likely to face in the role. The responses are three to five plausible ways to handle it, and the candidate ranks them from most to least effective, or picks the best and worst of the set.

A retail scenario might read: “A customer is angry that an item is out of stock and is raising their voice at the front counter. What do you do first?” The response options might range from “apologize and offer to check another location” to “tell them to calm down before you can help.” There’s no arithmetic, no trick, just a judgment call under realistic pressure. It’s one of several formats worth knowing if you’re building assessments into a screening process for the first time.

Two ways to score the same scenario

How a scenario gets scored splits SJTs into two families. “Would do” tests ask what the candidate would actually do, and score against a behavioral tendency key. “Should do” tests ask what the ideal response is, and score against a knowledge-based key, what a competent person in that job is expected to know is correct.

That distinction sounds academic. It’s the hinge the rest of this post turns on, because the two scoring approaches hold up very differently once a test leaves the lab and someone starts selling practice packs for it.

What the research actually says about validity

The number vendors quote traces back to a real meta-analysis. Michael McDaniel and colleagues pooled 102 validity coefficients across more than 10,000 people and found SJTs generalize with a criterion validity of about .34 for predicting job performance. A later re-analysis by McDaniel’s team in 2007 landed closer to .26. Either way, that’s a real, useful correlation, in the range of cognitive ability tests and comfortably ahead of an unstructured chat.

Selection methodValidity coefficient
Structured interview.42
Cognitive ability test.31
Situational judgment test.26 to .34
Personality trait (best case: conscientiousness).19
Unstructured interview.19

Structured interview, cognitive ability, conscientiousness, and unstructured interview figures are the corrected estimates from Sackett, Zhang, Berry & Lievens (2022), Journal of Applied Psychology. SJT figures are from McDaniel et al.’s meta-analyses above. The two research streams use different methods, so treat the comparison as a rough scale, not a single unified ranking.

Christian, Edwards, and Bradley’s 2010 meta-analysis in Personnel Psychology dug into what actually drove that number, and this is the part the marketing pages skip. Validity wasn’t uniform. It moved with what the scenario was measuring (tests built around leadership and teamwork scenarios scored highest), and it moved with scoring method. Knowledge-based (“should do”) and behavioral-tendency (“would do”) instructions produced similar raw validity on average, but only knowledge-based scoring held up under coaching, which turns out to matter more than the topline number suggests.

The validity number was never about one universal test

Here’s the part that gets lost between the meta-analysis and the sales page. McDaniel’s .34, and the later .26, are pooled averages across dozens of SJTs, each one custom-built by industrial psychologists for one specific job, with a scoring key derived from what actually predicted performance at that employer. The number describes that construction process. It does not certify every scenario bank that happens to call itself a situational judgment test.

SHRM’s 2025 benchmarking data found 56% of employers now use some form of pre-employment assessment, and 79% weigh those scores as highly as a resume. That’s a lot of trust riding on a label that, on its own, guarantees nothing about how the test underneath was built.

Why so many situational judgment tests are coachable

The coaching and faking research

If an SJT is genuinely measuring judgment, coaching shouldn’t move the needle much. It does. A study of an SJT used for medical school admissions in Belgium found a coaching effect of roughly half a standard deviation, candidates who received coaching scored meaningfully higher than otherwise-identical candidates who didn’t.

Faking moves the number even further. Research on response distortion found candidates instructed to fake scored about 0.89 standard deviations higher than candidates answering honestly. At a typical 25% selection ratio, that shift is large enough that an employer picking “the highest scorers” would end up hiring mostly fakers, not the strongest performers.

Why a shared answer key is the real problem

This is exactly what an entire practice-test industry has built a business on. Search for any well-known SJT and you’ll find timed drills, sample answer keys, and coaching guides, because a generic, reusable scoring key is, by definition, something you can study for. The same dynamic shows up with Big Five personality assessments: a fixed, published trait model is easier to answer strategically than one nobody outside the company has seen.

It’s also the reason knowledge-based scoring resisted coaching better in Christian et al.’s analysis. There’s a specific, learnable pattern to “should do” answers when the key is fixed and shared across every employer using that off-the-shelf test.

None of this makes SJTs a bad idea. It means the format alone tells you nothing about whether the score in front of you reflects a candidate’s judgment or their access to a $30 prep course.

What actually makes a situational judgment test hard to game

The lever isn’t the format. It’s whose answer key the test is scoring against. A portable, universal key, the kind sold across thousands of employers, is exactly what a practice site can crack once and resell. A key that only exists inside one company, built from how that team actually wants a situation handled, isn’t something a generic prep pack can prepare anyone for.

How Truffle’s SJT works

This is the design choice behind Truffle’s SJT, which went live in April as part of the full talent assessment suite. Candidates rank response options from best to worst for each scenario, the same format the research above is about. But the ranking they’re scored against isn’t a universal key.

You build scenarios from a library or write your own, and for each one you can accept an AI-suggested default ordering as a starting point or set your own ranking based on how your team actually handles that situation. The score that comes back measures alignment with your answer, not a portable one that shows up in a coaching guide.

That also means Truffle’s SJT isn’t chasing the “validated” label the way an off-the-shelf test might. It measures alignment between how a candidate says they’d handle a scenario and what your team already agreed the right move is. There’s no universal correct answer to memorize, because the correct answer is specific to you, which is the property the coaching research above says actually matters.

Where a situational judgment test fits in your screening process

An SJT is one layer, not a verdict. A resume tells you about credentials and work history, the easiest thing to embellish and still where you cut the obvious mismatches. A one-way interview shows you how someone actually communicates, something no scenario ranking can replicate. An SJT adds a third signal: how someone reasons through a situation your team already has an opinion about.

Used alone, any one of these tells you something. Combined into a single screening process, the evidence compounds, because each layer is catching a signal the others miss. That’s also why an SJT shouldn’t be the first or only gate on a role. Pair it with resume screening to cut obvious mismatches first, so candidates spend the ten minutes a scenario ranking takes only once they’ve cleared the basics.

The question worth asking before you trust any score

The next time a vendor tells you their SJT is validated, the useful follow-up isn’t “how validated.” It’s “whose answer key is this test scoring against, and who decided it was correct.” A test built around a portable, universal key is coachable by design, no matter what number is printed on the spec sheet. A test built around your own team’s judgment isn’t something a candidate can study for in advance, because there’s nothing generic to study.

That’s the real shift SJTs deserve: less time asking whether the format works, and more time asking who wrote the key you’re actually scoring people against.

Frequently asked questions about situational judgment tests

What does SJT stand for?

SJT stands for situational judgment test. It’s an assessment format that presents a candidate with a realistic work scenario and a set of possible responses, then asks them to rank or select the most and least effective options.

Are situational judgment tests reliable?

The format has real research behind it, with meta-analyses putting criterion validity around .26 to .34 for predicting job performance. But that number describes tests built and scored a specific way. A generic, off-the-shelf SJT with a widely shared answer key is more vulnerable to coaching and faking than one scored against a specific employer’s own preferred approach.

Can candidates cheat or fake a situational judgment test?

Research on faking found scores can shift by nearly a full standard deviation when candidates answer strategically instead of honestly, enough to change who gets hired at a typical selection ratio. Coaching has a similar, if smaller, effect. Tests scored against a portable, universal answer key are the easiest to prep for, since that’s exactly what practice sites teach.

How is a situational judgment test different from a personality test?

A personality assessment, like one built on the Big Five model, measures general traits such as conscientiousness or extraversion against a norm group. An SJT measures how someone reasons through a specific, job-relevant scenario, scored against what a particular employer considers the right approach, not a universal personality profile.

End of dispatch

Founder, Truffle

Sean began his career in leadership at Best Buy Canada before scaling SimpleTexting from $1MM to $40MM ARR. As COO at Sinch, he led 750+ people and $300MM ARR. A marathoner and sun-chaser, he thrives on big challenges.

More from Field Notes

Truffle is candidate screening software built for the AI age

Start free trial

7 days · 30 credits · no card required

Start typing to search 300+ pages on hiretruffle.com.