Skip to Content

Can candidates fake behavioural assessments with AI? What the evidence actually says

Written by
Ben Schwencke
Updated
decorative gradient bars

Ask a hiring team what worries them about behavioural assessment in 2026 and you'll usually get the same answer: candidates are using AI to game it. And if you've watched what's happened to CVs, cover letters and interview answers over the past couple of years, that worry is entirely reasonable.

But here's the thing. The published research on AI and assessment faking tells a more interesting story than the panic does, and it points at a completely different fix than most people assume. Not detection. Design.

Before you rip out your behavioural assessment, or bolt a lie-detector onto it, it's worth understanding what the evidence actually shows.

What a behavioural assessment measures, and the catch built into most of them

Quick grounding first, because the terminology gets used loosely. A behavioural assessment is a standardised, validated measure of how someone works: how they approach problems, deal with people, handle pressure, attend to detail. Where a cognitive ability test tells you whether a candidate can do the job, a behavioural assessment tells you how they'll go about it. The two predict different things, both predict performance, and together they give you a far more complete picture than either alone.

Now the catch. Most behavioural assessments are self-report, meaning the candidate answers questions about their own tendencies. That design is why they're efficient, deeply researched and genuinely predictive. It's also why they've always had one obvious vulnerability: they ask people to describe themselves honestly, in the middle of a competitive process, with a job on the line.

You can probably see where this is going.

Candidates were faking these long before ChatGPT

Occupational psychologists have a name for the tendency to answer as the person you think the employer wants: social desirability bias. It's one of the most studied phenomena in self-report measurement, and it has been for decades. Ask a candidate whether they stay calm under pressure, mid-application, and you should absolutely expect the answer to lean flattering. That's not cynicism; it's just people being people.

"A motivated candidate has always been able to shade a personality questionnaire in their favour. Anyone selling you an assessment that was previously immune to faking was overselling."

Ben Schwencke, Chief Psychologist

So the honest starting point is this: the question was never whether faking is possible. It's how much it distorts the result, and what the assessment's design does about it. Which brings us to the bit everyone's actually worried about.

What AI actually changed (it's not what you think)

Here's where the evidence gets genuinely counterintuitive.

A 2025 study by Robie and colleagues in the International Journal of Selection and Assessment tested exactly the scenario keeping hiring teams up at night: can ChatGPT outperform humans at faking a personality assessment while dodging detection? The answer, without coaching, was largely no. The model didn't consistently beat motivated human fakers.

Read that again, because it reframes the entire problem. AI hasn't handed candidates some dramatic new capability. Skilled faking was already achievable by anyone who thought carefully about what the employer wanted to hear.

What changed is access. Faking a behavioural assessment competently used to take insight, effort and time. Now it takes a prompt. The ceiling didn't move. The floor did.

And that's not a reassuring finding, to be clear. When a capability goes from "available to the few who work at it" to "available to everyone in ten seconds", you don't get a handful of suspiciously polished profiles. You get your whole candidate pool compressed towards the same flattering middle, where the genuine differences between people, the very thing you're paying to measure, get harder and harder to see.

You don't lose the ability to measure. You lose range. And range is the thing shortlisting runs on.

A separate study in Personality and Individual Differences makes the access point vividly: it looked at whether generative AI could be used to hack top scores on high-stakes personality assessments. The era of assuming your typical candidate can't fake competently is over.

Why "just detect the fakers" doesn't work

The instinctive response is to police it. Bolt on a lie-detection scale, flag the suspicious profiles, job done. It's a comforting idea, and vendors will happily sell it to you.

The research is not kind to it. The same Robie study put two established detection approaches (impression-management measures and overclaiming questionnaires) through their paces, and found real limitations in how reliably either identifies manipulated responses. Detection scales are, and always have been, blunt instruments.

There's a deeper problem too. The detection mindset turns assessment into an adversarial exercise and puts the burden of proof on the candidate. Think about what that does in practice: your best applicants get treated as suspects, your false positives become awkward (or legally uncomfortable) conversations, and the whole experience starts to feel like airport security. Candidates talk. Reviews get written.

Policing is a poor substitute for measurement. Fortunately, the same research points somewhere much more useful.

The three design decisions that actually hold up

The most useful finding in the Robie research wasn't about catching anyone. It was about construction: forced-choice formats proved more resistant to faking than single-stimulus questions. In plain English: when a candidate has to pick between two equally appealing statements rather than rate each one on an agree/disagree scale, there's no obviously "best" answer to aim a prompt at. The route to a flattering profile narrows dramatically.

That single finding points at three practical decisions, in ascending order of robustness.

It's the design that resists faking, not the detectionRate-each-statementCandidate scores each itemon an agree/disagree scaleA clear "best" answer existsfor every single questionMost exposedForced-choiceCandidate picks betweenequally appealing optionsNo obviously "right" answerto aim a prompt atMore resistantTask-basedCandidate demonstrates itinstead of describing itNothing to self-report,so nothing to embellishPurpose-built to resist AISame science throughout. The difference is what the format gives a candidate to aim at.

1. Choose the response format deliberately

A well-constructed forced-choice behavioural measure is substantially harder to game than a straightforward agree/disagree questionnaire. And here's the important bit: this is a property of the instrument, decided long before your candidate ever sits it. If your current assessment uses single-stimulus ratings throughout, that's worth knowing today, not after your next hiring round.

2. Measure behaviour, don't just ask about it

Where a candidate demonstrates something through an interactive task rather than describing it, self-report faking has nothing to act on. There's no statement to agree with, so there's nothing to embellish. This is exactly why we built MindmetriQ around interactive tasks: it measures general cognitive ability (g) to the same strict psychometric standards as a traditional test, in a shorter, more engaging format that candidates actually finish, and it's purpose-built to be AI-resistant. You stop the cheating without compromising the rigour, and without treating anyone like a suspect.

3. Use assessment as evidence, not verdict

A behavioural profile is at its most powerful when it structures the next conversation. Feed it into a structured interview, with the same questions for every candidate and probing the areas the profile raises, and you get two things at once: a stronger prediction, and a natural sense-check on any profile that doesn't match the person in front of you.

Notice what none of these three require: catching anyone.

The honest candidate is the one who loses

Almost every conversation about AI and assessment is framed around the employer's risk. But there's a second party in this transaction, and they have more at stake.

When faking is easy and widespread, the person who loses is the candidate who answered honestly. They get measured accurately, then ranked against profiles that weren't, and they slide down your shortlist for the sole reason that they told the truth. Sit with that for a second, because it's quietly one of the least fair things that can happen in a hiring process.

The point of AI-resistant assessment design isn't to catch dishonest candidates. It's to make sure honest ones aren't punished for their honesty.

And it's not evenly distributed, either. The candidates hurt most are the ones who are unfamiliar with the process, less coached, and less likely to treat an assessment as a game to be beaten. That cuts directly against every diversity and social-mobility goal your hiring strategy has. If fairness is the reason you assess in the first place (and it should be), then faking-resistance isn't a security feature. It's a fairness feature.

Five questions to ask your assessment provider

If you're reviewing how you assess behaviour, these five questions will tell you most of what you need to know, and they're much better questions than "can it detect AI?"

  1. What response format does the assessment use, and what's the evidence on its resistance to faking?
  2. Where does it measure behaviour through a task, rather than through self-description?
  3. What validation evidence supports it, and how recently was it gathered?
  4. How is it monitored for fairness across different candidate groups?
  5. If a candidate challenged a decision informed by this assessment, what evidence would we point to?

A provider who can answer all five is one you can defend a hiring decision with. That, not any bold claim about beating AI, is the standard worth holding assessments to.

Conclusion and next steps

Generative AI hasn't broken behavioural assessment. What it has ended is the era when faking could be treated as a marginal problem affecting a handful of unusually determined candidates. The evidence points firmly away from detection and towards design: deliberate response formats, task-based measurement where the stakes are highest, and assessment used as structured evidence rather than a verdict.

"Well-designed behavioural assessment remains one of the most predictive, most fair tools in hiring. The organisations that get this right won't be the ones policing candidates. They'll be the ones who chose better instruments."

Ben Schwencke, Chief Psychologist

The next steps are to audit the response format of whatever you currently use to assess behaviour, ask your provider the five questions above, and consider task-based measurement wherever faking would hurt most: typically high-volume and early-careers pipelines, where validated ability testing is doing the heavy lifting anyway. If you'd like to see how our behavioural assessments and MindmetriQ handle all of this, you can book a call with our team to talk it through, or explore our psychometric tests for more information.

author profile ben schwencke
Primary author

Ben Schwencke

Chief psychologist at Test Partnership. MSc in Organisational Psychology with over ten years experience in psychometric testing.