Contents
She was the best interview of the day. Articulate, composed, well-researched. She’d prepared answers for every question, made eye contact, asked thoughtful follow-ups. You walked out of the room thinking: finally, someone who gets it.
Four months later she can’t start a task without being told exactly what to do. She crumbles when a colleague pushes back on her work. The enthusiasm that made her sparkle in the interview disappeared somewhere around week six, and what’s left is someone who needs more management time than the person she replaced.
Nothing in the interview predicted this. That’s the problem. Not that you’re bad at interviewing. That the interview was built to measure something other than what you actually needed.
You were measuring expression: how well someone presents under evaluation conditions. What you needed to measure was capability: whether they’d sustain performance when the conditions stopped being a 45-minute conversation and started being a job.
This is the difference between hiring the person who interviews best and hiring the person who performs best. For graduates with no work history, it’s a major distinction.
At a glance: what actually predicts graduate performance
- Unstructured interviews: r = 0.20, barely above chance
- GPA: r = 0.34
- Years of experience: r = 0.06, close to worthless
- Structured interviews: r = 0.58, nearly three times an unstructured interview
- Cognitive ability tests: r = 0.51 to 0.65
- Cognitive ability plus integrity test: r = 0.65, the strongest combination, and none of these require a work history
What you’re probably using (and what the research says about it)
Personnel psychologists have been studying which hiring methods actually predict job performance since the 1920s. The definitive work is the Schmidt and Hunter meta-analysis, originally published in 1998 and updated in 2016. It synthesised thousands of studies across 85 years and ranked every major selection method by how well it predicts whether someone will succeed in the role.
Here’s what it found, put simply.
Unstructured interviews (the format most hiring managers default to: open-ended questions, conversational flow, gut-feel evaluation) have a predictive validity of r = .20. On a scale where 1.0 is perfect prediction and 0 is random chance, that’s barely above noise.
GPA performs slightly better at r = .34, but employers are already abandoning it. Usage dropped from 73% in 2019 to 42% in 2026, because hiring managers figured out what the research always said: grades don’t reliably track to performance.
Years of experience, the thing graduates don’t have, turns out to be nearly worthless as a predictor. The 2016 update revised it down to r = .06. Essentially zero.
So the three signals most commonly used to evaluate graduates (how they come across in conversation, what grades they got, and how much experience they have) are the three weakest predictors in the entire research literature.
What actually works?
Structured interviews hit r = .58. Same questions, same order, scored against consistent criteria. Nearly three times the predictive power of an unstructured interview, at zero additional cost.
Cognitive ability tests hit r = .51 to .65 depending on the study. The single strongest standalone predictor across all job types. And critically, they don’t require the candidate to have any prior work experience.
The best combinations are stronger still. Cognitive ability + structured interview reaches r = .63. Cognitive ability + integrity test reaches r = .65.
The methods that predict performance don’t require a work history. The methods that require a work history don’t predict performance. A clear signal that the standard approach is wrong.
If you’re building a graduate programme, the assessment layer deserves at least as much thought as the sourcing. You’ve chosen your recruiting platform. Now the question is what happens when the candidates arrive.
Why “tell me about a time when…” fails with graduates
Behavioural interview questions are the backbone of most corporate hiring processes. “Tell me about a time when you dealt with a difficult stakeholder.” “Describe a situation where you had to manage competing priorities.” The logic is sound in theory: past behaviour predicts future behaviour. If someone has handled conflict well before, they’ll probably handle it well again.
The logic breaks down the moment the candidate has no relevant past behaviour to offer.
The fabrication problem
A university group project becomes “leading a cross-functional team through a complex deliverable.” A disagreement with a housemate becomes “resolving a stakeholder conflict through active listening and compromise.” The interviewer can’t verify any of it. The candidate who tells the most compelling story wins. Not the candidate with the strongest underlying capability. The best storyteller.
The AI problem
In 2026, candidates can prepare polished, contextually specific behavioural answers in seconds. Eighty-two percent of hiring managers can’t reliably identify AI-generated content. Applications per recruiter have risen 412% since 2022. Two-thirds of UK employers now say candidates use AI to misrepresent their skills. The behavioural interview was already weak for graduates. AI has hollowed it out.
The expression problem
This one is subtler and more important. Even without fabrication, even without AI, a behavioural interview measures how well someone performs under interview conditions. It rewards confidence, fluency, and the ability to construct a narrative under pressure. Those are real skills. They’re just not the skills that determine whether a graduate hire succeeds.
What determines success? The survey data is consistent. When Intelligent.com asked nearly a thousand business leaders why recent graduate hires failed, the top reasons were: lack of motivation or initiative (50%), lack of professionalism (46%), poor communication (39%), struggles with feedback (38%), and inadequate problem-solving (34%). All dispositional. All about who someone is, not what stories they can tell about who they’ve been.
A behavioural interview can’t reliably surface motivation. Or resilience. Or how someone responds when the work stops being novel and starts being repetitive. Those qualities show up over months, not in a 45-minute conversation. And asking a graduate “tell me about a time you stayed motivated through boring work” invites exactly the kind of rehearsed, unfalsifiable answer the process is already drowning in.
How to assess what actually predicts performance in grads
Three practical, research-backed changes you can make to your graduate hiring process.
1. Replace behavioural questions with situational ones
Stop asking “tell me about a time when.” Start asking “what would you do if.”
Situational questions present a realistic scenario and ask the candidate to reason through it in real time. They don’t require past experience. They don’t reward preparation or storytelling. They reveal how someone thinks when they can’t fall back on a rehearsed answer.
Here are five, each designed to surface a different quality.
Problem-solving under ambiguity
“You’re two weeks into a project and you realise the brief you were given contains contradictory requirements. Your manager is on leave for a week. What do you do?”
What to listen for: Does the candidate tolerate the ambiguity or freeze? Do they identify who else could help, or wait for the manager? Do they attempt a partial solution or do nothing? A strong answer shows initiative without recklessness. A weak answer waits for permission.
Response to critical feedback
“Your manager reviews your first major deliverable and tells you it misses the point entirely. They’re direct about it. What’s your first response, and what do you do over the next 24 hours?”
What to listen for: Emotional regulation. Does the candidate describe a process for absorbing feedback, or do they immediately defend their work? The strongest candidates separate the emotional reaction from the professional response. The weakest ones can’t.
Initiative without supervision
“You’ve completed your assigned tasks for the week by Wednesday afternoon. Nobody has given you anything new. What do you do with the remaining two days?”
What to listen for: This one is binary. Some candidates describe asking for more work, looking for problems to solve, or starting something they’ve noticed needs doing. Others describe waiting. The answer tells you more about sustained motivation than any behavioural question can.
Prioritisation under competing demands
“Three people need something from you by end of day. Your manager wants a report. A colleague needs help with a client presentation. An intern is stuck and has asked for your time. You can’t do all three well. Walk me through how you decide.”
What to listen for: Prioritisation logic. Does the candidate assess urgency and impact, or default to hierarchy? Do they communicate proactively with the people they can’t help immediately? Weak answers try to do everything. Strong answers make a decision and own it.
Adaptability when the brief changes
“You’ve spent three days building a recommendation based on the criteria your team agreed to. In the review meeting, your manager changes two of the four criteria. Half your work no longer applies. How do you handle it?”
What to listen for: Flexibility without resentment. The best candidates treat the change as information, not as a personal affront. They salvage what’s relevant and rebuild what isn’t. They don’t sulk. And they don’t pretend they’re fine with it when they’re clearly not. Honest adaptability beats performed enthusiasm.
2. Structure the interview
The single highest-impact change you can make, and it’s free.
Use the same questions, in the same order, with the same scoring criteria, for every candidate. This alone moves you from unstructured (r = .20) to structured (r = .58). Nearly three times the predictive power.
How to do it:
Before the first interview, define 8 qualities that matter for the specific role. Not generic qualities (“good communicator”). Specific ones relevant to the work (“can prioritise independently when managing multiple stakeholder requests”).
Write one situational question per quality. Score each answer on a 1-5 scale with brief anchor descriptions for each level. Compare candidates on the scores. Not on how you felt about them. Not on who you’d rather have lunch with. On the scores.
This may feel mechanical at first, but it works for this very reason. The “gut feel” approach is what produces the confidence bias that got you the articulate hire who couldn’t handle feedback.
3. Add a capability layer for what interviews can’t reach
Structured situational interviews are a significant improvement. But they still measure how someone reasons through a hypothetical in a 45-minute window. Some qualities only reveal themselves over time: sustained motivation when the work is routine, stress tolerance under real (not hypothetical) pressure, openness to feedback when it’s about their actual work and not a scenario.
These are what the research calls capabilities: the underlying behavioural dispositions that determine whether skills get applied consistently when the context is demanding. Not “can you do this?” but “will you keep doing this when it’s hard?” That’s a different question, and it requires a different kind of measurement. (For more of the nuance, see our article on capability vs competency).
Capability assessment measures these dispositions directly, without relying on self-report or past examples. It doesn’t ask graduates to describe their motivation. It assesses whether they’re naturally wired for sustained effort, whether initiative comes at a cost, and whether adaptability is a strength or something they can perform temporarily before it drains them.
For hiring managers, this fills the gap that even the best interview can’t close. You get the reasoning and judgement from the situational interview. You get the underlying wiring from the capability assessment. Together, they answer both questions: “how does this person think?” and “how will this person sustain?”

Why Reece Group’s graduate hires ramp 66% faster
When the Reece Group applied capability assessment to their graduate programme, they saw a 66% reduction in graduate time to competence. Not because they found different candidates. Because they measured differently. The candidates who looked strong on capability data were the ones who ramped fastest, handled complexity earliest, and needed the least management intervention. The interview alone hadn’t been surfacing them reliably.
The Intelligent.com data makes the same point from the other direction. The same survey that found 75% of companies were unsatisfied with graduate hires also asked what hiring managers actually want: initiative (57%), positive attitude (53%), work ethic (52%), adaptability (51%). Every quality on that list is dispositional. Every one is measurable through capability assessment. None of them are reliably surfaced by an unstructured interview asking “tell me about a time when.”
The frustration hiring managers feel about graduate hires isn’t inevitable. It’s a measurement problem. They’re measuring expression and expecting capability. When you measure capability directly, the outcomes change.
The strongest way to assess graduates considers their inherent capability
The problem with assessing graduates isn’t that they’re hard to evaluate. It’s that the evaluation was designed for someone else. Someone with a decade of workplace behaviour to draw on. Someone whose references can vouch for how they handle pressure, feedback, and ambiguity. Someone who can answer “tell me about a time when…” with a real answer from a real job.
Graduates will never have that at the point you’re assessing them. No amount of better behavioural questions changes that.
The question that matters isn’t “what have you done?” It’s “what are you capable of sustaining?” Different question. Different tools. And 85 years of research says the tools that answer it are available, validated, and underused.
If you want to see what capability assessment looks like for graduate hiring, here’s where to start.
How many situational questions should we ask in one interview?
Enough to cover each of the core qualities you defined for the role before the interview, without running so long that candidate fatigue affects later answers. Most structured interviews land around one question per quality, so define your qualities first and let that number drive interview length, rather than the other way round.
How do we get interviewers who default to “tell me about a time when” to actually switch to situational questions?
Give them the exact wording in advance and ask them not to deviate from it. The core discipline of a structured interview is asking the same question the same way to every candidate, so treat the switch as a script change, not a style preference, and sit in on a few early interviews to check it is being followed.
Can situational questions be rehearsed or gamed the same way behavioural ones can?
Less easily. Because situational questions ask what a candidate would do in a scenario they have not seen before, rather than retell a story they can prepare in advance, they are harder to answer from a memorised script. They are not immune to preparation entirely, candidates can rehearse general frameworks, but they cannot rehearse the specific scenario the way they can rehearse a behavioural story.
Should the same situational questions be used for every role, or customised per team?
Customised per role. The five example questions in this guide test general qualities like initiative and adaptability, but the scenario itself should reflect the actual work. A finance graduate role and a client facing role should be tested with different scenarios, even if the underlying quality being assessed is the same.
How do we score answers consistently across multiple interviewers?
Agree on the scoring scale and the anchor descriptions for each quality before any interviews happen, not after. Where possible, have two interviewers score the same candidate independently and compare notes, since that is the fastest way to catch where your scoring criteria are too vague.
Does this approach work for career changers, not just recent graduates?
Yes. The core problem this guide addresses, that behavioural interviews reward people with relevant stories to tell, applies to anyone without directly relevant experience in the role they are applying for, not only recent graduates.
How long should a structured interview take compared to a traditional one?
Roughly the same length, since you are replacing one type of question with another rather than adding extra steps. What usually changes is where the time goes, less time spent following up on a rehearsed story, more time spent probing how the candidate reasons through the scenario in front of you.
What if a hiring manager insists on keeping their gut feel in the process?
Let them keep an overall impression, but do not let it override the scored criteria. The point of structured scoring is not to remove human judgement, it is to make sure the judgement is being applied to the same evidence for every candidate, so gut feel becomes one input recorded alongside the scores, not a silent override of them.
How do we know if our own interview questions are actually predictive over time?
Track scores against actual performance for the hires you make, not just against who you hired. If your highest scoring candidates consistently become your strongest performers a year in, the questions are working. If there is no relationship, or a negative one, that is a sign the scoring criteria or the questions themselves need revisiting.
Is capability assessment a replacement for the interview, or a complement to it?
A complement. The situational interview tells you how someone reasons through a scenario in the moment. Capability assessment tells you whether the underlying disposition, motivation, resilience, openness to feedback, is there to sustain that reasoning once the job stops being a short conversation. Neither replaces the other, used together they answer different questions.



