Interviews and simulations are often compared on surface features: one involves a person, the other software; one takes an hour, the other a few minutes. Those differences are real but secondary. The difference that determines what you can conclude from each is the type of evidence it produces.
An interview produces statements. A candidate recalls an episode, selects which details to include, and describes it in retrospect. A simulation produces actions: the decisions a candidate makes inside a live situation, in the order they make them, with the constraints of the role in force. These are not two grades of the same evidence. They are different kinds of thing, and they support different conclusions.
Section 01 Two kinds of evidence
The clearest way to see the distinction is to take one competency and assess it both ways. Consider service recovery - how a person handles a customer whose expectations have been broken. It appears in the competency framework of nearly every frontline role.
Assessing the same competency: service recovery
Interview - a statement
“Tell me about a time you handled an upset customer.”
What you receive- An episode the candidate selected, almost always one that ended well
- Recalled after the fact, with the outcome already known
- Compressed into roughly ninety seconds of narration
- Delivered with time to consider, and often rehearsed
- Unverifiable, and dependent on how well they tell a story
Simulation - an action
A guest reaches the desk. Her room was released this morning after an overbooking. Two people are queuing behind her. “I booked this three months ago. What are you going to do about it?”
What you receive- The words the candidate actually used, in the moment
- Whether they checked availability before offering anything
- Whether they acknowledged the guests waiting behind
- What they did after the guest rejected the first offer
- The order they addressed things in, under time pressure
Both approaches are aimed at the same competency. Only one of them observed it. The interview produced a description of service recovery that the candidate composed; the simulation produced an instance of service recovery that the candidate performed.
Four things follow from that difference, and the sections below take each in turn: the unit of evidence is an action rather than a statement, the performance is current rather than recalled, it occurs in context rather than in the abstract, and what gets analyzed is the behavior itself rather than an impression of a story.
Section 02 Actions instead of statements
A statement is a claim about behavior. An action is the behavior. Only one of them can be wrong about itself.
When a candidate says “I always check with the guest before making a change,” that sentence may be accurate, aspirational, or simply what they believe good service sounds like. Nothing in the interview can separate those possibilities, because the only evidence available is the sentence itself.
In a simulation the equivalent evidence is whether they checked. There is no gap between the claim and the conduct, because no claim was invited. This closes the most persistent weakness in interview-based assessment: the correlation between describing good practice and following it is considerably weaker than most hiring processes assume.
It also changes who is advantaged. Describing your own competence well is a distinct skill from possessing it. It correlates with education, with confidence, with practice at interviewing, and with speaking the assessor's first language. In most frontline roles it is not part of the job. Assessing actions removes that layer, and with it a source of disadvantage for candidates who work well but present modestly.
Section 03 Performance instead of recollection
An interview samples memory. A simulation samples performance.
Everything a candidate reports in an interview has passed through recall, and recall is reconstructive rather than archival. People remember outcomes more reliably than the reasoning that produced them, and once an episode's ending is known it reshapes the account of what led there. A decision that was uncertain at the time is remembered as deliberate. Hesitation disappears. The story acquires a coherence the event did not have.
None of that requires dishonesty. It is how memory works, and it means an interview answer is a reconstruction produced with the benefit of hindsight - even from a candidate scrupulously trying to be accurate.
The candidate also controls the sample. Asked for a time they handled something difficult, they will supply their best available case, which tells you about the top of their range rather than its typical value. A simulation presents a situation nobody selected, so what you observe is how this person handles this problem now - including the parts that did not go smoothly.
Section 04 Context: where judgment happens
Competence is not a property someone carries around. It appears, or fails to, inside specific conditions.
An interview necessarily strips out context. The candidate sits in a quiet room with nothing else demanding attention, unlimited time to compose a sentence, and complete information about how the story ends. Almost none of that resembles the role.
The work happens under different conditions: incomplete information, several things needing attention simultaneously, someone waiting, a policy that does not quite cover the case, and no opportunity to draft an answer. Judgment in frontline roles is largely the ability to act sensibly under exactly those conditions, and an interview removes every one of them before asking about it.
A simulation restores them. The queue is present. The information is partial. The guest interrupts. The candidate must decide what to handle first, knowing that attending to one thing means leaving another. That is why the same competency assessed in context and out of context produces different results - the context is not background to the skill, it is the setting the skill exists in.
Why this matters for prediction
A candidate who describes prioritization well in a quiet room and a candidate who prioritizes well with three things demanding attention are not necessarily the same person. The second is the one you are hiring.
Section 05 Analysis: what is examined
The difference in evidence produces a difference in analysis. Different material is available, so different things can be measured.
What an interview yields for analysis
The material is a narrative and the analyst is a person forming an impression of it. Even in a well-structured interview with defined criteria, the rater is judging how convincing an account was, frequently writing notes afterward rather than during, and unable to separate the quality of the underlying behavior from the quality of its telling. The record that survives is a summary judgment about a story.
What a simulation yields for analysis
The material is the candidate's own conduct, captured completely, which makes several distinct things measurable:
- The response itself - what they said or chose, assessed against defined criteria for that task, with the reasoning behind each score retained.
- Sequence and prioritization - the order things were addressed, which is invisible in a retrospective account but often the clearest signal of operational judgment.
- Accuracy - whether the discrepancy was noticed, the figure was right, the procedure was followed.
- Pace - how quickly they moved, and where they hesitated.
- Recovery - what they did when their first approach did not work, which is where the difference between candidates is usually largest.
- Comparison against a benchmark - performance measured against how your own strong performers complete the same tasks, rather than against an interviewer's internal standard.
Every element is traceable to a specific moment in the candidate's performance. When a hiring manager asks why one candidate scored above another, the answer refers to what each of them did, at which step, against which criterion - not to how each came across.
Section 06 Side by side
| Dimension | Job interview | Job simulation |
|---|---|---|
| Unit of evidence | A statement about behavior | The behavior itself |
| Time frame | Past, recalled after the outcome is known | Present, as the situation unfolds |
| Who selects the example | The candidate, from their best available cases | The role, from its everyday demands |
| Conditions | Quiet room, full attention, time to compose an answer | Partial information, competing demands, time pressure |
| What is analyzed | An impression of a narrative | Responses, sequence, accuracy, pace, and recovery |
| Verifiability | Low - accounts cannot usually be checked | Complete - every action is recorded |
| Advantages candidates who | Describe their work fluently and interview often | Perform the work well |
| Basis of comparison | Different conversations, reconciled afterward | Identical tasks, identical criteria |
| Best suited to assessing | Motivation, circumstances, expectations, two-way fit | Capability, judgment, and task performance |
Section 07 Where interviews remain essential
None of the above makes interviews obsolete. It locates them. There is a class of question where a statement is the appropriate evidence, because the subject matter is intention rather than capability - and for those, an interview is the correct instrument and a simulation is no substitute.
Intent and circumstances
Whether someone wants the role, can work the shift pattern, can travel to the site, what notice they owe, what they expect to earn, and where they hope this leads. These are matters only the candidate can report, and asking them directly is the right method.
The conversation runs both ways
Candidates are assessing you at the same time. An interview is where a manager answers questions, describes the team, addresses hesitations, and gives someone a reason to choose this employer. In tight labor markets that persuasive function is often the main event.
Cases the process did not anticipate
Any standardized method handles typical candidates well and unusual ones poorly. A career changer whose relevant experience comes from another sector, an applicant with an easily explained gap, someone whose situation falls outside the design assumptions - these need a person who can follow up, and a good interviewer adds real value there.
Complex, senior, and low-volume roles
Where a role is highly individual, where volumes are small enough that consistency is manageable by hand, or where the work cannot be meaningfully sampled in a few minutes, a well-structured interview remains the appropriate primary instrument.
Section 08 What the research adds
The selection literature approaches this from a different angle - measuring how strongly each method's results correlate with later job performance - and it is worth knowing what it says, including the parts that complicate a simple story.
Predictive validity of common selection methods
Operational validity coefficients, Sackett et al. (2022). Higher indicates a stronger relationship with subsequent job performance.
Job simulations are a digital form of the work sample test. Estimates vary between meta-analyses - Schmidt & Hunter (1998) reported .51 for structured interviews and .54 for work samples - and the correction methods behind them remain debated. Treat these as an ordering, not as precise measurements.
Two points are worth drawing out.
First, a well-structured interview is a strong predictor - in these estimates, ahead of work samples. Anyone claiming simulations simply out-predict interviews is overstating the evidence, and a buyer with an industrial-organizational psychologist on staff will say so.
Second, and more revealingly, the gap inside the interview category dwarfs the gap between categories. Structured interviews predict at roughly twice the rate of unstructured ones. What structure does is constrain the conversation toward job-relevant content and hold the criteria steady - in other words, it moves an interview closer to sampling behavior and further from evaluating rapport.
Read that way, the research supports the distinction this article makes rather than contradicting it. Methods perform better the more directly they engage with the work, and a simulation takes that principle to its conclusion by removing the verbal intermediary altogether. It also explains why simulations tend to be more consistent in practice: a structured interview only performs at .42 when it is genuinely structured, which depends on interviewer training and discipline holding across every site and shift. A simulation's structure is a property of the instrument, not of the person administering it.
Section 09 Using both: a working sequence
Because the two methods produce different evidence, the productive arrangement is sequential rather than competitive. Establish capability where behavior can be observed; spend interview time on the questions that require a person to answer them.
Simulation, offered to every applicant
Capability is assessed before any filtering, on actions rather than on a document. Candidates who would not have survived a résumé screen but can clearly do the work are identified here.
Shortlist on demonstrated performance
Prediction scores and task-level results, compared against your own benchmark, decide who advances. This replaces the stage that previously relied on inference from a CV.
Structured interview with the shortlist
With capability established, interview time goes to intent: motivation, availability, expectations, and the candidate's own questions. The conversation can also reference what the candidate actually did in the simulation, which produces a far more specific discussion than a hypothetical ever does.
Decision on combined evidence
Observed performance and reported intent together. Where the two disagree, that disagreement is informative and worth examining rather than averaging away.
A useful side effect
Interviewing someone whose work you have already seen changes the conversation. Instead of “tell me about a time you handled a difficult customer,” a manager can ask why the candidate made the specific choice they made at a specific moment. The answer to that question is very difficult to rehearse.
Key terms
- Self-report evidence
- Information a candidate provides about their own behavior, attitudes, or history. Appropriate for intent and circumstances; weaker for capability, since it cannot be separated from recall and self-presentation.
- Behavioral evidence
- A record of what a person actually did in a given situation. What a simulation collects, and what a work sample test is designed to produce.
- Work sample test
- Any method that has a candidate perform a representative sample of the actual job. Job simulations are a digital, scalable form of one.
- Structured interview
- An interview in which every candidate receives the same questions, answers are rated against defined criteria with anchored examples, and interviewers score independently before discussion.
- Predictive validity
- The strength of the relationship between a selection method's results and later job performance, expressed as a correlation coefficient. Values around .30 to .40 are considered strong in this field.
Section 10 Common questions
- Do simulations replace interviews?
- Isn't a behavioral interview already asking about actions?
- Can candidates game a simulation the way they prepare for interviews?
- The research says structured interviews predict better than work samples. How do you square that?
- Is a simulation the same as an AI interview?
- Are simulations fairer than interviews?
- What about roles where talking is the job?