An AI-conducted interview is not a recording session. You speak, the system transcribes what you said, and the question it asks next is chosen based on your answer. Preparing for that loop is a genuinely different exercise from preparing for a camera and a countdown timer, and most candidates walk in with the wrong training.
The boundary, drawn early
A one-way video interview gives you a fixed list of prompts. You read a question, a timer starts, you record, and the file goes into a pile for someone to watch later. Nothing in that process reacts to you. If your answer is thin, the next prompt is still the next prompt.
An AI-conducted interview reacts. You say "I rebuilt the reporting process," and the system asks what the process looked like before you touched it. You mention a team, and it asks how many people reported to you. The questions are generated live from your own words, and a score is assembled while you are still talking.
The practical consequence matters. For a pre-recorded interview, rehearsing a polished ninety-second block per topic works well. Here it fails, because the second question is unscripted and your block does not answer it. You need material you can move around, not a script.
What these systems actually read
Almost everything of consequence runs through the transcript. The speech-to-text layer converts your answer into text, and the scoring happens on that text. Which means the model of "impressing the machine" that people imagine, warm tone, confident energy, is largely beside the point.
What tends to get measured:
- Topic coverage. Whether your answer contains the concepts the role expects. If the job is about stakeholder management and your answer never names a stakeholder, the coverage is low regardless of how good the story was.
- Structure and completeness. Whether an identifiable situation, action and result can be located in what you said. A single blob of narration scores worse than the same content with visible parts.
- Answer length and response latency. Very short answers read as low effort. Very long ones drift off topic. Long silences before you start can be logged.
- Speech pace and clarity. Some systems score delivery: words per minute, filler density, how cleanly the transcription came out.
That last one deserves a flag. Scoring a candidate on voice or affect is contested, and it is increasingly regulated. Illinois' Artificial Intelligence Video Interview Act requires disclosure and consent for AI analysis of interview video. New York City's Local Law 144 requires an independent bias audit for automated employment decision tools used on city candidates. The EU AI Act puts recruitment systems in its high-risk category, with transparency and human-oversight obligations attached. At least one large vendor dropped facial analysis from its product after sustained criticism.
My opinion, since this is where the format is weakest: transcript scoring is defensible and voice-affect scoring is not. A transcript is a record of what you claimed and can be checked against your CV. An inference about your enthusiasm from your pitch contour is a guess dressed up as a measurement, and it punishes accents, nerves and anyone whose baseline delivery is quiet.
Structure is the scoring surface
Everyone tells you to use STAR. With a human, it is a suggestion, because a decent interviewer will reconstruct your story even if you tell it out of order. With a scoring system, the parts have to be findable. If the result never appears as a distinct statement, it cannot be credited.
So say the parts out loud. "The situation was that our onboarding took eleven days. What I did was rewrite the verification step. The result was six days by March." That sounds slightly mechanical to a human ear. It is unambiguous to a parser, and it costs you nothing.
Signpost your answers verbally. The scoring layer is looking for components it can identify, and it will not infer a result that you only implied.
Say the noun
This is the single habit that changes scores the most. These systems do not infer. If you say "the main CRM," the transcript contains "the main CRM," not Salesforce. If you say "a large team," it contains "a large team," not eleven engineers across two sites.
Name the tool. Name the metric. Name the unit. Say "we cut refund processing from four days to one" rather than "we made refunds much faster." Say "Python and dbt" rather than "the usual data stack." Pronouns are the other trap: "we handled it, and then they signed off" gives the transcript almost nothing to attribute to you.
When the follow-up goes after a gap
The adaptive follow-up is the part people are unprepared for. It usually does one of three jobs, and it helps to know which one you are getting.
- Clarification. It did not parse something. "You mentioned a migration, which system was being migrated?" Answer plainly and move on.
- Depth. Your answer was credible and it wants specifics. This is good news. Give the number, the timeline, the constraint you were working under.
- The gap probe. Your answer skipped something the role cares about, and the system noticed. "You described the rollout. What happened when it failed?"
The third one is where candidates panic and repeat their first answer in different words. Repetition is the worst response available, because it adds no new content and the transcript shows you filling time. If you genuinely do not have the experience, say so and pivot to the nearest real thing: "I have not run a migration at that scale. The closest is a cutover of about forty accounts, and what I learned there was..." That produces new, checkable content. A bluff produces content a human reviewer can later catch.
Silence and the missing human
There is nobody to read your pause. A human interviewer sees you thinking and waits. A system sees a gap in the audio stream, and it may cut you off or log the delay.
Fill the thinking time out loud. "Let me pick the best example for that" is two seconds of speech that stops the clock running on empty air. Similarly, say when you are done. A clear "that is the example I would give" prevents the awkward hang where you wait for a reaction that never arrives.
Do a dry run with your actual microphone before the day. Headset mics beat laptop mics by a wide margin for transcription accuracy, and transcription accuracy is your score. If you have an accent the system may struggle with, slow down slightly and lean harder on naming things explicitly, since a well-transcribed keyword survives where a garbled clause does not.
When it breaks, and what you can ask for
Assume something will go wrong once. Have the recruiter's email open in another window before you start, and if the session freezes or the audio drops, write immediately with the timestamp and what happened. A message sent within minutes is treated very differently from one sent the next day.
You are also entitled to ask questions before you consent. Ask whether a human reviews the output before any rejection. Ask what happens to the recording and how long it is kept. If you have a stammer, a hearing condition or anything else that interacts badly with automated scoring, ask for an accommodation in writing, early, and ask for a human-led alternative. In the EU, automated decisions producing legal or similarly significant effects carry a right to human intervention under Article 22 of the GDPR, and a hiring rejection is a reasonable candidate for that. Most employers, faced with a polite written request, will offer a call instead. The ones that refuse have told you something useful.
An evening of preparation that actually works
Pick six stories from your own history. For each one, write down the numbers: dates, headcount, the before value, the after value, the tools by name. If your career history is scattered across old profiles, a LinkedIn-to-CV tool like Postulit will pull the timeline into one page so you are not reconstructing dates from memory at eleven at night.
Then practise the hard part. Record yourself giving one story in ninety seconds, then answer a follow-up you invented specifically to attack its weakest point. Do that six times. You are not rehearsing answers, you are rehearsing recovery, and recovery is the only thing the adaptive format really tests.