No avatar, no calendar, and a conversation that pushes further when an answer stays vague. Here's how voice screening works, what the score actually means, and where it belongs in a hiring funnel.
AI voice interview: how voice-based screening works
An AI voice interview is a screening conversation where the candidate speaks with a voice-based system instead of typing answers or facing a camera. There's no avatar and no face on the screen. Just a microphone and a conversation that changes direction based on what the candidate says. The system transcribes the answers, matches them against the competencies the role requires, and leaves the hiring team a structured report.
That's the short version. The longer one is more interesting, because voice sits between two formats most teams already know: the live avatar interview on one end, the pre-recorded one-way video interview on the other.
Why voice became its own category
For a while, "AI interview" meant one of two things. Either the candidate recorded video answers to fixed questions, or they talked to a rendered avatar.
Both work. Both come with things worth thinking through.
The avatar format is technically impressive and some candidates genuinely enjoy it. For others, having a face on screen pulls part of their attention away from the conversation, and their energy goes to the image in front of them rather than to their answer. There's a preference on the employer side too. However advanced the technology gets, some companies aren't ready to put an AI avatar in front of a candidate, usually for employer brand reasons.
There's also a question of scale. A real-time avatar experience runs on heavier infrastructure. On a 40-person shortlist that difference isn't noticeable; at 800 candidates a quarter it becomes a line you plan for. Voice gives you the same depth of conversation at the widest part of the funnel, in a structure that scales more easily.
Voice removes the face and keeps the conversation. What's left resembles a call that never gets scheduled, never runs late, and can happen at 11pm on a Sunday if that's when the candidate is free.
How the conversation works in practice
The thing to understand is that a voice interview is dynamic. It isn't a list of questions read aloud.
A typical flow:
- The candidate gets a link, opens it in a browser, checks their mic, and starts whenever they want.
- The system asks an opening question tied to the role.
- The candidate answers out loud, and speech is transcribed almost in real time.
- Based on that answer, the system decides what to ask next. A thin answer gets a clarifying question: something like "which part of that project were you responsible for exactly?" An answer with a concrete example gets a follow-up that goes one level deeper.
- The loop continues for the length you've scoped, usually 8 to 20 minutes.
- The transcript, the audio, and a structured evaluation land in your hiring dashboard.
Step 4 is where the value sits. When a candidate says "I ran the migration," a static form takes that at face value. A dynamic interview asks how big the team was, what difficulties came up along the way, and what they'd do differently today. That's the difference between a claim and evidence.
What scoring means
Most of the confusion lives here, so let's be clear about it.
A voice interview system produces a role-fit score. It does that by mapping what the candidate said against the competencies defined for the position and rating each one with the supporting quote attached. Communication clarity, depth of relevant experience, problem-solving approach, motivation for the role.
The purpose of that score is to put a large candidate pool into a readable order and show why each person sits where they do. The final call belongs to the hiring team; the system produces the evidence behind the decision, not the decision itself.
The practical benefit follows from that. A recruiter looking at 300 applications can read 300 structured summaries in the time 20 phone screens used to take, and the candidates moving to the next step surface much faster.
There's a gain on the consistency side too, and it's worth stating plainly rather than in marketing language. Structured interviews predict job performance better than unstructured ones. That was true long before AI. What a voice system adds is that everyone is genuinely evaluated through the same structure, in the same way. Friday afternoon fatigue and the impression left by the previous candidate don't leak into the assessment.
Maintaining that consistency is something you stay on top of, not something you assume. Evaluation criteria are defined for the position up front, every score is reported with the quote that supports it, and results are reviewed regularly. Which competency produced which score is always visible, so the evaluation doesn't become a closed box.
Where it fits in the funnel
Most teams place voice screening right after the application, sometimes after a skills assessment.
The problem it solves sits at the widest point of the funnel. You post a role, 400 people apply, and roughly 120 look plausible on paper. Finding room for 120 phone screens is beyond most hiring calendars. So the list gets scanned quickly, familiar company and university names stand out, and good candidates without a recognizable logo on their CV can slip past.
Voice screening steps in exactly there: where the volume is too high for a human conversation and the decision is too important for a CV scan.
There's no ceiling on seniority. What matters is how clearly the position is defined. The more detail you give on the role's priorities, the competencies you're looking for, and what the team expects, the more accurate both the questions and the evaluation become. With that clarity in place, the format works for senior roles too. In practice most teams still reach for it where volume is highest, since at executive level the mutual conversation carries more weight, so voice usually sits as a first step ahead of the human interview. And when the candidate pool is already small, going straight to a conversation is often quicker than adding a step.
Candidate experience is the underrated part
Screening tools are bought by employers and used by candidates. That makes the candidate-side experience as decisive as the efficiency gain on the employer side.
A few concrete reasons the voice format lands well with candidates:
- No scheduling. On its own, this removes the most common drop-off point in early-stage hiring.
- No setup. It runs in the browser.
- No camera. A candidate who isn't comfortable on video, or who is joining from their car during a lunch break, isn't disadvantaged by their background.
- Without an evaluator watching them, many candidates say they speak more freely than they would in a first-contact interview.
- A candidate who feels the process treated them fairly will say so, whatever the outcome.
The ATS question
Conversations about interview automation eventually run into the same sentence: "we already use an ATS and we're not moving to a new system."
You don't have to. Voice interviewing is built to connect to your applicant tracking system, or to run on its own without connecting to anything. There are two ways to use it and neither requires an implementation project.
- Without a connection. You add candidates directly in the Coensio panel and send interview invitations from there. For trying a single role quickly, or for a team that doesn't use an ATS, this is the most practical route.
- With an ATS connection. Once you connect your applicant tracking system to Coensio, candidates who apply to a given position flow through automatically, matched to the role. No moving lists by hand and no separate setup for each posting.
Either way, invitations, completion tracking, and reports sit together in the Coensio panel. You don't have to decide up front: plenty of teams start without a connection and switch the ATS integration on later as volume grows.
In short, sending an interview link doesn't require an implementation project, a data migration, or a separate lift from your IT team.
Frequently asked questions
Is an AI voice interview the same as a phone screen?Close, but not the same. A phone screen implies a call and a scheduled slot. Voice interviews usually run in the browser and asynchronously, so the candidate starts when they choose rather than at a set time.
Can the AI ask follow-up questions?
Yes. That's the core difference from pre-recorded video interviews. The next question is shaped by the previous answer.
Does it replace the recruiter?
No. It produces evidence and a score. The hiring team decides who advances and who gets hired.
How long does a voice interview take?
Usually 8 to 20 minutes, depending on how many competencies you're measuring.
Is candidate data secure?
Yes. At Coensio, audio recordings and transcripts are processed in line with GDPR and KVKK requirements, with access permissions, retention periods, and data location defined at the outset.
Coensio's product family already covered technical assessments, case studies, AI Interview Assistant, and TalentRadar. AI Voice Interview is the newest addition: a dynamic voice interview that asks follow-up questions, built for teams that would rather not use an avatar or want a format that scales across a wider funnel. It scores role fit, produces a detailed report per candidate, and works alongside the ATS you already use.
