Why a Group-Meeting Company Runs a One-on-One Benchmark
August 30, 2026
End-of-turn detection is the floor under everything Baton does — and it is the one part of it a third party can check.
We build for meetings. Four people, five people, a room and two dial-ins, everybody talking over each other in the good way. So it is a fair question why one of the benchmarks we care most about — and the one LiveKit just published our result on — is made entirely of two-party customer-service phone calls.
The short answer: end-of-turn detection is the floor that everything else in Baton stands on, and eot-bench is a rigorous, public, third-party measurement of exactly that floor. We did not build it. We cannot tune it. That is the point.
Turn-taking asks two questions
When an AI is in a conversation, deciding whether to speak is really two decisions stacked on top of each other:
Has this person actually finished talking?
And if they have — is the next turn mine?
In a one-on-one, the second question is close to trivial. There are two of you. If they stopped, it is your turn; silence is the failure mode. That is why a conventional voice agent can get away with a voice-activity detector and a timeout: wait 700 milliseconds, hear nothing, start talking.
Add a third and fourth person and the second question becomes the hard one. Someone finished a sentence, and the right next speaker is almost never the AI. It is the person who was asked. It is the person who has been waiting to jump in. It is nobody, because the room is thinking. That is the problem Baton exists to solve, and it is why we had to build a polyadic benchmark to measure it — nothing that existed measured turn-taking beyond two participants.
The second question rests on the first
Here is the part that is easy to miss. All of that group reasoning — who was addressed, who is owed the floor, whether this is a side-bar — is worth nothing if you get the first question wrong.
If your system cannot tell a thinking pause from a finished thought, it does not matter how elegantly it decides whose turn is next. It will make that beautiful decision at the wrong moment, in the middle of somebody's sentence. Perfect floor management on top of a bad end-of-turn signal is just a well-reasoned interruption.
So end-of-turn detection is not a component we bolted onto Baton. It is Baton v1 — the first axis of the engine to ship, the one that answers has this speaker finished, continuously, as a probability, twenty times a second. The broader engine reasons about when to speak, interrupt, wait, or think more deeply. None of it works on a shaky floor.
Baton answers that question by fusing two signals that fail in different places. An acoustic model listens to the trailing seconds of audio — prosody, the shape of a falling intonation, the difference between a breath and a stop. A text model reads only the words a streaming transcriber has actually emitted by that instant, never the completed turn, because handing it the finished sentence would leak the future and flatter every number we report. When the text model is confident, it gets a vote; when it is genuinely undecided — about three moments in ten — the audio call stands on its own.
A false cut costs more in a group
Both benchmarks measure the same failure, but a group changes what it costs.
In a one-on-one, cutting someone off is rude and recoverable. They stop, you stop, they finish, the conversation moves on. Two people can repair a collision quickly because there are only two of them.
In a meeting, an early cut does more damage than an interruption. The AI does not just talk over the person speaking — it takes the floor from whoever was about to get it. A hand-off that was already in motion between two humans gets destroyed by a third voice that was never owed the turn, and the repair now involves everybody in the room. Worse, most of the time the correct behaviour was silence anyway. In our own polyadic evaluation set, eleven of fifteen meetings never open the floor to the agent at all. The frontier realtime models still fire in those meetings. The value is in the turn not taken.
That asymmetry decides which end of the trade-off we tune. Every end-of-turn model trades false cut-offs against latency: commit early and you interrupt more, wait longer and you feel sluggish. A dyadic product is pushed toward speed, because in a 1:1 a slow response is the thing users notice. We are pushed the other way, which is why we lead with false cut-offs rather than with milliseconds.
Where Baton lands
eot-bench scores every model on how often it cuts a speaker off who was not finished, inside a given latency budget, and on how quickly it commits when the turn really has ended. On LiveKit's board Baton places second of thirteen on both axes, behind only the model LiveKit's own authors built.
The axis a meeting feels first is dead air: how long everyone waits after a speaker has actually finished. Hold every model to the same 5% false-cut-off budget and rank on that:
The board's other axis holds latency to 300 ms and ranks by how often each model cuts a speaker off. Baton is second there too:
Data and second chart from LiveKit's leaderboard, read September 27, 2026. Click either to open the live board.
| Model | False cuts @300 ms | @600 ms | Latency @5% |
|---|---|---|---|
| LiveKit Turn Detector v1 | 9.9% | 4.5% | 543 ms |
| JoinIn AI Baton | 12.3% | 4.8% | 577 ms |
| Deepgram Flux | 12.9% | 9.9% | 1151 ms |
| ultraVAD | 27.7% | 11.9% | 899 ms |
| LiveKit v1-mini | 27.8% | 12.1% | 1070 ms |
| Smart Turn v3.2 | 35.2% | 14.8% | 1051 ms |
| Soniox | n/a | 5.5% | 647 ms |
| OpenAI GPT Realtime 2 | n/a | n/a | 1143 ms |
| VAD baseline | 55.6% | 21.7% | 1600 ms |
English rows as LiveKit publishes them in livekit/eot-bench: false cut-offs at 300 ms and 600 ms latency budgets, mean latency at a 5% cut-off budget. “n/a” means no operating point reached that budget. Baton's latency at the 10% budget is 350 ms. The prediction artifacts behind every row are committed in the harness, so any of this can be re-scored.
Three caveats we would rather state than have someone find.
We predicted this number before LiveKit measured it. The shipped model was chosen by lowest seed index rather than by benchmark score. Our own four-seed mean was 12.5%, our best single seed 12.2%, and we quoted the mean; LiveKit's independent run landed at 12.3%, inside that range. Every free choice in the scheme was written down and hashed before the confirmation model existed, so the headline is a prediction we kept rather than the best of forty arms.
The margin over Deepgram Flux at 300 ms is inside our own seed noise. Twelve-three against twelve-nine is not a real separation. The separation is on the other columns — 4.8% against 9.9% at the 600 ms budget, and 577 ms against 1151 ms on latency.
It is English-grade in English only. The acoustic half is built on whisper-small.en. Across all fourteen languages the mean is AUC 0.8762 and 25.1% false cuts — 12.5% in English against a 26.1% non-English mean. It works everywhere we measured it. It is not English-grade anywhere else.
What eot-bench cannot tell you
eot-bench is dyadic by construction: 367 of its 400 items have exactly two roles, and the median turn is a little over nine seconds of customer-service phone call. It is an excellent measurement of one thing and structurally silent about the other. It cannot tell you whether an agent knows a question was meant for someone else. It cannot tell you what happens when two people start at once. It cannot score the meeting where the right answer for sixty-five straight seconds is to say nothing, because in a conversation with two participants that behaviour is indistinguishable from being broken.
On a dyadic benchmark, holding the floor is actively penalised. We show a row we lose on our scoreboard for exactly that reason.
So we run both. eot-bench is the axis a third party can check, on a harness we do not control, against a field that includes the people who do this for a living — and it keeps us honest about the floor. Full-Duplex Polyadic, the benchmark we built and are publishing, covers the axis nobody else measures, because every voice benchmark we could find assumed two participants. A benchmark you build and then win carries no independent weight on its own. Paired with a third-party result on the part that is comparable, it means something.
More on how we got here in Why We Built Baton.
Trying it
Baton runs inside Pulse today, and is available as a standalone service in beta — take the models and run them in your own process, or call the hosted API. Both run the identical pipeline, so the choice is operational rather than about accuracy. It sits between your transcription and your model: bring your own LLM, keep your voice, keep your stack.
Beta access: request an API key at hello@joinin.ai. The full benchmark detail, the interactive demo, and the cost comparison live on the Baton page.



