Why We Built Baton
August 20, 2026
Every voice AI on the market is built for one user. Real conversations have more.
Put a frontier voice model into a real meeting — four people, overlapping speech, someone thinking mid-sentence — and it does the same three things every time. It talks over people. It fills the thinking pauses. And it answers questions that were never meant for it.
Not because it is unintelligent. Because it has no idea whose turn it is.
That is the problem Baton solves — the turn-taking engine we built to decide when an AI speaks, and when it stays out of the way.
Expensive Answers, Terrible Timing
We noticed this the first time we put an AI into a real customer meeting. The answers were good — and metered by the second, whether or not any of them were wanted. The timing was awful. It would interrupt a VP mid-thought to offer a correct fact nobody had asked for, then sit silent when someone actually turned to it.
That is not a knowledge problem or a voice-quality problem. It is a timing problem — and specifically a group timing problem. In a one-on-one, every pause is an invitation: there is nobody else whose turn it could be. Add a third and fourth person and that assumption breaks completely. Most of the silences in a room belong to somebody else.
When we measured it, the gap was not subtle.
Across the same 51 scripted meetings, the frontier realtime models spent 69 to 73 seconds talking over a human being. Baton spent 1.5 seconds.
Races Are Lost at the Hand-Off
A relay is won by the best team, not the fastest runner. And it is often won and lost at the hand-off — four quicker sprinters lose to a team that passes cleanly, because the exchange is the only moment where two runners have to agree, in motion, about who is carrying the thing.
A conversation works the same way. The floor is the baton. It moves from person to person, and the exchange is so quick and so practised that people barely notice doing it — a glance, a falling tone, a half-second of silence that means go ahead. Get it right and nobody notices. Get it wrong and everyone does.
That is exactly what an AI keeps failing at. It is a fast runner that grabs the baton whenever it sees an opening, and drops it when it is actually handed over.
We named it Baton because our AI is about the team — the interaction, and improving how AI and humans collaborate. The goal was never to make your AI faster. It is to have it take the baton at the right moment, and carry it only as far as its leg of the race.
What Baton Actually Does
Baton watches the conversation about twenty times a second and does two jobs.
The first is gating: is the floor open, and is it open to us? Most of the time the answer is no, and the correct behaviour is silence — the hardest turn in any conversation is the one you do not take.
The second is signal. Baton hands your LLM the things it cannot work out on its own: has this person actually finished talking, or are they just thinking? Were they talking to me, or to someone else in the room? What kind of interaction is this — a question, a side-bar, an interruption, a backchannel? Your model gets social context instead of a raw transcript, so it can respond like a participant rather than a search box.
Nobody was measuring any of this, so we had to build the benchmark before we could build the model. Every voice AI benchmark we could find assumed two participants — and on a benchmark like that, holding the floor is a penalty: staying quiet reads as failing to respond. Today Baton leads the polyadic benchmark we built to measure group turn-taking, and places second of eleven on the third-party end-of-turn benchmark we did not build.
It does not replace your stack
Baton is a small policy, not another large model. Keep your transcription, keep your model, keep your voice — it sits between them and decides the moment.
That shape is also why it is cheap, and this is the number that is actually measured in multiples. A realtime model is always on: it is an LLM consuming every second of your conversation, reasoning about it and billing by the second, whether or not it has anything to say. Baton wakes your LLM when the floor has actually opened. The conversation still gets there — with the context it needs — just at the moments that call for a response, rather than every second in between. Over the same audio that was 842 policy decisions and 4 LLM calls — around ~39× less than leaving a frontier model always on and metered.
And it is not only for meetings
Meetings are simply the hardest case we could find, and the one we live in. But the problem is not meetings — it is real-world interaction. Anywhere more than two people share a space, turns get taken: a customer call where someone puts you on hold to ask a colleague, a support desk with a supervisor listening in, a kiosk in a busy room, a car with three passengers, a clinician and a patient and a family member. Interruptions, side-bars, questions aimed at someone else.
And the rules are not the same in each. Human communication is remarkably adaptable — it bends to the group, the situation and the norms of the room. An interruption that is normal in a stand-up is rude in a clinic. A two-second pause means go ahead in one room and I am still thinking in another. That adaptability is exactly what a fixed timeout can never capture, and exactly what Baton is built to read.
Baton already runs inside Pulse, our real-time meeting assistant. As a standalone service it is coming soon, for teams building voice AI that has to live in a room with people.
See the benchmarks, and hear the same meeting with different agents dropped into it, on the Baton page →
