Baton smartly passes the floor to your AI.
The first turn-taking engine built for group conversation — and on the polyadic benchmark made to measure it, it beats both OpenAI and Google by 3×.
Drop a frontier voice model into a real meeting and it talks over people, fills thinking pauses, and can’t tell whether a question is for it or for someone else. That isn’t voice quality or knowledge — it’s timing in a group. Baton decides the one thing they all get wrong: when to hand the floor to your agent, and when to stay silent.
pol·y·ad·ic — involving more than two participants. A 1:1 is dyadic; a meeting is polyadic. Every voice AI on the market is built for the first case.
The floor goes to whoever’s turn it is. Your AI is just another voice in the room.
Good at two. In a different league at four.
In a 1:1 we trade blows with the frontier — real, but percentage points. Add a third and fourth person and the comparison stops being close. Same stack, same harness, deterministic scoring, no LLM judge. Full-Duplex Polyadic is a new benchmark of ours, releasing soon — nothing existing measured turn-taking beyond two participants.
Same stack, same harness, deterministic scoring, no LLM judge. OpenAI is gpt-realtime 2.1; Google is Gemini 3.1 Live. Hover the i marks for method notes and the row we lose.
Pending validationBenchmark results are pending independent validation and may change before product launch.
The hardest turn is the one you don’t take
Pick a meeting, then press play on each agent. You are hearing the identical room — same humans, same words, same timing — with only the agent swapped back in. The strip shows who held the floor: humans on top, the agent underneath. Red is the agent speaking over a human. In most of these the floor never opens, and the only correct behaviour is silence.
A child asks for the iPad
correct = stay silentEleven of the fifteen FDP-Showcase meetings never open the floor to the agent at all. In those, the frontier models still fire — Baton stays out.
Pending validationClips and verdicts come from our own harness; scoring is pending validation and may change before product launch.
Bring your own LLM. Keep your pipeline.
Baton decides whether to speak. Your stack decides everything else — audio arrives constantly, and your LLM only wakes when the floor actually opens.
- Any STTSoniox, Deepgram, Whisper, AssemblyAI — or whatever already runs in your stack.
- Any LLMGPT-5.1, Claude, Gemini, Llama, or your own fine-tune. Baton never replaces the brain.
- Any TTSElevenLabs, Cartesia, OpenAI, or Piper fully local.
- Any shapeCascade, hybrid, realtime-native, on-prem — no audio-native model required.
Baton colocates with your STT, wherever your audio already lands — no new hop, no new region, no audio crossing a boundary it wasn’t already crossing. Running it at the edge is an optional upgrade for the last few milliseconds, not a requirement.
Pending validationMeasured on our own runs. Figures are pending validation and may change before product launch.
Why it stays cheap
A realtime model is an LLM consuming every second of your meeting — reasoning about it, and billing for it, whether or not it has anything to say. Measured over the same 6.7 minutes of audio:
~39× less, on the same audio, at list pricing — and this figure is Baton plus estimated costs for the STT and LLM around it, not the policy alone. No GPU, no 7B audio model to serve, no per-second audio metering.
Pending validationCost figures are estimates from our own measurements at list pricing — pending validation and subject to change before product launch.
Request information about Baton
Building group voice AI, or evaluating turn-taking for your own agent? We’ll walk you through the benchmark, the integration, and how Baton drops into your existing STT/LLM/TTS stack.
