Real-time TTS for cluster-meet
R&D on a text-to-speech stack fast enough to speak inside a live meeting, where latency is the whole product.

cluster-meet is our in-house meeting platform. This was the research effort behind giving it a voice: text-to-speech that speaks inside the live call instead of rendering audio afterwards.
Doing it in real time changes what counts as good. Offline TTS is judged on how natural the result sounds, and you can take as long as you like getting there. In a meeting, the audio has to start before the moment it belongs to has gone past. That makes time-to-first-audio the number you’re really designing against, plus holding a steady stream once you’ve started. Naturalness is what you optimise inside that budget.
The work was exploratory. Which approaches hold up at conversational latency, and what the quality actually costs you when you get there.
innoscripta SE · 2026