Coda, Rime’s new speech synthesis model for enterprise conversations at scale

19/05/2026

Rime introduces Coda, a text-to-speech model for real-time conversational agents that reproduces the rhythm, pauses and intonation of natural conversation, built for high-concurrency environments.

Coda, Rime’s new speech synthesis model for enterprise conversations at scale

Coda is Rime's new speech synthesis model, available in public beta since May 19, 2026. The company developed it to meet the needs of businesses managing large volumes of phone-based customer interactions, a segment where Rime already operates with several clients at scale.

Unlike its previous model, Arcana v3, Coda uses a jointly trained dual-decoder architecture: one focused on semantic understanding of the text and another on acoustic reconstruction. This design captures elements of natural conversation — such as rhythm, pauses and intonation contour — that tend to be lost in models built for monologues or pre-recorded audio content. The system processes audio in 80-millisecond frames, each encoded into 15 acoustic codes, and delivers the signal via streaming.

Rime has also developed its own inference infrastructure, called TIGERSTRIPE, which runs the model both in the cloud and on-premises. The company states that this proprietary implementation outperforms standard industry solutions. The greatest improvements over Arcana are seen in multilingual contexts.

In comparative quality evaluations against ElevenLabs — both its Flash and v3 models — Coda scores higher across all tested languages, with the largest margin in Spanish. Compared to Arcana, the most notable gains are in German, Spanish and French.

The model is available through Rime's dashboard and API.

Key points

  • Coda is Rime's new speech synthesis model for real-time enterprise conversational agents
  • Reproduces natural rhythm, pauses and intonation through a dual semantic and acoustic decoder
  • Outperforms models built for monologues or pre-recorded audio
  • Includes proprietary inference infrastructure, available in cloud and on-premises deployments
  • In quality comparisons against ElevenLabs, scores higher across all languages tested, with the largest gap in Spanish
  • Compared to its previous model Arcana, the greatest gains are in German, Spanish and French

Related AI

Rime

Voice synthesis models

Company specialising in text-to-speech models for enterprise voice agents. Reproduces the breathing, pauses and rhythm of real speech, with pronunciation control that requires no model ...

Lastest news

★★★★★
Rate us on Google
This website uses technical, personalization and analysis cookies, both our own and from third parties, to facilitate anonymous browsing and analyze website usage statistics. We consider that if you continue browsing, you accept their use.