Thinking Machines presents an AI model that processes audio, video and text in real time

11/05/2026

Thinking Machines Lab has published a research preview of TML-Interaction-Small, an interaction model designed to collaborate with the user in real time, without waiting for conversation turns.

Thinking Machines presents an AI model that processes audio, video and text in real time

Unlike current language models, which work in turns —the user writes or speaks, the model responds, and the cycle repeats—, TML-Interaction-Small continuously processes audio, video and text while generating responses. This architecture, which Thinking Machines calls Interaction Models, allows the system to interrupt, pause, react to visual cues or speak at the same time as the user, in a way similar to how a conversation between people works.

The model works with 200-millisecond chunks of simultaneous input and output, eliminating the need to artificially detect when the user has finished speaking before responding. This enables capabilities that current systems cannot offer without additional components: simultaneous translation, real-time pronunciation correction or spontaneous interruptions depending on context.

The most distinctive capability is visual proactivity: reacting to what happens on screen or camera without the user saying anything. In internal tests of exercise repetition counting and temporal action localization in video, OpenAI and Google's real-time models scored near zero, while TML-Interaction-Small completed the tasks significantly.

The entire system, including the audio and image processing components, was trained from scratch jointly, without relying on pre-existing external encoders.

TML-Interaction-Small is a MoE (mixture of experts) model with 276 billion total parameters and 12 billion active per inference. For tasks requiring deeper reasoning, it delegates to a secondary model running asynchronously in the background, while the main model remains active in the conversation.

In public benchmarks provided by Thinking Machines, the model outperforms OpenAI and Google's real-time systems on interactivity metrics: it scores 77.8 on FD-bench v1.5 compared to 47.8 for GPT Realtime 2.0 in high-quality mode, and achieves a response latency of 0.40 seconds versus 1.18 seconds for OpenAI's model in minimal mode. On intelligence benchmarks such as Audio MultiChallenge, TML-Interaction-Small reaches 43.4%, above GPT Realtime 2.0's 37.6% in standard mode, though below the 48.5% that model achieves with extended reasoning enabled.

Thinking Machines plans to open a limited research preview in the coming months and release larger models throughout 2026.

Key points

  • Thinking Machines launches TML-Interaction-Small, its first real-time interaction model.
  • It processes audio, video and text simultaneously, without conversation turns.
  • It can interrupt, react to visual cues or speak while the user is speaking.
  • Its most distinctive capability: it responds to what happens on camera without being asked.
  • On interactivity benchmarks, it outperforms OpenAI and Google's real-time models.
  • On intelligence, it is competitive with GPT Realtime in standard mode.
  • For complex tasks, it delegates to a secondary model running in the background.
  • The entire system was trained from scratch, without external encoders.
  • A limited preview is planned soon, with larger models coming in 2026.

Videos

Related AI

Thinking Machines

AI research and development laboratory

AI research company focused on frontier models, multimodal systems and human-AI collaboration. Publishes open research and develops tools for model customisation and ...

Lastest news

★★★★★
Rate us on Google
This website uses technical, personalization and analysis cookies, both our own and from third parties, to facilitate anonymous browsing and analyze website usage statistics. We consider that if you continue browsing, you accept their use.