Reka has introduced Rho-1, a 19 billion parameter omni-reasoning model that collapses the traditional multimodal stack into a single neural network. Trained from scratch, this model handles text, images, video, and robot actions within one unified architecture — a direct alternative to the fragmented pipelines that currently dominate agentic AI. The research preview demonstrates a five-turn session where Rho-1 draws a lighthouse, generates a bounding box, animates the scene, edits the video, and answers questions about the changes, all without calling any external tool or specialist model.
The Problem with Today’s Multimodal Pipelines
Most systems that process multiple modalities rely on a central planner delegating tasks to separate image, video, and object detection models. Each handoff introduces latency, and each downstream model operates with a narrow, context-stripped view of the original input. Reka argues that this architecture is inherently inefficient — not just in speed but in consistency, because spatial and temporal information is lost every time the request is reformatted for a different model.
By treating everything — text, visual features, robotic actions — as tokens within a single context window, the model maintains a coherent internal representation throughout any interaction.
Rho-1 eliminates those handoffs completely. By treating everything — text, visual features, robotic actions — as tokens within a single context window, the model maintains a coherent internal representation throughout any interaction. The research team points to a simple demonstration: a user asked the model to draw a lighthouse, then put a box around it, then animate a drone flying toward it, then add a snowstorm, and finally explain what changed between the videos. All five turns used the same model, the same KV cache, and produced consistent lighting, geometry, and camera motion.
Architecture: Two Expert Streams, One Shared State
Rho-1’s design revolves around two native token formats. Discrete tokens carry text, symbolic reasoning, and high-level commands. Continuous tokens carry image latents, video frames, robot actions, and proprioception. Every transformer block contains two expert weight streams — one for understanding (processing language and visual parsing) and one for generation (denoising latents into pixels). Crucially, both streams share attention mechanisms and operate over the same KV cache.
When the model needs to produce pixels, the understanding stream emits a discrete handoff token. The generation stream then renders the output from the full accumulated state. Training combines next-token prediction for discrete sequences with flow matching for continuous outputs. This approach yields practical efficiencies: bounding boxes are emitted as coordinate tokens directly from the same network that generated the image, and video frames reuse the original image representation rather than re-encoding from scratch.
How Does Rho-1 Handle Video, Text, and Robot Actions in One Model?
Rho-1 uses a single shared KV cache that accumulates state across all modalities. The understanding stream processes incoming text, image latents, or robot proprioception as tokens. The generation stream produces outputs — whether words, pixels, or action commands — by attending to that accumulated state. This means that generating bounding boxes, animating video, or emitting robotic joint controls all happen within the same attention space, without any specialized submodels or external tools.
Performance: Base Model vs Distilled Variant
Reka reports that the base Rho-1 model generates video at a median of 0.79x real-time, with a watchable stream starting approximately six seconds into generation. In a measured comparison against an illustrative multi-agent pipeline (13.8 seconds to first clip), the base model returned its first clip in 7.0 seconds. A distilled variant cuts the denoising process from 99 steps to just 8, producing a 5.3-second clip in roughly one second, with minimal quality loss according to Reka’s internal tests. The same distilled model matched the fastest dedicated image models in image generation latency and was the quickest model tested to the first text token.
These figures come from vendor-run benchmarks, not independent evaluations, but they highlight a clear direction: reducing the number of denoising passes dramatically lowers latency without sacrificing output fidelity. The distilled variant is particularly interesting for real-time applications like gaming, simulation, or live video editing where speed is critical.
Continuous Rollouts and Robotics
Rho-1 is designed for streaming generation. New instructions enter through the understanding stream and update the model’s state mid-rollout, allowing for continuous video generation that can be steered in real time. Reka demonstrates this by forking a single opening scene into two continuations — “bank left” and “bank right” — without restarting the generation process.
For robotics, actions and future frames decode from the same latent state. In a LIBERO simulation episode, Rho-1 emits seven action channels as continuous tokens. To scale beyond the limited teleoperation logs that constrain most robot learning pipelines, Reka pairs Rho-1 with its Inverse Dynamics Model, which infers control signals from raw video. This creates a loop: the world model generates possible futures, the inverse dynamics model translates those futures into control commands, and the system can self-supervise large-scale robotic training without manual data collection.
How Rho-1 Compares to Other Multimodal Models
| Feature | Reka Rho-1 | ByteDance BAGEL | BAAI Emu3.5 | Google Genie 3 |
|---|---|---|---|---|
| Developer | Reka | ByteDance Seed | BAAI | Google DeepMind |
| Parameters | 19B | 14B total, 7B active (MoT) | 34B | Not disclosed |
| Inputs | Text, image, video, actions, proprioception | Text, image | Interleaved text and image | Text prompt, navigation inputs |
| Outputs | Text, image, video, actions, proprioception | Text, image | Interleaved text and image | Interactive video world |
| Native video generation | Yes, capped at 672×384 | No | No (image frames, not native clips) | Yes, 720p at 24 fps |
| Real-time steering | Yes, continuous rollouts | No | No | Yes, promptable world events |
| Robot actions | Native continuous action tokens | No | Embodied manipulation demos | Takes navigation actions, does not emit them |
| Open weights | No | Yes, Apache 2.0 | Yes, Apache 2.0 | No |
| Access today | Research preview via email | Hugging Face, GitHub | Hugging Face, emu.world app | Project Genie for Google AI Ultra subscribers |
| Source | Reka blog | GitHub | Hugging Face | DeepMind blog |
Genie 3 access per Project Genie; Emu3.5 parameter count per its Hugging Face model card.
Training Footprint and Current Limits
Rho-1 was trained on 320 H100 GPUs for approximately three months. The model’s native video generation is currently capped at 672×384 pixels, which limits its applicability for high-resolution applications. Reka has not released public weights, an API, or pricing — the offering is limited to a research preview accessible by requesting access via email.
The decision to keep weights closed, even as competitors like ByteDance Seed and BAAI release their models under Apache 2.0 licenses, may slow community adoption and independent benchmarking. For companies evaluating Rho-1, the lack of an API also means they must manage their own infrastructure, though the model’s relatively modest 19B parameter count makes it deployable on at least one high-end GPU.
Implications for Agentic AI and Robotics
Rho-1 represents a significant step toward unified multimodal intelligence. By removing the handoff between specialist models, it reduces latency, preserves context across modalities, and opens the door to truly interactive world models that can reason, generate, and act within a single loop. The robotics applications are particularly promising: combining world model rollouts with inverse dynamics for self-supervised learning could reduce the reliance on expensive teleoperation data, potentially accelerating progress in manipulation and navigation tasks.
However, the field is moving fast. Google’s Genie 3 already offers 720p video at 24 fps with real-time steering, albeit without the ability to emit robot actions. BAAI’s Emu3.5 and ByteDance’s BAGEL provide open-weight alternatives for text and image generation, even if they lack native video and action outputs. Rho-1’s unique advantage — a single model that directly outputs actions — may define a new category if Reka can match the resolution and framerate of competitors while maintaining the multimodal coherence that pipelines struggle to achieve.
The research preview period will test whether the architectural innovations translate into reliable, production-ready performance. If the distilled variant can maintain quality at near-real-time speeds, and if robotics teams can effectively pair it with the Inverse Dynamics Model, Rho-1 could become a foundational component for next-generation embodied AI systems. For now, it stands as a compelling proof that the multimodal stack can indeed be collapsed — and that the next frontier in AI may be not just bigger, but genuinely unified.
- How does Rho-1 handle video, text, and robot actions in one model?Rho-1 uses a single shared KV cache that accumulates state across all modalities, with two expert streams for understanding and generation.
- What is the problem with current multimodal pipelines?Current systems rely on a central planner delegating tasks to separate models, introducing latency and losing context.
- What are the key architectural features of Rho-1?Rho-1 has two native token formats (discrete and continuous) and two expert weight streams that share attention mechanisms and operate over the same KV cache.
- How does Rho-1 compare to Google's Genie 3?Genie 3 offers high-resolution video but cannot emit robot actions, while Rho-1 directly outputs actions in a unified model.
- What are the potential applications of Rho-1 in robotics?Rho-1 could reduce reliance on expensive teleoperation data by combining world model rollouts with inverse dynamics for self-supervised learning.