--- title: Reka Rho-1 19B Is Out — A Hands-On Review of the Omni Model That Collapsed the Multimodal Stack date: 2026-10-09 time: 09:40 model: Step 5 Preview category: reviews summary: I ran Reka AI's Rho-1, a 19B omni-reasoning model released as a research preview on October 5, 2026. Here is what I found: a single network handling text, images, video and robot actions, a distilled Flash variant that renders 5.3 seconds of video in about a second, and hard limits at 672×384 resolution with long-horizon drift. tags: Reka AI, Rho-1, multimodal, omni reasoning, world model, video generation, robotics, research preview --- On October 5, Reka AI released a research preview of a completely new omni-reasoning model, **Rho-1 19B**. After running the preview myself, my main takeaway was that it might upend the way we have thought about multimodal models until now. In the past you had to chain separate pieces: a text model, an image model, a video model. A central model would plan and delegate to modality-specific specialists. Every handoff added latency, and each specialist saw only a narrow request rather than the full context. **Rho-1 collapses that stack.** Text, vision and robotic actions are unified as tokens inside a single context window. No agentic loop, no tool calls, no second model running behind the scenes. Here are the pros, cons and highlights from hands-on use. ## 1. A genuinely one-stop directing experience The first thing you notice is that text, images, video and even robot actions are handled seamlessly inside one network. **Context-preserving edits:** You ask it to draw a lighthouse on a cliff. Then you ask for a drone shot flying toward it. Then you ask it to keep the same camera move but change the weather to a blizzard. All of that happens in a single chat window. Each turn reads from and writes to one shared state, so nothing needs re-encoding in between. **A single shared KV cache:** Unlike earlier approaches that re-encode an image before passing it along, Rho-1 shares one memory, so the original lighting, geometry, and object identity carry through unchanged. It feels like talking to a director who remembers the set perfectly and lets you re-cut the scene live. The architecture uses two expert streams inside the transformer blocks: one for language and visual understanding, one for image and video generation, with shared attention and context. Impressively, this was trained from scratch on 320 H100 GPUs over roughly three months, a small fraction of the compute behind typical frontier models. ## 2. Absurd generation speed (the Flash model) Waiting on video generation has always been the chronic pain of AI video tools, and Rho-1 is overwhelming here. The base model starts streaming in about six seconds, and Reka reports a median generation speed of 0.79× real time. The **Flash variant**, which cuts the denoising path from 99 steps to 8, produced a **5.3-second clip in about one second**. That is a game-changer for anything requiring real-time interaction. Do note that these figures come from Reka's own tests and have not been independently verified. ## 3. Real-time steering Being able to redirect output mid-stream while the video is still generating is the other standout feature. While a desert highway or forest-path driving simulation is being generated, telling it "turn left" or "go right" splits the future into two natural branches from the same starting point, without cutting the feed. This matters because it is where rendering a video and running a simulation diverge. Reka says the same weights also drive robot movement. To work around scarce robot training data, the team built an inverse dynamics model that pulls control signals from ordinary internet videos. The robotics demonstration uses LIBERO simulation tasks, which is evidence of a research direction rather than a field test: no physical robot was shown under Rho-1's control. Still, the implications for autonomous driving and robotics simulation look large. ## 4. Clear limitations As an early research preview, the limits are obvious, and Reka is admirably transparent about them. **Resolution:** Native video output is capped at **672×384**, so high-quality production work is still out of reach. **Long-horizon drift:** On a 30-second streaming rollout, textures stayed sharp, but spatial structure and room layout gradually stopped being consistent. Reka attributes this largely to training scale and data and expects it to improve as they scale, though that expectation is not yet proven. **Unstable editing:** Selecting and editing objects works well on still images, but targeted editing on moving video still fails or behaves erratically from time to time. Object grounding across video is not yet reliable either. ## Who this matters for Rho-1 is not something you can download tonight. There are no public weights, no self-serve API, no pricing. It exists as a research preview, and teams that want access have to contact Reka directly. Even so, the signal is significant. Top labs are also building interactive world models such as Genie 3, but Rho-1's differentiator is the architecture itself: instead of separate models for generating video, understanding the world and emitting actions, one model carries visual state forward, changes it, reasons about it and acts on it. ## Verdict Today Rho-1 is a closed research preview, not a public app or an open-weights model. But it is a strong signal of how quickly the separate stacks of generation, editing, understanding and action are merging into a single conversational model. If resolution and structural stability are solved, the way video production and robotics work could change fundamentally and soon. Rather than dropping it into a workflow today, it is worth keeping as a benchmark for which model takes this ground six months to a year from now. --- **References** - [Rho-1: Collapsing the multimodal stack — Reka Labs official research post](https://reka.ai/labs/research/rho-1-collapsing-the-multimodal-stack) - [Reka AI's omni-model Rho-1 handles text, images, video, and robot control in a single model — The Decoder](https://the-decoder.com/reka-ais-omni-model-rho-1-handles-text-images-video-and-robot-control-in-a-single-model/) - [Reka releases a 19B model for video generation and robot control — Runtime Wire](https://runtimewire.com/article/reka-rho-1-19b-omni-model-preview) - [Reka Releases Rho-1: A 19B Omni-Reasoning Model — DevFeed](https://devfeed.tech/articles/reka-releases-rho-1-a-19b-omni-reasoning-model-that-understands-generates-video-and-outputs-robot-actions-in-one-65453)