Skip to main content
World Today News
  • Home
  • News
  • World
  • Sport
  • Entertainment
  • Business
  • Health
  • Technology
Menu
  • Home
  • News
  • World
  • Sport
  • Entertainment
  • Business
  • Health
  • Technology
the-decoder.com: Reka AI Unveils Rho-1 Omni-Model for Text, Images, Video, and Robot Control

the-decoder.com: Reka AI Unveils Rho-1 Omni-Model for Text, Images, Video, and Robot Control

October 6, 2026 Rachel Kim – Technology Editor Technology

Reka AI Unveils Rho-1 Omni-Model for Text, Images, Video, and Robotics

Reka AI released a research preview of Rho-1, a 19-billion-parameter omni-model capable of processing and generating text, images, video, and robot control actions within a single neural network, the-decoder.com reported. Unlike traditional architectures that route tasks to separate specialized models, Rho-1 manages all modalities as tokens inside one shared context window without relying on tool calls or external systems.

The Tech TL;DR:

  • Unified Processing: Handles text, images, video, and robotic physical control signals natively within a 19-billion-parameter neural network context window.
  • Real-Time Adaptation: Generates continuous video in real time while responding dynamically to new instructions on the fly without system reboots.
  • Data Scaling: Trained on 320 H100 GPUs for roughly three months, utilizing a custom inverse dynamics model to translate standard internet videos into robotic control actions.

Single Context Architecture Eliminates Modular Routing

The core architectural distinction of Rho-1 lies in its unified sequence handling. Traditional multimodal deployments pipeline discrete models, passing outputs from a vision encoder to a language model, and then routing control variables to external execution scripts. Rho-1 bypasses this latency bottleneck by processing every modality as uniform tokens inside a single memory space. According to the-decoder.com, the exact same network weights responsible for predicting camera images are simultaneously utilized to drive physical robot movements.

To solve the historical scarcity of real-world robotic training sets, Reka AI engineered an inverse dynamics model. This internal module extracts actionable control signals directly from ordinary internet videos, expanding the dataset pool beyond specialized robotics lab recordings. The underlying model completed its training run across 320 H100 GPUs over a span of approximately three months.

Rho-1 Enables Real-Time Video Output and Adaptive Control

Production environments requiring low-latency video synthesis and adaptive robotic manipulation face significant hurdles with segmented AI pipelines. Rho-1 addresses this by producing continuous video output in real time. Operators can alter parameters or supply new instructions on the fly, and the network adjusts output generation instantaneously without halting or restarting execution pipelines.

This capability reflects a wider industry shift toward foundational world models capable of internalizing physical laws and spatial reasoning rather than relying on brittle, task-specific heuristics.

# Conceptual API Inference Call for Unified Token Stream
import reka

client = reka.Client()
response = client.chat.complete(
    model="rho-1",
    messages=[
        {"role": "user", "content": [
            {"type": "text", "text": "Execute sorting protocol and render continuous status stream."},
            {"type": "image", "url": "s3://factory-floor/bin-state-01.png"}
        ]}
    ],
    modalities=["text", "video", "robot_control"]
)
print(response.output_stream)

Gemini Omni Integrates Multimodal Generation and Natural Language Editing

Gemini Omni introduces native multimodal video generation, combining prompts, images, clips, audio, or templates into a single creative workflow. Instead of separating ideas by format, it understands different references as one connected creative instruction. Video remixing allows builders to work from videos they already have without restarting every time, making iteration faster and more practical when combining the girl walking by the sea clip with a product clip to create a cinematic TVC-style advertisement that blends lifestyle beauty shots with polished product visuals for a premium, elegant skincare commercial. Targeted scene editing supports precise edits inside an existing video, allowing creators to replace the spaghetti in both people’s plates with creamy pumpkin soup while keeping everything else the same, maintaining original composition, motion, and style. Gemini Omni Flash takes this further with native multimodal inputs, physics and world simulation that understands gravity and fluid dynamics, conversational video editing for multi-turn workflows, text and action synchronization that renders kinetic typography and explainer graphics, and digital avatar creation for personalized content at scale.

Share this:

  • Share on Facebook (Opens in new window) Facebook
  • Share on X (Opens in new window) X

Worth a look

  • SHIFTER: Håkan Silfvernagel leaves Sopra Steria to join Forte Data and AI team
  • Hyperliquid Strategies achieves 251 percent share price increase

Related

robotics, robots

Search:

World Today News

World Today News is your trusted source for global journalism — breaking headlines, in-depth analysis, and reporting from around the world.

Quick Links

  • Privacy Policy
  • About Us
  • Accessibility statement
  • California Privacy Notice (CCPA/CPRA)
  • Contact
  • Cookie Policy
  • Disclaimer
  • DMCA Policy
  • Do not sell my info
  • EDITORIAL TEAM
  • Terms & Conditions

Browse by Location

  • GB
  • NZ
  • US

Connect With Us

© 2026 World Today News. All rights reserved. Your trusted global news source directory.
For contact, advertising, copyright, issues email: office@world-today-news.com

Privacy Policy Terms of Service