The Qwen team has released three embodied AI models under a single banner, Qwen-Robot-Suite, marking a significant push to address the fragmentation that has long plagued robotics research. Rather than offering a single, monolithic solution, the suite comprises three distinct foundation models—Qwen-RobotManip, Qwen-RobotWorld, and Qwen-RobotNav—each built on a Qwen vision-language backbone and targeting a different core robotics problem: manipulation, world modeling, and navigation. The models collectively tackle the industry’s persistent challenge of data heterogeneity, where policies trained on one robot arm rarely transfer to another due to incompatible observation and action formats.
The suite represents a coordinated attempt to impose structure on a deeply fragmented field. All three models are built on Qwen vision-language backbones but diverge sharply in architecture and output. Qwen-RobotManip is a Vision-Language-Action (VLA) model for manipulation, built on Qwen3.5-4B. Qwen-RobotWorld is a language-conditioned video world model using a 60-layer MMDiT with a frozen Qwen2.5-VL encoder. Qwen-RobotNav is a navigation model built on Qwen3-VL, available in 2B, 4B, and 8B parameter sizes. Two of the three—RobotManip and RobotNav—ship with public GitHub repositories; RobotWorld is presented as a research paper only.
Qwen-Robot-Suite: Three Independent Models For a Fragmented Field
Qwen-Robot-Suite is not a single model. It is a suite of three independent foundation models, each designed to address the data fragmentation problem from a different angle. Robotics data is fragmented across hardware and tasks. Different robots use incompatible observation and action formats. A policy trained on one arm rarely transfers to another. The three research reports address this fragmentation in distinct ways.
RobotManip aligns action representations so manipulation data scales. RobotWorld uses language as a unified action interface for video prediction. RobotNav exposes a controllable observation interface for navigation tasks. The table below summarizes the core split between the three releases.
| Model | Problem | Backbone | Output |
|---|---|---|---|
| Qwen-RobotManip | Robotic manipulation | Qwen3.5-4B (Qwen-VL) | Continuous robot actions |
| Qwen-RobotWorld | Embodied world modeling | Frozen Qwen2.5-VL | Predicted future video |
| Qwen-RobotNav | Mobile navigation | Qwen3-VL (2B/4B/8B) | Waypoint trajectories |
Qwen-RobotManip: Alignment Unlocks Scale for Manipulation
Qwen-RobotManip is a Vision-Language-Action (VLA) foundation model built on Qwen-VL that predicts continuous robot actions. A VLA model takes camera views and a language instruction and outputs low-level robot actions. The core challenge is that manipulation data is heterogeneous by nature. Different robots record states and actions in incompatible formats. When demonstrations arrive with mismatched representations, scaling data produces interference. RobotManip solves this with a unified alignment framework.
The Unified Alignment Framework
The framework has three complementary mechanisms. First is a canonical state-action representation, an 80-dimensional vector with per-dimension binary masking. This vector holds two 29-dimensional per-arm blocks plus 22 reserved dimensions. Each block stores joint positions, end-effector pose, gripper state, and dexterous hand joints. Robots populate only the dimensions they have.
Second is a camera-frame delta pose parameterization. End-effector actions are expressed as deltas in the camera frame, making visually similar motions numerically proximate across embodiments. Third is an in-context policy adaptation mechanism that reads recent execution history as an implicit embodiment identifier, adjusting behavior at deployment time without parameter updates. A dual-stream co-training strategy runs alongside this, jointly optimizing manipulation data and a vision-language stream to prevent the backbone’s perception and reasoning from eroding.
The Data Engine
RobotManip assembles roughly 38,100 hours of manipulation data using only open-source datasets and human videos. No proprietary data collection was used. A human-to-robot synthesis pipeline produces most of this scale, converting egocentric hand demonstrations into robot trajectories and rendering across 15 robot platforms. This synthesis alone yields about 24,808 hours of demonstrations from approximately 1,933 hours of egocentric source data. Open-source robot datasets contribute over 11,000 hours.
The pipeline separates action alignment from visual alignment. Action alignment retargets hand keypoints to gripper poses. Visual alignment uses SAM3 masking, ProPainter inpainting, and MuJoCo inverse kinematics. A five-stage curation pipeline then filters the combined corpus, catching sudden changes, temporal misalignment, and extreme values. One check found 81% of episodes in a subset failed state-action alignment.
Benchmark Results
The research report argues standard benchmarks fail to measure generalization, noting that models without robot pretraining match pretrained ones on in-distribution tests. RobotManip therefore focuses on out-of-distribution (OOD) settings.
| Benchmark (OOD) | Prev. SOTA (π0.5) | Qwen-RobotManip |
|---|---|---|
| LIBERO-Plus | 84.4 | 91.4 |
| RoboTwin-C2R Hard | 47.9 | 69.4 |
| EBench | 27.1 | 45.6 |
| RoboCasa365 | 16.9 | 35.9 |
| RoboTwin-IF | 49.6 | 72.2 |
The largest reported gap is on cross-embodiment transfer, where RobotManip reaches 23.9% using camera-frame EEF actions, which is 3.2 times the 7.5% achieved by π0.5. The model also ranks first on the RoboChallenge Table30-v1 generalist track, scoring a 20% relative improvement over the prior best. Real-robot validation covers AgileX ALOHA, Franka, UR, and ARX platforms.
Qwen-RobotWorld: Language as a Universal Action Interface
Qwen-RobotWorld is a language-conditioned video world model that predicts future visual trajectories from a current observation. Natural language serves as the unified action interface. A world model learns environment dynamics. Given a current state and an action, it predicts the next state. RobotWorld represents states as video frames and actions as text. This is important because language is embodiment-agnostic. One instruction encodes the action sequence, goal, and constraints, and it works across a Franka gripper, an Aloha dual-arm system, or a humanoid.
The Double-Stream MMDiT Architecture
The model uses a 60-layer double-stream Multimodal Diffusion Transformer (MMDiT). An understanding stream processes a frozen Qwen2.5-VL encoder’s features. A generation stream processes video-VAE latents. The two streams interact via joint attention at every layer. Using an MLLM as the action encoder gives two advantages: it parses compositional instructions and constrains physically plausible transitions. The MMDiT has 20 billion parameters. The VAE adopts the Wan-VAE architecture. The context length supports up to 48,360 video tokens. A Scene2Robot mechanism reuses this backbone for cross-embodiment synthesis, processing scene, robot reference, and generation segments together to enable human-to-robot video transfer without robot-specific prompting.
The Embodied World Knowledge Dataset
Training uses the Embodied World Knowledge (EWK) dataset, which contains roughly 8.6 million video-text pairs spanning over 200 million observation frames. The corpus covers four embodied domains plus general video. Manipulation provides about 5.9 million samples across 20+ morphologies. Driving, navigation, and human-to-robot transfer fill out the rest. An action-language mapping framework standardizes everything, converting 20+ embodiment types and 500+ action categories into language. A hierarchical five-layer annotation pipeline produces the captions.
Benchmark Results
RobotWorld was evaluated on four established benchmarks, ranking first overall on two of them.
| Benchmark | Result | Ranking |
|---|---|---|
| EWMBench | 4.60 | 1st overall |
| DreamGen Bench | 4.95 | 1st overall |
| WorldModelBench | 8.99 | 1st open-source (3rd overall) |
| PBench | 0.804 | 1st open-source |
On EWMBench it leads motion fidelity with an HSD of 0.566, a 33% gain over the runner-up. Scene consistency reaches 0.914. On WorldModelBench it scores 1.00 on four physics-adherence categories: Newton’s laws, mass conservation, fluid dynamics, and gravity. Penetration scores 0.94, and instruction following scores 2.33 out of 3.0.
Qwen-RobotNav: A Controllable Interface for Navigation
Qwen-RobotNav is a scalable navigation model built on Qwen3-VL. It reframes multi-task navigation as observation context modeling. The model exposes a parameterized interface for external control. Navigation spans many task families—instruction following, point-goal navigation, object search, target tracking, and driving—each demanding a different strategy for consuming the visual stream. No fixed context strategy serves all tasks well.
The Parameterized Interface
RobotNav formulates all tasks as waypoint trajectory prediction, predicting eight waypoints, each with a 2D position and heading. A lightweight 4-layer MLP head produces these from the backbone. The interface has two configuration dimensions. Task modes select navigation behavior across VLN, PointNav, ObjNav, and Tracking. Observation parameters govern how visual history is encoded, including a visual token budget, temporal decay, and per-camera importance weights. Training-time randomization over all parameters ensures robustness. Camera identity and temporal order use natural-language tags, requiring zero architectural modification to Qwen3-VL. Supporting a new platform needs only a new prompt template.
The Agentic System
The interface makes RobotNav a building block for agentic systems. An upper-tier planner decomposes long-horizon goals into sub-goals. Qwen3.6-Plus serves as this planner in the system. The planner reconfigures RobotNav’s task mode mid-episode, while RobotNav serves as the reactive executor. The two tiers communicate exclusively through natural language. A two-level memory supports long-horizon reasoning: single-episode memory summarizes each rollout, and cross-episode memory accumulates durable conclusions like searched regions.
Benchmark Results
RobotNav was trained on 15.6 million samples. Navigation trajectory data forms 85% of this, while vision-language reasoning data fills the remaining 15%.
| Benchmark | Metric | Result |
|---|---|---|
| VLN-CE RxR (Val-Unseen) | Success Rate | 76.5% |
| VLN-CE R2R (Val-Unseen) | Success Rate | 72.1% |
| EVT-Bench | Tracking Rate | 90.0% |
| HM3Dv2 (ObjectNav) | Success Rate | 75.6% |
| NAVSIM | PDMS | 91.4 |
The agentic system sets new state-of-the-art on Embodied Question Answering, improving over the best prior method by 10.8% on HM-EQA and by 15.4% on EXPRESS-Bench while requiring 77% fewer navigation steps. The report shows performance improving from 2B to 8B parameters and states that joint multi-task training develops a shared spatial-planning substrate that transfers across task families.
What This Means for Developers
Each model maps to concrete deployment scenarios. For teams with a Franka arm and a handful of demonstrations, fine-tuning RobotManip on their own workspace is a direct path to improved performance on clutter and unseen states compared to training from scratch. For cross-embodiment skill transfer, a policy jointly fine-tuned on 6,000 CobotMagic and 130 ARX demonstrations and tested on four novel ARX tasks with zero target-task demonstrations achieved 55.0% success, over four times the best ablated variant.
RobotWorld serves as a synthetic data engine for VLA policies that need more training data than physical collection allows, and as a policy evaluation environment for running policies against generated trajectories before real hardware deployment. RobotNav, when paired with an upper-tier planner like Qwen3.6-Plus, forms an agentic system that outperforms prior methods on Embodied Question Answering. The same model handles autonomous driving as one task mode, reaching 91.4 PDMS on NAVSIM.
Comparison Table: The Three Models
The table below consolidates the technical details as a reference for choosing the right model.
| Attribute | RobotManip | RobotWorld | RobotNav |
|---|---|---|---|
| Task type | Manipulation (VLA) | Video world model | Navigation |
| Backbone | Qwen3.5-4B | Frozen Qwen2.5-VL | Qwen3-VL |
| Action interface | Camera-frame EEF / joint | Natural language | Waypoint trajectories |
| Training data | ~38,100 hours | 8.6M video-text pairs | 15.6M samples |
| Key architecture | DiT flow-matching head | 60-layer double-stream MMDiT | MLP action head |
| Headline result | 1st on RoboChallenge Table30-v1 | 1st on EWMBench, DreamGen | 76.5% SR on VLN-CE RxR |
| Output | Continuous actions | Predicted video | 8 waypoints (x, y, θ) |
| Public repo | Yes (GitHub) | Blog only | Yes (GitHub) |
Implementation Note: The Canonical Action Vector
The RobotManip action representation is the mechanism that lets different robots share one model. The per-dimension binary mask ensures gradients flow only through semantically populated entries, preventing spurious supervision on absent degrees of freedom. The same masking principle appears in the flow-matching loss, where each sample contributes equally regardless of how many dimensions are active, stopping robots with more populated slots from dominating optimization.
Who Should Try This Now
The Qwen-Robot-Suite is available now for researchers and practitioners who need to move beyond the fragmentation problem in embodied AI. RobotManip and RobotNav are available on GitHub, making them immediately useful for teams working on manipulation or navigation tasks. RobotWorld, while available only as a research paper, presents a compelling architecture for anyone building synthetic data pipelines or video-based policy evaluation tools.