Robbyant Launches LingBot-VA 2.0 Causal Video-Action Model

Robbyant's LingBot-VA 2.0 is a native causal video-action foundation model that redefines robot manipulation.

By Central
LingBot-VA 2.0 pretrains a causal DiT with a semantic visual-action tokenizer for embodied AI.
Highlights
  • LingBot-VA 2.0 is the first embodied-native foundation model for generalist robot manipulation.
  • The model uses a causal Diffusion Transformer with a sparse mixture-of-experts video stream.
  • It demonstrates few-shot adaptation and reactive control capabilities in real-world robot tasks.

Robbyant, the embodied AI unit operating within Ant Group, has released LingBot-VA 2.0, which the team describes as the first embodied-native foundation model purpose-built for generalist robot manipulation. Unlike prior approaches that retrofit video generators for control tasks, LingBot-VA 2.0 pretrains the entire stack for embodiment from the ground up. This distinction matters because it directly addresses structural mismatches that have limited earlier video-action models in real-world robotic applications.

Why a Video-Action Foundation Model Needs to Be Native, Not Adapted

Most existing video-action models repurpose components designed for digital content creation: a reconstruction-oriented VAE and a bidirectional video-diffusion backbone with an action module bolted on. Robbyant identifies three fundamental limitations with this approach. Pixel-reconstruction latents preserve appearance but carry limited physical structure information. Iterative denoising over video tokens is too slow for closed-loop control. And generic video objectives never teach the model how actions reshape the world.

A fourth mismatch is structural: backbones trained with bidirectional attention operate on the full sequence simultaneously, but control unfolds strictly forward in time. LingBot-VA version 1.0 finetuned a bidirectional stack into a causal model. Version 2.0 pretrains a causal Diffusion Transformer (DiT) natively, eliminating the architectural compromise from the start.

Version 1: A Tokenizer That Puts States and Actions in One Space

The first stage of the system replaces the compression-only VAE with a semantic visual-action tokenizer. Following the RepWAM approach, the tokenizer adds two objectives to reconstruction. Semantic alignment pulls visual latents toward a frozen Perception Encoder teacher. A latent-action objective extracts compact transition variables between consecutive latents: an inverse dynamics model predicts each latent action, and a forward dynamics model decodes it into a transport map plus residual.

Because world states and actions now share one latent space, unlabeled web video carries action-relevant supervision that the model can exploit during training.

Version 2: A Causal DiT With a Sparse MoE Video Stream

Built on top of that shared latent space, version 2 pretrains a causal DiT that retains the Mixture-of-Transformers layout from version 1.0. A video expert and an action expert share one causal self-attention module, but each owns a separate feed-forward pathway. The two streams scale asymmetrically. The video expert replaces its dense feed-forward network with a sparse mixture-of-experts routed layer containing 128 routed SwiGLU experts with top-8 routing and one shared expert, using auxiliary-loss-free Loss-Free Balancing. The action expert keeps a dense FFN at hidden dimension 768.

The video backbone totals roughly 13.0 billion parameters, with about 1.9 billion active. Adding the action expert and multi-chunk prediction heads, training covers about 15.3 billion parameters, of which roughly 2.5 billion activate per token at inference. Training uses a rectified-flow objective with a hybrid Muon plus AdamW optimizer.

Training Objectives: Multi-Chunk Prediction and Co-Training

Two innovations shape what the model learns beyond architecture. Multi-chunk prediction (MCP) addresses the problem of myopic supervision: teacher forcing supervises only the next chunk, so the model can cut loss by copying appearance rather than learning dynamics. MCP attaches three lightweight modules that predict the next three chunks. In ablation testing, it matched the baseline’s 45,000-step accuracy in 20,000 steps, a 2.3x training speedup.

Five objectives are co-trained rather than staged: text-to-image, text-to-video, text-image-to-video-action, in-context learning, and human-robot co-training. Sampling follows a coarse-to-fine schedule, from appearance grounding through video-action control, keeping every objective alive to avoid forgetting earlier priors.

Hierarchical Planning and Foresight Reasoning

Above the low-level policy sits a VLM planner, LoRA-finetuned with a frozen vision tower. It emits structured JSON containing completion status, instruction, generation instruction, and local scene description, running at about 2 Hz behind an asynchronous shared buffer. The policy reads this output at each chunk boundary, so planner latency never blocks execution.

Foresight Reasoning addresses the serial bottleneck that plagues real-time deployment. Rather than waiting for model inference to complete before executing a command, the system runs prediction and execution as asynchronous streams. While the robot executes chunk a_t, the video expert imagines its outcome and the action expert decodes a_{t+1} from that prediction. Each returning observation is encoded into the true latent, overwriting the imagined one to prevent drift. A forward-dynamics grounding loss trains the video expert specifically for this role.

Benchmark Performance on RoboTwin 2.0

Evaluation covers both simulation and real hardware. On RoboTwin 2.0, models train on 2,500 clean plus 25,000 randomized demonstrations across 50 tasks. LingBot-VA 2.0 achieves a 93.8 percent success rate on clean demonstrations and 93.4 percent on randomized demonstrations, for an average of 93.6 percent. This compares favorably to the previous version (92.2 percent average), Motus (87.9 percent), pi0.5 (79.8 percent), and X-VLA (72.9 percent).

On the inference speed front, a cascade of optimizations brings chunk latency from 927 milliseconds down to 142 milliseconds and asynchronous control frequency from 35 Hz up to 225 Hz. The optimizations include consistency distillation, low-precision compiled execution with FP8 TensorRT engines, long-horizon attention optimization with a paged and ragged KV cache using FlashInfer attention, and runtime overhead reduction.

Deployment Shapes: Few-Shot, Cross-Task, and Reactive Control

Beyond benchmarks, four deployment capabilities stand out. The model adapts from 10 to 15 demonstrations for few-shot onboarding, with real-world evaluation using 20 teleoperated demonstrations per task from a single multi-task checkpoint. In-context learning lets a human demonstration video replace the text instruction, enabling the policy to execute unseen compositional tasks after finetuning on four seen tasks. Human-robot co-training retargets hand poses into the robot action space, with each hand becoming a virtual parallel gripper across an egocentric corpus of 65,400 episodes. And reactive control scenarios such as air hockey and conveyor belt handling demonstrate the policy’s ability to anticipate moving objects.

What This Means for Embodied AI Development

LingBot-VA 2.0 represents a deliberate architectural departure from the prevailing practice of adapting generative video models for robotic control. By pretraining a causal DiT with a semantically aligned tokenizer and sparse MoE video stream from scratch, Robbyant has demonstrated that native embodiment design yields measurable gains in both success rate and inference efficiency. The asynchronous Foresight Reasoning loop, in particular, offers a practical template for closing the latency gap between high-capacity foundation models and real-time robotic control.

For developers and researchers working on robot manipulation, the full technical report and project page are available through Robbyant’s technology portal. The key takeaway is that the next generation of generalist robot policies may not come from finetuning foundation models built for other domains, but from building foundation models designed for embodiment from the first line of code.

Share This Article