NVIDIA AI Releases Molt, 8.6K-Line Agentic RL Framework

NVIDIA's NeMo team unveils Molt, a compact 8.6K-line agentic RL framework that prioritizes researcher velocity and training correctness.

By Central
Molt combines Ray, vLLM, and NVIDIA AutoModel in a single asynchronous loop for efficient agentic RL training.
Highlights
  • Molt is an 8.6K-line agentic RL framework from NVIDIA's NeMo team, released under Apache 2.0.
  • Molt composes Ray, vLLM, and NVIDIA AutoModel without forking them, benefiting from upstream updates.
  • Molt ensures token-exact correctness invariants, avoiding training on tokens not generated by the policy.

Agentic reinforcement learning research has long been defined by constant algorithm modification — new estimators, new pipeline stages, new rollout schemes. In mainstream frameworks, every change threads through layers of trainer, distributed backend, and rollout glue. That cost lands on the researcher at every iteration, making rapid experimentation a luxury few can afford. NVIDIA’s NeMo team has released Molt, a PyTorch-native agentic RL framework with an unusual design target: the codebase should be compact enough for a researcher to hold in their head, and for an AI coding assistant to read and reason about in its entirety. The stated footprint is roughly 8.6K lines of RL code.

What is Molt? An 8.6K-Line Agentic RL Framework Built for Researcher Velocity

Molt is a PyTorch-native agentic reinforcement learning framework from NVIDIA’s NeMo team, released under Apache 2.0. Its primary design goal is researcher velocity through extreme code compactness. The framework measures roughly 8.6K lines of RL code, traced by following the import graph from each framework’s RL entry point. For comparison, the same method counts about 62K lines for verl, 25K for slime, and 7.2K for OpenRLHF. Molt ships with launch codes, Slurm scripts, and a prebuilt container, making it immediately deployable on supported hardware.

Agentic RL is reinforcement learning applied to agents that interact with environments over multiple turns — think multi-step tool use, code execution, or vision-language tasks. Unlike standard RL, the agent’s policy is typically a large language model that generates tokens, making the training loop both computationally expensive and architecturally complex. Molt targets this exact complexity by composing three existing components rather than forking them: Ray for placement and asynchronous queues, vLLM for rollout, and NVIDIA AutoModel with FSDP2 for training. Because none of the three is forked, upstream improvements arrive as a container pin rather than a rebase.

Three Components, One Asynchronous Loop

Molt’s runtime architecture is built around three components working in a single asynchronous loop. An agent pool, a set of vLLM engines behind a request router, and a single trainable policy actor. A streaming pool keeps prompt groups in flight so the engines never drain while the actor trains. Partial rollout pauses the engines, broadcasts actor shards over NCCL directly to each engine, and resumes retained requests instead of discarding them.

The training loop itself is straightforward. The agent — a plain Python program — sends prompts and observations to the vLLM engines. The engines return token IDs and log-probabilities through a request router that keeps every request of one rollout on the engine holding its prefix cache. The Ray queue emits a batch when enough groups finish, and the trainer consumes the batch for an optimizer step. After the step, weight refit broadcasts the updated actor shards back to the engines via NCCL, bypassing the request router entirely. This design allows the engines to continue serving inference requests even while the trainer is computing gradients, decoupling training throughput from heavy-tailed generation latency.

Is Molt Deployable? Hardware Requirements and Real-World Constraints

Molt ships under Apache 2.0 with launch codes, Slurm scripts, and a prebuilt container. The research paper positions it as research infrastructure, not a production training service, and hardware is the real gate. The shipped recipes assume 2 nodes of 8 H100 GPUs, split with 8 GPUs for training and 8 for rollout.

That puts Molt in reach of frontier and frontier-adjacent labs, well-funded AI startups doing post-training, enterprise AI research groups in finance, healthcare, and robotics that train agents against proprietary environments, and academic labs with multi-node H100 or H200 access. Applications include multi-turn tool-use agents, code-execution agents, vision-language environments (the shipped geo3k recipe), LLM-as-judge reward loops, and on-policy distillation onto a smaller student.

The hardware requirement is significant but not exceptional for contemporary AI research. Two nodes of 8 H100 GPUs is a substantial investment, but it is the minimum configuration required to run Molt’s asynchronous loop effectively. The framework is designed to scale up from this baseline, and the report states the full asynchronous loop has been run end to end on a 700B MoE model at expert parallelism 256.

The Agent Is an Ordinary Program

One of Molt’s most distinctive design choices is that the agent is an ordinary Python program. An RL run names one Python module that exports an AgentRunnercodecodecodecode. Everything else is ordinary code, including the reward function. This means researchers can write reward functions as graders, sandboxed tools, LLM-as-judge, or a vision-language environment — whatever Python can express.

Two forms are supported. With Envcodecodecodecode, the framework owns the LLM loop in a Gymnasium-aligned step()codecodecodecode, providing a familiar interface for researchers coming from standard RL frameworks. With ChatAgentcodecodecodecode, the user owns the loop through a stock OpenAI or Anthropic SDK. Molt launches a loopback server that speaks both wire protocols, and every request decodes server-side into one token-exact accumulation. When a long-horizon agent compacts its context and rewrites the prefix, the server seals the current segment and opens a fresh one automatically.

This design is a deliberate departure from frameworks that require researchers to adapt their code to a specific agent interface or reward structure. By making the agent an ordinary program, Molt reduces the cognitive overhead of translation between the researcher’s mental model and the framework’s expectations. The reward function, in particular, can be any Python code — there is no special reward API to learn.

Never Train on a Token You Did Not Generate

Three correctness invariants organize Molt’s design around a fundamental principle: never train on a token you did not generate. Token identity means sampled token IDs define the trajectory, not a retokenized transcript. Policy-version semantics means trainable tokens keep their behavior-policy log-probabilities, and asynchronous use is corrected per token behind a sequence-level gate. Forward consistency means the rollout and actor must agree on model semantics.

The token identity invariant is particularly important. When a trainer re-tokenizes a transcript, the token boundaries can move even though the text is identical. The report calls this class of failure “quiet” — nothing crashes, the update is simply wrong. Molt avoids this entirely by passing token IDs throughout the pipeline. Prompts enter as token IDs, completions return as token IDs with per-token log-probabilities, and text never passes through a tokenizer mid-episode.

For mixture-of-experts policies, forward consistency matters most. The rollout and training routers select experts independently, and small numerical differences can flip top-k choices. Molt applies rollout routing replay, where vLLM returns its per-token expert IDs and the training forward replays them. This ensures that the gradient is computed with respect to the exact same expert selection that was used during generation, eliminating a subtle source of bias in MoE training.

Benchmarking Molt: Throughput and Comparison with Slime

The Molt technical report includes benchmark results comparing Molt with slime, another agentic RL framework. The measurements were taken using a Qwen3-30B-A3B model in bf16 with a 16,384-token context and 8,192-token response cap, running on 2 nodes of 8 H100 GPUs in a fully asynchronous and disaggregated configuration.

For seconds per optimizer step, Molt measured 119.4 ± 2.3 seconds, while slime measured 109.5 ± 10.3 seconds. For tokens per GPU-second, Molt achieved 461 tok/GPU/s versus slime’s 502 tok/GPU/s. The report emphasizes reading the error bars, not the means. The slime cross-run spread of 102 to 121 seconds overlaps Molt’s band, so the report claims no superiority in either direction.

The benchmark checkpoint also exposes an upstream distributed-MoE forward mismatch, so these rows measure throughput only, not convergence. This is a critical distinction — throughput benchmarks can show one framework as faster while hiding differences in training stability, sample efficiency, or final policy quality. The report is transparent about this limitation.

What matters more than raw throughput is Molt’s design for researcher velocity. In a field where the bottleneck is often the time to implement and test a new idea, a codebase that is 8.6K lines versus 62K lines represents a significant reduction in cognitive load. A researcher can hold the entire framework in their head, and an AI coding assistant can read and reason about it in its entirety. This is the design target, and it is what sets Molt apart from larger frameworks.

Scale Knobs: Expert Parallelism and Configuration Flexibility

Molt is designed to scale through configuration, not migration. The launch script that trains a dense 4B model expresses a DeepSeek-V3-class MoE as a flag. The framework supports expert parallelism from 1 to 256, with corresponding configurations for small MoE actors, vLLM tensor parallelism combined with data parallelism, and mid-size to large MoE layouts.

The scale knobs include expert parallelism through the --fsdp.ep_sizecodecodecodecode flag, multi-token prediction speculative decoding to reduce generation time per step, and optimizer CPU offload to reduce peak GPU memory. For the Qwen3.6-35B-A3B multimodal MoE recipe, enabling MTP speculative decoding drops generation time per step from 329 seconds to 64 seconds, moving the workload from generation-bound to training-bound. Offloading Adam states to host memory reduces peak GPU memory from 64.7 GB to 46.4 GB, with a modest 18% increase in policy training time.

The interactive explainer included in the report demonstrates these scale knobs visually, showing how expert parallelism maps onto GPU layouts and how different configurations affect throughput and memory usage. The report provides concrete numbers for each configuration, allowing researchers to make informed trade-offs based on their available hardware and training objectives.

What Molt Means for Agentic RL Research

Molt represents a deliberate bet on code compactness as a primary design virtue. In a field where frameworks have grown increasingly complex, with layered abstractions, distributed backends, and extensive glue code, Molt strips away everything that is not directly about the RL loop. The result is a framework that is easier to understand, easier to modify, and easier to reason about — for both humans and AI coding assistants.

The implications for agentic RL research are significant. Current frameworks require substantial upfront investment in learning the framework itself before any research can begin. With Molt, a researcher can read the entire codebase in a few hours and start modifying it immediately. This lowers the barrier to entry for agentic RL research and makes it feasible for smaller labs and individual researchers to contribute.

The choice to compose rather than fork existing components is also strategically sound. By using Ray, vLLM, and NVIDIA AutoModel as composed components, Molt benefits from upstream improvements without requiring maintenance of forked code. When vLLM adds support for a new architecture or optimization, Molt users get it with a container pin update rather than a manual rebase. This keeps the framework lean and focused on its core value proposition: the RL loop itself.

Molt’s focus on token-exact correctness invariants — never training on a token you did not generate — sets a quality bar that not all frameworks meet. The quiet failures that arise from retokenization or policy-version drift are real problems in agentic RL, and Molt is one of the first frameworks to address them systematically. This may prove to be its most durable contribution, as the research community increasingly recognizes the importance of training correctness in RL from AI feedback and agentic applications.

For researchers evaluating frameworks, Molt offers a clear value proposition: if you need to understand and modify every line of your RL training loop, and you have access to multi-node H100 hardware, Molt is worth serious consideration. Its compact codebase, clean architecture, and principled approach to training correctness make it a strong foundation for agentic RL research. The report’s transparent benchmarking, including the honest admission that throughput is comparable to slime within error bars, builds trust and sets the right expectations. Molt is not claiming to be faster — it is claiming to be more understandable, and that may be exactly what the field needs.

Share This Article