PixVerse R2: Scaling Real-Time Omni World Models
PixVerse R2 scales real-time audiovisual world models with Omni Causal AR, persistent world state, multimodal control, and ultra-few-step acceleration.
Overview
PixVerse R1 introduced the world’s first publicly launched, general-purpose real-time audiovisual world model. It moved video generation beyond one-off output toward a continuously evolving world that responds to input, updates its state, and produces synchronized video and audio in real time.
We see the future of world models advancing along two complementary axes. The first is a scalable foundation for streaming video generation, supported by a training and acceleration framework that preserves quality, consistency, and efficiency as data, tasks, controls, and temporal horizons expand. The second is a richer input-and-interaction layer, where multimodal and agent-driven workflows introduce more expressive control signals and structured spatial context.
R2 was rebuilt around this roadmap. Omni Causal AR continuously scales this streaming-video foundation across data, modalities, tasks, control signals, and temporal horizons, while Real-Time Acceleration carries those capabilities into ultra-few-step real-time operation without relearning them. In parallel, a unified multimodal interface allows text, references, audio, actions, and agent-generated controls to update the same running world, bringing model scaling and richer interaction into one system.
Overall Framework
Conventional long-video and real-time generation frameworks typically rely on a chain of training stages: converting a bidirectional model into an autoregressive model, constructing a task-specific distillation teacher from a bidirectional model, applying DMD distillation, and then introducing self-rollout DMD training to narrow the train-inference gap. Repeated capability transfer across these stages can degrade generation quality and long-horizon stability.
R2 consolidates this pipeline into two scalable processes: continuously pretraining Omni Causal AR as a unified world model, then distilling the same causal foundation directly into a real-time model. Fewer stage boundaries reduce repeated task alignment and model handoffs, allowing new data, tasks, control signals, and model capacity to translate more directly into real-time world-modeling performance.

Figure 1. A conventional five-stage pipeline versus R2’s two-process framework: continuously scale world modeling in Omni Causal AR, then distill the same causal foundation directly for real-time inference.
We introduce R2, the first unified scaling architecture for real-time audiovisual world models. We develop continuous multimodal control, persistent causal world state, explicit failure-state replay, and ultra-few-step same-backbone acceleration as one system.
Omni Causal AR
Omni Causal AR is R2’s scalable causal world-modeling backbone. It converts a full spatiotemporal generative prior into a continuously pretrained world-state transition model that advances in one temporal direction and predicts the next synchronized audio-video segment from available history and live control.
As training expands from isolated clips to continuous long video and multi-turn interaction trajectories, one model learns the causal chain connecting historical world state, current interactive input, and future world evolution.
Multimodal Interaction and Synchronized Output
Four types of input enter one running world: text prompts, multimodal references, audio, and actions such as WASD or continuous controls. They can arrive while the model is generating, alter the next world state, and drive synchronized video and audio. The runtime loop is continuous: generate, receive input, update the world, and keep generating.

Figure 2. Four live input families—text, multimodal references, audio, and actions—enter one running causal world and drive the next synchronized video-audio state.
Dynamic Chunk Generation
Control signals operate at different temporal scales. A text prompt or reference may describe a complete event. A WASD action requires immediate, fine-grained feedback. Audio carries semantic content, rhythm, and synchronization constraints simultaneously. Using a fixed chunk length therefore forces the model to trade responsiveness against expressive completeness.
Dynamic Chunk Generation segments video and audio according to the semantic boundaries and duration of the active control signal. A maximum chunk size bounds computation and stabilizes training. Short chunks improve responsiveness to actions and local instructions, while longer chunks preserve event structure, motion, and audiovisual coherence. This allows heterogeneous tasks and controls to share one causal generation interface without changing the model architecture.
Hybrid Teacher Forcing / Diffusion Forcing
A long-running autoregressive model faces a persistent train-inference gap. During training, the historical context is often cleaner than the history produced by the model during deployment. Training only on clean history can produce a strong next-step model that remains fragile to small accumulated errors. Training only on perturbed history can weaken image quality, motion, and control response.
R2 combines teacher forcing on clean histories with diffusion forcing on noisy histories. Clean histories preserve generation quality, while noisy histories improve robustness to self-generated deviations. Together, they narrow the train-inference gap and slow error accumulation over long rollouts.
Causal Attention Mask and Relative Temporal RoPE
Within this hybrid strategy, R2 jointly constrains causal visibility and temporal position. The attention mask specifies which historical states each query can access. Under Diffusion Forcing, the model attends only to the history represented by Sink Memory and Rolling History. Under Teacher Forcing, a clean query attends to clean causal history, while a noisy query attends to the preceding clean history and its own noisy state. Future states and other noisy states remain masked.
Relative Temporal RoPE assigns bounded temporal coordinates inside the same active attention window. If a video block contains V frames, appears in slot s, and the frame offset is r, its temporal position is p = sV + r. As the rolling window advances, real time continues moving forward, but the sink, recent history, and current chunk remain mapped to a stable and bounded range of relative coordinates. The joint mask-and-position design preserves the causal boundary while reducing pressure from absolute-position extrapolation.

Figure 3. Teacher Forcing protects the clean-history quality ceiling; Diffusion Forcing trains recovery from imperfect histories. Both branches share bounded causal visibility and relative temporal coordinates.
Multi-Timescale Memory
Long-running world modeling cannot simply retain every historical state. Continuously growing history increases compute and memory cost, while aggressive truncation can lose character identity, scene setup, and key events. R2 separates memory by temporal function:
- Sink Memory preserves persistent anchors such as character identity, environment, style, and world rules.
- Rolling History tracks recent motion, pose, camera behavior, and environmental changes so that local state transitions remain continuous.
- Object KV Cache reuses and compresses object-level historical representations that remain relevant to future evolution.
Three parallel memory channels keep one world alive under a bounded budget. Sink Memory preserves persistent identity and rules; Rolling History carries recent dynamics; Object KV Cache retains object-level state that still matters to future evolution.
Error Bank
Long-horizon drift often begins with a small error that is written into history and inherited by later generations. Error Bank stores representative failure states and replays them during training alongside normal histories. The model learns not only how to continue from ideal states, but how to recognize and recover after deviation has already entered the world.
Error-aware replay turns deployment failures into reusable pretraining signals. In internal stage evaluations, the long-horizon brightness-drift metric decreased from 0.201 to 0.129, a 35.8% reduction. Twenty of 29 long-sequence samples showed improvement, and spurious motion was reduced in all five no-motion samples. Error Bank therefore creates a feedback loop from rollout failures to targeted replay and improved recovery.
Real-Time Acceleration
Accelerate, do not relearn. R2 accelerates an already capable causal world model rather than relearning long-horizon dynamics in a separate few-step generator. The continuously pretrained Omni Causal AR model provides both student ODE initialization and the foundation of the distillation teacher, aligning teacher, student, and autoregressive runtime around the same causal trajectory.
This shared starting point lets the acceleration stage focus on the actual cost of ultra-few-step generation: preserving control strength, teacher distribution, visual realism, long-context state, and high-resolution structure under a real-time budget.
DDMD with Adversarial Regularization
Under an ultra-few-step budget, condition alignment, distribution matching, and realism are related but distinct optimization objectives. Condition alignment requires accurate responses to prompts, references, audio, and actions. Distribution matching preserves the teacher prior. Adversarial supervision anchors the student output to real data.
Following DDMD, R2 constructs separate directions for condition alignment and distribution matching at tCA and tDM:

R2 further introduces adversarial regularization derived from DMD2:

Three optimization forces converge in one student. Condition alignment strengthens control response, distribution matching preserves the teacher prior, and adversarial regularization anchors realism under a small sampling budget.
Block-Sparse Attention
R2 partitions the spatiotemporal token sequence into computation blocks. Block-level relevance estimates which key-value blocks matter most to each query block. Exact softmax attention is then computed only over the selected connections, while lower-contribution links are omitted.
The routing is not random pruning. It concentrates computation on dependencies connecting active multimodal controls, recent motion, and persistent world state. This converts dense attention into structured computation over a smaller set of blocks, reducing matrix operations and memory traffic.
R2 reaches over 90% attention sparsity while preserving visual quality, motion continuity, condition following, and long-horizon stability across the four reported dimensions in the current internal evaluation. The model is trained to adapt to sparse attention patterns, releasing critical budget for lower latency, higher resolution, and longer continuous generation.
Pyramid Ultra-Few-Step Distillation
High-resolution generation spends substantial computation on global structure that can often be established at lower resolution. R2 organizes the generation path into one or two low-resolution stages followed by a high-resolution stage. The lower-resolution stages establish scene layout, subject motion, and camera structure. The high-resolution stage restores texture, edges, and local detail.
This coarse-to-fine pyramidal design assigns global structure and motion to lower-resolution stages and reserves high-resolution computation for high-frequency detail. It avoids redundant high-resolution work, enabling fewer refinement steps and lower latency while preserving structural coherence and visual detail.
Conclusion and Outlook
Conclusion
PixVerse R1 established real-time audiovisual world models as a practical paradigm. PixVerse R2 makes that paradigm scalable. One continuously pretrained Omni Causal AR backbone compounds data, modalities, tasks, controls, and temporal horizons; three memory channels preserve persistent state; Error Bank learns from generated failures; and Real-Time Acceleration compresses the same causal foundation instead of relearning the world.
The architecture is already producing measurable gains: over 90% attention sparsity while preserving the four reported quality dimensions in the current internal evaluation, and 35.8% lower long-horizon brightness drift in Error Bank stage evaluations. R1 made AI video interactive; R2 makes real-time audiovisual worlds more persistent, controllable, and scalable.
Current Frontier and Scaling Outlook
R2 establishes a path for scaling visual quality, motion, world consistency, and temporal horizons without giving up real-time interaction. The next phase expands model capacity, high-quality audiovisual and interaction data, training horizons, task coverage, and capability-preserving acceleration—raising both the ceiling of Omni Causal AR and the efficiency with which that capability reaches real-time operation.
The direction is clear: more capable worlds, longer-running state, richer controls, and stronger audiovisual expression inside the same scalable causal system.
References
- PixVerse R1 Explained: Real-Time AI Video World Model
- Decoupled DMD: CFG Augmentation as the Spear, Distribution Matching as the Shield
- DMD2: Improved Distribution Matching Distillation for Fast Image Synthesis
- Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- Pyramidal Flow Matching for Efficient Video Generative Modeling