Research Portfolio · Toward AGI

One World Model, Asked to Predict Every Modality

Ziyang Ye  叶子扬
Research Assistant · CUHK-Shenzhen & SLAI (Shenzhen Loop Area Institute)
B.Eng. Software Engineering · Jilin University
World Models · Generative Models · Embodied AI
Overview

A Single Predictive World Model

A world model is, at bottom, one question asked over and over: given what has happened and what I do next, what comes next? The thing it predicts can be raw pixels, a compressed latent, or an embodied action — but it is the same question, and I believe it can live in one model. Video world models and embodied agents are not two research agendas; they are the same world model, asked to predict a different modality.

My main, completed work is on the pixel-and-latent side. LIVE (ICML 2026) is a cycle-consistency training trick that keeps a long-horizon interactive video world model stable far past 256 frames, in real time, where autoregressive rollouts normally drift and collapse. The insight is the group's, and I worked on making it run: supervising whether a rollout can be reversed back to where it started, which bounds error accumulation without distilling from a teacher.

On the action side, two projects extend the same model to predict what to do. GeometryWAM is a world-action model — a Video-DiT that imagines the scene and an Action-DiT that controls embodied action — using 3D geometry as a training-time supervision signal that costs nothing at inference. AgenticVLA puts a reasoning VLM cerebrum over a VLA cerebellum that executes, folding cognitive operations like counting, waiting, and conditionals into the VLA's own action space rather than bolting them on from outside.

The bet underneath all of it: pixels, latents, and actions are just modalities one model can be asked to predict, and learning to predict all three in a single model is a path toward AGI. Near term, I am scaling LIVE onto a 14B-scale world model with memory; further out, toward worlds many agents can share.

The map

One predictive model asked to forecast three modalities — LIVE on pixels and latents, GeometryWAM and AgenticVLA on action. Same question, different output.

Modality · Pixels & Latents

Predicting the World as It Looks

Ask the world model to predict pixels and latents and it becomes an interactive video world. The hard part is keeping it stable over a long horizon — my main project here is LIVE.

LIVE · ICML 2026

LIVE — Stable Long-Horizon Interactive Video Worlds

LIVE keeps an interactive video world model coherent over a long horizon. The model generates frames autoregressively under a stream of control (camera or action); small errors compound, so after a few hundred frames the world drifts and collapses. LIVE holds it stable far past 256 frames, in real time, where prior approaches fall apart.

stable past 256 frames real-time · >20 FPS SOTA on 3 benchmarks
LIVE — stable long-horizon rollouts and flat error curve
Stability. A baseline degrades into artifacts by the 256th frame while LIVE holds structure; the error-vs-rollout curve stays flat, and it generalizes across indoor, Minecraft, and game-engine scenes. (ICML 2026)

How it works — cycle-consistency

Supervising the long rollout directly is ill-posed: once it drifts, there is no ground-truth frame to compare against. LIVE instead supervises whether the rollout can be reversed back to where it started — roll forward from a short prompt under the control stream (no gradient), reverse the rollout and its controls in time, then reconstruct the known starting frames with a diffusion loss. Reconstruction only succeeds if the forward error stayed small, so the objective bounds error accumulation — and it needs no teacher model to distill from, unlike prior long-horizon methods.

LIVE — comparison of autoregressive training paradigms: Teacher Forcing, Diffusion Forcing, Self-Forcing, and LIVE
Where LIVE sits. Teacher Forcing trains on ground-truth context, so the model never sees its own errors (a train/test mismatch); Diffusion Forcing injects noise into the context but still cannot model real rollout error; Self-Forcing distills a teacher at the sequence level, leaving drift unbounded. LIVE rolls forward, then reverses to recover the start under a frame-level diffusion loss — bounding error through the cycle-consistency objective. (Fig. 2)
LIVE training pipeline — forward rollout, reverse, reconstruct
Method. Forward rollout → temporally reverse the rollout and its controls → a flow-matching loss reconstructs the starting frames. Right: the reverse-forcing attention pattern.

Results & what's next

On RealEstate10K, LIVE leads PSNR at every horizon and stays real-time. At 256 frames its FID is ≈ 10 while the baselines blow up to ≈ 60 / ≈ 90; FVD is lowest across RealEstate10K, UE Engine, and Minecraft.

What's next. Scaling this onto a 14B-scale world model with memory, and further out toward worlds many agents can share (multi-agent / multiplayer, in the spirit of NVIDIA's Gamma-World).

LIVE — RealEstate10K results table
RealEstate10K. LIVE leads PSNR / LPIPS / SSIM across the 0–64 / 0–128 / 0–200 / ≥256-frame buckets, in the real-time bracket.
Modality · Action

Predicting What to Do in the World

Ask the same model to predict action and it becomes an embodied agent — GeometryWAM grounds it in 3D; AgenticVLA folds reasoning into the action space.

GeometryWAM

GeometryWAM — Grounding the World in 3D

A world-action model with a Video-DiT backbone that imagines the scene and an Action-DiT that controls embodied action — the same model that predicts what the world will look like also decides what to do in it. During training it is supervised by geometry (per-frame depth / 3D structure), so the representation learns the spatial structure a policy needs.

The design is deliberately asymmetric: the geometry tokens attend to video and action, but video and action never read geometry back. Because that dependency is one-directional, the whole geometry branch can be dropped at inference with zero overhead and no train/test gap. Geometry is a teacher that goes home before the exam.

No separate 3D encoder is needed — the geometry is tokenized by the same VAE as the video, so it adds no new architecture.

Video-DiT backbone + Action-DiT control 3D supervision · free at inference shared VAE · no extra params
GeometryWAM tri-modal attention mask
Attention. Over z (latent frame), a (action chunk), and g (depth latent): geometry rows read video and action, but no video/action row reads a geometry column — so deleting g leaves the policy's computation identical.
AgenticVLA

AgenticVLA — A Cerebrum for the Body

Today's VLAs grasp and wipe well, then fail at anything that needs reasoning over time — "wait five seconds," "scoop exactly three," "if the water boils, turn off the heat." These are failures of cognition about action, not dexterity. AgenticVLA gives the body a brain.

Cerebrum — VLM reasons · what & when
A reasoning VLM reads goal, scene, and memory, and decides what to do and when.
↓   drives high-frequency control   ↓
Cerebellum — VLA executes · how
A VLA does the high-frequency low-level control — installed as a skill inside the VLM's repertoire, not an external tool.

The distinctive move is to install "VLA-as-a-skill" inside the VLM so that cognitive operations — counting, waiting, conditional branching — are unified into the VLA's own action space rather than wrapped around it with an external clock and counter. The aim is to do this with the least new data.

VLM cerebrum · VLA cerebellum count · wait · branch → in the action space
Convergence

Why One Model for Every Modality

Pixels, latents, and actions are not separate problems that happen to share a name. They are different answers to one question — what comes next? A predictor that has learned all three together shares one representation of how the world moves, so imagining a scene, compressing it, and acting in it draw on the same learned dynamics rather than three systems stitched at the seams.

That is the through-line connecting LIVE, GeometryWAM, and AgenticVLA. Each project pushes on one modality; the longer game is to hold all three in one model — first scaling LIVE onto a 14B-scale world model with memory, then toward shared, multi-agent worlds.

One model, asked to predict every modality of experience — that is the bet, and a path toward AGI.
Contact

Open to research collaboration & graduate study

Ziyang Ye (叶子扬)
Research Assistant · CUHK-Shenzhen & SLAI (Shenzhen Loop Area Institute)
B.Eng. Software Engineering · Jilin University