A Single Predictive World Model
A world model is, at bottom, one question asked over and over: given what has happened and what I do next, what comes next? The thing it predicts can be raw pixels, a compressed latent, or an embodied action — but it is the same question, and I believe it can live in one model. Video world models and embodied agents are not two research agendas; they are the same world model, asked to predict a different modality.
My main, completed work is on the pixel-and-latent side. LIVE (ICML 2026) is a cycle-consistency training trick that keeps a long-horizon interactive video world model stable far past 256 frames, in real time, where autoregressive rollouts normally drift and collapse. The insight is the group's, and I worked on making it run: supervising whether a rollout can be reversed back to where it started, which bounds error accumulation without distilling from a teacher.
On the action side, two projects extend the same model to predict what to do. GeometryWAM is a world-action model — a Video-DiT that imagines the scene and an Action-DiT that controls embodied action — using 3D geometry as a training-time supervision signal that costs nothing at inference. AgenticVLA puts a reasoning VLM cerebrum over a VLA cerebellum that executes, folding cognitive operations like counting, waiting, and conditionals into the VLA's own action space rather than bolting them on from outside.
The bet underneath all of it: pixels, latents, and actions are just modalities one model can be asked to predict, and learning to predict all three in a single model is a path toward AGI. Near term, I am scaling LIVE onto a 14B-scale world model with memory; further out, toward worlds many agents can share.
One predictive model asked to forecast three modalities — LIVE on pixels and latents, GeometryWAM and AgenticVLA on action. Same question, different output.
Predicting the World as It Looks
Ask the world model to predict pixels and latents and it becomes an interactive video world. The hard part is keeping it stable over a long horizon — my main project here is LIVE.
LIVE — Stable Long-Horizon Interactive Video Worlds
LIVE keeps an interactive video world model coherent over a long horizon. The model generates frames autoregressively under a stream of control (camera or action); small errors compound, so after a few hundred frames the world drifts and collapses. LIVE holds it stable far past 256 frames, in real time, where prior approaches fall apart.
How it works — cycle-consistency
Supervising the long rollout directly is ill-posed: once it drifts, there is no ground-truth frame to compare against. LIVE instead supervises whether the rollout can be reversed back to where it started — roll forward from a short prompt under the control stream (no gradient), reverse the rollout and its controls in time, then reconstruct the known starting frames with a diffusion loss. Reconstruction only succeeds if the forward error stayed small, so the objective bounds error accumulation — and it needs no teacher model to distill from, unlike prior long-horizon methods.
Results & what's next
On RealEstate10K, LIVE leads PSNR at every horizon and stays real-time. At 256 frames its FID is ≈ 10 while the baselines blow up to ≈ 60 / ≈ 90; FVD is lowest across RealEstate10K, UE Engine, and Minecraft.
What's next. Scaling this onto a 14B-scale world model with memory, and further out toward worlds many agents can share (multi-agent / multiplayer, in the spirit of NVIDIA's Gamma-World).
Predicting What to Do in the World
Ask the same model to predict action and it becomes an embodied agent — GeometryWAM grounds it in 3D; AgenticVLA folds reasoning into the action space.
GeometryWAM — Grounding the World in 3D
A world-action model with a Video-DiT backbone that imagines the scene and an Action-DiT that controls embodied action — the same model that predicts what the world will look like also decides what to do in it. During training it is supervised by geometry (per-frame depth / 3D structure), so the representation learns the spatial structure a policy needs.
The design is deliberately asymmetric: the geometry tokens attend to video and action, but video and action never read geometry back. Because that dependency is one-directional, the whole geometry branch can be dropped at inference with zero overhead and no train/test gap. Geometry is a teacher that goes home before the exam.
No separate 3D encoder is needed — the geometry is tokenized by the same VAE as the video, so it adds no new architecture.
AgenticVLA — A Cerebrum for the Body
Today's VLAs grasp and wipe well, then fail at anything that needs reasoning over time — "wait five seconds," "scoop exactly three," "if the water boils, turn off the heat." These are failures of cognition about action, not dexterity. AgenticVLA gives the body a brain.
The distinctive move is to install "VLA-as-a-skill" inside the VLM so that cognitive operations — counting, waiting, conditional branching — are unified into the VLA's own action space rather than wrapped around it with an external clock and counter. The aim is to do this with the least new data.
Why One Model for Every Modality
Pixels, latents, and actions are not separate problems that happen to share a name. They are different answers to one question — what comes next? A predictor that has learned all three together shares one representation of how the world moves, so imagining a scene, compressing it, and acting in it draw on the same learned dynamics rather than three systems stitched at the seams.
That is the through-line connecting LIVE, GeometryWAM, and AgenticVLA. Each project pushes on one modality; the longer game is to hold all three in one model — first scaling LIVE onto a 14B-scale world model with memory, then toward shared, multi-agent worlds.
Open to research collaboration & graduate study
Ziyang Ye (叶子扬)
Research Assistant · CUHK-Shenzhen & SLAI (Shenzhen Loop Area Institute)
B.Eng. Software Engineering · Jilin University