Pixel-space world models have recently emerged as a promising approach for GUI agents, predicting future screens prior to execution and subsequently re-perceiving them for action selection. However, this two-stage paradigm faces two key challenges in GUI environments: (L1) the high computational cost of pixel reconstruction and (L2) the need to reconstruct large and diverse appearance-level details beyond what is required for action selection. In practice, this screen-generation objective can induce an identity shortcut problem, where future-state prediction largely preserves the current state rather than capturing action-induced changes. In this work, we reformulate GUI world modeling by shifting the prediction target from full-screen reconstruction to action-induced future information in latent space. Concretely, we leverage a large multimodal model as the latent world model, whose shared representations support two dedicated heads for both action selection and predictive transition modeling. The action head directly translates these latent representations into action selection, while the transition head captures action-conditioned future information. To mitigate the identity shortcut, we further introduce an action-aware contrastive objective that encourages the transition representations to capture discriminative action-induced changes. Extensive experiments across offline and online GUI settings show that our method achieves up to 31.9% and 40.6% relative gains in success rate over pixel-space world models, respectively, while substantially reducing inference latency by up to 98.9%.
Figure 2 highlights two mismatches in pixel-space GUI world modeling. (a) Rendering future screens adds substantial runtime, while Latent2World retains a stronger success–efficiency trade-off. (b) GUI transitions show lower similarity and much wider variation than natural-video transitions, making exact future-screen reconstruction an unnecessarily difficult target.
(a) Performance–efficiency trade-off of GUI world models. (b) Transition-similarity distributions measured with DINOv2 and SigLIP2.
These observations motivate predicting decision-relevant future information in latent space instead of reconstructing the full next screen.
Figure 3 exposes the identity shortcut . (a) Even when the ground-truth next screen changes drastically, pixel-space predictors can remain close to the current GUI. (b) As transition size increases, shortcut severity rises and downstream SR drops for pixel-space world models, while Latent2World keeps both trends substantially more stable.
(a) Identity-shortcut example. (b) Shortcut severity and downstream success rate across weak, medium, and strong GUI transitions.
On strong transitions , shortcut severity reaches 98% for Diffusion and 66% for Code2World, versus 41% for Latent2World; the corresponding SRs are 50.0, 48.5, and 61.5. Larger GUI changes therefore amplify identity preservation in pixel space, whereas latent prediction better preserves action-relevant performance.
Latent2World learns action selection and action-conditioned future prediction in one shared multimodal latent space, so predictive supervision directly shapes the representation used for decision making.
(a) Training: the Action Head scores candidates while the Transition Head learns the action-induced next-state latent with an InfoNCE objective. (b) Inference: the world model re-scores low-confidence decisions directly in latent space, without rendering a future screen.
The true next state is the positive target; the current state and memory-bank transitions are negatives. This explicitly teaches the latent to represent what the action changes .
A confidence gate invokes the world model only when needed. The transition head is discarded at inference, avoiding the render-and-re-perceive loop.
We report Online task-level evaluation on AndroidWorld and Offline step-level evaluation on AndroidControl and AITZ. Online is shown by default; use the buttons or swipe on mobile to switch views.
Across five backbones on AndroidWorld, Latent2World improves Overall SR over the corresponding base agent by +4.5 to +12.5 points , while avoiding the large runtime cost of pixel-space future generation.
| Backbone | Setting | Easy SR | Medium SR | Hard SR | Overall SR | Tokens/Task | Mean Steps | Run Time |
|---|---|---|---|---|---|---|---|---|
| GPT-5.6 | Base | 68.33 | 50.00 | 16.67 | 54.46 | 363.8K | 13.12 | 127.4 |
| Code2World | 68.33 | 52.94 | 22.22 | 56.25 | 525.4K | 13.71 | 370.9 | |
| Diffusion2Image | 68.33 | 47.06 | 16.67 | 53.57 | – | 13.21 | 2320.3 | |
| Latent (Ours) | 73.33 | 58.82 | 27.78 | 61.61 | 329.4K | 12.20 | 141.4 | |
| Gemini-3.6-Flash | Base | 70.00 | 52.94 | 33.33 | 58.93 | 204.6K | 11.42 | 168.6 |
| Code2World | 71.67 | 52.94 | 22.22 | 58.04 | 310.0K | 12.55 | 482.3 | |
| Diffusion2Image | 68.33 | 52.94 | 22.22 | 55.75 | – | 11.68 | 2490.3 | |
| Latent (Ours) | 73.33 | 58.82 | 38.89 | 63.39 | 200.4K | 11.12 | 228.1 | |
| GUI-Owl-7B | Base | 25.00 | 5.88 | 0.00 | 15.18 | 226.1K | 16.67 | 159.5 |
| Code2World | 26.67 | 5.88 | 5.56 | 16.69 | 421.64K | 18.12 | 659.4 | |
| Diffusion2Image | 26.67 | 5.88 | 0.00 | 16.07 | – | 16.81 | 1260.2 | |
| Latent (Ours) | 33.33 | 11.76 | 5.56 | 22.32 | 229.0K | 16.99 | 205.3 | |
| Qwen3-VL-8B | Base | 41.67 | 14.71 | 5.56 | 27.68 | 105.9K | 11.21 | 125.8 |
| Code2World | 38.33 | 20.59 | 11.11 | 28.57 | 288.4K | 14.10 | 519.9 | |
| Diffusion2Image | 53.33 | 14.71 | 5.56 | 33.93 | – | 13.59 | 789.9 | |
| Latent (Ours) | 55.00 | 26.47 | 16.67 | 40.18 | 129.3K | 13.40 | 223.8 | |
| Qwen3-VL-32B | Base | 56.67 | 32.35 | 11.11 | 41.96 | 320.8K | 13.29 | 703.6 |
| Code2World | 58.33 | 35.29 | 11.11 | 43.75 | 597.3K | 16.84 | 1453.6 | |
| Diffusion2Image | 55.00 | 26.47 | 5.56 | 38.39 | – | 17.81 | 3154.6 | |
| Latent (Ours) | 60.00 | 38.23 | 16.67 | 46.42 | 385.4K | 16.26 | 960.0 |
On AndroidControl and AITZ, Latent2World improves action-type prediction across all evaluated backbones and increases step-level SR in most backbone–dataset settings with modest overhead.
| Backbone | Setting | AC Type | AC Ground. | AC SR | AITZ Type | AITZ Ground. | AITZ SR | Tokens/Task | Run Time |
|---|---|---|---|---|---|---|---|---|---|
| GPT-5.6 | Base | 71.92 | 60.46 | 59.84 | 59.84 | 47.99 | 41.30 | 3585.57 | 1.14 |
| Code2World | 70.34 | 59.61 | 58.22 | 58.60 | 47.03 | 39.24 | 12634.21 | 9.31 | |
| Diffusion2Image | 68.31 | 55.01 | 57.94 | 56.40 | 46.80 | 38.59 | – | 135.20 | |
| Latent (Ours) | 75.57 | 58.75 | 58.52 | 64.78 | 48.21 | 46.10 | 4260.90 | 1.46 | |
| Gemini-3.6-Flash | Base | 76.02 | 60.78 | 57.25 | 65.14 | 49.67 | 45.19 | 2001.71 | 8.03 |
| Code2World | 75.80 | 57.61 | 55.50 | 64.37 | 48.08 | 44.15 | 9557.18 | 21.20 | |
| Diffusion2Image | 74.75 | 59.89 | 55.69 | 64.08 | 47.61 | 43.51 | – | 194.51 | |
| Latent (Ours) | 80.36 | 60.50 | 61.13 | 66.19 | 45.80 | 47.25 | 2658.46 | 8.49 | |
| GUI-Owl-7B | Base | 73.23 | 41.65 | 40.40 | 60.73 | 35.56 | 34.78 | 1120.64 | 2.14 |
| Code2World | 68.52 | 36.08 | 31.80 | 60.03 | 33.59 | 33.45 | 3890.48 | 9.00 | |
| Diffusion2Image | 66.39 | 37.72 | 33.11 | 59.01 | 34.76 | 33.20 | – | 100.88 | |
| Latent (Ours) | 74.16 | 43.55 | 41.96 | 60.99 | 35.23 | 35.06 | 1181.68 | 2.22 | |
| Qwen3-VL-8B | Base | 75.94 | 58.07 | 55.56 | 67.57 | 49.23 | 44.52 | 1391.27 | 4.70 |
| Code2World | 75.21 | 56.80 | 54.04 | 67.21 | 48.72 | 43.90 | 8170.46 | 14.88 | |
| Diffusion2Image | 74.29 | 56.73 | 54.96 | 67.26 | 49.30 | 43.31 | – | 171.68 | |
| Latent (Ours) | 77.15 | 58.95 | 57.09 | 68.35 | 50.40 | 45.62 | 1702.17 | 4.83 |
AC = AndroidControl. SR requires both the predicted action type and argument to be correct.
After removing shared current-state information, the t-SNE view shows how each method organizes action-induced future representations . Latent2World forms the clearest Click, Scroll, and Type clusters around their action centroids.
Action-specific t-SNE structure for Base, Diffusion, Code2World, and Latent2World. Latent (Ours) shows the strongest separation around the corresponding action centroids.
@misc{yu2026latentworldaction,
title = {Latent World Action Model for GUI Agents},
author = {Seungjun Yu and Hojun Choi and Seojeong Park and Jaeyo Shin and Hyunjung Shim},
year = {2026},
note = {Preprint}
}