Latent World Action Model for GUI Agents

Preprint
* Equal contribution.
† Corresponding author.

Abstract

Pixel-space world models have recently emerged as a promising approach for GUI agents, predicting future screens prior to execution and subsequently re-perceiving them for action selection. However, this two-stage paradigm faces two key challenges in GUI environments: (L1) the high computational cost of pixel reconstruction and (L2) the need to reconstruct large and diverse appearance-level details beyond what is required for action selection. In practice, this screen-generation objective can induce an identity shortcut problem, where future-state prediction largely preserves the current state rather than capturing action-induced changes. In this work, we reformulate GUI world modeling by shifting the prediction target from full-screen reconstruction to action-induced future information in latent space. Concretely, we leverage a large multimodal model as the latent world model, whose shared representations support two dedicated heads for both action selection and predictive transition modeling. The action head directly translates these latent representations into action selection, while the transition head captures action-conditioned future information. To mitigate the identity shortcut, we further introduce an action-aware contrastive objective that encourages the transition representations to capture discriminative action-induced changes. Extensive experiments across offline and online GUI settings show that our method achieves up to 31.9% and 40.6% relative gains in success rate over pixel-space world models, respectively, while substantially reducing inference latency by up to 98.9%.

Why Latent World Modeling?

Figure 2 highlights two mismatches in pixel-space GUI world modeling. (a) Rendering future screens adds substantial runtime, while Latent2World retains a stronger success–efficiency trade-off. (b) GUI transitions show lower similarity and much wider variation than natural-video transitions, making exact future-screen reconstruction an unnecessarily difficult target.

Performance-efficiency gains and GUI transition variability

(a) Performance–efficiency trade-off of GUI world models. (b) Transition-similarity distributions measured with DINOv2 and SigLIP2.

These observations motivate predicting decision-relevant future information in latent space instead of reconstructing the full next screen.

Mitigating the Identity Shortcut

Figure 3 exposes the identity shortcut . (a) Even when the ground-truth next screen changes drastically, pixel-space predictors can remain close to the current GUI. (b) As transition size increases, shortcut severity rises and downstream SR drops for pixel-space world models, while Latent2World keeps both trends substantially more stable.

Identity shortcut example and shortcut severity across transition sizes

(a) Identity-shortcut example. (b) Shortcut severity and downstream success rate across weak, medium, and strong GUI transitions.

On strong transitions , shortcut severity reaches 98% for Diffusion and 66% for Code2World, versus 41% for Latent2World; the corresponding SRs are 50.0, 48.5, and 61.5. Larger GUI changes therefore amplify identity preservation in pixel space, whereas latent prediction better preserves action-relevant performance.

Latent World-Action Model

Latent2World learns action selection and action-conditioned future prediction in one shared multimodal latent space, so predictive supervision directly shapes the representation used for decision making.

Latent world-action model training and inference

(a) Training: the Action Head scores candidates while the Transition Head learns the action-induced next-state latent with an InfoNCE objective. (b) Inference: the world model re-scores low-confidence decisions directly in latent space, without rendering a future screen.

Training
Action scoring + action-aware future prediction

The true next state is the positive target; the current state and memory-bank transitions are negatives. This explicitly teaches the latent to represent what the action changes .

Inference
Single-pass latent rescoring

A confidence gate invokes the world model only when needed. The transition head is discarded at inference, avoiding the render-and-re-perceive loop.

Experimental Results

We report Online task-level evaluation on AndroidWorld and Offline step-level evaluation on AndroidControl and AITZ. Online is shown by default; use the buttons or swipe on mobile to switch views.

Swipe left / right to switch between Online and Offline
Online Evaluation

AndroidWorld

Across five backbones on AndroidWorld, Latent2World improves Overall SR over the corresponding base agent by +4.5 to +12.5 points , while avoiding the large runtime cost of pixel-space future generation.

Backbone Setting Easy SR Medium SR Hard SR Overall SR Tokens/Task Mean Steps Run Time
GPT-5.6 Base 68.33 50.00 16.67 54.46 363.8K 13.12 127.4
Code2World 68.33 52.94 22.22 56.25 525.4K 13.71 370.9
Diffusion2Image 68.33 47.06 16.67 53.57 – 13.21 2320.3
Latent (Ours) 73.33 58.82 27.78 61.61 329.4K 12.20 141.4
Gemini-3.6-Flash Base 70.00 52.94 33.33 58.93 204.6K 11.42 168.6
Code2World 71.67 52.94 22.22 58.04 310.0K 12.55 482.3
Diffusion2Image 68.33 52.94 22.22 55.75 – 11.68 2490.3
Latent (Ours) 73.33 58.82 38.89 63.39 200.4K 11.12 228.1
GUI-Owl-7B Base 25.00 5.88 0.00 15.18 226.1K 16.67 159.5
Code2World 26.67 5.88 5.56 16.69 421.64K 18.12 659.4
Diffusion2Image 26.67 5.88 0.00 16.07 – 16.81 1260.2
Latent (Ours) 33.33 11.76 5.56 22.32 229.0K 16.99 205.3
Qwen3-VL-8B Base 41.67 14.71 5.56 27.68 105.9K 11.21 125.8
Code2World 38.33 20.59 11.11 28.57 288.4K 14.10 519.9
Diffusion2Image 53.33 14.71 5.56 33.93 – 13.59 789.9
Latent (Ours) 55.00 26.47 16.67 40.18 129.3K 13.40 223.8
Qwen3-VL-32B Base 56.67 32.35 11.11 41.96 320.8K 13.29 703.6
Code2World 58.33 35.29 11.11 43.75 597.3K 16.84 1453.6
Diffusion2Image 55.00 26.47 5.56 38.39 – 17.81 3154.6
Latent (Ours) 60.00 38.23 16.67 46.42 385.4K 16.26 960.0
Offline Evaluation

AndroidControl & AITZ

On AndroidControl and AITZ, Latent2World improves action-type prediction across all evaluated backbones and increases step-level SR in most backbone–dataset settings with modest overhead.

Backbone Setting AC Type AC Ground. AC SR AITZ Type AITZ Ground. AITZ SR Tokens/Task Run Time
GPT-5.6 Base 71.92 60.46 59.84 59.84 47.99 41.30 3585.57 1.14
Code2World 70.34 59.61 58.22 58.60 47.03 39.24 12634.21 9.31
Diffusion2Image 68.31 55.01 57.94 56.40 46.80 38.59 – 135.20
Latent (Ours) 75.57 58.75 58.52 64.78 48.21 46.10 4260.90 1.46
Gemini-3.6-Flash Base 76.02 60.78 57.25 65.14 49.67 45.19 2001.71 8.03
Code2World 75.80 57.61 55.50 64.37 48.08 44.15 9557.18 21.20
Diffusion2Image 74.75 59.89 55.69 64.08 47.61 43.51 – 194.51
Latent (Ours) 80.36 60.50 61.13 66.19 45.80 47.25 2658.46 8.49
GUI-Owl-7B Base 73.23 41.65 40.40 60.73 35.56 34.78 1120.64 2.14
Code2World 68.52 36.08 31.80 60.03 33.59 33.45 3890.48 9.00
Diffusion2Image 66.39 37.72 33.11 59.01 34.76 33.20 – 100.88
Latent (Ours) 74.16 43.55 41.96 60.99 35.23 35.06 1181.68 2.22
Qwen3-VL-8B Base 75.94 58.07 55.56 67.57 49.23 44.52 1391.27 4.70
Code2World 75.21 56.80 54.04 67.21 48.72 43.90 8170.46 14.88
Diffusion2Image 74.29 56.73 54.96 67.26 49.30 43.31 – 171.68
Latent (Ours) 77.15 58.95 57.09 68.35 50.40 45.62 1702.17 4.83

AC = AndroidControl. SR requires both the predicted action type and argument to be correct.

Action-Specific Latent Structure

After removing shared current-state information, the t-SNE view shows how each method organizes action-induced future representations . Latent2World forms the clearest Click, Scroll, and Type clusters around their action centroids.

t-SNE visualization of action-induced future representations

Action-specific t-SNE structure for Base, Diffusion, Code2World, and Latent2World. Latent (Ours) shows the strongest separation around the corresponding action centroids.

BibTeX

@misc{yu2026latentworldaction,
  title  = {Latent World Action Model for GUI Agents},
  author = {Seungjun Yu and Hojun Choi and Seojeong Park and Jaeyo Shin and Hyunjung Shim},
  year   = {2026},
  note   = {Preprint}
}