Verify2Act: Critic-Guided Latent World Models for Verifying Language-Conditioned Manipulation Plans

Chrisantus Eze Christopher Crick
Oklahoma State University
Under review

A vision-language model proposes candidate plans, a latent world model imagines each one, and a critic checks the imagined outcome against the goal before the robot moves.

Overview

Abstract

Vision-Language Models (VLMs) can propose multi-step manipulation plans, but they have no grounded model of how a scene will change, and the pixel-level world models that could check those plans are slow and drift over long horizons. We present Verify2Act, a modular verification framework that separates semantic action proposal from latent physical verification. A VLM proposes candidate plans, a latent world model (V2A-WM) imagines each of them in the DINOv2 feature space, and a dual-head critic checks the imagined rollouts for temporal consistency and goal satisfaction before anything is executed. V2A-WM extends recent flow-matching world models with cross-attention action grounding and a causal temporal history window. We deploy the full system on a low-cost robot arm, training the world model and critic only on data generated in a digital twin. Across six language-conditioned tasks, Verify2Act reaches 83.3% task success, compared with 61.7% for a latent world model without our architectural changes, 53.3% for the VLM alone, and 46.7% for a pixel-diffusion world model. It verified every plan it executed and incurred no planning-induced failures: all of its failed episodes were caused by low-level grasp or placement slips. In simulation, Verify2Act obtains the highest five-task chain success on CALVIN among the compared baselines with 30× fewer reflective VLM calls than pixel-diffusion verification, and the highest success on cluttered nut assembly.

83.3%real-robot task success over six tasks (VLM alone: 53.3%)
0planning-induced failures: every executed plan was verified
30×fewer reflective VLM calls than pixel-diffusion verification on CALVIN

Method

Verify2Act treats the VLM only as a source of candidate plans. Each candidate is rolled out by V2A-WM in the frozen DINOv2 latent space, and a dual-head critic scores the imagined rollout: one head checks that consecutive imagined states are temporally consistent, the other scores the final imagined state against the language goal. Plans that fail are sent back to the VLM for re-planning; only an accepted plan reaches the robot.

Verify2Act architecture: a VLM proposes plans, V2A-WM imagines them in latent space, and a dual-head critic accepts or rejects them.
Verify2Act overview. Stage 1: a VLM proposes candidate multi-step plans from the camera observation and the language instruction. Stage 2: V2A-WM rolls out each candidate in the frozen DINOv2 latent space. Stage 3: the critic scores each rollout, and rejected plans trigger re-planning.

V2A-WM

V2A-WM builds on the flow-matching cores of DINO-WM and RLA-WM with two changes: cross-attention action grounding between DINOv2 patch features and CLIP tokens, which ties language commands to visual object patches, and a causal temporal history window with validity masks, which preserves visual memory across gripper occlusions.

V2A-WM training and inference diagram.
V2A-WM training and inference.

Sim-to-real through a digital twin

For the robot, the world model and the critic are trained only on data generated in a digital twin of the workspace (20k episodes). At test time, a real2sim step aligns the camera frame with the twin before the plan is imagined.

One planning call on the robot, showing the camera frame, the real2sim render, the imagined outcome of each candidate and its goal score.
One planning call on the robot (goal: “Put the red block to the left of the blue block”). After real2sim alignment of the camera frame, V2A-WM imagines each candidate plan and the goal head scores the imagined state. The first candidate is rejected on the basis of its imagined outcome; the second is accepted and executed.

Real-World Execution

A Yahboom DOFBOT with an arm-mounted camera and four colored blocks. Robot footage is sped up 4× or 6×, as labeled in each clip.

Rejecting a plan before it runs

Goal: “Clear all cool-colored blocks into the bin and leave the yellow block.” All three proposed candidates also bin the red block and are rejected by the goal head; the re-planned candidate is accepted.

Verify2ActBin clearing

The verified plan bins green and blue and leaves red and yellow on the table.

VLM-OnlyBin clearing, no verification

The same task without a world model or critic: the first proposed plan is executed and the red block goes into the bin.

Verify2ActSpatial relation

Goal: “Put the red block to the left of the blue block.” Plan accepted with goal score 0.99.

Verify2ActStacking

Goal: “Stack the blue block on top of the yellow block.” Plan accepted with goal score 1.00.

Failure caseCorrect plan, failed execution

The plan is correct and verified, but the low-level skill fails to complete the stack. All of Verify2Act's failed episodes were of this kind.

More rollouts

The bin-clearing task on two other layouts, unedited apart from the speed-up.

Verify2ActBin clearing, second layout

Green and blue are binned; red and yellow stay on the table.

Verify2ActBin clearing, third layout

Green and blue are binned; red and yellow stay on the table.

These demonstrations were filmed without the letter-size sheet used in the paper's evaluation. The planning panels were computed from the recorded camera frames and the plans the VLM proposed during the filmed runs.

Results

Real robot

All methods share the same VLM (GPT-4o), prompts and low-level skills, and differ only in how a proposed plan is verified.

Task success on the robot (%). Avg.: mean over the six tasks, 10 episodes each.
MethodBin-2Bin-coolBin-3Left-ofRight-ofStackAvg.
VLM-Only80402060804053.3
Diffusion-WM8040040804046.7
RLA-WM100503080605061.7
Verify2Act (Ours)100906090808083.3
Bar charts of task success per task and of episode outcomes per method.
Results on the robot. (a) Task success per task and averaged over tasks. (b) Outcome of the episodes of each method: success, execution failure (the plan was correct but a skill failed), or planning failure (the executed plan could not satisfy the goal).
Planning versus execution on the robot. Plan correct: executed plans that satisfy the goal if all skills succeed. Verified: executed plans accepted by the critic. Top-1: offline plan ranking on 42 real scenes (chance ≈ 6%). “—”: no critic.
VLM-OnlyDiffusion-WMRLA-WMVerify2Act
Plan correct (%)93.393.395.0100
Planning failures (%)6.76.75.00.0
Execution failures (%)43.346.733.316.7
Step success (%)69.261.573.082.5
Verified plans (%)——61.8100
Top-1, real2sim-aligned (%)——71.497.6
Top-1, camera frame (%)——23.871.4
VLM calls / episode1.002.001.841.13
Planning time / call (s)3.430.028.424.0
Episode time (s)81107126109

Simulation: CALVIN

CALVIN ABCD→D. SRk: fraction of sequences completing at least k consecutive subtasks. Avg. Len.: mean subtasks completed (max 5). Refl. Calls: reflective VLM queries triggered by critic rejections. Averaged over 3 seeds.
MethodSR1SR2SR3SR4SR5Avg. Len.Refl. Calls
MCIL37.32.70.40.00.00.38—
HULC88.973.358.747.538.33.06—
GR-194.989.684.478.973.14.21—
VLM-Only (GPT-4o)100.096.091.083.070.04.400
DINO-WM97.093.087.076.066.04.190
RLA-WM100.095.090.081.069.04.3516
Diffusion-WM100.096.091.079.066.04.32444
V2A-WM (AdaLN)100.096.089.081.071.04.3727
V2A-WM (Non-Causal)99.094.090.084.073.04.4033
Verify2Act (Ours)100.098.088.082.075.04.4315

Simulation: cluttered nut assembly

Cluttered nut assembly (Robosuite). SR: episode success. NCR: nut completion rate. Prec.: critic-accepted plan success. Rej.: critic rejection rate. Repl.: re-plans per episode. Averaged over 3 seeds.
MethodSRNCRPrec.Rej.Repl.
VLM-Only (GPT-4o)17.037.9——0.0
DINO-WM20.036.946.69.21.5
RLA-WM43.067.676.269.817.6
Diffusion-WM55.069.178.951.89.0
V2A-WM (AdaLN)30.047.355.155.07.1
V2A-WM (Non-Causal)31.051.158.149.15.4
Verify2Act (Ours)58.070.361.849.010.3

Citation

@misc{eze2026verify2act,
  title  = {Verify2Act: Critic-Guided Latent World Models for Verifying
            Language-Conditioned Manipulation Plans},
  author = {Eze, Chrisantus and Crick, Christopher},
  year   = {2026},
  note   = {Under review}
}