PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents

Vighnesh Subramaniam1,2,*, Boris Katz1, Brian Cheung1,3, Chun-Liang Li2, Tomas Pfister2, Yale Song2
1 MIT CSAIL 2 Google Cloud AI Research 3 UCSF
*Correspondence: vsub851@mit.edu
Paper Code

Closed-loop generation inside the denoising process.

Overview of the PreviewDiff method.
PreviewDiff turns one open-loop denoising trajectory into a critic-guided tree over intermediate latent continuations.
Preview
Decode partial clean-latent estimates at selected checkpoints, without waiting for a full generation.
Critique
Ask a multimodal judge to score prompt satisfaction and propose compact correction notes.
Branch
Re-noise locally and continue under corrected conditioning, keeping the search near promising partial scenes.

Abstract

Diffusion models produce striking images and videos, but they still struggle with compositional details such as object counts, attribute binding, spatial relations, and temporally grounded actions. Best-of-N sampling can spend more compute at test time, but it is open-loop: it only chooses among completed outputs and cannot repair a promising trajectory before it fails.

PreviewDiff is a training-free test-time search method that converts diffusion sampling into multimodal, critic-guided search over intermediate latents. At selected denoising checkpoints, it decodes a partial preview, asks a multimodal judge to score and critique it, and uses the resulting natural-language feedback to branch over semantic prompt edits and locally re-noised latent continuations. Across image and video benchmarks, PreviewDiff improves over budget-matched Best-of-N and scalar-search baselines.


What changes relative to Best-of-N?

Best-of-N selection

Samples many complete outputs, scores each final image or video, and returns the best one. This is simple and strong, but the verifier only acts after mistakes have already been baked into the final sample.

PreviewDiff search

Uses the same kind of multimodal signal earlier. It previews partial denoising states, asks for a critique, branches over targeted semantic fixes plus local latent restarts, and prunes before spending full rollout compute.

  1. State: an intermediate clean-latent estimate, checkpoint time, positive/negative conditioning, and feedback history.
  2. Action: a compact positive fix note, optional avoid note, restart depth, and branch seed.
  3. Reward: a multimodal prompt-satisfaction score for the decoded preview.
  4. Tree policy: beam-pruned expansion with max backup over descendant values.
Conceptual diagram of PreviewDiff search states, critic feedback, and latent continuation.
The multimodal model is used both as a value estimator and as a policy-like source of semantic correction actions.

Results Summary

We find gains across image generation with SDXL, SD-3.5, and FLUX.2, and video generation with LTX-Video and Wan 2.2. The headline pattern is consistent: spending critic compute inside denoising beats spending it only on final-sample selection.

Scaling curves comparing PreviewDiff and Best-of-N for image and video generation.
PreviewDiff improves as denoiser-step budget increases and stays above budget-matched Best-of-N for both image and video generation.
Fixed-budget results

Image generation

ModelMethodGenEval2 Gemini P@4GenEval2 Qwen-3 ValGenEval2 Qwen-3 P@4T2I-CoReBench Gemini P@4T2I-CompBench++ Gemini P@4T2I-CompBench++ BLIP VQA
SDXLPreviewDiff2.2921.97651.9752.0153.4740.4844
SDXLBest-of-N2.0901.55231.7351.7132.9630.3976
SDXLEvoSearch2.103—1.61961.75433.0130.4521
SDXLParticle-Sample1.902—1.75441.91543.0030.4255
SD-3.5PreviewDiff1.9651.5881.8921.7652.0110.2566
SD-3.5Best-of-N1.7541.38541.7131.5451.9130.2011
FLUX.2 9BPreviewDiff3.1122.6751.9832.7813.7580.7851
FLUX.2 9BBest-of-N2.7752.5041.8752.6743.5100.7256

Video generation

ModelMethodT2V-CompBench Gemini P@4T2V-CompBench Qwen3-VL P@4LLaVA-EvalNarrLV Gemini P@4NarrLV Qwen-2.5 EvalV-Bench 2.0 Gemini P@4
LTX-Video 9BPreviewDiff3.0211.9870.45232.16762.173.162
LTX-Video 9BBest-of-N2.7161.4110.39221.54455.923.123
LTX-Video 9BEvoSearch2.8141.9010.37132.01261.143.071
LTX-Video 9BVideo-T12.8611.9540.43292.11561.873.124
Wan 2.2 A14BPreviewDiff3.5412.0190.63113.41171.223.554
Wan 2.2 A14BBest-of-N3.3621.8750.60112.87468.163.292
Wan 2.2 A14BVideo-T13.4651.9010.62663.229—3.509

Interactive PreviewDiff search tree

Prompt:

Move the slider to see the tree unfolding!

PreviewDiff tree animation frame.
Decode roots

The controls will scrub through the tree-building process.

Explore a root-to-branch-to-final path

You can across root previews and click a child node. The diagram redraws the tree edges and highlights the chosen continuation.

root 0
Root preview · checkpoint 5
Selected root preview. root 0 seed 98052 · cp5

Select a root to update its child branches.

Child branches · checkpoint 12
Final rollout
Final PreviewDiff output for the selected example.
Final PreviewDiff output

Selected child branch.


PreviewDiff search walkthrough

This paper figure follows one search from intermediate root previews through critic feedback and branch scoring to the selected final generation.

PreviewDiff qualitative walkthrough figure showing root previews, feedback, child scores, and final branch.
One complete PreviewDiff search, showing root previews, critic feedback, prompt edits, child scores, and the chosen final branch.

PreviewDiff vs Best-of-N qualitative examples

Use the slider to compare PreviewDiff with one Best-of-N sample.

Example

1 / 15

PreviewDiff example.
PreviewDiff
Best-of-N example.
Best-of-N

PreviewDiff vs Best-of-N video examples

01 · Warm streetlights on snow

1 / 10
Prompt

In a snowy nighttime setting, the streetlights cast a warm, orange glow over the quiet street.

PreviewDiff
Best-of-N

The two videos restart together when the example changes. Use “Play all” or “Restart” to compare motion from the same point in time.


What matters in the search?

The paper’s ablations show that search width produces the largest gains, while depth and semantic variants provide complementary improvements. Earlier checkpoint intervention is useful because it leaves enough denoising time for a correction to affect the sample.

Ablations over checkpoint timing, tree depth, semantic variants, and search width.
Search-axis ablations for images and videos.
Checkpoint timing and base model scaling results.
Earlier two-checkpoint schedules perform best, and gains remain across larger base models.

Human evaluation

Human raters compare PreviewDiff, Best-of-N, and EvoSearch on content and style. The strongest gains are on content, which matches the method’s focus on prompt adherence: objects, attributes, counts, relations, and actions.

Human evaluation of PreviewDiff compared with Best-of-N and EvoSearch.
PreviewDiff is preferred more often than the baselines, especially for content correctness.