TILT: Model-Intrinsic Reward Alignment
for Compositional Diffusion

Debottam Dutta, Jianchong Chen, Jaehoon Hahm, Romit Roy Choudhury

Under Review

Abstract

Diffusion models often struggle with compositional prompts, producing samples in which some concepts dominate while others are missing or weakly represented. TILT is a training-free framework that poses eventual pure-mode sampling as a reward for intermediate-time alignment. The reward is intrinsic to the model, yields a tractable target under a variational approximation, and can be applied across modalities. Experiments on vision and audio benchmarks show improved compositional alignment while preserving sample quality compared with prior baselines.

Main Question

How should intermediate generation be steered so that the denoised samples at t=0 are drawn from the pure modes?

Pure Modes

A pure mode is a region of the final clean-sample distribution that represents the full composition. Its samples express the requested concepts together, without one dominating or disappearing.

Model-Intrinsic Reward

TILT rewards intermediate states that lead to pure modes at the final sample. The reward comes from the diffusion model itself, contrasting joint and constituent conditions without an external reward model.

TL;DR: Instead of heuristically correcting intermediate diffusion states, TILT treats landing in a pure compositional mode at the final sample as a model-intrinsic reward and derives test-time guidance toward that target.

Text-to-Image

Qualitative comparisons across compositional prompts.

Prompt: a green school bus and a red bag

SDXL
SDXL: a green school bus and a red bag
CFG++
CFG++: a green school bus and a red bag
CO3
CO3: a green school bus and a red bag
DoS
DoS: a green school bus and a red bag
SuperDiff
SuperDiff: a green school bus and a red bag
R2F
R2F: a green school bus and a red bag
TILT (Ours)
TILT (Ours): a green school bus and a red bag

Prompt: The soft yellow duckling swam next to the sleek black swan.

SDXL
SDXL: The soft yellow duckling swam next to the sleek black swan.
CFG++
CFG++: The soft yellow duckling swam next to the sleek black swan.
CO3
CO3: The soft yellow duckling swam next to the sleek black swan.
DoS
DoS: The soft yellow duckling swam next to the sleek black swan.
SuperDiff
SuperDiff: The soft yellow duckling swam next to the sleek black swan.
R2F
R2F: The soft yellow duckling swam next to the sleek black swan.
TILT (Ours)
TILT (Ours): The soft yellow duckling swam next to the sleek black swan.

Prompt: a metallic jewelry and a wooden spoon

SDXL
SDXL: a metallic jewelry and a wooden spoon
CFG++
CFG++: a metallic jewelry and a wooden spoon
CO3
CO3: a metallic jewelry and a wooden spoon
DoS
DoS: a metallic jewelry and a wooden spoon
SuperDiff
SuperDiff: a metallic jewelry and a wooden spoon
R2F
R2F: a metallic jewelry and a wooden spoon
TILT (Ours)
TILT (Ours): a metallic jewelry and a wooden spoon

Prompt: an oval coffee table and a square end table

SDXL
SDXL: an oval coffee table and a square end table
CFG++
CFG++: an oval coffee table and a square end table
CO3
CO3: an oval coffee table and a square end table
DoS
DoS: an oval coffee table and a square end table
SuperDiff
SuperDiff: an oval coffee table and a square end table
R2F
R2F: an oval coffee table and a square end table
TILT (Ours)
TILT (Ours): an oval coffee table and a square end table

Prompt: a photo of a pizza right of a banana

SDXL
SDXL: a photo of a pizza right of a banana
CFG++
CFG++: a photo of a pizza right of a banana
CO3
CO3: a photo of a pizza right of a banana
DoS
DoS: a photo of a pizza right of a banana
SuperDiff
SuperDiff: a photo of a pizza right of a banana
R2F
R2F: a photo of a pizza right of a banana
TILT (Ours)
TILT (Ours): a photo of a pizza right of a banana

Prompt: a photo of a person and a sink

SDXL
SDXL: a photo of a person and a sink
CFG++
CFG++: a photo of a person and a sink
CO3
CO3: a photo of a person and a sink
DoS
DoS: a photo of a person and a sink
SuperDiff
SuperDiff: a photo of a person and a sink
R2F
R2F: a photo of a person and a sink
TILT (Ours)
TILT (Ours): a photo of a person and a sink

Prompt: a photo of a surfboard and a suitcase

SDXL
SDXL: a photo of a surfboard and a suitcase
CFG++
CFG++: a photo of a surfboard and a suitcase
CO3
CO3: a photo of a surfboard and a suitcase
DoS
DoS: a photo of a surfboard and a suitcase
SuperDiff
SuperDiff: a photo of a surfboard and a suitcase
R2F
R2F: a photo of a surfboard and a suitcase
TILT (Ours)
TILT (Ours): a photo of a surfboard and a suitcase

Prompt: a photo of a red train and a purple bear

SDXL
SDXL: a photo of a red train and a purple bear
CFG++
CFG++: a photo of a red train and a purple bear
CO3
CO3: a photo of a red train and a purple bear
DoS
DoS: a photo of a red train and a purple bear
SuperDiff
SuperDiff: a photo of a red train and a purple bear
R2F
R2F: a photo of a red train and a purple bear
TILT (Ours)
TILT (Ours): a photo of a red train and a purple bear

Text-to-Audio

Listen to SCORE-CLAP and SCORE-TILT for the same prompt.

Prompt: A series of doors sliding open as gusts of wind blows and glass clanks

SCORE-CLAP
SCORE-TILT (Ours)
Spectrogram comparison for A series of doors sliding open as gusts of wind blows and glass clanks

Prompt: Water running softly in the background followed by several shrill beeps and a man speaking

SCORE-CLAP
SCORE-TILT (Ours)
Spectrogram comparison for Water running softly in the background followed by several shrill beeps and a man speaking

Prompt: Wind noise while a water vehicle is traveling across water, and a man talks

SCORE-CLAP
SCORE-TILT (Ours)
Spectrogram comparison for Wind noise while a water vehicle is traveling across water, and a man talks

Prompt: A gun firing several times followed by a revolver chamber spinning and metal clanking as a man talks and a person grunts

SCORE-CLAP
SCORE-TILT (Ours)
Spectrogram comparison for A gun firing several times followed by a revolver chamber spinning and metal clanking as a man talks and a person grunts

Prompt: A man speaks as rain pitter-patters and thunder rumbles

SCORE-CLAP
SCORE-TILT (Ours)
Spectrogram comparison for A man speaks as rain pitter-patters and thunder rumbles
Method and pure-mode intuition

Compositional diffusion models can drift toward regions where a joint prompt behaves too similarly to one constituent concept, causing concept dominance or omission. Prior correctors suppress these collisions locally at intermediate noise levels.

TILT instead defines the desired outcome in data space: a final clean sample should lie in a pure joint mode. The resulting reward is intrinsic to the pretrained model, contrasting joint-condition likelihood against constituent-condition likelihoods.

Pure-mode reward and non-commutativity intuition
Intermediate correction need not commute with denoising to the final data distribution; TILT aligns generation with the eventual pure-mode objective.
Concept dominance

The reward is computed from the diffusion model itself rather than an external vision-language or preference model. Under the paper's variational approximation, it becomes a contrast between the model's joint and single-concept denoising behavior.

This makes the same principle usable without training a modality-specific reward model.

↑ Back to Top