ORCA

Hunting compositional failures in text-to-image diffusion

Arshia Hemmat1,2,3,4*, Amirhossein Vahidi1,3*, Amitis Shidani5,6, Mohammad Vali Sanian1,7,8, Hesam Asadollahzadeh1,9, Aryan Yazdan Parast9, Mohammad Lotfollahi1,2,3,4

1Wellcome Sanger Institute  2Cambridge Stem Cell Institute  3Cambridge Centre for AI in Medicine  4Department of Medicine, University of Cambridge  5University of Oxford  6Apple  7University of Helsinki  8Institute for Molecular Medicine Finland  9University of Melbourne

*Equal contribution

The paper in 68 seconds. No sound needed: everything is captioned.Download MP4

Hunt the failure

You are the evaluator. Each round shows a prompt and the picture a model drew. Decide whether it matches, and if not, what went wrong. These are the same checks GenEval runs, and the same failures ORCA was built to hunt.

Scenes are drawn by this page to illustrate GenEval's failure categories. They are not samples from ORCA or any other model.

Why image models lose the plot

Every object can look perfect and the picture still be wrong. The usual suspect is the text encoder. CLIP, trained contrastively, reads a prompt almost like a bag of words: it knows which concepts appear, not how they bind. SD3 and FLUX added a T5 encoder to keep that structure, yet the failures persist. Our argument: the information is there, but the denoising objective never directly rewards tying it to what the image looks like.

One extra loss, during training only

ORCA pulls a middle block of the diffusion transformer toward a compact summary of the real image, through a readout chosen by the prompt. At inference it is switched off: zero extra parameters, memory or FLOPs.

ORCA diagram: the T5 minus CLIP residual chooses a projection K of the block-8 hidden state h; the projected readout is pulled toward the top-64 DINOv2 principal components of the real training image.
Default setting on DiT-L/2: block 8 of 24, rank 64, loss weight 1.0.

Residual text signal

Delta y equals W times the T5 embedding minus the CLIP embedding

A learned map from T5 into CLIP space. What is left over is what T5 knows that CLIP does not.

Low-rank visual target

z of x equals the top-n principal components of the centred DINOv2 features

Top principal components of frozen DINOv2 features. Fixed before training, so the target cannot collapse.

Prompt-chosen readout

K of Delta y is the QR orthonormalisation of an MLP applied to Delta y

A small MLP turns the residual into an orthonormal basis, so the prompt decides which slice of the hidden state is read out.

L ORCA equals the squared norm of stop-gradient z of x minus K transpose h T

Added to the usual diffusion loss with a single weight. Performance is robust over about one order of magnitude of that weight.

Half the training, better images

MS-COCO at 256×256, three diffusion-transformer backbones, the same optimiser and schedule for every method. Only the auxiliary loss changes.

16.65FID on DiT-L/2 at 200K steps, vs 24.01 vanilla and 20.05 REPA at 400K
0.291GenEval on DiT-L/2 at 200K steps, vs 0.247 vanilla and 0.275 REPA at 400K
≈3×vanilla's GenEval score on colour attribution and position, same 150K-step budget
ORCA Vanilla REPA REG

Where the gains land

GenEval per task on DiT-L/2, all methods at 150K steps.

Vanilla REPA ORCA

Inside ORCA

The loss only sees the part of the hidden state inside the prompt-chosen subspace; everything orthogonal to it is left free. Scrub through training to watch that subspace fill up and line up with the visual target.

0K
Energy of h inside the subspace
58%
Cosine with the visual target
0.52

Rank 64 inside a 768-dimensional state. Values follow the measured curves in the appendix, which level off at about 80K steps.

Different prompts, different subspaces

Mean principal angle between the readout subspaces of two prompts. The further apart the meaning, the further apart the subspaces.

    Colour swap Object change Relation added Several changes

    Cite

    BibTeX
    @inproceedings{hemmat2026orca,
      title     = {{ORCA}: Hunting Compositional Failures in
                   Text-to-Image Diffusion},
      author    = {Hemmat, Arshia and Vahidi, Amirhossein and
                   Shidani, Amitis and Sanian, Mohammad Vali and
                   Asadollahzadeh, Hesam and Yazdan Parast, Aryan and
                   Lotfollahi, Mohammad},
      booktitle = {Advances in Neural Information Processing Systems},
      year      = {2026},
      eprint    = {2610.09841},
      archivePrefix = {arXiv}
    }

    arXiv abstractPDFCode on GitHub