Hunt the failure
You are the evaluator. Each round shows a prompt and the picture a model drew. Decide whether it matches, and if not, what went wrong. These are the same checks GenEval runs, and the same failures ORCA was built to hunt.
Scenes are drawn by this page to illustrate GenEval's failure categories. They are not samples from ORCA or any other model.
Why image models lose the plot
Every object can look perfect and the picture still be wrong. The usual suspect is the text encoder. CLIP, trained contrastively, reads a prompt almost like a bag of words: it knows which concepts appear, not how they bind. SD3 and FLUX added a T5 encoder to keep that structure, yet the failures persist. Our argument: the information is there, but the denoising objective never directly rewards tying it to what the image looks like.
One extra loss, during training only
ORCA pulls a middle block of the diffusion transformer toward a compact summary of the real image, through a readout chosen by the prompt. At inference it is switched off: zero extra parameters, memory or FLOPs.
Residual text signal

A learned map from T5 into CLIP space. What is left over is what T5 knows that CLIP does not.
Low-rank visual target

Top principal components of frozen DINOv2 features. Fixed before training, so the target cannot collapse.
Prompt-chosen readout

A small MLP turns the residual into an orthonormal basis, so the prompt decides which slice of the hidden state is read out.
Half the training, better images
MS-COCO at 256×256, three diffusion-transformer backbones, the same optimiser and schedule for every method. Only the auxiliary loss changes.
Where the gains land
GenEval per task on DiT-L/2, all methods at 150K steps.
Inside ORCA
The loss only sees the part of the hidden state inside the prompt-chosen subspace; everything orthogonal to it is left free. Scrub through training to watch that subspace fill up and line up with the visual target.
Rank 64 inside a 768-dimensional state. Values follow the measured curves in the appendix, which level off at about 80K steps.
Different prompts, different subspaces
Mean principal angle between the readout subspaces of two prompts. The further apart the meaning, the further apart the subspaces.
Cite
@inproceedings{hemmat2026orca,
title = {{ORCA}: Hunting Compositional Failures in
Text-to-Image Diffusion},
author = {Hemmat, Arshia and Vahidi, Amirhossein and
Shidani, Amitis and Sanian, Mohammad Vali and
Asadollahzadeh, Hesam and Yazdan Parast, Aryan and
Lotfollahi, Mohammad},
booktitle = {Advances in Neural Information Processing Systems},
year = {2026},
eprint = {2610.09841},
archivePrefix = {arXiv}
}