NeurIPS 2026 Main track
ORCA: Hunting Compositional Failures in Text-to-Image Diffusion
Arshia Hemmat*, Amirhossein Vahidi*, Amitis Shidani, Mohammad Vali Sanian, Hesam Asadollahzadeh, Aryan Yazdan Parast, Mohammad Lotfollahi

Ask a text-to-image model for "a red cube on a blue sphere" and you will often get a blue cube, or a red sphere, or the cube underneath. These failures are predictable: attributes bind to the wrong objects, spatial relations invert, and scenes with several objects lose count.
We show the model already has the information it needs to get this right. What it lacks is a training signal that tells it to line that information up with what it draws.
Why this was hard
The obvious diagnosis is that the text encoder throws compositional structure away. That is true of CLIP, whose contrastive training compresses a sentence into something close to a bag of concepts, and it is exactly why recent architectures add a T5 encoder alongside it. But the failures survive the extra encoder. So the information is not missing.
Our argument is that it is misaligned. T5 keeps the structure of "red cube on blue sphere", but in a representation space shaped by language modelling rather than by vision, and the denoising objective never directly rewards the model for connecting the two. Supplying that connection means answering two questions at once: which part of the text signal carries the composition, and what visual signal should it be matched against.
How it works
ORCA, for Orthogonal Residual Compositional Alignment, answers both with one auxiliary loss added to ordinary diffusion training.
On the text side, we do not use T5 directly. We use the residual between T5 and CLIP: a learned linear map \(W\) carries the T5 embedding into CLIP's space, and \(\Delta y = W z_y^T - z_y^C\) is whatever T5 knows that CLIP cannot linearly express. That residual is where the compositional structure CLIP discards ends up living.
On the visual side, the target comes from a frozen self-supervised encoder. We take DINO features of the training image and keep only the top principal components, giving a low-rank visual target \(z(x)\). The cross-modal information the model needs turns out to be concentrated in that low-rank subspace, and we prove a bound to match: the cross-modal information recoverable at a given rank is limited by the spectral mass of the visual encoder's covariance in its top components.
The two meet through a small predictor. An MLP \(g_\theta\) maps the residual \(\Delta y\) to a matrix, and a QR decomposition turns it into an orthogonal basis \(K(\Delta y)\). That basis is prompt-dependent, so each prompt selects its own readout subspace. We take the hidden state \(h_T\) from an intermediate transformer block and train with
\[\mathcal{L}_{\mathrm{ORCA}} = \left\lVert \mathrm{sg}[z(x)] - K(\Delta y)^\top h_T \right\rVert^2\]
alongside the usual diffusion loss, where \(\mathrm{sg}\) stops gradients flowing into the target. The predictor exists only during training. At inference the model is unchanged, so the improvement costs nothing at generation time.
Across three diffusion-transformer backbones, DiT-B/2, DiT-L/2 and U-ViT-L, ORCA improves both FID and GenEval over the vanilla model and over REPA. On DiT-L/2 it reaches FID 16.65 and GenEval 0.291 at 200K training steps, beating the strongest baseline trained for 400K steps at half the cost. Put the other way, it reaches the vanilla model's 400K-step FID in 4.8 times fewer steps, and its GenEval in 3.6 times fewer. The largest gains land exactly where the failures were: attribute binding and spatial relations.
What it does not do
ORCA is a training-time method. It improves models trained with the loss; it does nothing for a model that has already been trained without it, short of further training. The bound we prove also cuts both ways: a low-rank target can only carry the cross-modal information that sits in the top components of the visual encoder, so whatever compositional structure lives outside that subspace is out of reach by construction. Raising the rank buys more of it, at the cost of a less focused target.
Accepted to the NeurIPS 2026 main track. Paper on arXiv, and the full project page with the explainer video.