Scrub the denoising trajectory
A diffusion LM does not commit to an answer once. It refines the whole sequence over many denoising steps, and factuality moves during that process. Drag the slider on a real trace from the paper and watch where the answer goes and where the attention lands.
Question.
Gold answer:
Traces are the real denoising outputs reported in the paper's appendix. The attention graph is a schematic rendering of the structure described there, not a plot of raw attention weights.
Factuality is a property of the trajectory
Detectors built for auto-regressive models can be pointed at a diffusion LM, and several stay competitive. What they cannot see is the denoising trace, where the evidence actually accumulates. TDGNet turns each step into a sparsified attention graph, updates a memory per token, and reads out over the whole path.
No single step is enough
Static snapshots peak in the middle of the trajectory (0.66 AUROC) and fall off at both ends. Reading every step gives 0.68: the signal is distributed, not localised.
Structure carries the signal
Remove the attention edges and the model collapses to a per-token classifier at 0.40 to 0.49 AUROC. The relational structure, not the isolated hidden states, separates factual from hallucinated.
It holds across architectures
Four diffusion LMs from 7B to 16B, dense and Mixture-of-Experts. TDGNet reaches 0.76 on LLaDA 1.5 and 0.74 on LLaDA 2.1-mini, despite its block-local attention.
Three stages over the graph sequence
At every denoising step, the head-averaged attention map is thresholded into a sparse directed graph over tokens. The detector walks that sequence in denoising order.
-
Spatial aggregation
A message-passing network pools over each token's in-neighbourhood, so a token's representation reflects what it is attending to at that moment, not just its own hidden state.
m̄ᵢ⁽ᵗ⁾ = mean ψ(hⱼ, hᵢ, eⱼᵢ) -
Temporal memory
A GRU carries a persistent memory per token across steps, so the detector can tell steady grounding apart from a token that briefly looked fine and then drifted.
sᵢ⁽ᵗ⁾ = GRU(m̄ᵢ⁽ᵗ⁾, sᵢ⁽ᵗ⁺¹⁾) -
Trajectory readout
Temporal attention weights the steps that matter instead of averaging them, then pooling gives one score for the response or one score per token.
zᵢ = Σₜ αᵢ⁽ᵗ⁾ · Linear(sᵢ⁽ᵗ⁾)
Results
Response-level detection across Math, CommonsenseQA, HotpotQA and TriviaQA, plus token-level localisation and cost.
| Method | Math | CSQA | HotpotQA | TriviaQA | Average |
|---|---|---|---|---|---|
| LLaDA-8B-Instruct | |||||
| Semantic Entropy | 0.68 | 0.64 | 0.61 | 0.66 | 0.65 |
| Lexical Similarity | 0.51 | 0.53 | 0.54 | 0.54 | 0.53 |
| LN-Entropy | 0.70 | 0.59 | 0.55 | 0.55 | 0.60 |
| Perplexity | 0.67 | 0.60 | 0.51 | 0.54 | 0.58 |
| EigenScore | 0.56 | 0.54 | 0.56 | 0.59 | 0.56 |
| TSV | 0.72 | 0.61 | 0.55 | 0.50 | 0.60 |
| CCS | 0.59 | 0.56 | 0.60 | 0.54 | 0.57 |
| TDGNet | 0.72 | 0.65 | 0.64 | 0.72 | 0.68 |
| Dream-7B-Instruct | |||||
| Semantic Entropy | 0.59 | 0.50 | 0.68 | 0.69 | 0.62 |
| Lexical Similarity | 0.71 | 0.68 | 0.71 | 0.67 | 0.69 |
| LN-Entropy | 0.57 | 0.59 | 0.52 | 0.53 | 0.55 |
| Perplexity | 0.52 | 0.57 | 0.51 | 0.54 | 0.54 |
| EigenScore | 0.68 | 0.55 | 0.63 | 0.70 | 0.64 |
| TSV | 0.68 | 0.71 | 0.43 | 0.50 | 0.58 |
| CCS | 0.59 | 0.56 | 0.64 | 0.60 | 0.60 |
| TDGNet | 0.72 | 0.72 | 0.74 | 0.74 | 0.73 |
Cite
Swap in the ACL Anthology entry once the proceedings are published. Until then, this entry is the stable one.
@inproceedings{hemmat2026tdgnet,
title = {{TDGNet}: Hallucination Detection in Diffusion Language
Models via Temporal Dynamic Graphs},
author = {Hemmat, Arshia and Torr, Philip and
Chen, Yongqiang and Yu, Junchi},
booktitle = {Proceedings of the 2026 Conference on Empirical Methods
in Natural Language Processing (EMNLP)},
year = {2026},
address = {Budapest, Hungary},
publisher = {Association for Computational Linguistics}
}