Spectral Detection of Report Template Leakage in Radiology Visual Question Answering Corpora
- Authors
-
-
Hamza Bennani
Department of Computer Science, Polydisciplinary Faculty of Larache, Abdelmalek Essaadi University, Route de Rabat, Larache 92000, Morocco
Author
-
Mehdi Elami
Department of Information Technology, Higher School of Technology, Sultan Moulay Slimane University, Mghila Airport Road, Beni Mellal 23000, Morocco
Author
-
Anas Cherkaoui
Department of Computer Engineering, National School of Applied Sciences, Ibn Zohr University, Agadir 80000, Morocco
Author
-
- Abstract
-
Radiology visual question answering datasets are commonly derived from imaging reports, structured labels, and question templates. This construction can accelerate dataset creation, but it may also introduce statistical artifacts that make the answer predictable from wording patterns rather than from the intended clinical content. This paper develops an empirical leakage audit for report-derived radiology question answering corpora. The study examines whether template identifiers, lexical operators, and report-origin signatures encode answer information beyond the clinical finding being queried. A simulated multi-source corpus of 52,800 binary chest-imaging questions is generated with controlled prevalence, template reuse, phrase inheritance, and annotation noise. Three leakage diagnostics are evaluated: conditional mutual information, spectral separability of question-template embeddings, and a sparse logistic probe trained without image variables. The proposed spectral residual test isolates label-aligned directions in the question-template matrix after removing finding and institution effects. Across 200 Monte Carlo replications, high-leakage corpora showed a mean residual spectral ratio of 0.417, compared with 0.092 under leakage-controlled generation. The sparse lexical probe achieved 82.6% accuracy in high-leakage settings while receiving no image or finding-severity information, and permutation testing rejected conditional label independence in 96.5% of high-leakage runs. A leakage-reduced resampling procedure lowered probe accuracy to 56.8% while preserving finding prevalence and question coverage. These results show that dataset artifacts can be measured before model training and that corpus-level spectral tests can reveal answer-bearing language regularities that ordinary label summaries miss.
- References
- Downloads
- Published
- 2026-04-04
- Section
- Articles
- License
-
Copyright (c) 2026 authors

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
