DrawingVQA benchmarks multi-depth visual-textual reasoning on real construction drawings — a domain where getting it wrong has physical consequences
DrawingVQA is a real-world benchmark for multi-depth visual-textual reasoning over construction drawings. Technical drawings combine dense symbolic notation, cross-referenced sheets and spatial reasoning — a combination general multimodal benchmarks do not test.
Domain benchmarks like this matter more than their citation counts suggest. General multimodal evaluation has drifted toward natural images and screenshots, which are the easy case: the symbols are conventional and the reasoning is shallow. A construction drawing requires following a reference across sheets and holding a spatial model while doing it.
It is also a domain with a real error cost. A model that misreads a drawing does not produce a bad caption, it produces a bad building. Benchmarks anchored to consequences tend to produce more honest numbers than benchmarks anchored to preference.
Kimbodo — AI Research & Papers — July 20, 2026 → · arXiv — Multiagent Systems and multimodal benchmark listings →