New photogrammetry self-validation protocol can't detect 100m errors even at 1.00 confidence
Confidence scores saturate at near-perfect while true errors swing up to 14x
Behnam Asadi's new paper on arXiv formalizes a track-leakage-free hold-out self-validation protocol for photogrammetric reconstruction. The method withholds a deterministic subset of images and re-localizes each against only 3D points seen by at least two retained images—creating a track-level barrier so no view is tested against structure it helped create. The held-out observations still enter the bundle adjustment that built the trusted structure, yielding an mAA confidence score. Across operational GNSS-referenced captures plus ETH3D, EuRoC, and IMC 2025 benchmarks, the protocol reports four key findings: (i) The protocol is computationally well-posed—good reconstructions score millidegree self-consistency. (ii) Thresholded self-consistency saturates and does not track absolute accuracy: confidence stays near 1.00 while true RTK error swings up to 14x within a capture, and per-capture correlation is sign-unstable (-0.57 to +0.98). (iii) It flags gross failure only when the failure destroys internal consistency—a fragmenting model drops confidence, but a single self-consistent globally-distorted model evades detection. Three of four captures gave one model wrong by 55–106 m at confidence 1.00. (iv) The same dichotomy holds on independent ETH3D and IMC 2025 ground truth.
This is a negative/characterization result: a ground-truth-free self-consistency signal is shown to saturate and be blind to coherent global distortion. The protocol measures internal geometric consistency, not absolute accuracy—neither a substitute for control-point assessment nor a general gross-failure gate. Asadi releases the protocol and degradation harness for the community. For professionals relying on photogrammetric inspection without ground truth, this serves as a critical warning: high self-consistency does not guarantee accuracy. Automated inspection pipelines that trust such self-validation risk missing large-scale distortion errors of 100+ meters.
- Track-leakage-free hold-out protocol achieves millidegree self-consistency on good reconstructions
- Confidence saturates at 1.00 even when true RTK error is up to 14x larger, with sign-unstable cross-capture correlation
- Globally distorted models escaped detection: three of four captures had 55-106m errors at 1.00 confidence
Why It Matters
High self-confidence in photogrammetric models does not guarantee accuracy—external control points remain essential.