03 / PROOF
Proof, not promises.
Measured results, provider evidence, and the failure that changed the number we report.
Repaired diagnostic result
- Terra recall
- 88.9%
- Repaired deterministic
- 58.7%
- Difference
- +30.2pp
- 95% CI
- [13.7, 46.9]
POST-HOC SEGMENTATION-REPAIRED DIAGNOSTIC
REAL RAZORPAY TEST MODE PROOF
A read-only check against the real provider.
Recorded August 29, 2026. No communication was sent, no money moved, and no live lookup occurs in the public fixture demo.
- PAYMENT
- CAPTURED
- ORDER
- PAID
- VERDICT
- BLOCK
WHAT BROKE / F012
The evaluation instrument had a segmentation defect.
The frozen baseline kept bare URLs inside semantic units, against the written ontology. IoU matching turned real detections into misses. We preserved the original comparator, built a segmentation-only diagnostic repair, and reduced the reported advantage from +55.6pp to +30.2pp.
The smaller number is the public headline.LIMITATIONS
The confirmatory claim remains unmade.
- The planned fresh confirmatory corpus was not executed before submission.
- Pilot power was 0.62 against the pre-registered target of 0.80.
- The stricter lower-CI-bound interpretation does not clear +0.15.
- Only GPT-5.6 Terra was evaluated as the semantic classifier.
FALSE-POSITIVE COST
Blocking can also be wrong.
A false positive means PREFLIGHT blocks a legitimate recovery intervention. The cost is delayed or lost recovery, plus unnecessary human review. No rupee value is claimed.