03 / PROOF

Proof, not promises.

Measured results, provider evidence, and the failure that changed the number we report.

Repaired diagnostic result

Terra recall
88.9%
Repaired deterministic
58.7%
Difference
+30.2pp
95% CI
[13.7, 46.9]

POST-HOC SEGMENTATION-REPAIRED DIAGNOSTIC

REAL RAZORPAY TEST MODE PROOF

A read-only check against the real provider.

Recorded August 29, 2026. No communication was sent, no money moved, and no live lookup occurs in the public fixture demo.

PAYMENT
CAPTURED
ORDER
PAID
VERDICT
BLOCK

WHAT BROKE / F012

The evaluation instrument had a segmentation defect.

The frozen baseline kept bare URLs inside semantic units, against the written ontology. IoU matching turned real detections into misses. We preserved the original comparator, built a segmentation-only diagnostic repair, and reduced the reported advantage from +55.6pp to +30.2pp.

The smaller number is the public headline.

LIMITATIONS

The confirmatory claim remains unmade.

  • The planned fresh confirmatory corpus was not executed before submission.
  • Pilot power was 0.62 against the pre-registered target of 0.80.
  • The stricter lower-CI-bound interpretation does not clear +0.15.
  • Only GPT-5.6 Terra was evaluated as the semantic classifier.

FALSE-POSITIVE COST

Blocking can also be wrong.

A false positive means PREFLIGHT blocks a legitimate recovery intervention. The cost is delayed or lost recovery, plus unnecessary human review. No rupee value is claimed.