Watch it work: a specialist gets it wrong, and the close still can't complete
The strongest evidence for this architecture isn't a passing test — it's what happened on a fresh database no agent had ever seen, during the M8 end-to-end run.
AP gets a real judgment call wrong
AP's live duplicate-detector run found the seeded $14,200 CloudScale Hosting duplicate pair — and, on its own reasoning, dismissed it as a false positive.
"Two separate legitimate billing items."— AP Agent, live run, dismissing the actual seeded anomaly
Controller catches it anyway
Controller reruns its own independent duplicate scan rather than trusting AP's — precisely so a specialist's dismissal can't quietly become the last word. It flagged the identical pair, filed a correction request, and refused to mark the period ready once the combined unresolved exposure ($28,032.95) crossed the $25,000 materiality threshold.
The close blocks — correctly
Gate 1 failed. Not a bug: the system did exactly what it was designed to do when a specialist's judgment and an independent re-check disagree.
AP corrects itself on the very next pass
Re-run against the same database, AP saw Controller's correction request, correctly reversed the duplicate, and was explicit about what it found and why.
The paid duplicate represents a genuine double-payment needing vendor collections follow-up outside AP's scope.— AP Agent, second live run
Gate 1 passes, a human approves, the period closes
Controller's next pass found trial balance and subledgers tying out cleanly and zero duplicate candidates remaining. A human ran orchestrator/approve.py --approved-by "Jordan Ellis, VP Finance", and period_status flipped to closed for the first time against this database.