Delta Sentinel
An auditor for claimed agent improvements that checks variance, grader tampering, contamination and overfitting, with a recorded self-audit that rejected its own claimed gain.

Problem
A reported benchmark improvement can come from noise, leaked tasks or a changed grader. The claim needs evidence beyond a single headline score.
Approach
Built four audit probes, paired statistical comparisons and an offline replay interface in FastAPI and Vue. A synthetic fixture lineup exercises the intended failure modes; two adapters connect the audit machinery to supported evaluation setups.
Impact
The recorded self-audit returned NOT_SUPPORTED for its author’s claimed improvement. Five purpose-built fixtures exercised the probes correctly, which establishes fixture behavior rather than real-agent performance or broad benchmark coverage.
Key Metrics
Technologies
Links
My Role
Defined the audit thesis, interfaces and evidence rules, directed AI-assisted implementation across separate lanes, and investigated calibration and the rejected self-audit.