AI/ML • 2026

Delta Sentinel

An auditor for claimed agent improvements that checks variance, grader tampering, contamination and overfitting, with a recorded self-audit that rejected its own claimed gain.

Delta Sentinel — Evidence auditing. Conceptual illustration of paired evidence records under an audit lens, with an unresolved amber mismatch.
AI-generated conceptual illustration.

Problem

A reported benchmark improvement can come from noise, leaked tasks or a changed grader. The claim needs evidence beyond a single headline score.

Approach

Built four audit probes, paired statistical comparisons and an offline replay interface in FastAPI and Vue. A synthetic fixture lineup exercises the intended failure modes; two adapters connect the audit machinery to supported evaluation setups.

Impact

The recorded self-audit returned NOT_SUPPORTED for its author’s claimed improvement. Five purpose-built fixtures exercised the probes correctly, which establishes fixture behavior rather than real-agent performance or broad benchmark coverage.

Key Metrics

NOT_SUPPORTED
Recorded self-audit verdict

Technologies

PythonFastAPINumPySciPyVue 3PiniaECharts

Links

My Role

Defined the audit thesis, interfaces and evidence rules, directed AI-assisted implementation across separate lanes, and investigated calibration and the rejected self-audit.