AI/ML • 2026

FaultLab

A team-built, AI-assisted lab for testing LLM agents under tool failures, with fixed-code checks and fresh challenges before a proposed repair can replace the baseline.

FaultLab — Agent reliability. Conceptual illustration of isolated test worlds, an amber tool fault and a teal gate leading to evidence records.
AI-generated conceptual illustration.

Problem

An agent that receives no response from a tool can repeat a completed action or report success without evidence. A passing retry alone does not establish that a recovery rule is dependable.

Approach

Co-developed a Python/FastAPI and React lab with isolated SQLite worlds, injected tool faults and eight fixed-code checks. Saved regression evidence can be reviewed separately from fresh tests; configuration checks, model-call limits and saved request IDs keep reruns bounded and retries tied to the same execution.

Impact

The historical September 14 campaign matched 143 of 143 Weave traces to local records. Through the September 30 continuation, zero generated policies were accepted: incomplete fault exposure and lab errors remained visible, with the original baseline retained.

Key Metrics

8 checks
Code referee
143 of 143
Historical trace verification · Sep 14
0 · baseline retained
Generated policies accepted · Sep 30

Technologies

PythonFastAPISQLiteReactTypeScriptW&B Weave

Links

My Role

Co-developed with Aryan Bhusari as Team Gatekeeper, followed by continued AI-assisted engineering. The implementation and experiment results describe team capabilities; individual ownership of every later component is not established.

Team Size: 2 people