FaultLab
A team-built, AI-assisted lab for testing LLM agents under tool failures, with fixed-code checks and fresh challenges before a proposed repair can replace the baseline.

Problem
An agent that receives no response from a tool can repeat a completed action or report success without evidence. A passing retry alone does not establish that a recovery rule is dependable.
Approach
Co-developed a Python/FastAPI and React lab with isolated SQLite worlds, injected tool faults and eight fixed-code checks. Saved regression evidence can be reviewed separately from fresh tests; configuration checks, model-call limits and saved request IDs keep reruns bounded and retries tied to the same execution.
Impact
The historical September 14 campaign matched 143 of 143 Weave traces to local records. Through the September 30 continuation, zero generated policies were accepted: incomplete fault exposure and lab errors remained visible, with the original baseline retained.
Key Metrics
Technologies
Links
My Role
Co-developed with Aryan Bhusari as Team Gatekeeper, followed by continued AI-assisted engineering. The implementation and experiment results describe team capabilities; individual ownership of every later component is not established.
Team Size: 2 people