Full-Stack • 2026

Earshot

A WebRTC voice capture agent with typed tools and an interface that shows disagreement between speech recognition and model-written customer names.

Earshot — Voice to records. Conceptual illustration of a voice waveform becoming typed records, with an amber mismatch visible between the two paths.
AI-generated conceptual illustration.

Problem

Turning a conversation into structured records introduces two sources of uncertainty: what the recognizer heard and what the model decided to write.

Approach

Built a realtime voice interface with seven typed tools, closed-vocabulary validation and structured errors that guide the model’s next call. The public MVP visualizes tool actions and offers two repair paths when the recognized name and tool argument disagree.

Impact

Name disagreement appeared in six of seven manually observed sessions. That small observation motivated a visible correction interaction; it is not an accuracy benchmark or a measured reduction in CRM work. The expanded 2.0 implementation is user-confirmed but remains outside the linked public main branch.

Key Metrics

7 typed tools
Public MVP tool interface
6 of 7
Name divergence · manual sessions

Technologies

Next.jsReactTypeScriptWebRTCOpenAI Realtime APIZodWeb Audio

Links

My Role

Designed and built the voice/tool interaction with AI assistance and wrote the phased 2.0 specification. Collaboration and handoff are recorded in the public repo; the user-confirmed 2.0 implementation is unpublished.