01
Problem
The initial prototype could draft useful recommendations, but every edge case required informal judgment. Operators lacked confidence in when to trust the output, when to escalate, and which source material had influenced a response.
- No explicit routing between auto-draft, human review, and escalation.
- Source documents mixed stable policy, working notes, and one-off exceptions.
- Quality feedback lived in conversations instead of reusable evaluation fixtures.