Historical case study
From HR voice agent to evidence-based assistance
The original prototype conducted recruitment screening calls and produced scores and a recommendation. The pipeline worked technically; that evaluation was also its regulatory risk.
Original scope
A voice agent asked five application questions. A second model scored communication, motivation, and role fit and returned Hire, Maybe, or Reject.
Core finding
A recruitment evaluation does not become low-risk through friendlier wording. Recommendations, rankings, and person scores influence decisions and should be treated as high-risk functionality.
EU AI Act, Annex III ↗Robust redesign
- predefined job-related criteria instead of open personality judgments
- transcript correction and source evidence for every extracted claim
- abstention when evidence is missing instead of an invented score
- no emotion, language, personality, or origin inference
- no automatic rejection and a mandatory human decision
- versioning, audit logs, quality metrics, and documented risk management
Product decision
The public portfolio demo now uses a different conversation scenario. The active system reviews product concepts. The HR artifacts remain in the monorepo but are isolated from every live route.