Aug 2024 - Present
Remote, USA
Software Engineer (AI/ML team) · Highmark Inc.
Call-center voice training with an LLM on the other end of the line, the judge that grades it, and the registry and evals that decide which agents reach production.
- Shipped a real-time AI call-center training platform (React/FastAPI/WebSockets) where an LLM plays trainer-configured customer personas over the phone, replacing $1M+/year of consultant-led training.
- Designed the rubric-based LLM-as-judge that scores each completed roleplay, tuning it against coach-scored transcripts until its scores tracked the coaches' before trainees saw any feedback.
- Built an enterprise AI-agent lifecycle control plane governing versioned releases from staging to production via a centralized registry, durable evaluation runs, policy-gated GitLab promotions, and cross-environment observability.
- Architected a multi-agent evaluation harness combining specialized reviewers, deterministic checks, and owner-defined and hidden tests to catch quality, safety, latency, and cost regressions before promotion.
- Created a team-shared, access-controlled knowledge workspace on Vertex AI Search, implementing a RAG pipeline with metadata-filtered retrieval, configurable chunking, and citation-grounded answers over user-uploaded documents.
React · FastAPI · WebSockets · LLM-as-Judge · Multi-Agent Evaluation · RAG · Vertex AI Search · GitLab