Skip to content

Resume

Simon Xu

Software Engineer (AI/ML)

  • Production AI Systems
  • MS CSE Georgia Tech
  • US citizen, no sponsorship needed
  • Open to relocation
  • Applied AI · Full-Stack AI · Forward Deployed Engineer

01

Experience

Aug 2024 - Present
Remote, USA

Software Engineer (AI/ML team) · Highmark Inc.

Call-center voice training with an LLM on the other end of the line, the judge that grades it, and the registry and evals that decide which agents reach production.

  • Shipped a real-time AI call-center training platform (React/FastAPI/WebSockets) where an LLM plays trainer-configured customer personas over the phone, replacing $1M+/year of consultant-led training.
  • Designed the rubric-based LLM-as-judge that scores each completed roleplay, tuning it against coach-scored transcripts until its scores tracked the coaches' before trainees saw any feedback.
  • Built an enterprise AI-agent lifecycle control plane governing versioned releases from staging to production via a centralized registry, durable evaluation runs, policy-gated GitLab promotions, and cross-environment observability.
  • Architected a multi-agent evaluation harness combining specialized reviewers, deterministic checks, and owner-defined and hidden tests to catch quality, safety, latency, and cost regressions before promotion.
  • Created a team-shared, access-controlled knowledge workspace on Vertex AI Search, implementing a RAG pipeline with metadata-filtered retrieval, configurable chunking, and citation-grounded answers over user-uploaded documents.

React · FastAPI · WebSockets · LLM-as-Judge · Multi-Agent Evaluation · RAG · Vertex AI Search · GitLab

May 2024 - Aug 2024
Remote, USA

Software Engineer Intern (AI/ML team) · Highmark Inc.

The internal LLM gateway, a policy Q&A bot that cites its sources, and the multimodal assistant that became the company's GenAI platform.

  • Developed a Retrieval-Augmented Generation system using Flask, integrating embedding models and Gemini to answer employee queries on internal policies with cited sources; surveyed users reported 70% less search time.
  • Engineered the internal LLM API gateway exposing a single Gemini endpoint to engineering teams company-wide, with schema-validated JSON output, per-consumer cost attribution, and circuit-breaking on provider errors.
  • Built and scaled an internal multimodal AI assistant on Vertex AI (Gemini), enabling employees to process text, PDFs, images, and voice through a React and FastAPI stack; serves 30K users and processes 6M+ prompts annually.

React · FastAPI · Flask · Vertex AI · Gemini · RAG · Embedding Models

02

Education

Master of Science in Computational Science & Engineering

Georgia Institute of Technology

2022-2024 · December 2024

Bachelor of Arts in Applied Economics

University of California - Santa Barbara

2019-2021 · March 2021

  • Microsoft AI Hackathon participant
  • Career Essentials in Generative AI by Microsoft and LinkedIn (issued April 2024)

03

Skills

Languages
PythonTypeScriptJavaScriptGoSQLAgentic AI Coding (Claude Code, Codex, Cursor)
AI / LLM
RAGAgent Harness EngineeringLLM-as-JudgeAgent EvalsFine-tuningEmbedding Models
Frameworks
PyTorchHugging Face TransformersPEFTvLLMLangChainTemporalMCPFastAPIReact
Data / Infra
PostgreSQL/pgvectorRedisDockerKubernetesCI/CDAWSGCPOpenTelemetry

04

Selected projects

Proxy LoopJul 2026 - Present

  • Fine-tuned Qwen3-8B via SFT + 4-bit QLoRA on 1.5K provenance-tracked agent trajectories synthesized and filtered by a Claude Sonnet teacher; lifted held-out task completion 58%→67% and cut false completion 6%→2%.
  • Architected a Fast/Slow agent harness separating low-latency voice interaction from long-horizon planning and tool use, with shared state, retries, approval gates, and completion checks across phone, browser, and email.
  • Implemented durable long-running agent workflows with Temporal + PostgreSQL, persisting approval waits and scheduled follow-ups while using idempotent retries and callback deduplication to recover across worker restarts.

Python · Qwen3-8B · QLoRA · PEFT · vLLM · Temporal · PostgreSQL

ConductorApr 2026 - Jun 2026

  • Engineered a Go-based distributed agent runtime with event-sourced recovery and idempotent tool execution, allowing long-running LLM workflows to resume after crashes without replaying completed side effects.
  • Built a Kubernetes-native execution layer with PostgreSQL leases and fencing tokens to coordinate distributed workers, handling pod failures, graceful shutdowns, and rolling updates while preserving single-writer execution.
  • Added end-to-end OpenTelemetry + Prometheus observability across model and tool calls, validating throughput, p95/p99 latency, and failure recovery through chaos and load testing.

Go · Kubernetes · PostgreSQL · Event Sourcing · OpenTelemetry · Prometheus

EngramJan 2026 - Mar 2026

  • Reduced repeated context tokens by approximately 40% through cross-session state handoffs.
  • Consolidated raw sessions into durable short- and long-term memory with stale-memory cleanup and conflict detection.
  • Validated retrieval reliability with Recall@5, MRR, and traces for ranking regressions and injected context.

Python · MCP · Event-Driven Architecture · Agent Memory · BM25 · Evaluation

FigBrainApr 2025 - May 2025

  • Converted natural-language prompts into FigJam components at approximately 85% first-try accuracy.
  • Reduced API costs by approximately 28% and lowered latency with PostgreSQL-backed memory, caching, and model routing.

Python · TypeScript · LangGraph · MCP · Figma API · PostgreSQL

↵ ask · esc close

Keyboard shortcuts

⌘ K / Ctrl K
Ask the portfolio AI
/
Ask the portfolio AI
g h
Go to Home
g p
Go to Projects
g r
Go to Resume
g a
Go to Ask
g m
Go to Music
t
Toggle light / dark
?
This sheet
h i
Say hi