Systems I have built and evaluated.

Selected work across LLM measurement, data infrastructure, online experiments, realtime agents, and research software. Each case study records the constraints, design decisions, and failures behind the implementation.

ENGINEERING CASE STUDIES
01

Problem Taxonomy Research System

DATA + LLM EVALUATION · 2025–PRESENT

The pipeline recovered 30.1M archived pages, extracted structured measurements for 88,336 startups, built human gold standards, optimized LLM classifiers with DSPy/GEPA, and organized embeddings into 265 interpretable themes.

02

Negotiation Task Reader

HUMAN-AI EXPERIMENT SYSTEM · 2026

A six-condition platform that controls the evidence layer while varying AI behavior. Includes provenance-aware retrieval, structured outputs, SSE streaming, experiment assignment, Postgres persistence, and a research dashboard.

03

CareCorgi Agent + Voice

AGENT + VOICE INFRASTRUCTURE · 2025

FastAPI and Google ADK backend with function-calling care tools, Firestore sessions, Vertex AI Memory Bank, plus a Twilio-to-Gemini Live audio path with transcoding, backpressure, barge-in, and reconnect handling.

DEPLOYED RESEARCH TOOLS
04

HBS Interview Workbench

REACT · TYPESCRIPT · VERCEL · IN USE

Diarized transcription, LLM paragraph segmentation, source search, and structured evidence review for an HBS research team conducting an ongoing 331-interview program, including a localized Chinese-startup workflow.

05

Annotation Infrastructure

NEXT.JS · PRISMA · POSTGRES · IN USE

Balanced assignment, reservations, resumable sessions, quality checks, and exports for human coding. The resulting 800+ annotations support inter-rater reliability and gold-label inputs to model evaluation.

RESEARCH ARCHIVE

Earlier human-AI research

These earlier projects cover controlled experiments, computational analysis, and human evaluation of AI systems.