Negotiation Task Reader
Six-condition human-AI experiment platform with citation-grounded assistance, schema validation, parallel evidence scans, SSE streaming, model fallbacks, and revision-safe behavioral logging.
AI researcher and founding engineer. I work on what happens after the base model: evaluation, data pipelines, and human-AI systems.
The Problem Taxonomy project turns licensed company data and historical web archives into structured measurements. Every stage is recoverable and testable, from WARC retrieval through human-label evaluation.
I built the data recovery, LLM measurement, annotation, evaluation, orchestration, and analysis layers. The pipeline uses Athena to reduce the Common Crawl index before fetching WARC byte ranges, then applies structured LLM extraction, human-gold evaluation, and interpretable representation learning.
Constraint: 883 GiB of archived HTML, probabilistic model outputs, a licensed dataset that cannot be exposed, and a paper that must reproduce every number.
Six-condition human-AI experiment platform with citation-grounded assistance, schema validation, parallel evidence scans, SSE streaming, model fallbacks, and revision-safe behavioral logging.
Memory-enabled caregiving agent and realtime phone service spanning ADK function calling, Vertex AI Memory Bank, Firestore, Twilio Media Streams, audio transcoding, bounded queues, and VAD-gated barge-in.
An HBS research team uses this tool for diarized transcription, LLM paragraph segmentation, source search, and structured evidence review; it supports an ongoing 331-interview program localized for Chinese startups.
Research infrastructure for balanced assignment, resumable coding, and gold-set construction. More than 800 annotations across psychological frameworks feed agreement metrics and DSPy evaluation inputs.
eval_first()I build codebooks, independent human labels, inter-rater reliability, held-out metrics, ablations, and error analysis into the pipeline from the start.
resume_from(checkpoint)Pointer tables, bounded concurrency, content-addressed outputs, checkpoint rows, manifests, and versioned prompts turn expensive runs into inspectable state.
preserve(provenance)Evidence scope, model fallbacks, interaction revisions, completion semantics, and audit logs sit alongside structured-output schemas.
Designing AI for context analysis in humanitarian frontline negotiations.
Human evaluation of large language model based chatbots in a high-stakes support context.
A controlled experiment on reducing biased decision-making on dating websites.