
LLM Response Quality Classification
A reproducible four-label classification workflow for evaluating accuracy, completeness, reasoning quality, and hallucination risk.
- 150+
- reviewed outputs
- 1.00
- macro F1 demo
- 2
- baseline families
Seattle, WA · U.S. Citizen
Data Scientist · Applied AI · Analytics Engineering
I build validated data pipelines, statistical and machine learning workflows, and LLM evaluation systems that turn messy inputs into reproducible decisions.
Selected work

A reproducible four-label classification workflow for evaluating accuracy, completeness, reasoning quality, and hallucination risk.

Configuration-driven ingestion with raw snapshots, validation, rejected-record handling, curated customer metrics, and SQL publishing.

A leakage-aware equity panel using historical index membership, liquidity controls, walk-forward model comparison, factor regimes, placebo tests, moving-block bootstrap, Reality Check, and locked holdout evaluation. The final analysis reports an exploratory signal rather than claiming persistent alpha.
Experience
Early-Stage Data & AI Startup
Built Python and SQL workflows, a three-layer data lake pattern, model-selection pipelines, RAG evaluation, and repeatable technical handoff assets.
Feelie
Standardized 500K+ conversation records and connected DistilBERT emotion classification, LLM response generation, and rubric-based output review.
UW Engineering Services
Built validated ETL and SQL reporting workflows, improved query performance by 40%, and automated 200+ hours of annual reporting work.
UW Department of Civil Engineering
Processed 1.2M+ sensor records per month with timestamp alignment, missing-value treatment, noise filtering, and feature engineering; built congestion-monitoring datasets, applied clustering to identify recurring traffic patterns and high-risk locations, and evaluated ARIMA and LSTM forecasting approaches for short-term traffic prediction.
Capabilities
Model selection, feature engineering, cross-validation, hypothesis testing, regression, classification, clustering, and time-series forecasting.
XGBoost · Random Forest · Logistic Regression · PCA · ARIMA · LSTM
RAG workflows, response-quality classification, retrieval evaluation, prompt testing, embeddings, hallucination review, and model evaluation.
PyTorch · scikit-learn · TensorFlow · LangChain · OpenAI API
ETL development, data validation, layered data lakes, schema design, SQL optimization, analytics datasets, and reporting automation.
Python · SQL · pandas · AWS · Tableau · Power BI
Education
University of Washington · Seattle, WA
Scientific computing, machine learning for large-scale data, database systems, probability, and statistical modeling.
University of Washington · Seattle, WA
Data structures, algorithms, numerical analysis, statistical methods, and computational mathematics.