Seattle, WA · U.S. Citizen

Jiangbin Huang

Data Scientist · Applied AI · Analytics Engineering

I build validated data pipelines, statistical and machine learning workflows, and LLM evaluation systems that turn messy inputs into reproducible decisions.

1.2M+
sensor records processed monthly
40%
SQL query performance improvement
200+
manual reporting hours saved annually
150+
LLM responses evaluated

Selected work

Projects I can walk through

Support ticket routing project visualization

Customer support tooling

Support ticket routing

I wanted a small project that went beyond a notebook: a message goes through an API, the prediction is saved, and uncertain cases are left for a reviewer. It runs locally with FastAPI, Streamlit, and SQLite.

31K
training rows
8.5K
test rows
0.60
review cutoff
Read the project notes
LLM Response Quality Classification project visualization

Applied AI / NLP evaluation

LLM Response Quality Classification

I built a preference-classification pipeline around 135K public Arena votes, then tested grouped splits, soft labels, pairwise scoring, and a frozen temporal holdout.

135K
human votes
0.285
best dev macro F1
27K
locked test rows
Read the project notes
Regional Energy Data Lake project visualization

Energy data engineering / forecasting

Regional Energy Data Lake

I joined hourly prices, generation, weather, and archived forecast runs for DE-LU, France, and Austria. The raw files stay available, while the curated panel is used to compare forecast errors with price movement.

30
forecast runs
3
market zones
21,600
Gold rows
Read the project walkthrough
Point-in-Time S&P 500 Return Modeling project visualization

Quantitative research / point-in-time validation

Point-in-Time S&P 500 Return Modeling

I built a stock panel with historical index membership, walk-forward evaluation, transaction costs, placebo tests, and a locked holdout. The result is exploratory evidence, not a claim of persistent alpha.

704
historical members
8
locked OOS quarters
0.04
locked OOS R²
Read the project walkthrough

Experience

Work that involved data, models, and review

Co-Founder & Data Scientist

Early-Stage Data & AI Startup

Built Python and SQL workflows, a three-layer data lake pattern, model-selection pipelines, RAG evaluation, and repeatable technical handoff assets.

Data / AI Researcher

Feelie

Standardized 500K+ conversation records and connected DistilBERT emotion classification, LLM response generation, and rubric-based output review.

Data Analyst Intern

UW Engineering Services

Built validated ETL and SQL reporting workflows, improved query performance by 40%, and automated 200+ hours of annual reporting work.

Research Assistant

UW Department of Civil Engineering

Processed 1.2M+ sensor records per month with timestamp alignment, missing-value treatment, noise filtering, and feature engineering; built congestion-monitoring datasets, applied clustering to identify recurring traffic patterns and high-risk locations, and evaluated ARIMA and LSTM forecasting approaches for short-term traffic prediction.

Capabilities

What I use on actual projects

Statistics & Machine Learning

I use fixed effects, regression, clustering, cross-validation, and time-series models when the question needs an interpretable comparison.

scikit-learn · statsmodels · XGBoost · ARIMA · LSTM

LLM and response review

I build small evaluation sets, write labels that a reviewer can apply, and check where a model is incomplete, ungrounded, or simply wrong.

PyTorch · scikit-learn · LangChain · OpenAI API

Data pipelines and reporting

I use Python and SQL to move raw files into tables that can be checked, rerun, and explained to someone who did not write the pipeline.

Python · SQL · pandas · Parquet · AWS · Tableau

Education

Applied mathematics foundation

Master of Science, Applied Mathematics

University of Washington · Seattle, WA

Scientific computing, machine learning for large-scale data, database systems, probability, and statistical modeling.

Bachelor of Science, Applied Mathematics

University of Washington · Seattle, WA

Data structures, algorithms, numerical analysis, statistical methods, and computational mathematics.