Seattle, WA · U.S. Citizen

Jiangbin Huang

Data Scientist · Applied AI · Analytics Engineering

I build validated data pipelines, statistical and machine learning workflows, and LLM evaluation systems that turn messy inputs into reproducible decisions.

1.2M+
sensor records processed monthly
40%
SQL query performance improvement
200+
manual reporting hours saved annually
150+
LLM responses evaluated

Selected work

Production-oriented projects

LLM Response Quality Classification project visualization

Applied AI / NLP evaluation

LLM Response Quality Classification

A reproducible four-label classification workflow for evaluating accuracy, completeness, reasoning quality, and hallucination risk.

150+
reviewed outputs
1.00
macro F1 demo
2
baseline families
View technical work on GitHub
Three-Layer Data Lake Pipeline project visualization

Data engineering / analytics

Three-Layer Data Lake Pipeline

Configuration-driven ingestion with raw snapshots, validation, rejected-record handling, curated customer metrics, and SQL publishing.

3
data layers
98.4%
valid demo rows
3
reason-coded rejects
View technical work on GitHub
Point-in-Time S&P 500 Return Modeling project visualization

Quantitative research / point-in-time validation

Point-in-Time S&P 500 Return Modeling

A leakage-aware equity panel using historical index membership, liquidity controls, walk-forward model comparison, factor regimes, placebo tests, moving-block bootstrap, Reality Check, and locked holdout evaluation. The final analysis reports an exploratory signal rather than claiming persistent alpha.

704
historical members
8
locked OOS quarters
0.04
locked OOS R²
View full research walkthrough

Experience

From data contracts to model review

Co-Founder & Data Scientist

Early-Stage Data & AI Startup

Built Python and SQL workflows, a three-layer data lake pattern, model-selection pipelines, RAG evaluation, and repeatable technical handoff assets.

Data / AI Researcher

Feelie

Standardized 500K+ conversation records and connected DistilBERT emotion classification, LLM response generation, and rubric-based output review.

Data Analyst Intern

UW Engineering Services

Built validated ETL and SQL reporting workflows, improved query performance by 40%, and automated 200+ hours of annual reporting work.

Research Assistant

UW Department of Civil Engineering

Processed 1.2M+ sensor records per month with timestamp alignment, missing-value treatment, noise filtering, and feature engineering; built congestion-monitoring datasets, applied clustering to identify recurring traffic patterns and high-risk locations, and evaluated ARIMA and LSTM forecasting approaches for short-term traffic prediction.

Capabilities

Methods backed by working pipelines

Statistics & Machine Learning

Model selection, feature engineering, cross-validation, hypothesis testing, regression, classification, clustering, and time-series forecasting.

XGBoost · Random Forest · Logistic Regression · PCA · ARIMA · LSTM

Applied AI & LLM Systems

RAG workflows, response-quality classification, retrieval evaluation, prompt testing, embeddings, hallucination review, and model evaluation.

PyTorch · scikit-learn · TensorFlow · LangChain · OpenAI API

Data Engineering & Analytics

ETL development, data validation, layered data lakes, schema design, SQL optimization, analytics datasets, and reporting automation.

Python · SQL · pandas · AWS · Tableau · Power BI

Education

Applied mathematics foundation

Master of Science, Applied Mathematics

University of Washington · Seattle, WA

Scientific computing, machine learning for large-scale data, database systems, probability, and statistical modeling.

Bachelor of Science, Applied Mathematics

University of Washington · Seattle, WA

Data structures, algorithms, numerical analysis, statistical methods, and computational mathematics.