- Design assessments and use performance data to improve instructional outcomes.
- Communicate complex quantitative concepts to diverse learners.
Mathematics graduate building end-to-end data systems — from rigorous statistical modeling to production data pipelines on the cloud. I turn messy data into things that hold up under proof.
I work the full distance between theory and production. A math background gives me the modeling rigor; hands-on work with the modern data stack gives me the engineering. Below is the same person, viewed two ways.
Designing pipelines and warehouses that move and reshape data reliably — orchestration, modeling, and cloud infrastructure. Comfortable from SQL optimization down to bootloaders and paging.
End-to-end predictive modeling grounded in statistics — feature work, model selection, and honest evaluation. I care about what a metric actually means, not just that it went up.
A statistical engine for experiment analysis built from first principles — Welch's t-test, Mann-Whitney U, Bonferroni and Benjamini-Hochberg corrections, and exact noncentral-t power analysis — behind a Streamlit interface where a product manager uploads an experiment CSV and gets a full report with effect sizes, bootstrap CIs, and minimum-detectable-effect diagnostics. Every guarantee is verified by Monte Carlo simulation, not assumed.
End-to-end ML pipeline comparing Logistic Regression, Random Forest, and XGBoost on telecom customer records, with SHAP interaction analysis to surface the real drivers of churn. Deployed as a live Streamlit app.
Compared regression tree, OLS, and random forest models on Queens co-op and condo transactions. Used permutation importance to rank predictors and handled high-missingness features with median imputation plus missingness indicators.
A 9-script PostgreSQL analytics pipeline over 100K+ Olist orders, with reusable views for revenue, retention, delivery, and seller performance. Surfaced a clear link between delivery speed and review scores in a 5-page Power BI dashboard.
A production-grade modern data stack pipeline streaming 100K+ Olist e-commerce orders through Kafka into Snowflake, orchestrated by Airflow DAGs. dbt models the raw data into a dimensional schema with 21 passing tests and full generated documentation.
Open to Data Engineering and Data Science roles. The fastest way to reach me is email.