I'm a statistics and computer science graduate who builds the pipelines, models, and visualizations that help teams and institutions see clearly and move with confidence — across public-sector, corporate, and consulting work alike.
I work at the intersection of statistics, software, and applied analytics. Across internships and research roles, the through-line has stayed the same: take messy, multi-source data and turn it into something a decision-maker can stand behind.
That's meant orchestrating ingestion pipelines from federal sources to study post-COVID migration, building file-validation systems that keep tens of thousands of records honest, and putting analyses in front of audiences like a county Chief Data Officer and an Economic Development Authority.
I care about the last mile of data work — the dashboard a director actually opens, the chart that reframes a debate, the validation routine that means nobody has to second-guess the numbers. Rigorous underneath, clear on the surface.
Designed and deployed an evaluation dashboard measuring investment-program effectiveness across income distributions, so leadership could pinpoint target communities and monitor annual impact. Built an executive-facing analytics dashboard tracking call volume, categorization, and average handling times, and automated data workflows with Python, SQL, and REST API integrations in a CI/CD environment — eliminating manual overhead on credential and permissions management.
Integrated and orchestrated ingestion pipelines from the FAO, U.S. Customs and Border Protection, and the U.S. Census — cleaning, standardizing, and transforming raw data with repeatable validation and deduplication routines to study post-COVID migration. Developed predictive and statistical analyses connecting food insecurity and migration across the Northern Triangle, and communicated findings through clear analytic narratives for policy-focused stakeholders.
Optimized data pipelines and front-end rendering for the Social Impact Data Commons platform using D3.js, improving accessibility of large datasets. Developed regex-based processing workflows to classify and organize 1,000+ repository files, enforcing schema compliance and quality-assurance standards across the platform.
Manipulated and analyzed a 166,000-row dataset in Python and R, improving data quality for more robust statistical analysis. Built a reusable web-scraping tool that cut data-collection time by 50%, ran regression analysis on economic variables related to altruistic behavior, and presented findings to local policymakers.
Engineered a Python-based automated pipeline to ingest, validate, and track billing and invoice records, improving billing efficiency by 10% and reducing processing errors. Built a reporting dashboard and analytical reports on monthly net-collection rates, lifting collections to 90% and increasing financial visibility for leadership.
Two command-line Python tools that convert between addresses and coordinates at scale. The geocoder batches large CSVs into 10,000-row chunks against the Census API and geocodes roughly 333 rows/second; the reverse geocoder uses the OpenStreetMap Nominatim API to fill addresses from latitude/longitude, validated against Google's API to within ~0.001° accuracy. Built with resumable runs, docstring-documented modules, and LaTeX documentation for reuse.
Read the report (PDF)A full-stack PHP/MySQL web app that surfaces the right song for a given mood or context. Designed a normalized relational schema (songs, lyrics, artists, moods, contexts, and linking tables for many-to-many tagging), then implemented triggers, stored procedures, and secured the app layer with prepared statements, input validation, and CSRF protection. Hosted on the UVA CS MySQL server with portable deployment to local or cloud MySQL.
View on GitHubAnalyzed ~11K Gallup survey responses merged with 50+ years of U.S. nuclear generation data to model how favorability toward nuclear energy shifts by political affiliation, age, and education over time. Built time-segmented logistic regression, decision tree, neural network, and SVM classifiers across three eras; the decision tree performed best, with political identity emerging as the strongest and most consistent predictor of nuclear support.
A CSV validation engine that enforces data integrity across the Social Impact Data Commons. Applies repeatable routines for validation, deduplication, and business-rule checks to keep 10,000+ data files consistent and accurate across many attributes and value ranges — catching schema violations and bad values before they reach the platform.
View on GitHubCombined Virginia DOE school-quality data with Census American Community Survey economics to quantify how income, parental education, work hours, and chronic absenteeism relate to elementary reading outcomes. Produced bivariate maps, correlation analysis, and LOESS-smoothed trends across every city elementary school, then translated the findings into five cost-conscious, actionable recommendations for the nonprofit.
Read the report (PDF)A multiple-regression study across all 50 states identifying which socioeconomic and policy factors best predict infant mortality rate. Ran stepwise variable screening, multicollinearity checks, nested F-tests, weighted least squares, and full regression-assumption diagnostics in SAS to land on a defensible, interpretable model — finding median income, racial composition, and abortion legislation to be meaningful predictors.
Read the report (PDF)Built a predictive classification model analyzing business diversity in Fairfax County, tuning SVM, Decision Tree, and Probit models with NLP techniques to track minority business ownership and cut classification error by 12%. Presented findings to the County's Chief Data Officer and Economic Development Authority.
View live projectI'm looking for data analyst, engineer, and scientist roles where rigorous work meets real impact. Happy to talk.