Andrew Amore

Data science leader with a foundation in software engineering, statistical modeling, and production machine learning.

About Me

Andrew Amore

I build data-driven systems that turn complex information into actionable insights. My work spans scalable data pipelines, statistical modeling, and deployment/maintenance of production machine learning systems.

I hold an M.S. in Statistical Science from Duke University and graduated Cum Laude from The Ohio State University with a background in mechanical engineering, computer science, and business. I've held professional positions supporting national defense, analytical consulting, pharmaceutical research and insurance. This breadth of experience, in combination with a diverse educational background, lets me bridge the gap between research and real-world deployment.

I am passionate about mentoring teams, fostering a culture of analytical rigor, and leading initiatives that translate data into strategic advantage. I am always looking for opportunities to add value and drive impact in any team role.

Skills & Technologies

Statistical Modeling Bayesian & Frequentist Methods Machine Learning Production ML Systems Data Pipelines Python Cloud (AWS) Docker

Portfolio

A selection of projects spanning statistical modeling, machine learning, and data engineering.

Dec 2022

Information Retrieval in Natural Language Processing

Modern Question Answering (QA) systems consist of two components: readers and retrievers. Retrievers reduce the passage search space for answer extraction and limit the overall accuracy of QA methods. Conventional retrievers consume large amounts of resources, reducing their viability to large corporations or well funded institutions.

NLP Information Retrieval Transfer Learning DPR
Read More
Oct 2022

A Box Office Analysis: Factors Influencing Film Profitability

Movie production companies are interested in understanding what factors contribute to a financially successful film. Box office data from all 2019 film releases was collected from Kaggle and enhanced with additional features to address this question while controlling for potential confounders using a hierarchical model.

Statistics Hierarchical Model Regression Data Visualization
Read More
Apr 2022

Bayesian Phylogenetics

An immune response produces antibodies to remove infectious material in the human body. A "slow" evolutionary process, called clonal selection, uses natural selection to identify a binding match between antibody and antigen receptors. One identification strategy uses an ancestral tree, called phylogenetic tree reconstruction, to infer candidate antibodies.

Bayesian Statistics MCMC Phylogenetics Simulation
Read More
Apr 2022

Standard Error Estimation for Clustered Data

In causal inference, experimental data is often collected from groups of individuals which form clusters. When estimating a statistical quantity from grouped data it is important for statisticians to take the clustering structure into account when performing standard error (SE) estimation. Literature suggests many conventional SE estimation methods ignore or underestimate grouping and which can cause severe downward bias in parameter estimation.

Causal Inference Bootstrap Monte Carlo Standard Errors
Read More
Dec 2019

NFL Big Data Bowl

The NFL wants to reduce the risk of non-contact injury to professional athletes. This analysis statistically evaluates the influence of playing surface on non-contact injury risk using hypothesis testing and logistic regression. Geospatial data was recorded every 0.1 seconds (10 Hz) for different positional players across 267,000+ plays and modeled in GCP.

GCP Computer Vision ResNet-50 Logistic Regression Geospatial
Read More