Andrew Amore
Data science leader with a foundation in software engineering, statistical modeling, and production machine learning.
About Me

I build data-driven systems that turn complex information into actionable insights. My work spans scalable data pipelines, statistical modeling, and deployment/maintenance of production machine learning systems.
I hold an M.S. in Statistical Science from Duke University and graduated Cum Laude from The Ohio State University with a background in mechanical engineering, computer science, and business. I've held professional positions supporting national defense, analytical consulting, pharmaceutical research and insurance. This breadth of experience, in combination with a diverse educational background, lets me bridge the gap between research and real-world deployment.
I am passionate about mentoring teams, fostering a culture of analytical rigor, and leading initiatives that translate data into strategic advantage. I am always looking for opportunities to add value and drive impact in any team role.
Skills & Technologies
Portfolio
A selection of projects spanning statistical modeling, machine learning, and data engineering.
Information Retrieval in Natural Language Processing
Modern Question Answering (QA) systems consist of two components: readers and retrievers. Retrievers reduce the passage search space for answer extraction and limit the overall accuracy of QA methods. Conventional retrievers consume large amounts of resources, reducing their viability to large corporations or well funded institutions.
A Box Office Analysis: Factors Influencing Film Profitability
Movie production companies are interested in understanding what factors contribute to a financially successful film. Box office data from all 2019 film releases was collected from Kaggle and enhanced with additional features to address this question while controlling for potential confounders using a hierarchical model.
Bayesian Phylogenetics
An immune response produces antibodies to remove infectious material in the human body. A "slow" evolutionary process, called clonal selection, uses natural selection to identify a binding match between antibody and antigen receptors. One identification strategy uses an ancestral tree, called phylogenetic tree reconstruction, to infer candidate antibodies.
Standard Error Estimation for Clustered Data
In causal inference, experimental data is often collected from groups of individuals which form clusters. When estimating a statistical quantity from grouped data it is important for statisticians to take the clustering structure into account when performing standard error (SE) estimation. Literature suggests many conventional SE estimation methods ignore or underestimate grouping and which can cause severe downward bias in parameter estimation.
NFL Big Data Bowl
The NFL wants to reduce the risk of non-contact injury to professional athletes. This analysis statistically evaluates the influence of playing surface on non-contact injury risk using hypothesis testing and logistic regression. Geospatial data was recorded every 0.1 seconds (10 Hz) for different positional players across 267,000+ plays and modeled in GCP.
