CV

Massachusetts Bay Transportation Authority (MBTA)

2024-present

As a data engineer in the Data Resource Group, I divide my time between "building/maintaining data tools" and "making sure we're asking the right questions about our data tools".

  • Created Python libraries for analyzing and updating Informatica ETL pipelines, enabling the team to identify and fix problems before they impact stakeholders, and saving hundreds of hours of debugging time
  • Created Python libraries to connect data pipelines between SQL Server, Informatica, and Tableau, surfacing high-impact operations insights to MBTA stakeholders as well as the general public
  • Implemented source version control workflows, development guidelines, automated build testing, and instruction manuals for data pipelines, which greatly increased team capacity through cross-training
  • Established and promoted Kanban and Scrum methodologies to enable the team to more easily prioritize immediate and long-term initiatives
  • Conducted user research and user testing across the organization, surfacing data-related challenges and opportunities for collaboration

Adobe

2018-2023

My job title was "machine learning engineer", but my actual work focused on data, software engineering, and computational linguistics. I worked on a variety of products in Adobe Acrobat, Adobe Sign, and Adobe Express.

• Optimized AWS Athena queries to run 2x faster while also making the queries more readable and maintainable

• Developed engineering team guidelines to emphasize psychological safety and strong engineering practices

• Used full-stack web app libraries and design systems to build interactive prototypes for concept validation and usability testing

• Applied and introduced user research and usability testing techniques to multiple internal and external features, from ideation and concept validation to 100% rollout

• Built data visualization pipelines for quantitative and qualitative error analysis

• Architected, implemented, and deployed inference and evaluation libraries for state-of-the-art machine learning models

• Used software engineering and linguistics techniques to speed up PyTorch inference code by 10x, reducing compute cost by 80x

• Led company-wide seminars and workshops on computational linguistics, natural language processing, the intersection of linguistics and UX, and the application of quantitative and qualitative research methods

• Developed self-study and coaching materials to help machine learning scientists learn full-stack web development and prototyping skills

Patents

Model-based semantic text searching

Machine-learning tool for generating segmantation and topic metadata for documents

Publications

SemEval-2020 Task 6: Definition Extraction from Free Text with the DEFT Corpus

DEFT: a corpus for definition extraction in free- and semi-structured text

McLean Hospital

2017-2018

As a Research Data Analyst in the Psychosis Neurobiology Laboratory, I used natural language processing and data science techniques to support research in schizophrenia and bipolar disorder.

• Architected and implemented natural language processing pipelines for electronic health records, designed to support a variety of supervised and unsupervised techniques while prioritizing data privacy

• Organized annotation efforts by subject-matter experts

• Built and maintained relationships with clinicians, researchers, and data managers to support ongoing research collaborations

• Built data visualization pipelines for quantitative and qualitative error analysis

• Presented experiments at research workshops

Publications

Assessing the Efficacy of Clinical Sentiment Analysis and Topic Extraction in Psychiatric Readmission Risk Prediction

Analysis of Risk Factor Domains in Psychosis Patient Health Records

Saigon South International School

2011-2016

As a Middle/High School Mandarin Teacher and Modern World Language Department Chair, my teaching practice emphasized the complex and ever-changing nature of multilingualism and multiliteracy.

• Taught middle and high-school Mandarin from beginning to advanced level

• Designed trauma-informed assessment practices emphasizing deliberative practice, student voice and choice, and “resubmit and revise” opportunities

• Mentored student-led language and literacy programs for community members

• Developed teaching materials that drew on students’ unique linguistic backgrounds and life experiences

• Created innovative lesson plans for understanding Chinese characters in terms of their constituent parts

• Developed school-wide language policy that emphasized inclusion, multilingualism, and multiliteracy

Software languages and tools

By subfield, in descending order of familiarity

Languages

Python

Javascript / TypeScript

SQL

Rust

Java

R

Haskell

Scala

Backend tools

Django

Flask

FastAPI

Express

Actix

Rocket

NLP tools

spaCy

gensim

NLTK

Doccano

Prodi.gy

Stanford CoreNLP

Frontend tools

React

Vue

Svelte

Tailwind

Data engineering tools

pandas

scikit-learn

seaborn

matplotlib

SQL Server

Informatica

Tableau

AWS Athena

Spark

Airflow

DevOps tools

Docker

Kubernetes

Kafka

Splunk