CV
Massachusetts Bay Transportation Authority (MBTA)
2024-present
As a data engineer in the Data Resource Group, I divide my time between "building/maintaining data tools" and "making sure we're asking the right questions about our data tools".
- Created Python libraries for analyzing and updating Informatica ETL pipelines, enabling the team to identify and fix problems before they impact stakeholders, and saving hundreds of hours of debugging time
- Created Python libraries to connect data pipelines between SQL Server, Informatica, and Tableau, surfacing high-impact operations insights to MBTA stakeholders as well as the general public
- Implemented source version control workflows, development guidelines, automated build testing, and instruction manuals for data pipelines, which greatly increased team capacity through cross-training
- Established and promoted Kanban and Scrum methodologies to enable the team to more easily prioritize immediate and long-term initiatives
- Conducted user research and user testing across the organization, surfacing data-related challenges and opportunities for collaboration
Adobe
2018-2023
My job title was "machine learning engineer", but my actual work focused on data, software engineering, and computational linguistics. I worked on a variety of products in Adobe Acrobat, Adobe Sign, and Adobe Express.
• Optimized AWS Athena queries to run 2x faster while also making the queries more readable and maintainable
• Developed engineering team guidelines to emphasize psychological safety and strong engineering practices
• Used full-stack web app libraries and design systems to build interactive prototypes for concept validation and usability testing
• Applied and introduced user research and usability testing techniques to multiple internal and external features, from ideation and concept validation to 100% rollout
• Built data visualization pipelines for quantitative and qualitative error analysis
• Architected, implemented, and deployed inference and evaluation libraries for state-of-the-art machine learning models
• Used software engineering and linguistics techniques to speed up PyTorch inference code by 10x, reducing compute cost by 80x
• Led company-wide seminars and workshops on computational linguistics, natural language processing, the intersection of linguistics and UX, and the application of quantitative and qualitative research methods
• Developed self-study and coaching materials to help machine learning scientists learn full-stack web development and prototyping skills
Patents
Model-based semantic text searching
Machine-learning tool for generating segmantation and topic metadata for documents
Publications
SemEval-2020 Task 6: Definition Extraction from Free Text with the DEFT Corpus
DEFT: a corpus for definition extraction in free- and semi-structured text
McLean Hospital
2017-2018
As a Research Data Analyst in the Psychosis Neurobiology Laboratory, I used natural language processing and data science techniques to support research in schizophrenia and bipolar disorder.
• Architected and implemented natural language processing pipelines for electronic health records, designed to support a variety of supervised and unsupervised techniques while prioritizing data privacy
• Organized annotation efforts by subject-matter experts
• Built and maintained relationships with clinicians, researchers, and data managers to support ongoing research collaborations
• Built data visualization pipelines for quantitative and qualitative error analysis
• Presented experiments at research workshops
Publications
Analysis of Risk Factor Domains in Psychosis Patient Health Records
Saigon South International School
2011-2016
As a Middle/High School Mandarin Teacher and Modern World Language Department Chair, my teaching practice emphasized the complex and ever-changing nature of multilingualism and multiliteracy.
• Taught middle and high-school Mandarin from beginning to advanced level
• Designed trauma-informed assessment practices emphasizing deliberative practice, student voice and choice, and “resubmit and revise” opportunities
• Mentored student-led language and literacy programs for community members
• Developed teaching materials that drew on students’ unique linguistic backgrounds and life experiences
• Created innovative lesson plans for understanding Chinese characters in terms of their constituent parts
• Developed school-wide language policy that emphasized inclusion, multilingualism, and multiliteracy
Software languages and tools
By subfield, in descending order of familiarity
Languages
Python
Javascript / TypeScript
SQL
Rust
Java
R
Haskell
Scala
Backend tools
Django
Flask
FastAPI
Express
Actix
Rocket
NLP tools
spaCy
gensim
NLTK
Doccano
Prodi.gy
Stanford CoreNLP
Frontend tools
React
Vue
Svelte
Tailwind
Data engineering tools
pandas
scikit-learn
seaborn
matplotlib
SQL Server
Informatica
Tableau
AWS Athena
Spark
Airflow
DevOps tools
Docker
Kubernetes
Kafka
Splunk