Pipeline Parallelism Theory
A theoretical study of convergence and delayed updates in randomized PipeDream-style pipeline parallelism.
Hi, I’m
Machine Learning ResearcherPhD Candidate
I work on efficient optimization and compression methods for large language models, including pruning, sparse fine-tuning, quantization, and pipeline parallelism.

Research areas
Four connected directions across model compression, adaptation, and distributed training.
Methods for removing redundant model parameters while preserving model quality.
02Training a small, carefully selected subset of model parameters efficiently.
03Reducing model precision and memory requirements while controlling quality degradation.
04Optimization and convergence analysis for training models across pipeline stages with delayed updates.
Selected work
Research implementations and open-source tools built around practical constraints.
A theoretical study of convergence and delayed updates in randomized PipeDream-style pipeline parallelism.
Methods for adapting language models by updating only a small, carefully selected subset of parameters.
A method for pruning large language models using second-order information and coordinated weight compensation.
An open-source, local-first tool for automatically editing recorded speech.
Publications
Peer-reviewed papers and recent preprints in optimization and efficient machine learning.
Writing
Project updates and technical writing from ongoing work.
Why I built an open-source, local-first tool for cleaning retake-heavy narration.