Hi, I’m

Ivan Ilin

Machine Learning ResearcherPhD Candidate

I work on efficient optimization and compression methods for large language models, including pruning, sparse fine-tuning, quantization, and pipeline parallelism.

Open to machine learning research internships

ivan.ilin@kaust.edu.sa
View CV
Ivan Ilin seated beside a studio microphone
Portrait of Ivan Ilin.
StatusPhD Candidate
AffiliationKAUST · Computer Science
Google Scholar
153Citations
5h-index
4i10
GitHub
622Stars
449Followers
26Repos

Here are some of my favorite projects, from research implementations to open-source tools.

QK-Wanda

We couple query and key pruning by extending Wanda scores with opposite-projection activation energies and sharing the pruning budget across Q and K.

First Theory for PipeDream

Randomized PipeDream captures PipeDream's stale-weight behavior in an analyzable block-SGD model, revealing how pipeline depth affects convergence.

Pruning LLMs with Thanos

A block-wise algorithm for pruning large language models using second-order information and coordinated weight compensation.

Super-Tuning for LLMs

We introduce Super, which selects a sparse trainable support using activation-aware pruning scores, and Supra, a matched-budget sparse-plus-LoRA adapter.

Selected publications

2026
PreprintarXiv

QK-Wanda: Coupling Queries and Keys for Unstructured Pruning

Ivan Ilin, Peter Richtárik

2026
PreprintarXiv

Demystifying Pipeline Parallelism: First Theory for PipeDream

Ivan Ilin, Peter Richtárik

2025
PreprintarXiv

Thanos: A Block-wise Pruning Algorithm for Efficient Large Language Model Compression

Ivan Ilin, Peter Richtárik

2024
Conference paperNeurIPS 2024Published

PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression

Vladimir Malinovskii, Denis Mazur, Ivan Ilin, Denis Kuznedelev, Konstantin Burlachenko, Kai Yi, Dan Alistarh, Peter Richtárik

Personal blog

Research milestones, project updates, thoughts, and talks from recent years.

See other posts