Hi, Alex 👌
alexdremov.me
- 🔥 Managing Thousands of DL Experiments and Staying Sane
- 🌮 Rethinking Quantization-Aware Training: Why Your QAT Length is Probably Wrong
- 🔥 Understanding Flash Attention: Writing the Algorithm from Scratch in Triton
- 🚀 Speed Up PyTorch With Custom Kernels. But It Gets Progressively Darker
- ❤️ Simple Ways to Speed Up Your PyTorch Model Training
- Compute-Optimal Quantization-Aware Training — Aleksandr Dremov, David Grangier, Angelos Katharopoulos, Awni Hannun
ICLR 2026 - Training dynamics of the cooldown stage in warmup-stable-decay learning rate scheduler — Aleksandr Dremov, Alexander Hägele, Atli Kosson, Martin Jaggi
TMLR, J2C Certification (ICLR 2026)
Bachelor of Science — Graduated with honors in PSAMI Informatics and Computational Technologies.
Master of Science — Graduated from EPFL Data Science program





