Research

Publications.

My work on efficient training, language model architectures, and training objectives reflects a growing focus on ML systems. Earlier work in NLP and computer vision is included below.

Citation information on Google Scholar

arXiv2026

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training

Akhmed Sakip, Erland Hilman Fuadi, Omar Sayedelahl, Zonghang Li, Jianshu She, Alham Fikri Aji, Steve Liu, Eric Xing, Qirong Ho

Jointly adapting batch size and parallelism to improve useful training progress per unit of time.

ICML 2026 poster for Predicting the Order of Upcoming Tokens Improves Language Modeling, showing the token order prediction architecture and results.
ICML2026
Findings of ACL2026

Softpick: No Attention Sink, No Massive Activations with Rectified Softmax

Zayd Muhammad Kawakibi Zuhri, Erland Hilman Fuadi, Alham Fikri Aji

Rethinking attention normalization to eliminate attention sinks and massive activations, with implications for quantization and sparsity.

arXiv2026

LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation

Patrick Amadeus Irawan, Erland Hilman Fuadi, Shanu Kumar, Alham Fikri Aji, Yova Kementchedjhieva

Recovering language capabilities after multimodal adaptation through selective distillation from the original language model.

NAACL2024

COPAL-ID: Indonesian Language Reasoning with Local Culture and Nuances

Haryo Akbarianto Wibowo, Erland Hilman Fuadi, Made Nindyatama Nityasya, Radityo Eko Prasojo, Alham Fikri Aji

Evaluating commonsense reasoning grounded in Indonesian culture, in both standard and colloquial Indonesian.