COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training
Jointly adapting batch size and parallelism to improve useful training progress per unit of time.
Research
My work on efficient training, language model architectures, and training objectives reflects a growing focus on ML systems. Earlier work in NLP and computer vision is included below.
Citation information on Google ScholarJointly adapting batch size and parallelism to improve useful training progress per unit of time.
An auxiliary training objective that learns the relative order of future tokens using a ranking loss.
Rethinking attention normalization to eliminate attention sinks and massive activations, with implications for quantization and sparsity.
Recovering language capabilities after multimodal adaptation through selective distillation from the original language model.
Evaluating commonsense reasoning grounded in Indonesian culture, in both standard and colloquial Indonesian.
Combining self-supervised image transformations with a gating mechanism for supervised image classification.
Comparing classification approaches for recognizing distracting activities during driving.