Applied Scientist 2 · Microsoft
Divyanshu Aggarwal
Applied scientist and NLP researcher working on efficient language-model adaptation and distillation, multilingual and continual learning, and evaluation at scale.
Professional experience
Applied Scientist 2
Microsoft India · India Applied Sciences
Aug 2025 — Present
Building efficient language-model capabilities for product applications, from adaptation and training-data design through large-scale evaluation.
- Finetuning large and small language models to improve application intelligence across new tasks and domains.
- Exploring on-policy distillation of frontier models into cost-efficient, purpose-built small language models.
- Supporting a strong research culture through paper curation, reading groups, and technical presentations.
Research Fellow
Microsoft Research India · Mentored by Dr. Sunayana Sitaram
Sep 2023 — Jul 2025
Researched how pretrained language models can gain multilingual capabilities efficiently without sacrificing existing performance.
- Developed active-forgetting and modular-learning approaches for cross-lingual transfer and language adaptation.
- Studied catastrophic forgetting, multilingual parameter-efficient finetuning, and broad multilingual evaluation.
- Ran distributed pretraining and post-training experiments on clusters spanning more than 100 GPUs.
AI Researcher
American Express AI Labs · Advanced NLP
Aug 2022 — Aug 2023
Applied language models to customer-care and text-understanding workflows while building scalable training-data systems.
- Finetuned LLaMA models for response generation and built text-classification workflows with DeBERTa and RoBERTa.
- Combined abstractive summarization and classification for high-precision analysis of call-log data.
- Designed data curation and preprocessing pipelines on petabyte-scale Hadoop infrastructure.
Publications
ACL 2026 Findings
Improving Cross Lingual Transfer by Pretraining with Active Forgetting
EMNLP 2025
Improving Consistency in LLM Inference using Probabilistic Tokenization
NAACL 2025 Findings
MAPLE: Multilingual Evaluation of Parameter Efficient Finetuning of Large Language Models
ACL 2024 Findings
MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks
NAACL 2024
Evaluating Inter-Bilingual Semantic Parsing for Indian Languages
NLP4ConvAI @ ACL 2023
IndicXNLI: Evaluating Multilingual Inference for Indian Languages
EMNLP 2022
XInfoTabS: Evaluating Multilingual Tabular Natural Language Inference
FEVER @ ACL 2022
A Review of Deep Learning Techniques for Protein Function Prediction
IEEE INCET 2021
Fine-tuning Distributional Semantic Models for Closely-Related Languages
VarDial @ EACL 2021
* Equal contribution