Shamil Chollampatt

I am a Staff Machine Learning Researcher at Zoom, where I lead multilingual LLM development and have taken multilingual LLMs from research to production at scale. Before Zoom, I was a machine translation research scientist at Rakuten. I have a PhD in natural language processing (NLP) from National University of Singapore (NUS), and I was a recipient of the ISEP Scholarship (formerly NGS scholarship).

My areas of expertise include multilingual LLM post-training, multilinguality evaluation, machine translation, and speech translation.


Email

shamil [dot] cm [at] gmail [dot] com

Bio

Education

National University of Singapore
Aug 2014 - Sep 2019
Ph.D. in Natural Language Processing

National Institute of Technology Calicut
Jul 2009 - Jun 2013
Bachelor of Technology in Computer Science

Work

Zoom
Feb 2023 - Present
Machine Learning Researcher (Staff)
Jul 2021 - Jan 2023
Research Scientist (Senior)

Rakuten
Aug 2019 - Jul 2021
Research Scientist

National University of Singapore
Aug 2018 - Aug 2019
Research Fellow

Oracle
Jun 2013 - Jun 2014
Member Technical Staff

Publications

  • Shamil Chollampatt, Minh Quang Pham, Sathish Indurthi, Marco Turchi. 2025. Cross-lingual Evaluation of Multilingual Text Generation. In COLING 2025. pdf
  • Sathish Indurthi, Wenxuan Zhou, Shamil Chollampatt, Ravi Agrawal, Kaiqiang Song, Lingxiao Zhao, Chenguang Zhu. 2024. Improving Multilingual Instruction Finetuning via Linguistically Natural and Diverse Datasets. In Findings of EMNLP 2024. pdf
  • Sathish Indurthi, Shamil Chollampatt, Ravi Agrawal, Marco Turchi. 2023. CLAD-ST: Contrastive Learning with Adversarial Data for Robust Speech Translation. In EMNLP 2023. pdf
  • Minh-Quang Pham, Sathish Indurthi, Shamil Chollampatt, Marco Turchi. 2023. Select, Prompt, Filter: Distilling Large Language Models for Summarizing Conversations. In EMNLP 2023. pdf
  • Shamil Chollampatt, Raymond Hendy Susanto, Liling Tan, Ewa Szymanska. 2020. Can Automatic Post Editing Improve NMT? In EMNLP 2020. pdf code bib
    @InProceedings{chollampatt2020apenmt,
      author    = {Chollampatt, Shamil and Susanto, Raymond Hendy and Tan, Liling and Szymanska, Ewa},
      title     = {Can Automatic Post Editing Improve NMT?},
      booktitle = {Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing},
      year      = {2020},
      month     = {November},
      address   = {Online}
    }
  • Raymond Susanto, Shamil Chollampatt, and Liling Tan. 2020. Lexically Constrained Neural Machine Translation with Levenshtein Transformer. In ACL 2020. pdf code bib
    @InProceedings{susanto2020clevt,
      author    = {Susanto, Raymond and Chollampatt, Shamil and Tan, Liling},
      title     = {Lexically Constrained Neural Machine Translation with Levenshtein Transformer},
      booktitle = {Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics},
      year      = {2020},
      month     = {July},
      address   = {Online}
    }
  • Shamil Chollampatt, Weiqi Wang, and Hwee Tou Ng. 2019. Cross-Sentence Grammatical Error Correction. In ACL 2019. pdf code bib
    @InProceedings{chollampatt2019csgec,
      author    = {Chollampatt, Shamil and Wang, Weiqi and Ng, Hwee Tou},
      title     = {Cross-Sentence Grammatical Error Correction},
      booktitle = {Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics},
      year      = {2019},
      month     = {July},
      address   = {Florence, Italy}
    }
  • Shamil Chollampatt and Hwee Tou Ng. 2018. Neural Quality Estimation of Grammatical Error Correction. In EMNLP 2018. pdf code bib
    @InProceedings{chollampatt2018nqegec,
      author    = {Chollampatt, Shamil and Ng, Hwee Tou},
      title     = {Neural Quality Estimation of Grammatical Error Correction},
      booktitle = {Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing},
      year      = {2018},
      month     = {November},
      address   = {Brussels, Belgium}
    }
  • Shamil Chollampatt and Hwee Tou Ng. 2018. A Reassessment of Reference-Based Grammatical Error Correction Metrics. In COLING 2018. pdf code bib
    @InProceedings{chollampatt2018reassess,
      author    = {Chollampatt, Shamil and Ng, Hwee Tou},
      title     = {A Reassessment of Reference-Based Grammatical Error Correction Metrics},
      booktitle = {Proceedings of the 27th International Conference on Computational Linguistics},
      year      = {2018},
      month     = {September},
      address   = {Santa Fe, New Mexico, USA}
    }
  • Shamil Chollampatt and Hwee Tou Ng. 2018. A Multilayer Convolutional Encoder-Decoder Neural Network for Grammatical Error Correction. In AAAI 2018. pdf arXiv code bib
    @InProceedings{chollampatt2018mlconv,
      author    = {Chollampatt, Shamil and Ng, Hwee Tou},
      title     = {A Multilayer Convolutional Encoder-Decoder Neural Network for Grammatical Error Correction},
      booktitle = {Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence},
      year      = {2018},
      month     = {February},
      address   = {New Orleans, Louisiana, USA}
    }
  • Shamil Chollampatt and Hwee Tou Ng. 2017. Connecting the Dots: Towards Human-level Grammatical Error Correction. In BEA Workshop @ EMNLP 2017. pdf code bib
    @InProceedings{chollampatt2017smtgec,
      author    = {Chollampatt, Shamil and Ng, Hwee Tou},
      title     = {Connecting the Dots: Towards Human-level Grammatical Error Correction},
      booktitle = {Proceedings of the 12th Workshop on Innovative Use of NLP for Building Educational Applications},
      year      = {2017},
      month     = {September},
      address   = {Copenhagen, Denmark}
    }
  • Shamil Chollampatt, Duc Tam Hoang, and Hwee Tou Ng. 2016. Adapting Grammatical Error Correction Based on the Native Language of Writers with Neural Network Joint Models. In EMNLP 2016. pdf bib
    @InProceedings{chollampatt2016adaptngec,
      author    = {Chollampatt, Shamil and Hoang, Duc Tam and Ng, Hwee Tou},
      title     = {Adapting Grammatical Error Correction Based on the Native Language of Writers with Neural Network Joint Models},
      booktitle = {Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing},
      year      = {2016},
      month     = {November},
      address   = {Austin, Texas, USA}
    }
  • Shamil Chollampatt, Kaveh Taghipour, and Hwee Tou Ng. 2016. Neural Network Translation Models for Grammatical Error Correction. In IJCAI 2016. pdf arXiv bib
    @InProceedings{chollampatt2016ngec,
      author    = {Chollampatt, Shamil and Taghipour, Kaveh and Ng, Hwee Tou},
      title     = {Neural Network Translation Models for Grammatical Error Correction},
      booktitle = {Proceedings of the 25th International Joint Conference on Artificial Intelligence},
      year      = {2016},
      month     = {July},
      address   = {New York, USA}
    }
  • Duc Tam Hoang, Shamil Chollampatt, and Hwee Tou Ng. 2016. Exploiting N-best Hypotheses to Improve an SMT Approach to Grammatical Error Correction. In IJCAI 2016. pdf arXiv bib
    @InProceedings{hoang2016nbestgec,
      author    = {Hoang, Duc Tam and Chollampatt, Shamil and Ng, Hwee Tou},
      title     = {Exploiting N-Best Hypotheses to Improve an {SMT} Approach to Grammatical Error Correction},
      booktitle = {Proceedings of the 25th International Joint Conference on Artificial Intelligence},
      year      = {2016},
      month     = {July},
      address   = {New York, USA}
    }

Patents

Granted

  • Distilling Language Models. US 12,737,560 B2 (Granted Sep 2026). pdf
  • Fine-tuning Language Models Using Target Language Vocabulary and Parallel Data for Machine Translation. US 12,711,331 B1 (Granted Aug 2026). pdf
  • Recommendation and Summarization of Unread Messages. US 12,641,047 B2 (Granted May 2026). pdf
  • Priority-based Scheduling of Translation Requests. US 12,511,500 B2 (Granted Dec 2025). pdf
  • Expanding Online Chat Communications Based on Chat Context. US 12,381,839 B1 (Granted Aug 2025). pdf

Pending Applications

  • Language Capability Evaluation of Large Language Models. US 2025/0298995 A1 (Pending).
  • Multilingual Dataset Collection for Large Language Model Training. US 2025/0384226 A1 (Pending).
  • Contrastive Learning with Adversarial Data for Robust Speech Translation. US 2024/0419927 A1 (Pending).
  • Providing Multistream Machine Translation During Virtual Conferences. US 2023/0351123 A1 (Pending).
  • Providing Real-Time Translation During Virtual Conferences. US 2023/0351124 A1 (Pending).