Vivek Iyer

vivekiyer.jpg

I am a final-year PhD student in NLP at the University of Edinburgh and an Apple AI/ML PhD Fellow, supervised by Dr. Alexandra Birch.

My research focuses on multilinguality, multimodality and personalization, motivated by a desire to adapt language models for the long tail of users, languages, modalities and cultures. My work has been published at EMNLP, NAACL, EACL and Interspeech. On the industry side, I have done research internships post-training LLMs at Meta and Apple, and earlier at Naver Labs Europe and IBM Research. For more, see my career summary or publications.

I’m currently on the industry job market for Research Scientist roles! Do reach out if you are hiring or think there is a fit.

news

Mar 10
2026
My work on Spectrum, an omnilingual, cross-modal LLM reasoning over 1500 languages in a shared, modality-agnostic latent space, is now published in the Omnilingual SONAR technical report 🌍
Nov 1
2025
Our paper on XL-Suite, an evaluation benchmark and synthetic data generation method for cross-lingual open-ended generation, has been accepted to Findings of EMNLP 2025. ACL logo 🎉
Sep 1
2025
Started a research internship at Meta (FAIR) in Paris, working on Spectrum, an omnilingual speech-text language model 🇫🇷
Mar 3
2025
Started a research internship at Apple in Pittsburgh, on personalization of reward models 🇺🇸
Dec 5
2024
Selected as an Apple Scholar in AI/ML — one of 21 recipients selected across the globe! 🎓
Sep 20
2024
[EMNLP 2024] 2 long papers on low-resource LLM-MT (Spotlight) and cultural transcreation of menus accepted at WMT 2024! See you all in Miami :us: :sunny:
Jul 11
2024
Gave an invited talk at IBM Research (slides) as part of the “Papers We Wrote” program which features presentations from their top-performing interns. :memo: :star:
Jun 21
2024
[NAACL 2024] Presented a poster on our submission on adapting LLMs for very low-resource MT, ranked #3 at AmericasNLP shared task, at NAACL 2024 in Mexico City. :sunrise: :mexico:

noteworthy publications

  1. arXiv
    Omnilingual SONAR: Cross-Lingual and Cross-Modal Sentence Embeddings Bridging Massively Multilingual Text and Speech
    Omnilingual SONAR Team, João Maria Janeiro, Pere-Lluís Huguet Cabot, Ioannis Tsiamas, Yen Meng,  Vivek Iyer, Guillem Ramírez, Loic Barrault, Belen Alastruey, Xiang "Tony" Cao, Yu-An Chung, Marta R. Costa-Jussa, David Dale, Kevin Heffernan, Jaehyeong Jo, Artyom Kozhevnikov, Alexandre Mourachko, Christophe Ropers, Holger Schwenk, and Paul-Ambroise Duquenne
    2026
  2. EMNLP
    XL-Suite: Cross-Lingual Synthetic Training and Evaluation Data for Open-Ended Generation
    Vivek Iyer, Pinzhen Chen, Ricardo Rei, and Alexandra Birch
    In Findings of the Association for Computational Linguistics: EMNLP 2025 Nov 2025
  3. WMT (Spotlight)
    Quality or Quantity? On Data Scale and Diversity in Adapting Large Language Models for Low-Resource Translation
    Vivek Iyer, Bhavitvya Malik, Pavel Stepachev, Pinzhen Chen, Barry Haddow, and Alexandra Birch
    In Proceedings of the Ninth Conference on Machine Translation Nov 2024
  4. Interspeech
    mHuBERT-147: A Compact Multilingual HuBERT Model
    Marcely Zanon Boito,  Vivek Iyer, Nikolaos Lagos, Laurent Besacier, and Ioan Calapodescu
    In Proceedings of INTERSPEECH 2024 Sep 2024