Adnan El Assadi

Research Engineer at Tyce

I work on making language models more capable and efficient, especially on long contexts and long-horizon tasks, and on evaluations that show where today’s models still fall short.

I’ve studied when LLMs are worth their cost (The Embedder’s Dilemma) and where models still trail humans (HUME), and I’ve led large-scale benchmarks for audio (MAEB) and video (MVEB). I’m also a core maintainer of MTEB, a widely used open-source AI evaluation framework with 25M+ downloads.

On the agent side, I’ve contributed to SWE-Marathon, which tests whether agents can sustain ultra-long-horizon software work, and to Harbor, which provides infrastructure and a curated benchmark suite for evaluating agents at scale.

I’m currently at Tyce, where I build agent harnesses for construction workflows. Previously, I was a Research Associate in the Abudayyeh-Gootenberg Lab at Harvard Medical School, where I fine-tuned LLMs with reinforcement learning for scientific objectives and built the RL environments to train them.

A few other things:

  • I graduated from Carleton University as valedictorian with a perfect GPA, earning a B.C.S. Honours in Computer Science and a minor in Mathematics.
  • I started my undergrad at Koç University in Istanbul, studying Computer Engineering and Mathematics, and worked on Turkish LLMs with Deniz Yuret at the KUIS AI Center before transferring to Carleton.

If you’d like to collaborate, or have an opportunity you think I’d be a good fit for, email me at adnanassadi56@gmail.com.

news

Sep 03, 2026 New paper: Harbor Adapters and Harbor-Index (NeurIPS 2026).
Aug 13, 2026 New paper: The Embedder’s Dilemma: LLMs Are Better, but at What Cost? (COLM 2026).
Jun 12, 2026 New paper: MVEB: Massive Video Embedding Benchmark (Findings of EMNLP 2026).
Jun 05, 2026 New paper: SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work? (NeurIPS 2026).
Feb 17, 2026 New paper: MAEB: Massive Audio Embedding Benchmark (NeurIPS 2026).

selected publications

  1. MAEB: Massive Audio Embedding Benchmark
    Adnan El Assadi, Isaac Chung, Chenghao Xiao, and 15 more authors
    In Advances in Neural Information Processing Systems (NeurIPS), 2026
  2. SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?
    Rishi Desai, Jesse Hu, Joan Cabezas, and 23 more authors
    In Advances in Neural Information Processing Systems (NeurIPS), 2026
  3. Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation
    Lin Shi, Haowei Lin, Zixuan Zhu, and 123 more authors
    In Advances in Neural Information Processing Systems (NeurIPS), 2026
  4. MVEB: Massive Video Embedding Benchmark
    Adnan El Assadi, Roman Solomatin, Isaac Chung, and 13 more authors
    In Findings of the Association for Computational Linguistics: EMNLP 2026, 2026
  5. The Embedder’s Dilemma: LLMs Are Better, but at What Cost?
    Adnan El Assadi, Niklas Muennighoff, and Jinhyuk Lee
    In Conference on Language Modeling (COLM), 2026
  6. HUME: Measuring the Human-Model Performance Gap in Text Embedding Tasks
    Adnan El Assadi, Isaac Chung, Roman Solomatin, and 2 more authors
    In International Conference on Learning Representations (ICLR), 2026