publications
2026
- MAEB: Massive Audio Embedding BenchmarkAdnan El Assadi, Isaac Chung, Chenghao Xiao, and 15 more authorsIn Advances in Neural Information Processing Systems (NeurIPS), 2026
We introduce the Massive Audio Embedding Benchmark (MAEB), a large-scale benchmark covering 30 tasks across speech, music, environmental sounds, and cross-modal audio-text reasoning in 100+ languages. We evaluate 50+ models and find that no single model dominates across all tasks: contrastive audio-text models excel at environmental sound classification (e.g., ESC50) but score near random on multilingual speech tasks (e.g., SIB-FLEURS), while speech-pretrained models show the opposite pattern. Clustering remains challenging for all models, with even the best-performing model achieving only modest results. We observe that models excelling on acoustic understanding often perform poorly on linguistic tasks, and vice versa. We also show that the performance of audio encoders on MAEB correlates highly with their performance when used in audio large language models. MAEB is derived from MAEB+, a collection of 98 tasks. MAEB is designed to maintain task diversity while reducing evaluation cost, and it integrates into the MTEB ecosystem for unified evaluation across text, image, and audio modalities. We release MAEB and all 98 tasks along with code and a leaderboard at https://github.com/embeddings-benchmark/mteb.
@inproceedings{elassadi2026maeb, title = {MAEB: Massive Audio Embedding Benchmark}, author = {El Assadi, Adnan and Chung, Isaac and Xiao, Chenghao and Solomatin, Roman and Jha, Animesh and Chand, Rahul and Singh, Silky and Wang, Kaitlyn and Khan, Ali Sartaz and Nasser, Marc Moussa and Fong, Sufen and He, Pengfei and Xiao, Alan and Munot, Ayush Sunil and Shrivastava, Aditya and Gazizov, Artem and Muennighoff, Niklas and Enevoldsen, Kenneth}, booktitle = {Advances in Neural Information Processing Systems (NeurIPS)}, year = {2026}, } - SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?Rishi Desai, Jesse Hu, Joan Cabezas, and 23 more authorsIn Advances in Neural Information Processing Systems (NeurIPS), 2026
AI agents are increasingly expected to complete long-horizon workflows that require sustained progress over hours, millions of tokens, and complex environments. Yet current agent benchmarks largely evaluate short-form tasks, such as single pull requests, small tickets, or 5-10 minute exercises, limiting our ability to measure agents’ capabilities in planning, long-context understanding, and memory use. We introduce SWE-Marathon, a benchmark of 20 long-horizon tasks spanning software engineering and adjacent technical domains. Each task consists of a unique executable environment, a human-written reference solution, and a multi-layer verification suite. Logged agent attempts average 27.2M total tokens, making SWE-Marathon substantially longer-horizon than existing SWE and command-line agent benchmarks. Current frontier coding agents solve fewer than 30% of tasks. Failures often arise from poor self-verification, self-reported infeasibility, and premature termination. We also observe reward-hacking behavior in 13.8% of rollouts, where agents attempt to exploit the environment or verifier to bypass the intended workflow. SWE-Marathon includes adversarial review of test suites and execution environments, as well as multi-layer checks designed to prevent shortcut solutions. We release SWE-Marathon, evaluation code, and agent trajectories at https://swe-marathon.org/.
@inproceedings{desai2026swemarathon, title = {SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?}, author = {Desai, Rishi and Hu, Jesse and Cabezas, Joan and Harsola, Neel and Shukla, Pratyush and Chaim, Roey Ben and El Assadi, Adnan and Kamath, Omkaar Mukund and Faldu, Fenil and Hebbar, Prannay and Sun, Jiankai and Li, Yiyuan and Srinivasan, Pramod and Gupta, Ishan and Settles, Christopher and Wang, Daniel and Chen, Derek and Raja, Pranav and Liu, Albert and Šuppa, Marek and Sasikumar, Nevasini and Kong, Luyang and Quintanilla, Erik and Li, Xiangyi and Bercovich, Ivan and Dillmann, Steven}, booktitle = {Advances in Neural Information Processing Systems (NeurIPS)}, year = {2026}, } - Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic EvaluationLin Shi, Haowei Lin, Zixuan Zhu, and 123 more authorsIn Advances in Neural Information Processing Systems (NeurIPS), 2026
Evaluating agents on the growing number of agentic benchmarks is challenging because they often require complex environments and agent integrations. We introduce Harbor Adapters, a unified evaluation infrastructure for agentic benchmarks. Our work makes three contributions. First, we develop benchmark adapters that port more than 80 benchmarks to evaluate arbitrary agents, and validate them through rigorous code review and parity experiments. Second, we conduct a large-scale evaluation of 8 models spanning capability tiers across 54 benchmarks; every model is run with Terminus-2 and with one of 3 native harnesses. This enables a broader analysis of agent capabilities and failure modes than was previously possible. Third, we introduce Harbor-Index, a curated set of 82 difficult, diverse, and high-quality tasks spanning 29 benchmarks, refined from the adapted suite through difficulty filtering, AI and human audit, and an audit-and-fix loop. Harbor-Index preserves the challenge and breadth of large-scale agentic evaluations while being affordable to run; no evaluated model-harness configuration exceeds 30% pass rate, and the strongest (GPT-5.5 with Codex) reaches 28.0%. We release the adapters, evaluation results, in-depth analysis, and Harbor-Index as open-source artifacts to support more reliable and comprehensive evaluation of language-model agents.
@inproceedings{shi2026harbor, title = {Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation}, author = {Shi, Lin and Lin, Haowei and Zhu, Zixuan and Zhou, Xiaoyue and Li, Xiang and Lin, Xiangning and Deng, Yaxuan and Xu, Han and Li, Yuangang and Li, Shanda and Chen, Zizhao and Xing, Hanwen and Raj, Harsh and Chen, Bo and Shi, Quan and Dillmann, Steven and Gao, Yipeng and Khanna, Puneesh and Lu, Ruofan and Zhou, Chao Beyond and Yang, Michael and Zhang, Robert and Chai, Siyuan and Chang, Jiayu and Chen, Yizhao and Chen, Xiaokun and Dai, Yiwei and Yang, Wenting and Liu, Hange and Liu, Minghao and Wang, Zihan and El Assadi, Adnan and Stroebl, Benedikt and Buchanan, E. Kelly and Meng, Han and He, Junwei and Yu, Longxuan and Shayanfar, Radin and Lee, Yukyung and Dong, Zhikang and Hart, Allen G and Wei, Anjiang and Kashyap, Anurag and Khatua, Arpandeep and Zheng, Audrey Jixin and Ma, Chengrui and Heineman, David and Chen, Dubing and Trinh, Hai-Anh and Fang, Haishuo and Zhang, Hefan and Shen, Hui and Sugiura, Issa and Sun, Jiankai and Gao, Jiechao and Lin, Junhong and Li, Junnan and Yang, Kai and Hsiung, Lei and Wang, Maoyu and Tang, Mengze and Omi, Nabil and Raoof, Negin and Edwards, Nicholas and Guo, Octavia and Mastromichalakis, Orfeas Menis and Ji, Pengliang and Hejman, Przemysław and Qi, Qi and Lin, Qunshu and Zhuang, Richard and Yang, Rui and Zheng, Ruichen and Marten, Ryan and Fazliani, Shaghayegh and Hou, Shizheng and Jiang, Sicong and Li, Sijie and Yuan, Boqin and Glass, Michael and Bian, Song and Zhuo, Terry Yue and Wu, Tianqing and Tang, Tom and Zhao, Wanjia and Xuan, Weihao and Liang, Wenhua and Liu, Xian and Lan, Xin and Zhang, Xuan and Zhao, Xuandong and Tang, Yanchuan and Jiang, Yifan and Li, Yijiang and Guan, Yitong and Li, Yizhi and Liu, Yonghui and Tang, Yuheng and Yujun and Mao and Zhao, Yunfei and Wang, Yuxin and Tang, Yuxuan and Tang, Zhenheng and Li, Zhifei and Wang, Ziruo and She, Ziyu and Liu, Kaiyuan and Chaabane, Iheb and Tang, Yuxin and Li, Xiangyi and GNVV, Satya Sai Srinath Namburi and Zheng, Xinyue and Konwinski, Andy and Li, Boxuan and Chen, Leon Liangyu and Dimakis, Alex and Carlini, Nicholas and Vosoughi, Soroush and Koyejo, Sanmi and He, Di and Guha, Etash and Feuer, Benjamin and Merrill, Mike and Schmidt, Ludwig and Shaw, Alex}, booktitle = {Advances in Neural Information Processing Systems (NeurIPS)}, year = {2026}, } - MVEB: Massive Video Embedding BenchmarkAdnan El Assadi, Roman Solomatin, Isaac Chung, and 13 more authorsIn Findings of the Association for Computational Linguistics: EMNLP 2026, 2026
We introduce the Massive Video Embedding Benchmark (MVEB), a 23-task benchmark for video embeddings spanning classification, zero-shot classification, clustering, pair classification, retrieval, and video-centric question answering. We evaluate 33 models and find that no single model dominates: MLLM-based embeddings lead on classification, clustering, pair classification, and QA; multimodal binding leads on retrieval and zero-shot classification; generative MLLMs without contrastive adaptation collapse on cross-modal tasks. Paired video-only vs. audio+video evaluations show that audio’s contribution depends on dataset annotation provenance: audio helps when labels were produced from both modalities and hurts when they were produced from visuals alone, a six-point gap consistent across model families. MVEB is derived from MVEB+, a 184-task pool, and is designed to maintain task diversity while reducing evaluation cost. It integrates into the MTEB ecosystem for unified evaluation across text, image, audio, and video. We release MVEB and all 184 tasks along with code and a leaderboard at https://github.com/embeddings-benchmark/mteb.
@inproceedings{elassadi2026mveb, title = {MVEB: Massive Video Embedding Benchmark}, author = {El Assadi, Adnan and Solomatin, Roman and Chung, Isaac and Xiao, Chenghao and Shah, Deep and Dey, Manan and Sudhakar, Shriya and Bugaud, Zacharie and Siblini, Wissam and Munot, Ayush Sunil and Devavarapu, Yashwanth and Ireddi, Rakshitha and Yang, Michelle and Kardos, Márton and Muennighoff, Niklas and Enevoldsen, Kenneth}, booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026}, year = {2026}, } - The Embedder’s Dilemma: LLMs Are Better, but at What Cost?Adnan El Assadi, Niklas Muennighoff, and Jinhyuk LeeIn Conference on Language Modeling (COLM), 2026
Should you replace your text-embedding pipeline with a large language model? We answer this with a controlled, cost-aware comparison of ten LLMs across six families and 26 embedding models (118M to 14B parameters) on 37 tasks spanning classification, semantic textual similarity (STS), clustering, pair classification, and retrieval. In aggregate the two paradigms are effectively tied: the best LLM (Gemini 3.1 Pro, 77.6) and the best embedding model (77.2) differ by 0.4 points. Their strengths differ by task: LLMs lead on reasoning-heavy retrieval, embedding models lead on classification, and the two match on clustering, STS, and pair classification. Reaching that parity is expensive. An LLM costs up to 1,431x more than an embedding model of comparable quality (USD 154 vs. USD 0.11 per benchmark pass), and the open LLMs tested process tokens 2.5 to 736x more slowly on the same GPU. Reasoning tokens account for 28 to 81% of LLM inference cost; lower reasoning budgets preserve or improve retrieval quality for most models in our ablation. The Pareto frontier contains the leading embedding models and one LLM, Gemini 3.1 Pro. These results support a division of labour: use embedding models for similarity, classification, and clustering, and reserve LLMs for reasoning-intensive retrieval. Our code, datasets, and results are publicly available at https://github.com/embeddings-benchmark/embedders-dilemma.
@inproceedings{elassadi2026embedders, title = {The Embedder's Dilemma: LLMs Are Better, but at What Cost?}, author = {El Assadi, Adnan and Muennighoff, Niklas and Lee, Jinhyuk}, booktitle = {Conference on Language Modeling (COLM)}, year = {2026}, } - HUME: Measuring the Human-Model Performance Gap in Text Embedding TasksAdnan El Assadi, Isaac Chung, Roman Solomatin, and 2 more authorsIn International Conference on Learning Representations (ICLR), 2026
Comparing human and model performance offers a valuable perspective for understanding the strengths and limitations of embedding models, highlighting where they succeed and where they fail to capture meaning and nuance. However, such comparisons are rarely made, as human performance on embedding tasks is difficult to measure. To fill this gap, we introduce HUME: Human Evaluation Framework for Text Embeddings. While frameworks like MTEB provide broad model evaluation, they lack reliable estimates of human performance, limiting the interpretability of model scores. We measure human performance across 16 MTEB datasets spanning reranking, classification, clustering, and semantic textual similarity across linguistically diverse high- and low-resource languages. Humans achieve an average performance of 77.6% compared to 80.1% for the best embedding model, though with substantial variation: models reach high performance on some datasets while struggling on notably low-resource languages. Our human annotations also reveal multiple dataset issues. We additionally benchmark nine LLMs as annotators on reranking, classification, and STS tasks, finding that they fall short of human performance (76.1% vs. 81.2%) despite offering scalability advantages. We provide human performance baselines, insights into task difficulty patterns, and an extensible evaluation framework that enables a more meaningful interpretation of results and informs the development of both models and benchmarks. Our code, dataset, and leaderboard are publicly available at https://github.com/embeddings-benchmark/mteb.
@inproceedings{elassadi2026hume, title = {HUME: Measuring the Human-Model Performance Gap in Text Embedding Tasks}, author = {El Assadi, Adnan and Chung, Isaac and Solomatin, Roman and Muennighoff, Niklas and Enevoldsen, Kenneth}, booktitle = {International Conference on Learning Representations (ICLR)}, year = {2026}, }