We are releasing HLE-Diamond, a refined subset of Humanity’s Last Exam (HLE) question collection, following a year-long process of cleaning and refinement with input from research communities.
HLE-Diamond consists of 1,000 questions.
Main results. We compare the performance of current models on HLE-Diamond without tools.
GPT-6 Astra60.6%
Claude Opus 5.555.0%
Claude Fable 5.151.3%
Claude Opus 538.6%
Gemini 3.8 Flash34.3%
GPT-6 Sol33.8%
GPT-5.6 Sol31.2%
Muse Spark 1.325.4%
Grok 4.723.4%
All models are evaluated with reasoning high.
Dataset. HLE-Diamond consists of 500 reasoning and 500 knowledge questions, measuring reasoning and expert knowledge, respectively.
GPT-6 Astra
Reasoning75.6%Knowledge45.6%Claude Opus 5.5
Reasoning63.2%Knowledge46.8%Claude Fable 5.1
Reasoning62.0%Knowledge40.6%Claude Opus 5
Reasoning47.0%Knowledge30.2%Gemini 3.8 Flash
Reasoning38.6%Knowledge30.0%GPT-6 Sol
Reasoning44.2%Knowledge23.4%GPT-5.6 Sol
Reasoning38.8%Knowledge23.6%Muse Spark 1.3
Reasoning31.6%Knowledge19.2%Grok 4.7
Reasoning32.4%Knowledge14.4%
Evaluation with tools. HLE-Diamond questions are designed to be answerable in a closed-book setting, testing both reasoning and expert knowledge. Since HLE is also used to evaluate the capabilities of agentic systems, we outline our recommended settings for evaluating HLE-Diamond with tools here.
GPT-6 Astra
Without tools60.6%With tools (web+code)82.9%Claude Opus 5.5
Without tools55.0%With tools (web+code)73.9%Claude Fable 5.1
Without tools51.3%With tools (web+code)72.4%Claude Opus 5
Without tools38.6%With tools (web+code)69.1%Gemini 3.8 Flash
Without tools34.3%With tools (web+code)60.3%GPT-6 Sol
Without tools33.8%With tools (web+code)64.9%GPT-5.6 Sol
Without tools31.2%With tools (web+code)56.5%Muse Spark 1.3
Without tools25.4%With tools (web+code)55.5%
Acknowledgement. We extend our deepest gratitude to all participating question contributors and experts involved in creating and refining the dataset, and to the researchers whose inputs across HLE-Rolling shaped HLE-Diamond.
For any inquiries, please contact agibenchmark@safe.ai.
Citation
@article{phan2025lastexam,
title = {A benchmark of expert-level academic questions to assess {AI} capabilities},
author = {{Center for AI Safety} and {Scale AI} and {HLE Contributors Consortium}},
journal = {Nature},
volume = {649},
pages = {1139--1146},
year = {2026},
doi = {10.1038/s41586-025-09962-4},
eprint = {2501.14249},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2501.14249}
}