Humanity's Last Exam
HLE-Diamond Logo

Introducing HLE-Diamond

Center for AI Safety&Scale AI
Hugging FaceDatasetload_dataset("cais/hle-diamond")

Center for AI Safety and Scale AI

We are releasing HLE-Diamond, a refined subset of Humanity’s Last Exam (HLE) question collection, following a year-long process of cleaning and refinement with input from research communities.

HLE-Diamond consists of 1,000 questions.

Main results. We compare the performance of current models on HLE-Diamond without tools.

HLE-Diamond
  • GPT-6 Astra60.6%
  • Claude Opus 5.555.0%
  • Claude Fable 5.151.3%
  • Claude Opus 538.6%
  • Gemini 3.8 Flash34.3%
  • GPT-6 Sol33.8%
  • GPT-5.6 Sol31.2%
  • Muse Spark 1.325.4%
  • Grok 4.723.4%

All models are evaluated with reasoning high.

Dataset. HLE-Diamond consists of 500 reasoning and 500 knowledge questions, measuring reasoning and expert knowledge, respectively.

Reasoning and Knowledge Partitions
ReasoningKnowledge
  • GPT-6 Astra
    Reasoning75.6%
    Knowledge45.6%
  • Claude Opus 5.5
    Reasoning63.2%
    Knowledge46.8%
  • Claude Fable 5.1
    Reasoning62.0%
    Knowledge40.6%
  • Claude Opus 5
    Reasoning47.0%
    Knowledge30.2%
  • Gemini 3.8 Flash
    Reasoning38.6%
    Knowledge30.0%
  • GPT-6 Sol
    Reasoning44.2%
    Knowledge23.4%
  • GPT-5.6 Sol
    Reasoning38.8%
    Knowledge23.6%
  • Muse Spark 1.3
    Reasoning31.6%
    Knowledge19.2%
  • Grok 4.7
    Reasoning32.4%
    Knowledge14.4%

Evaluation with tools. HLE-Diamond questions are designed to be answerable in a closed-book setting, testing both reasoning and expert knowledge. Since HLE is also used to evaluate the capabilities of agentic systems, we outline our recommended settings for evaluating HLE-Diamond with tools here.

HLE-Diamond with tools
Without toolsWith tools (web+code)
  • GPT-6 Astra
    Without tools60.6%
    With tools (web+code)82.9%
  • Claude Opus 5.5
    Without tools55.0%
    With tools (web+code)73.9%
  • Claude Fable 5.1
    Without tools51.3%
    With tools (web+code)72.4%
  • Claude Opus 5
    Without tools38.6%
    With tools (web+code)69.1%
  • Gemini 3.8 Flash
    Without tools34.3%
    With tools (web+code)60.3%
  • GPT-6 Sol
    Without tools33.8%
    With tools (web+code)64.9%
  • GPT-5.6 Sol
    Without tools31.2%
    With tools (web+code)56.5%
  • Muse Spark 1.3
    Without tools25.4%
    With tools (web+code)55.5%

Acknowledgement. We extend our deepest gratitude to all participating question contributors and experts involved in creating and refining the dataset, and to the researchers whose inputs across HLE-Rolling shaped HLE-Diamond.

For any inquiries, please contact agibenchmark@safe.ai.

Citation

@article{phan2025lastexam,
      title = {A benchmark of expert-level academic questions to assess {AI} capabilities},
      author = {{Center for AI Safety} and {Scale AI} and {HLE Contributors Consortium}},
      journal = {Nature},
      volume = {649},
      pages = {1139--1146},
      year = {2026},
      doi = {10.1038/s41586-025-09962-4},
      eprint = {2501.14249},
      archivePrefix = {arXiv},
      primaryClass = {cs.LG},
      url = {https://arxiv.org/abs/2501.14249}
}