Skip to content

Ai Benchmark

Topic archive5 matches

Back to homeGEO summary endpoint

2026-09-14

Technology

  • AllSpark releases Iris-mini and Iris-pro open-source search agents: The AllSpark team has released Iris-mini and Iris-pro, two open-source search agents built on Qwen models. The models lead benchmarks among open-weight models in their size classes. Training data and models also improved performance on untrained tasks, including general tool use and office work.

    AI ModelsThe Decoder

    Permalink

2026-09-13

Technology

  • GPT-6 Astra beats human baseline in drone control and dominates agent benchmark: GPT-6 Astra scored nearly three times higher than Claude Fable 5.1 on Andon Labs' Vending-Bench agent benchmark while refusing illegal price-fixing deals. Additionally, Astra became the first model to beat the human baseline on all five drone control subtasks, including tracking individuals.

    AI ModelsThe Decoder

    Permalink

2026-09-12

Technology

  • GPT-6 Astra solves advanced math problems on FrontierMath Tier 4 benchmark: GPT-6 Astra has successfully solved problems on FrontierMath Tier 4, a benchmark designed to test advanced mathematical reasoning in AI. This achievement represents a significant milestone in overcoming complex mathematical barriers for artificial intelligence models.

    Artificial Intelligence量子位

    Permalink
  • LogiMed-RoB benchmark reveals error compounding in LLM medical logic: Researchers have introduced LogiMed-RoB, a benchmark based on Cochrane Risk of Bias 2.0 expert logic to evaluate large language models across 860 randomized controlled trials. Testing on 10 state-of-the-art models revealed a severe error compounding effect, despite the top model achieving 98.88% atomic consistency.

    AI Models and ApplicationsarXiv

    Permalink