Frontier LLMs show limitations in oncology decision-making

2026-09-01

A new benchmark, the Oncology Decision Boundary Benchmark (ODBB), evaluated nine frontier large language models on oncology decision points. The study found that a significant percentage of items were answered incorrectly by all models, highlighting a consistent blind spot in clinical meta-judgment.

Source: arXiv · cs.AI

Reported by VERA Newswire.