Frontier LLMs show limitations in oncology decision-making
2026-09-01
A new benchmark, the Oncology Decision Boundary Benchmark (ODBB), evaluated nine frontier large language models on oncology decision points. The study found that a significant percentage of items were answered incorrectly by all models, highlighting a consistent blind spot in clinical meta-judgment.
Source: arXiv · cs.AI
Reported by VERA Newswire.