Naive Bayes benchmarks favorably against LLMs for text classification

2026-09-25

New research from arXiv benchmarks Complement Naive Bayes against large language models (LLMs) for text classification tasks. The study indicates that while LLMs excel in zero-data scenarios, Naive Bayes performs comparably or better with labeled data, offering significantly higher throughput.

VERA Brief

AI-generated. Grounded in the article and its cited sources.

New research benchmarks Naive Bayes against large language models for text classification. While LLMs excel in zero-data scenarios, Naive Bayes performs comparably or better with labeled data, offering significantly higher throughput.

Key facts

  • Naive Bayes (NB) performs comparably or better than large language models (LLMs) on text classification tasks when labeled data is available.
  • LLMs show superiority in zero-data regimes, but this advantage may be prone to contamination.
  • Naive Bayes offers significantly higher throughput and lower energy consumption per sample compared to LLMs, especially on commodity CPUs.
  • The choice between Naive Bayes and LLMs for text classification is task-dependent.
  • Naive Bayes can reach parity with LLMs for topic classification at approximately 10^4 labels.

Source: arXiv · cs.LG

Reported by VERA Newswire.

More from September 2026 in The Record.