LLM Reliability Assessed: Prompt Wording Impacts Answers Despite Consistent Accuracy
2026-07-28
New research indicates that large language models' responses can vary significantly based on question phrasing, even when the meaning remains the same. This inconsistency suggests underlying knowledge may be present but not reliably accessed.
Source: arXiv · cs.AI
Reported by VERA Newswire.