Study audits vision-language models for bias in museum archives

2026-09-25

A new study on arXiv.org examines vision-language models (VLMs) for societal bias using artwork metadata from the Metropolitan Museum of Art. The research developed a quantitative audit framework to assess CLIP models, finding no statistically significant gender effect in initial evaluations.

VERA Brief

AI-generated. Grounded in the article and its cited sources.

A study on arXiv.org audited vision-language models for societal bias using artwork metadata from the Metropolitan Museum of Art. The research developed a quantitative audit framework to assess CLIP models, finding no statistically significant gender effect in initial evaluations.

Key facts

  • A study examined vision-language models for societal bias using artwork metadata from the Metropolitan Museum of Art.
  • Researchers audited Contrastive Language-Image Pretraining (CLIP) models using metadata from 1,500 objects.
  • The study analyzed 743 attributed works, with 534 attributed to males and 209 to females.
  • Initial evaluations showed no statistically significant gender effect for OpenAI CLIP or OpenCLIP.
  • High residual embedding variance suggests global zero-shot valuation metrics are a coarse measurement instrument.

Source: arXiv · cs.LG

Reported by VERA Newswire.

More from September 2026 in The Record.