Optimizing VLM Reward Models Using Demonstrations at Test Time
2026-06-02
A new method called Demo2Reward uses expert demonstrations to optimize Vision-Language Model (VLM) reward models for robotics. The technique aims to reduce false positives without requiring additional model training.
Source: arXiv · cs.LG
Reported by VERA Newswire.