New AI Method Addresses Challenges in Long-Horizon Agent Training
2026-07-28
Researchers have introduced Progress-conditioned Group Policy Optimization (ProGPO) to improve training for AI agents on complex, long-horizon tasks. The method aims to overcome sampling imbalances that hinder learning from sparse rewards.
Source: arXiv · cs.LG
Reported by VERA Newswire.