当前语言下暂无文章。
主题
智能体训练
关于后训练目标、探索机制与可靠智能体学习的持续整理。
Why Online RFT Falls Short of RLVR
Negative samples help preserve exploration and correct unstable reasoning paths that positive-only RFT can reinforce.
阅读文章Understanding GSPO from the Objective Level
GSPO moves clipping to the sequence level to reduce instability caused by routing changes in MoE training.
阅读文章