All the articles with the tag "PPO".
PPO 完整推导:策略梯度、actor-critic、GAE、裁剪目标与 KL 惩罚,一步步搭出 RLHF 里最经典的强化学习对齐算法。