RLVER Checkpoints trained via RLVER, the first RLVR framework to boost LLM empathy. RLVER/PPO-non-thinking 8B • Updated Jul 9, 2025 • 53 • 1 RLVER/GRPO-thinking 8B • Updated Jul 9, 2025 • 20 RLVER/PPO-thinking 8B • Updated Jul 9, 2025 • 36 RLVER/GRPO-non-thinking 8B • Updated Jul 9, 2025 • 195
RLVER Checkpoints trained via RLVER, the first RLVR framework to boost LLM empathy. RLVER/PPO-non-thinking 8B • Updated Jul 9, 2025 • 53 • 1 RLVER/GRPO-thinking 8B • Updated Jul 9, 2025 • 20 RLVER/PPO-thinking 8B • Updated Jul 9, 2025 • 36 RLVER/GRPO-non-thinking 8B • Updated Jul 9, 2025 • 195