Commit History

Overhaul trainer: TRL GRPO with env-backed reward, Qwen2.5-0.5B 4bit+LoRA, slim PyTorch CUDA base, heartbeat HTTP for HF Spaces health probe
d597642
verified

anugrah55 commited on

initial commit
af50ed7
verified

anugrah55 commited on