Upload PPO-aligned TinyLlama-1.1B model using MARS reward model on HHRLHF

a75499e verified 16 days ago

353 Bytes

	# payelb/HHRLHF_TinyLlama-1.1B_aligned_with_MARS_RM

	Base model: TinyLlama/TinyLlama-1.1B-Chat-v1.0

	Alignment dataset: Anthropic/hh-rlhf

	Reward model: payelb/HHRLHF_roberta-base_1k_fixed_MARS

	Method: PPO alignment with LoRA adapters.

	Notes:
	- Reward normalization and clipping enabled
	- KL control enabled
	- pad_token_id/eos_token_id explicitly set