openrlhf-training
par firecrawl
Cadre RLHF haute performance avec accélération Ray+vLLM. Utilisé pour l'entraînement PPO, GRPO, RLOO, DPO de grands modèles (7B-70B+). Construit sur Ray, vLLM, ZeRO-3. 2×…
npx skills add https://github.com/firecrawl/ai-research-skills --skill openrlhf-training