openrlhf-training
โดย firecrawl
เฟรมเวิร์ก RLHF ประสิทธิภาพสูงพร้อมการเร่งด้วย Ray+vLLM ใช้สำหรับการฝึก PPO, GRPO, RLOO, DPO ของโมเดลขนาดใหญ่ (7B-70B+) สร้างบน Ray, vLLM, ZeRO-3 2×…
npx skills add https://github.com/firecrawl/ai-research-skills --skill openrlhf-training