inference-server

Khởi động và kiểm tra máy chủ suy luận prime-rl. Sử dụng khi được yêu cầu chạy suy luận, khởi động vLLM, kiểm tra mô hình hoặc khởi chạy máy chủ suy luận.

npx skills add https://github.com/huggingface/prime-rl --skill inference-server

Inference Server

Starting the server

Always use the inference entry point — never vllm serve or python -m vllm.entrypoints.openai.api_server directly. The entry point runs setup_vllm_env() which configures environment variables (LoRA, multiprocessing) before vLLM is imported.

# With a TOML config
uv run inference @ path/to/config.toml

# With CLI overrides
uv run inference --model.name Qwen/Qwen3-0.6B --model.max_model_len 2048 --model.enforce_eager

# Combined
uv run inference @ path/to/config.toml --server.port 8001 --gpu-memory-utilization 0.5

SLURM scheduling

The inference entrypoint supports optional SLURM scheduling, following the same patterns as SFT and RL.

Single-node SLURM

# inference_slurm.toml
output_dir = "/shared/outputs/my-inference"

[model]
name = "Qwen/Qwen3-8B"

[parallel]
tp = 8

[slurm]
job_name = "my-inference"
partition = "cluster"
uv run inference @ inference_slurm.toml

Multi-node SLURM (independent vLLM replicas)

Each node runs an independent vLLM instance. No cross-node parallelism — TP and DP must fit within a single node's GPUs.

# inference_multinode.toml
output_dir = "/shared/outputs/my-inference"

[model]
name = "PrimeIntellect/INTELLECT-3-RL-600"

[parallel]
tp = 8
dp = 1

[deployment]
type = "multi_node"
num_nodes = 4
gpus_per_node = 8

[slurm]
job_name = "my-inference"
partition = "cluster"

Dry run

Add dry_run = true to generate the sbatch script without submitting:

uv run inference @ config.toml --dry-run true

Custom endpoints

The server extends vLLM with:

  • /v1/chat/completions/tokens — accepts token IDs as prompt input (used by multi-turn RL rollouts)
  • /update_weights — hot-reload model weights from the trainer
  • /load_lora_adapter — load LoRA adapters at runtime
  • /init_broadcaster — initialize weight broadcast for distributed training

Testing the server

curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen3-0.6B",
    "messages": [{"role": "user", "content": "Hi"}],
    "max_tokens": 50
  }'

Key files

  • src/prime_rl/entrypoints/inference.py — entrypoint with local/SLURM routing
  • src/prime_rl/inference/server.py — vLLM env setup
  • src/prime_rl/configs/inference.pyInferenceConfig and all sub-configs
  • src/prime_rl/inference/vllm/server.py — FastAPI routes and vLLM monkey-patches
  • src/prime_rl/templates/inference.sbatch.j2 — SLURM template (handles both single and multi-node)
  • configs/debug/infer.toml — minimal debug config

Thêm skills từ huggingface

cpu-kernels
huggingface
Cung cấp hướng dẫn về việc viết, tối ưu hóa và đánh giá hiệu năng các nhân CPU C++ với SIMD intrinsics (AVX2/AVX512) cho hệ sinh thái nhân của Hugging Face. Bao gồm…
official
generate-openenv-env
huggingface
Tạo môi trường OpenEnv từ một trường hợp sử dụng cụ thể (ví dụ: "tạo môi trường cho thư viện textarena"). Sử dụng khi được yêu cầu thiết kế hoặc triển khai một…
official
hf-mcp
huggingface
Sử dụng Hugging Face Hub qua các công cụ máy chủ MCP. Tìm kiếm mô hình, bộ dữ liệu, Spaces, bài báo. Lấy chi tiết kho lưu trữ, tải tài liệu, chạy tác vụ tính toán và sử dụng Gradio…
official
trl-training
huggingface
Huấn luyện và tinh chỉnh các mô hình ngôn ngữ transformer bằng TRL (Transformers Reinforcement Learning). Hỗ trợ huấn luyện SFT, DPO, GRPO, KTO, RLOO và Reward Model…
official
deploy-hf
huggingface
Triển khai môi trường OpenEnv lên Hugging Face Spaces. Sử dụng khi được yêu cầu triển khai, đẩy lên Hugging Face, hoặc cập nhật một space.
official
hf-space-recovery
huggingface
Diagnose and recover failing or stuck Hugging Face Space deployments for OpenEnv environments. Use when deploying envs from `envs/` to the Hub (`openenv`…
official
pre-submit-pr
huggingface
Xác thực các thay đổi trước khi gửi pull request. Chạy các kiểm tra toàn diện bao gồm lint, kiểm thử, đánh giá sự phù hợp và phân tích RFC. Sử dụng trước khi tạo…
official
example-skill
huggingface
Kỹ năng fixture mẫu để kiểm thử hành động
official