serving-llms-vllm
bởi firecrawl
Phục vụ các LLM với thông lượng cao bằng cách sử dụng PagedAttention và xử lý theo lô liên tục của vLLM. Sử dụng khi triển khai các API LLM sản xuất, tối ưu hóa suy luận…
npx skills add https://github.com/firecrawl/ai-research-skills --skill serving-llms-vllm