serving-llms-vllm
by firecrawl
Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference…
npx skills add https://github.com/firecrawl/ai-research-skills --skill serving-llms-vllm