serving-llms-vllm
por firecrawl
Atende LLMs com alta taxa de transferência usando o PagedAttention e o batching contínuo do vLLM. Use ao implantar APIs de LLM em produção, otimizando inferência…
npx skills add https://github.com/firecrawl/ai-research-skills --skill serving-llms-vllm