LLM을 위한 캘리브레이션 데이터 없이 수행하는 반이차 양자화. 캘리브레이션 데이터셋 없이 모델을 4/3/2비트 정밀도로 양자화할 때 사용하며, 빠른 속도를 제공합니다…
npx skills add https://github.com/firecrawl/ai-research-skills --skill hqq-quantization