Home / 🧱 AI Foundation Stack / ⚑ Inference / Serving

OnlyTerp/turboquant

First open-source implementation of Google TurboQuant (ICLR 2026) -- near-optimal KV cache compression for LLM inference. 5x compression with near-zero quality loss.

🧱 AI Foundation Stack βœ“ Commercial OK β˜… 76Python

Commercial license

βœ“ Commercial OKMIT

ε―ε•†η”¨οΌŒι€šεΈΈεͺιœ€δΏη•™θ‘—δ½œζ¬Šθ²ζ˜Ž/授權撝款

Topics

attentioncompressiondeep-learninggoogle-researchiclrkv-cachekv-cache-compressionllmllm-inferencemachine-learningmemory-optimizationpytorch
Ad