Home / π§± AI Foundation Stack / β‘ Inference / Serving
OnlyTerp/turboquant
First open-source implementation of Google TurboQuant (ICLR 2026) -- near-optimal KV cache compression for LLM inference. 5x compression with near-zero quality loss.
Commercial license
β Commercial OKMIT
ε―εη¨οΌιεΈΈεͺιδΏηθδ½ζ¬θ²ζ/ζζ¬ζ’ζ¬Ύ
Topics
attentioncompressiondeep-learninggoogle-researchiclrkv-cachekv-cache-compressionllmllm-inferencemachine-learningmemory-optimizationpytorch
Ad