Home / π§± AI Foundation Stack / β‘ Inference / Serving
quantumaikr/quant.cpp
LLM inference with 7x longer context. Pure C, zero dependencies. Lossless KV cache compression + single-header library.
Commercial license
β Commercial OKApache-2.0
ε―εη¨οΌιεΈΈεͺιδΏηθδ½ζ¬θ²ζ/ζζ¬ζ’ζ¬Ύ
Topics
delta-compressionembeddableggufkv-cachellmllm-inferencepure-cquantizationtransformerturboquant
Ad