Home / 🧱 AI Foundation Stack / ⚑ Inference / Serving

quantumaikr/quant.cpp

LLM inference with 7x longer context. Pure C, zero dependencies. Lossless KV cache compression + single-header library.

🧱 AI Foundation Stack βœ“ Commercial OK β˜… 394C

Commercial license

βœ“ Commercial OKApache-2.0

ε―ε•†η”¨οΌŒι€šεΈΈεͺιœ€δΏη•™θ‘—δ½œζ¬Šθ²ζ˜Ž/授權撝款

Topics

delta-compressionembeddableggufkv-cachellmllm-inferencepure-cquantizationtransformerturboquant
Ad