Home / 🧱 AI Foundation Stack / πŸ€– Agent Frameworks / Orchestration

huawei-csl/KVarN

KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag.

🧱 AI Foundation Stack βœ“ Commercial OK β˜… 440Python

Commercial license

βœ“ Commercial OKApache-2.0

ε―ε•†η”¨οΌŒι€šεΈΈεͺιœ€δΏη•™θ‘—δ½œζ¬Šθ²ζ˜Ž/授權撝款

Topics

agentic-aikv-cachellmllm-inferencelong-contextquantizationvllm
Ad