Home / 🧱 AI Foundation Stack / ⚑ Inference / Serving

jmaczan/tiny-vllm

Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM

🧱 AI Foundation Stack βœ“ Commercial OK β˜… 770C++

Commercial license

βœ“ Commercial OKApache-2.0

ε―ε•†η”¨οΌŒι€šεΈΈεͺιœ€δΏη•™θ‘—δ½œζ¬Šθ²ζ˜Ž/授權撝款

Topics

aiattentionbatchingcoursecppcudahpcinferencellmllm-inferencepagedattentiontiny-vllm
Ad