Home / 🧱 AI Foundation Stack / ⚑ Inference / Serving

avifenesh/bw24

From-scratch Rust+CUDA inference engine, bit-exact by construction β€” NVFP4, MoE, MTP speculative decoding, tuned against measured limits of one RTX 5090 Laptop (sm_120a).

🧱 AI Foundation Stack βœ“ Commercial OK β˜… 278Rust

Commercial license

βœ“ Commercial OKMIT

ε―ε•†η”¨οΌŒι€šεΈΈεͺιœ€δΏη•™θ‘—δ½œζ¬Šθ²ζ˜Ž/授權撝款

Topics

blackwellcudaggufgpu-kernelsllama-cppllm-inferencemoenvfp4rustspeculative-decoding
Ad