Home / π§± AI Foundation Stack / β‘ Inference / Serving
avifenesh/bw24
From-scratch Rust+CUDA inference engine, bit-exact by construction β NVFP4, MoE, MTP speculative decoding, tuned against measured limits of one RTX 5090 Laptop (sm_120a).
Commercial license
β Commercial OKMIT
ε―εη¨οΌιεΈΈεͺιδΏηθδ½ζ¬θ²ζ/ζζ¬ζ’ζ¬Ύ
Topics
blackwellcudaggufgpu-kernelsllama-cppllm-inferencemoenvfp4rustspeculative-decoding
Ad