Home / 🧱 AI Foundation Stack / ⚑ Inference / Serving

hec-ovi/vllm-awq4-qwen

vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-stream, vision, tool calling, 256K context, OpenAI-compatible, Docker. Matches DGX Spark FP8+DFlash+MTP at a third of the cost. No CUDA.

🧱 AI Foundation Stack βœ“ Commercial OK β˜… 40Python

Commercial license

βœ“ Commercial OKUnlicense

ε―ε•†η”¨οΌŒι€šεΈΈεͺιœ€δΏη•™θ‘—δ½œζ¬Šθ²ζ˜Ž/授權撝款

Topics

27bamd-strix-haloawqdflashdockergfx1151llm-inferencemultimodal-llmopenai-apiqwen3rdna35rocm
Ad