Home / ⚡ Inference / Serving

⚡ Inference / Serving

77 View all · Daily curated AI / LLM open-source intelligence, with plain-language notes and license checks.

mistralai/Leanstral-1.5-119B-A6B

· vllm, base_model:mistralai/Leanstral-2603, base_model:finetune:mistralai/Leanstral-2603

🧱 AI Foundation Stack License unclear

poolside/Laguna-M.1

text-generation · vllm, safetensors, laguna

🧱 AI Foundation Stack License unclear

zeraix/zeraix

Open-source local AI workspace — advancing on-device inference.

🧱 AI Foundation Stack ✓ Commercial OK ★ 86

Jia-Ethan/codex-keysmith

Version-independent Codex instruction deployment with dry-run, backups, hook isolation, and recovery.

🧱 AI Foundation Stack ✓ Commercial OK ★ 1,135

avifenesh/bw24

From-scratch Rust+CUDA inference engine, bit-exact by construction — NVFP4, MoE, MTP speculative decoding, tuned against measured limits of one RTX 5090 Laptop (sm_120a).

🧱 AI Foundation Stack ✓ Commercial OK ★ 278

re4/LibreCode

LibreCode - A Ollama cursor like coding / Reversing Interface

🧱 AI Foundation Stack License unclear ★ 385,613

OpenCPIL/prima.cpp

[Official] prima.cpp: Fast 30-70B LLM inference on heterogeneous and everyday home devices

🧱 AI Foundation Stack ✓ Commercial OK ★ 63

0xSero/glm-5.2-sm120

GLM-5.2-NVFP4-REAP-469B serving on SM120 (4× RTX PRO 6000 Blackwell) — one-command vLLM launch recipe, 250K context, DeepSeek Sparse Attention + MTP speculative decode

🧱 AI Foundation Stack License unclear ★ 150

cyyself/OpenTihui

on-device LLM for iOS with keyboard shortcuts

🧱 AI Foundation Stack ✓ Commercial OK ★ 52

chiennv2000/orthrus

Fast, lossless LLM inference via dual-view diffusion decoding.

🧱 AI Foundation Stack ✓ Commercial OK ★ 460

walkinglabs/modern-llm-notebook

A hands-on course for building modern LLMs from scratch in PyTorch, with 26 runnable Jupyter Notebooks covering tokenizers, attention, MoE, RLHF, inference, evaluation, and distillation.

🧱 AI Foundation Stack License unclear ★ 128

hec-ovi/vllm-awq4-qwen

vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-stream, vision, tool calling, 256K context, OpenAI-compatible, Docker. Matches DGX Spark FP8+DFlash+MTP at a third of the cost. No CUDA.

🧱 AI Foundation Stack ✓ Commercial OK ★ 40

AlexsJones/llmfit

Hundreds of models & providers. One command to find what runs on your hardware.

🧱 AI Foundation Stack ✓ Commercial OK ★ 27,845

ahammadmejbah/Awesome-Datasets-Hub

A curated collection of datasets for Large Language Models (LLMs), covering medical AI, NLP, multimodal learning, instruction tuning, reasoning, code generation, and evaluation benchmarks.

🧱 AI Foundation Stack License unclear ★ 145

jundot/omlx

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

🧱 AI Foundation Stack ✓ Commercial OK ★ 16,432

amitshekhariitbhu/llm-internals

Learn LLM internals step by step - from tokenization to attention to inference optimization.

🧱 AI Foundation Stack ✓ Commercial OK ★ 1,331

devnen/qwen3.6-windows-server

One-click Qwen3.6-27B inference on Windows. 158 tok/s on RTX 5090, 72 tok/s on RTX 3090. Native, no WSL, no Docker, no telemetry.

🧱 AI Foundation Stack License unclear ★ 222

0xNyk/council-of-high-intelligence

18 AI personas deliberate your hardest decisions across multiple LLM providers. Aristotle, Feynman, Kahneman, Torvalds & more — structured multi-round deliberation with genuine model diversity. One command: /council

🧱 AI Foundation Stack ✓ Commercial OK ★ 1,483

Arthur-Ficial/apfel

The free AI already on your Mac. CLI tool, OpenAI-compatible server, and interactive chat — all on-device via Apple Intelligence. No API keys, no cloud, no downloads.

🧱 AI Foundation Stack ✓ Commercial OK ★ 6,169

Andyyyy64/whichllm

Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.

🧱 AI Foundation Stack ✓ Commercial OK ★ 5,468

PMZFX/intel-arc-pro-b70-benchmarks

Benchmark results and performance data for the Intel Arc Pro B70 GPU (Xe2/Battlemage) - LLM inference, video generation, dual-GPU scaling.

🧱 AI Foundation Stack ✓ Commercial OK ★ 79

jmaczan/tiny-vllm

Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM

🧱 AI Foundation Stack ✓ Commercial OK ★ 770

Dynamis-Labs/spectralquant

SpectralQuant: Calibrated Eigenbasis Rotation and Water-Filled Bit Allocation for KV-Cache Compression

🧱 AI Foundation Stack ✓ Commercial OK ★ 197

quantumaikr/quant.cpp

LLM inference with 7x longer context. Pure C, zero dependencies. Lossless KV cache compression + single-header library.

🧱 AI Foundation Stack ✓ Commercial OK ★ 394

tinyBigGAMES/VindexLLM

VindexLLM is a pure Delphi, GPU-powered LLM inference engine that uses Vulkan compute shaders to run GGUF models entirely on the GPU. It performs full transformer inference without relying on Python, CUDA, or other external runtimes, requiring only vulkan-1.dll, which is typically included with modern GPU drivers.

🧱 AI Foundation Stack License unclear ★ 70

NightMean/OlliteRT

Turn your Android phone into an OpenAI-compatible LLM inference server — Fully local, private and Open Source

🧱 AI Foundation Stack ✓ Commercial OK ★ 136

brontoguana/krasis

Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware

🧱 AI Foundation Stack License unclear ★ 469

Kaden-Schutt/hipfire

RDNA-native LLM inference engine in Rust.

🧱 AI Foundation Stack License unclear ★ 483

AnkitNayak-eth/llmBench

llmBench is a high-depth benchmarking tool designed to measure the raw performance of local LLM runtimes (Ollama, llama.cpp) while providing deep hardware intelligence.

🧱 AI Foundation Stack ✓ Commercial OK ★ 45

mukel/gemma4.java

Fast Gemma 4 inference in pure Java

🧱 AI Foundation Stack ✓ Commercial OK ★ 65

psmarter/CUDA-Practice

CUDA编程练习项目-Hands-on CUDA kernels and performance optimization, covering GEMM, FlashAttention, Tensor Cores, CUTLASS, quantization, KV cache, NCCL, and profiling.

🧱 AI Foundation Stack ✓ Commercial OK ★ 156

Zyora-Dev/zse

The inference engine the open-source world built for itself.

🧱 AI Foundation Stack License unclear ★ 153

Mcourtyard/m-courtyard

M-Courtyard: Local AI Model Fine-tuning Assistant for Apple Silicon. Zero-code, zero-cloud, privacy-first desktop app powered by Tauri + React + mlx-lm.

🧱 AI Foundation Stack License unclear ★ 137

matt-k-wong/mlx-flash

Flash weight streaming for MLX: run massive models larger than your RAM on Apple Silicon.

🧱 AI Foundation Stack ✓ Commercial OK ★ 119

OnlyTerp/turboquant

First open-source implementation of Google TurboQuant (ICLR 2026) -- near-optimal KV cache compression for LLM inference. 5x compression with near-zero quality loss.

🧱 AI Foundation Stack ✓ Commercial OK ★ 76

BenChaliah/Tensa-Lang

TensaLang is a Tensor-first programming language, compiler, and runtime that let you write the Model’s inference engine (e.g. LLMs) and sampling in high level language, then compile it through MLIR to Multiple targets (e.g. CPU, CUDA, ROCm)

🧱 AI Foundation Stack License unclear ★ 74

TheToughCrane/nano-kvllm

This project aims to provide a high effective KV cache manage framework for llm inference and improve memory utilization and inference speed.

🧱 AI Foundation Stack ✓ Commercial OK ★ 62

OpenDCAI/Flash-MinerU

Ray-powered accelerator for MinerU, turning PDF → Markdown into a scalable, cluster-ready data infrastructure. 基于 Ray 的 MinerU 加速层,将 PDF → Markdown 构建为可扩展、面向集群的数据基础设施。

🧱 AI Foundation Stack License unclear ★ 54

Aaryan-Kapoor/ModelGate-Hackathon

🏆Winning Project | ModelGate is a contract-aware AI control plane that ingests customer contracts, extracts SLA/privacy/routing constraints, and generates an OpenAI-compatible endpoint that automatically routes every request to the optimal model. Simple queries go to cheap models. Complex queries go to premium ones.

🧱 AI Foundation Stack ✓ Commercial OK ★ 51

AbdelStark/attnres

Rust implementation of Attention Residuals from MoonshotAI/Kimi

🧱 AI Foundation Stack ✓ Commercial OK ★ 55

openmed-labs/synthvision

Synthetic medical VQA pipeline: 119K images annotated by frontier VLMs, cross-validated at 93% agreement, fine-tuned on 3 model families (2-3B params)

🧱 AI Foundation Stack License unclear ★ 39

Yog-Sotho/LLM-fine-tuner

Powerful no-code LLM fine-tuner: upload data → train → deploy in minutes. Unsloth 2-5× acceleration · QLoRA/DPO/RLHF/PPO/ORPO · Reward Model training · GGUF export · vLLM inference · BLEU/ROUGE/BERTScore · full CLI · Heretic Mode to unlock full model potential

🧱 AI Foundation Stack ✓ Commercial OK ★ 27

eullm/eullm

Open-source platform for creating, distributing and running sovereign EU-compliant LLMs. Verticalize any model for your domain, language and brand. AI Act ready.

🧱 AI Foundation Stack ✓ Commercial OK ★ 28