prism-ml/Ternary-Bonsai-27B-gguf
text-generation · llama.cpp, gguf, conversational
Home / ⚡ Inference / Serving
77 View all · Daily curated AI / LLM open-source intelligence, with plain-language notes and license checks.
text-generation · llama.cpp, gguf, conversational
image-text-to-text · gguf, llama.cpp, quantized
text-generation · transformers, gguf, text-generation
image-text-to-text · gguf, unsloth, fine tune
image-text-to-text · gguf, gemma4, unsloth
image-text-to-text · transformers, gguf, llama.cpp
text-generation · gguf, llama.cpp, quantized
any-to-any · transformers, gguf, gemma4
text-generation · mlx, safetensors, qwen3_5
text-generation · gguf, gemma4, coding
· gguf, deepseek_v4, unsloth
text-generation · transformers, gguf, glm_moe_dsa
text-generation · transformers, safetensors, gguf
text-generation · gguf, llama.cpp, qwen
image-text-to-text · hermes, gguf, uncensored
· transformers, gguf, qwen
· vllm, base_model:mistralai/Leanstral-2603, base_model:finetune:mistralai/Leanstral-2603
text-generation · vllm, safetensors, laguna
· transformers, gguf, reasoning
text-generation · peft, safetensors, gguf
image-text-to-text · transformers, gguf, unsloth
image-text-to-text · gguf, uncensored, qwen3.6
image-text-to-text · gguf, unsloth, fine tune
· gguf, qwen3, security
· gradio, build-small-hackathon, thousand-token-wood
text-generation · gguf, gemma4, uncensored
Open-source local AI workspace — advancing on-device inference.
Version-independent Codex instruction deployment with dry-run, backups, hook isolation, and recovery.
From-scratch Rust+CUDA inference engine, bit-exact by construction — NVFP4, MoE, MTP speculative decoding, tuned against measured limits of one RTX 5090 Laptop (sm_120a).
text-generation · transformers, gguf, llama
· jax, safetensors, needle
LibreCode - A Ollama cursor like coding / Reversing Interface
Simulate AI Personas in Social Feeds with Real-Time Chat Bots 2026
[Official] prima.cpp: Fast 30-70B LLM inference on heterogeneous and everyday home devices
· docker, region:us
GLM-5.2-NVFP4-REAP-469B serving on SM120 (4× RTX PRO 6000 Blackwell) — one-command vLLM launch recipe, 250K context, DeepSeek Sparse Attention + MTP speculative decode
on-device LLM for iOS with keyboard shortcuts
LLM Optimizer for NUMA, and monitor LLM system
Fast, lossless LLM inference via dual-view diffusion decoding.
A hands-on course for building modern LLMs from scratch in PyTorch, with 26 runnable Jupyter Notebooks covering tokenizers, attention, MoE, RLHF, inference, evaluation, and distillation.
Pure Rust Inference Engine
vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-stream, vision, tool calling, 256K context, OpenAI-compatible, Docker. Matches DGX Spark FP8+DFlash+MTP at a third of the cost. No CUDA.
Hundreds of models & providers. One command to find what runs on your hardware.
A curated collection of datasets for Large Language Models (LLMs), covering medical AI, NLP, multimodal learning, instruction tuning, reasoning, code generation, and evaluation benchmarks.
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
Learn LLM internals step by step - from tokenization to attention to inference optimization.
One-click Qwen3.6-27B inference on Windows. 158 tok/s on RTX 5090, 72 tok/s on RTX 3090. Native, no WSL, no Docker, no telemetry.
18 AI personas deliberate your hardest decisions across multiple LLM providers. Aristotle, Feynman, Kahneman, Torvalds & more — structured multi-round deliberation with genuine model diversity. One command: /council
The free AI already on your Mac. CLI tool, OpenAI-compatible server, and interactive chat — all on-device via Apple Intelligence. No API keys, no cloud, no downloads.
List of Permanent Free LLM API (API Keys)
Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.
Granite Switch — Build AI models like you build software
Benchmark results and performance data for the Intel Arc Pro B70 GPU (Xe2/Battlemage) - LLM inference, video generation, dual-GPU scaling.
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
SpectralQuant: Calibrated Eigenbasis Rotation and Water-Filled Bit Allocation for KV-Cache Compression
LLM inference with 7x longer context. Pure C, zero dependencies. Lossless KV cache compression + single-header library.
VindexLLM is a pure Delphi, GPU-powered LLM inference engine that uses Vulkan compute shaders to run GGUF models entirely on the GPU. It performs full transformer inference without relying on Python, CUDA, or other external runtimes, requiring only vulkan-1.dll, which is typically included with modern GPU drivers.
Turn your Android phone into an OpenAI-compatible LLM inference server — Fully local, private and Open Source
Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware
RDNA-native LLM inference engine in Rust.
Method for Long Context RLMs using verifiable Lambda Calculus
llmBench is a high-depth benchmarking tool designed to measure the raw performance of local LLM runtimes (Ollama, llama.cpp) while providing deep hardware intelligence.
Fast Gemma 4 inference in pure Java
CUDA编程练习项目-Hands-on CUDA kernels and performance optimization, covering GEMM, FlashAttention, Tensor Cores, CUTLASS, quantization, KV cache, NCCL, and profiling.
The inference engine the open-source world built for itself.
M-Courtyard: Local AI Model Fine-tuning Assistant for Apple Silicon. Zero-code, zero-cloud, privacy-first desktop app powered by Tauri + React + mlx-lm.
Flash weight streaming for MLX: run massive models larger than your RAM on Apple Silicon.
First open-source implementation of Google TurboQuant (ICLR 2026) -- near-optimal KV cache compression for LLM inference. 5x compression with near-zero quality loss.
TensaLang is a Tensor-first programming language, compiler, and runtime that let you write the Model’s inference engine (e.g. LLMs) and sampling in high level language, then compile it through MLIR to Multiple targets (e.g. CPU, CUDA, ROCm)
This project aims to provide a high effective KV cache manage framework for llm inference and improve memory utilization and inference speed.
Ray-powered accelerator for MinerU, turning PDF → Markdown into a scalable, cluster-ready data infrastructure. 基于 Ray 的 MinerU 加速层,将 PDF → Markdown 构建为可扩展、面向集群的数据基础设施。
TurboQuant 3-bit KV-cache quantization for llama.cpp
🏆Winning Project | ModelGate is a contract-aware AI control plane that ingests customer contracts, extracts SLA/privacy/routing constraints, and generates an OpenAI-compatible endpoint that automatically routes every request to the optimal model. Simple queries go to cheap models. Complex queries go to premium ones.
Rust implementation of Attention Residuals from MoonshotAI/Kimi
Synthetic medical VQA pipeline: 119K images annotated by frontier VLMs, cross-validated at 93% agreement, fine-tuned on 3 model families (2-3B params)
Powerful no-code LLM fine-tuner: upload data → train → deploy in minutes. Unsloth 2-5× acceleration · QLoRA/DPO/RLHF/PPO/ORPO · Reward Model training · GGUF export · vLLM inference · BLEU/ROUGE/BERTScore · full CLI · Heretic Mode to unlock full model potential
Open-source platform for creating, distributing and running sovereign EU-compliant LLMs. Verticalize any model for your domain, language and brand. AI Act ready.