nvidia/Nemotron-3-Embed-1B-BF16
sentence-similarity · sentence-transformers, safetensors, ministral3
Home / 📚 RAG / Retrieval / Vector DB
52 View all · Daily curated AI / LLM open-source intelligence, with plain-language notes and license checks.
sentence-similarity · sentence-transformers, safetensors, ministral3
text-generation · transformers, safetensors, gpt_oss
· license:cc-by-4.0, arxiv:2203.02155, region:us
Local-first ETL/ELT studio: a drag-and-drop visual pipeline designer that compiles to SQL and runs on DuckDB. Tiny desktop app, no servers, git-friendly workspaces.
Transformer architecture explained step by step - the full architecture, every attention variant, positional embeddings, and every layer inside a Transformer.
· task_categories:text-retrieval, task_categories:feature-extraction, language:en
High-performance Knowledge Graph engine for AI, LLMs, and GraphRAG — built for the next generation of intelligent applications.
A bio-inspired cognitive memory engine — a new paradigm for Graph RAG.
Native AI inference for PHP 8.3+ - run ONNX, GGUF (llama.cpp) and RubixML models directly in your PHP process via FFI. Chat, streaming, embeddings, RAG and vector search. No Python, no HTTP microservices, no Docker sidecars.
An enterprise AI workspace for model routing, multimodal chat, files, tools, billing, identity, and operations.
Selfhost modern LLM stacks. Run the whole fleet from your terminal
如影随形 · 本地优先桌面工作日志:富文本记录、语义检索、工作台总结/问答;模型自配,数据留在本机。 Local-first desktop work journal—rich logs, semantic search, AI summary & Q&A. BYO models, data stays yours.
World's fastest and most compact embedded vector database: exact by default, multimodal, local-first, and GPU-accelerated
Repo Explainer — turn any GitHub repo into a visual explainer page. Pipeline + 5 live examples.
Turn any codebase, with its docs, SQL schemas, configs, and PDFs, into a queryable knowledge graph. A /graphify skill for Claude Code, Cursor, Codex, and Gemini CLI: local deterministic AST parsing, every edge explained, no vector store.
Fine-tuned local RAG models. First: agmind-rag-splitter-ru — a Russian context-aware document splitter (T-lite-it-2.1 LoRA, distillation, GGUF, AMD Vulkan). 100% valid JSON / boundary-F1@±1 0.821.
The best-benchmarked open-source AI memory system. And it's free.
Train linear embedding adapters with triplet loss to align retrieval embeddings with your queries (RAG).
Deka — Aligning Human Intuition with Semantic Space
使用社交软件聊天记录结合向量数据库让AI更好的扮演对方的角色,在不微调模型的情况下可以达到可观的效果。把曾经的美好,续成往后的陪伴。
A complete collection of RAG interview questions, answers (505 questions & 41 RAG types), system design scenarios, architecture patterns, and production-ready concepts.
Independent Autistic Intelligence — a cyber brain for your AI. It never forgets a detail, remembers exactly what you said, and learns how you work over time. Free, local, works with Cursor, Claude Code, Codex, OpenClaw, Hermes and more. MIT.
A vector index built on TurboQuant, written in Rust with Python bindings
Fast search engine on object storage, with full text search, vectors, and SQL, natively on Parquet.
Voice notes for iPhone and macOS - 100% Rust, Dioxus, local-first (SQLite + LanceDB + RIG)
PostgreSQL-compatible SQL, graph, and vector database built from scratch in Rust.
基于 Streamlit、LangChain 与 Chroma 的轻量级 RAG 学习项目,支持本地知识库上传、检索增强问答与聊天式交互。
Run, quantize, and fine-tune LLMs on Apple Silicon. Pure Rust, no Python, no CUDA, no ONNX
Enhanced LanceDB memory plugin for OpenClaw — Hybrid Retrieval (Vector + BM25), Cross-Encoder Rerank, Multi-Scope Isolation, Management CLI
Drop-in prompt compression for production LLM apps. Cut your token bill 40-60% without changing your code. Python SDK, LLMLingua-2, MIT.
Enterprise-grade (40m+ LOC) codebase intelligence, zero-setup, local & private Plugin/Skill/Extension or MCP: hybrid semantic search, polyglot dependency graphs, symbol-level impact analysis & call-flow, interactive HTML viewer, cross-project & branch-aware search, DB/API/infra knowledge. 61% less tokens, 84% fewer calls, 37x faster. Cloud in beta.
Fast, local-first web content extraction for LLMs. Scrape, crawl, extract structured data — all from Rust. CLI, REST API, and MCP server.
Prismer Cloud
Privacy-first document intelligence engine — parse PDFs, DOCX, PPTX, XLSX & CSV into AI-ready chunks for RAG pipelines. Includes HITL review, 3-layer memory chat, and a production FastAPI server.
SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35B-A3B FP8 ~240 tok/s, 27B-int4 hybrid GDN+Mamba, Gemma4 26B/31B AWQ, 256K ctx. 321 patches: TurboQuant k8v4 KV, MTP/DFlash spec-decode, FULL cudagraph, hybrid GDN. vLLM pin dev424 + Control Center GUI.
一套为研究生和学术研究者设计的完整AI Prompt库 📖 包含内容: ✨ 40+ 精心设计的AI Prompt ✨ 论文选题系统方法(生成、评估、论证) ✨ 论文查找快速方案(8个不同方案) ✨ 文献综述框架和工具 ✨ Excel自动评估表格 ✨ 3个完整的论证模板 🚀 核心优势: ⚡ 节省时间 50-70%(选题3-5天而不是2-3周) 🎯 科学方法(基于系统的5维度评估体系) 💡 即插即用(所有Prompt直接复制可用) 📚 全流程覆盖(从选题到出版的完整方案) 🎓 适用人群: 👨🎓 硕士研究生 | 博士研究生 | 本科毕业设计 | 学术研究者 | 内容创作者
A pure-Python context management layer for LLM systems — retrieval, re-ranking, memory decay, and token-budget enforcement in one pipeline.
Dataset and benchmark for RAG on company internal documents.
Turn documents into high-quality instruction datasets with grounding, quality filtering, deduplication, and provenance.
VectorRAG.Net is a .NET-native high-performance vector database library for semantic search and RAG (Retrieval-Augmented Generation). Core search is based on Random Hyperplane LSH candidate generation with exact rerank by dot/cosine.
Graph RAG with pure vector search, achieving SOTA performance in multi-hop reasoning scenarios.
🚀 Engram-PEFT: An unofficial implementation of DeepSeek Engram. Inject high-capacity conditional memory into LLMs via sparse retrieval PEFT without increasing inference FLOPs / DeepSeek Engram 架构的非官方实现。通过参数高效微调 (PEFT) 为大语言模型注入超大规模条件记忆,支持稀疏更新且不增加推理开销。
The highest-scoring AI memory system ever benchmarked that isn't reliant on LLM reranking. And it's free & burns less tokens.
Self-hostable RAG platform - document ingestion, embedding, and vector search behind a simple REST API
RAG pipeline security testing toolkit - 27 techniques across 6 kill chain phases, mapped to MITRE ATLAS
用 Kubernetes 管理大型語言模型的運維工具。
Vectorless, Reasoning-Based Retrieval-Augmented Generation (RAG)
[knowledge-rag] - Drop docs, search instantly from Claude Code — 12 MCP tools, 20 format parsers, hybrid search + reranking. Zero servers, zero API keys, 100% local.
Shared memory MCP server — persistent, searchable, cross-client Claude, Opencode
A local-first Chrome extension that passively captures ChatGPT, Gemini, Claude, Grok, Perplexity conversations into a private memory graph. Features in-browser Hybrid RAG (Vector + BM25), semantic search, and 100% privacy via WebAssembly and IndexedDB. No servers, no API keys.
A reusable, open-source AI infrastructure designed for communities and organizations. It understands structured data, processes official announcements, and answers questions accurately by centralizing important information.
RadCrew.org website with LLM-powered chatbot