Home / 📊 Eval / Observability / Safety

📊 Eval / Observability / Safety

14 View all · Daily curated AI / LLM open-source intelligence, with plain-language notes and license checks.

LiquidAI/ifstruct-v1.0

· benchmark:official, benchmark:eval-yaml, task_categories:text-generation

🧱 AI Foundation Stack License unclear

Rapidata/svg-benchmark

· task_categories:text-to-image, task_categories:image-classification, task_categories:reinforcement-learning

🧱 AI Foundation Stack ✓ Commercial OK

ScaleAI/SWE-bench_Pro

· benchmark:official, benchmark:eval-yaml, size_categories:n<1K

🧱 AI Foundation Stack License unclear

cais/hle

· benchmark:official, license:mit, size_categories:1K<n<10K

🧱 AI Foundation Stack License unclear

gaia-benchmark/GAIA

· language:en, size_categories:n<1K, format:parquet

🧱 AI Foundation Stack License unclear

Idavidrein/gpqa

· benchmark:official, benchmark:eval-yaml, task_categories:question-answering

🧱 AI Foundation Stack ✓ Commercial OK

MadsLorentzen/ai-job-search

The job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.

🧱 AI Foundation Stack ✓ Commercial OK ★ 22,495

ibm-research/ScarfBench

· task_categories:text-generation, arxiv:2605.06754, region:us

🧱 AI Foundation Stack License unclear

MME-Benchmarks/Video-MME-v2

Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding

🧱 AI Foundation Stack License unclear ★ 369