smallevals — CPU-fast, GPU-blazing fast offline retrieval evaluation for RAG systems with tiny QA models.
-
Updated
Dec 4, 2025 - Python
smallevals — CPU-fast, GPU-blazing fast offline retrieval evaluation for RAG systems with tiny QA models.
Local decision recorder and offline policy evaluation for coding agents. Advisory tooling with explicit evidence and authority boundaries.
Lightweight Python library for interactive demo and inspection of recommender systems in Streamlit.
Support-ticket dedup & resolution finder — pgvector semantic retrieval evaluated with Netflix XP-style offline replay (recall@k / MRR). Shares a domain-agnostic similarity engine with its sibling repo 'clause'. Next.js 15 · Supabase · HF embeddings.
Local reproduction of Dream-RSI: model weights stay frozen while an executable search policy evolves, via persistent discovery trees and an offline replay engine that scores candidate policies at zero extra inference cost.
Offline prototype for intent-aware queue adaptation in music recommendation systems
Recommend Signal — temporal offline evaluation for recommendation policies, with explicit causal boundaries.
CTR/ranking fundamentals practice with feature crossing, Logistic Regression baselines, AUC/LogLoss/nDCG notes and reproducible evaluation scripts.
Algorithm Intern Candidate | Recommendation / Search Retrieval / Ranking | PyTorch + Faiss | Offline Evaluation / Negative Sampling / Badcase Analysis
可复现的中文离线内容推荐应用原型,覆盖合成行为数据、动态兴趣画像、双路召回、个性化排序、多样性重排、推荐解释与离线评估。
Offline finance reconciliation agent with evidence-backed matching, unposted review proposals and independent ledger-state checks.
轻量级终端 AI Agent Harness:手写状态机、能力审批、Checkpoint 恢复与增量审计;仅 1 个直接运行依赖,500+ 项测试与 10 项离线 Evals。
Offline security investigation agent with scoped evidence, cited dispositions, simulated review and independent state verification.
Local contract-review workflow with cited proposals, amendment coverage, exact approvals and authored outcome evaluation.
Offline RAG retrieval-quality harness. Recall@k, nDCG, MRR, chunking diagnostics, regression diffs. No LLM-as-judge required. CI-friendly.
Python and Java support workflow with reviewed returns, versioned policy, SQLite state and independent offline outcome tests.
End-to-end joke recommendation system with offline evaluation, FastAPI serving, Docker, and model artifacts.
Budgeted high-fidelity evaluation for symbolic regression
Production-ready recommender system suite: serving API, pipelines, algorithm SDK, and evaluation tooling.
Neural Thompson Sampling contextual bandit for personalized Type 2 Diabetes therapy selection — training pipeline, offline policy evaluation (IPS/SNIPS/DM/DR), safety gates, drift monitoring, and LLM-generated clinical explanations.
To associate your repository with the offline-evaluation topic, visit your repo's landing page and select "manage topics."