Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.
-
Updated
Oct 2, 2026 - Python
Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.
General Assembly's 2015 Data Science course in Washington, DC
Evidence-backed use cases, prompts, integrations, evaluations, and safety notes for OpenAI GPT-6 Astra.
Machine Learning notebooks for refreshing concepts.
Awesome Jev: a source-backed field guide to TypeSafe's System One model, with SDKs, live demos, agent tools, and independent evaluations.
Local Interpretable Model-Agnostic Explanations (R port of original Python package)
魔搭紫皮书|ModelScope Cookbook:面向开发者的开源模型应用实战指南,覆盖模型选型、推理、微调、评测、RAG、Agent 与 AIGC,从跑通第一个模型到构建实际应用。
🔍 Minimal examples of machine learning tests for implementation, behaviour, and performance.
an MLOps/LLMOps platform
AI-powered NBA game outcome predictor that uses advanced team stats and trend-based features to forecast winners and track model performance
A reproducible local LLM benchmarking platform with versioned datasets, deterministic scoring, and auditable reports.
Community registry of LLM serving-path traps that produce confidently wrong measurements: templates, tool parsers, reasoning fields, quant kernel paths, CUDA toolchains, KV allocation, eval harnesses, versioning. Symptom-first, with the check that catches each.
CloudCV GSoC Ideas
UBC ARBERT and MARBERT Deep Bidirectional Transformers for Arabic
Customers in the telecom industry can choose from a variety of service providers and actively switch from one to the next. With the help of ML classification algorithms, we are going to predict the Churn.
Evaluate visual models on your own images, JSON Schema, and production constraints with LangGraph, Pareto analysis, and a no-key Replay demo.
Measure and visualize machine learning model performance without the usual boilerplate.
Flexible tool for bias detection, visualization, and mitigation
A High-level Scorecard Modeling API | 评分卡建模尽在于此
An in-depth analysis of audio classification on the RAVDESS dataset. Feature engineering, hyperparameter optimization, model evaluation, and cross-validation with a variety of ML techniques and MLP
To associate your repository with the model-evaluation topic, visit your repo's landing page and select "manage topics."