Projects with this topic
-
Measure any tennis API against a published, reproducible method — response time, observed availability — and be told when your sample is too small to support a number.
Updated -
Zambo: the cross-AI execution layer. 136 native MCP tools over one endpoint, with a verifiable execution receipt for every call (AER-1 open draft). Free: generous free tier, no account. Day Pass $0.99 USDC on Base for 24h; Zambo Pass $29/month ($19/month for verified $ZAMBO holders). https://zambo.dev
Updated -
-
Open dataset: bitHuman avatar render speed (× real time) by device, from docs.bithuman.ai/performance. Code Apache-2.0, data CC BY 4.0.
Updated -
-
Raw benchmark results and statistical analysis.
Updated -
Объектив · Исследование распознавания техники: эталон, прогоны пайплайнов, метрики, лидерборд, выводы
Updated -
Evaluation harness for measuring how well AI models perform on SysML v2 modeling tasks.
Updated -
Public evidence datasets, claim atlases, and reproducible benchmarks for ClaimBound.
Updated -
An Extensible Benchmark Framework for Real-Time Applications
Documentation: https://rt-bench.gitlab.io/rt-bench/
Updated -
-
-
Helm Charts for various benchmarks
Updated -
The repository contains the code used in an extensive benchmark of co-occurence based inference methods to recover the interaction structure of microbial communities from metabarcoding data (16S rDNA-seq data)
Updated -
Empirical validation of C4 geometric defense against 16 Agents of Chaos. 550 adversarial prompts. 4 defense systems. 96.7% block rate. LLM validation on GPT-4o-mini + Mistral 7B. MIT.
Updated -
Kevlar Benchmark: OWASP Top 10 for Agentic Apps (AI-Agents) 2026 a Red Team Benchmark.
Updated -
A benchmark library with statistical analysis and plotting capabilities in C++. https://cppstatbench.musicscience37.com/
Updated -
Red Team AI Benchmark: Evaluating LLMs for authorized offensive-security tasks. Red Team AI Benchmark is a CLI model-evaluation benchmark. It measures how LLMs understand and respond to red-team questions and security scenarios; it is not a tool for carrying out those activities. Version 2 uses a rubric-based dataset instead of judging answers only against one golden response.
Updated -
Mini benchmarking suite — performance testing utilities.
Updated -
Automated LLM Benchmarking on GPU - tokens/sec, latency percentiles, VRAM profiling, multi-format support (HuggingFace, GGUF, GPTQ)
Updated