Projects with this topic
Sort by:
-
Empirical validation of C4 geometric defense against 16 Agents of Chaos. 550 adversarial prompts. 4 defense systems. 96.7% block rate. LLM validation on GPT-4o-mini + Mistral 7B. MIT.
Updated -
Sapient Eval is an open-source, AGPL-3.0+ compliant benchmarking framework for evaluating AI models across accuracy, speed, efficiency, and reasoning performance. Built around a modular, spec-driven architecture, it enables users to define industry-specific evaluation standards, run reproducible tests on locally hosted or remote models, and compare results transparently to establish measurable, empirical performance across synthetic intelligence systems. https://roxanneardary.com/sapient-eval/
Updated