Projects with this topic
-
LLM hallucination detector: highlights the words an AI answer was unsure about, from OpenAI-compatible token logprobs, and shows what it almost said instead. Terminal, HTML or Markdown report, CI gate that fails on low confidence. Rust CLI and library, runs locally.
Updated -
10 hands-on demos for agentic AI security: blind verification, AIBOM, eval invariants, authority confinement, recon/malware/LFI/SSRF/scan — offline only.
Updated -
Diff and version LLM outputs: semantic diff, a version store and lineage tracking across prompt iterations, for audit trails. Rust.
Updated -
Single-header C++17: A/B test prompts or models with Welch's t-test, Cohen's d and custom scorers. Part of llm-cpp.
Updated -
Single-header C++17: run a prompt N times, measure consistency, compare models or prompts and score responses. Part of llm-cpp.
Updated -
Production-grade toolkit for evaluating RAG (Retrieval-Augmented Generation) pipelines.
Updated -
Model Verification Layer is a modular system for evaluating, comparing, and validating the behavior of large language models using structured benchmarks, logic consistency checks, cross-model consensus analysis, and policy-aware constraints. It provides a transparent framework for understanding how different AI models perform under identical conditions, enabling more reliable model selection, safer deployment, and user-adaptive decision making. https://roxanneardary.com/model-verification-layer/
Updated -
Case study of a pre-alpha Algebra I tutoring platform focused on hybrid AI/deterministic architecture, multi-provider LLM evaluation, safety controls, and measured cost/latency tradeoffs.
Updated -
A small, transparent experiment testing whether language models distinguish solvable prompts from prompts containing missing or contradictory information, and whether their stated confidence tracks correctness
Updated -
Most LLM reasoning benchmarks come from Western math, English logic, or code. Bazi-Bench tests multi-step rule-following inference in a different formal system: traditional Chinese Ba Zi. Frozen tables, Python reference impl, gold CoT cases, all mechanically verifiable.
Updated