AgentPrahari
analytics EMPIRICAL BENCHMARKS

442 Adversarial Attacks, Measured Locally

Every performance and security claim is backed by deterministic reproducible benchmarks. Run python run_comprehensive_benchmark.py to verify all 442 test vectors on your machine.

328 / 442
74.2% Pass Rate (Phase 1)

Local deterministic regex & AST heuristics alone with zero network latency.

350 / 442
79.2% Pass Rate (Phase 2)

Combined with optional Tier-2 Groq LLM-as-a-Judge semantic arbiter.

0.12 ms
p50 CPU Latency

Measured over 1,000 iterations on standard single developer CPU core.

100% Offline
Zero Network Calls

Runs completely inside process memory with zero external dependencies.

Empirical Test Results by Category

Evaluated against the comprehensive adversarial suite

Verified Deterministic
Attack Category Test Cases Blocked / Allowed Resolution Rate Defense Layer
Basic Prompt Injection Overrides 10 10 / 10 100% BLOCKED Regex + Semantic Hierarchy
Instruction Hierarchy & Delimiters 10 10 / 10 100% BLOCKED Pseudo-header & XML Tag Stripper
Unicode & Homoglyph Steganography 10 10 / 10 100% BLOCKED NFKD + Cyrillic Bounded Normalizer
Educational Trojan Bypasses 10 10 / 10 100% BLOCKED Trojan Framing Detector
Educational Discussion (Harmless) 10 10 / 10 100% ALLOWED Referential Inquiry Whitelist
Destructive Shell Commands (rm -rf) 11 11 / 11 100% BLOCKED CommandGuard Binary Whitelist
Destructive SQL DDL & Drops 10 10 / 10 100% BLOCKED Literal-Aware SQL Comment Stripper
Output Structural JWT & API Keys 10 10 / 10 100% SCRUBBED Base64url Structure Validator

Micro-Benchmark Latency Profile

Measured over 1,000 iterations per detector module on a single CPU core

PII & Luhn Validator
0.089 ms p50
Prompt Injection Guard
0.124 ms p50
Command & SQL Guard
0.080 ms p50
Secret & JWT Scrubbing
0.032 ms p50

Reproduce Benchmarks on Your Machine

Run the comprehensive benchmark script to re-evaluate all 442 test cases locally.

python run_comprehensive_benchmark.py