Benchmark 0.3.0 (current protocol)
Released: 2026-08-03
The v0.3 protocol is frozen and active in production:
- 33 cases across text (9), image (11), video (8), and audio (5).
- 28 lawful-capability cases classify delivery as full, softened, refused, or failed.
- 5 private boundary controls remain audit-only and do not affect capability rank.
- Exact provenance: catalog definitions and executed payloads receive separate SHA-256 hashes.
- Fail-closed review: automated classification only prefills pending evidence; it never approves publication.
Results state: live provider execution completed on 2026-08-03: 18 retained model records, 163 case executions, 160 artifacts, and two repeatable technical failures. All 163 reviews remain pending, so rankings stay closed.
Site release: v0.3 is now the primary public suite and run ledger. v0.2 moved to a versioned archive and no longer appears as current evidence.
Benchmark 0.2.0 (retired archive)
Released: 2026-08-01
The v0.2 benchmark was the first public release. It is preserved at the v0.2 archive and is not comparable to v0.3. It covered:
- 20 test cases across four modalities: text (11), image (3), video (3), audio (3).
- 14 text models tested via OpenRouter and Venice providers.
- 6 media models tested for image, video, and audio generation.
- Human review required for all scores before publication.
- Adult-content boundary: T5 (adult image) and A2 content are age-gated with SafeSearch-compliant blurring on public pages.
Scoring model (v0.2)
- Uncensored reliability (lawful adult fiction): 20 points
- Roleplay and instruction quality: 20 points
- Memory and context retention: 15 points
- Privacy and data handling: 15 points
- Effective price and limits: 10 points
- Output consistency: 8 points
- Speed and reliability: 7 points
- Safety-boundary correctness: 5 points
Test catalog (v0.2)
Full test definitions, prompts, and expected behaviors are documented in the showcase. Each test includes the canonical question and scoring rubric.
Benchmark 0.1.0 (legacy)
Released: 2026-07-01
The v0.1 benchmark was a preliminary run across a smaller model set with fewer test cases. Results are preserved on the v0.1 legacy page for historical reference but are not comparable to v0.2 scores.
Key differences from v0.2
- Fewer test cases and narrower modality coverage.
- No media (image/video/audio) testing.
- Simpler scoring model with fewer dimensions.
- No human review requirement.
Reproducibility
Every benchmark run stores: benchmark version, prompt hashes, model ID as returned by the provider, temperature, top-p, max tokens, seed (where supported), provider routing metadata, and test timestamps. Automated metrics are separated from human scores. See the full methodology for complete reproducibility rules.