Benchmark 0.2.0 (current)
Released: 2026-08-01
The v0.2 benchmark is the first public release. It covers:
- 20 test cases across four modalities: text (11), image (3), video (3), audio (3).
- 14 text models tested via OpenRouter and Venice providers.
- 6 media models tested for image, video, and audio generation.
- Human review required for all scores before publication.
- Adult-content boundary: T5 (adult image) and A2 content are age-gated with SafeSearch-compliant blurring on public pages.
Scoring model (v0.2)
- Uncensored reliability (lawful adult fiction): 20 points
- Roleplay and instruction quality: 20 points
- Memory and context retention: 15 points
- Privacy and data handling: 15 points
- Effective price and limits: 10 points
- Output consistency: 8 points
- Speed and reliability: 7 points
- Safety-boundary correctness: 5 points
Test catalog (v0.2)
Full test definitions, prompts, and expected behaviors are documented in the showcase. Each test includes the canonical question and scoring rubric.
Benchmark 0.1.0 (legacy)
Released: 2026-07-01
The v0.1 benchmark was a preliminary run across a smaller model set with fewer test cases. Results are preserved on the v0.1 legacy page for historical reference but are not comparable to v0.2 scores.
Key differences from v0.2
- Fewer test cases and narrower modality coverage.
- No media (image/video/audio) testing.
- Simpler scoring model with fewer dimensions.
- No human review requirement.
Upcoming changes
v0.3 (planned)
- Product-level testing for hosted chat platforms (SpicyChat, CrushOn, etc.).
- Expanded image benchmark with real provider output.
- Longitudinal comparison: v0.2 → v0.3 score deltas tracked per model.
- Directory filters for price, privacy, memory, and media support.
Reproducibility
Every benchmark run stores: benchmark version, prompt hashes, model ID as returned by the provider, temperature, top-p, max tokens, seed (where supported), provider routing metadata, and test timestamps. Automated metrics are separated from human scores. See the full methodology for complete reproducibility rules.