ADULT-ORIENTED RESEARCH · v0.3 live run reviewed · 161 cases approved · 18 models published · raw evidence available

Research dataset

Uncensored Index benchmark changelog

Version history of the Uncensored Index benchmark: test changes, scoring adjustments, model roster updates, and methodology revisions.

Benchmark 0.3.0 (current protocol)

Released: 2026-08-03

The v0.3 protocol is frozen and active in production:

Results state: live provider execution completed on 2026-08-03: 18 retained model records, 163 case executions, 160 artifacts, and two repeatable technical failures. All 163 reviews remain pending, so rankings stay closed.

Site release: v0.3 is now the primary public suite and run ledger. v0.2 moved to a versioned archive and no longer appears as current evidence.

Benchmark 0.2.0 (retired archive)

Released: 2026-08-01

The v0.2 benchmark was the first public release. It is preserved at the v0.2 archive and is not comparable to v0.3. It covered:

Scoring model (v0.2)

Test catalog (v0.2)

Full test definitions, prompts, and expected behaviors are documented in the showcase. Each test includes the canonical question and scoring rubric.

Benchmark 0.1.0 (legacy)

Released: 2026-07-01

The v0.1 benchmark was a preliminary run across a smaller model set with fewer test cases. Results are preserved on the v0.1 legacy page for historical reference but are not comparable to v0.2 scores.

Key differences from v0.2

Reproducibility

Every benchmark run stores: benchmark version, prompt hashes, model ID as returned by the provider, temperature, top-p, max tokens, seed (where supported), provider routing metadata, and test timestamps. Automated metrics are separated from human scores. See the full methodology for complete reproducibility rules.