This publication covers adult-oriented AI systems. Adult media samples are age-gated and blurred by default. Indexable pages stay text-only.

Research dataset

Uncensored Index benchmark changelog

Version history of the Uncensored Index benchmark: test changes, scoring adjustments, model roster updates, and methodology revisions.

Benchmark 0.2.0 (current)

Released: 2026-08-01

The v0.2 benchmark is the first public release. It covers:

Scoring model (v0.2)

Test catalog (v0.2)

Full test definitions, prompts, and expected behaviors are documented in the showcase. Each test includes the canonical question and scoring rubric.

Benchmark 0.1.0 (legacy)

Released: 2026-07-01

The v0.1 benchmark was a preliminary run across a smaller model set with fewer test cases. Results are preserved on the v0.1 legacy page for historical reference but are not comparable to v0.2 scores.

Key differences from v0.2

Upcoming changes

v0.3 (planned)

Reproducibility

Every benchmark run stores: benchmark version, prompt hashes, model ID as returned by the provider, temperature, top-p, max tokens, seed (where supported), provider routing metadata, and test timestamps. Automated metrics are separated from human scores. See the full methodology for complete reproducibility rules.