ADULT-ORIENTED RESEARCH · v0.3 live run reviewed · 161 cases approved · 18 models published · raw evidence available
RETIRED / 0.2.0

The old twenty-test baseline.

Preserved exactly because deleting inconvenient history is worse than labeling it. This version used different cases, scoring, and review rules. Its 16 reviewed run records are not current model claims.

Why it was retired

v0.2 mixed a general capability benchmark with uncensored-specific claims, used weaker artifact retention, and published scores under a different review standard. v0.3 replaces the catalog rather than pretending those numbers still answer the question.