Evidence hierarchy
No lower-ranked source overrides a higher-ranked source. When sources conflict, we publish both and explain the discrepancy.
- Live benchmark result JSON: timestamped, hashed, provider-identified. This is the source of truth for every score.
- Provider API metadata: returned during the test run — model ID, routing, token counts, cost.
- First-party documentation: model cards, provider docs, published policies.
- Direct product testing: sign-up, purchase, use, cancellation — recorded with dates and evidence.
- Authenticated search data: for content planning only, never for scoring.
- Community discussion: used only as a pain-point signal to identify what to test.
- Third-party and affiliate claims: clearly labeled and never presented as test evidence.
Ranking rules
- Untested products and models are not ranked. A model profile may list architectural specs while awaiting a live test run, but no score or rank is published without evidence.
- Sponsorship never changes score or position. If a commercial relationship cannot be separated from editorial judgment, we reject it.
- Fixture output is not evidence. Synthetic test data used to validate the site pipeline is never presented as a model's real response.
- Vendor access does not guarantee favorable coverage. We test products we pay for, on free tiers, or with vendor-supplied access that is disclosed.
- Material corrections are logged and dated. See the corrections page for the full record.
- Scores expire after 45 days unless the model and provider route are unchanged and a fresh test confirms the score.
Reviewer standards
- All subjective scores (roleplay quality, character consistency, prose quality) require at least two independent human reviewers.
- Reviewers use fixed rubrics. Score changes greater than 15 points require a third reviewer.
- Reviewers do not know the model identity during scoring (blind review where technically feasible).
- Automated metrics (latency, token counts, cost, error rates) are computed programmatically and not subject to human adjustment.
Adult-content boundary
Testing includes consensual adult fictional scenarios using fictional consenting adults. Public samples remain non-explicit. Tests involving prohibited categories are excluded from publication and do not retain or publish harmful completions.
"Uncensored" means the model delivers on lawful adult fiction requests without refusal. The CLIMAX Benchmark scores models on their ability to follow instructions and produce requested output.
Excluded from publication
- Content involving or sexualizing minors.
- Nonconsensual intimate imagery.
- Real-person sexual deepfakes.
- Face-swap, undress, or nudify instructions.
- Jailbreak instructions designed to defeat safety systems.
- Explicit image galleries or embedded adult generators.
- User uploads, public comments, or anonymous ratings.
Commercial disclosure
Every page that contains an affiliate link discloses the relationship before the first commercial link. Affiliate links use rel="sponsored nofollow". Sponsored placements are labeled as such and excluded from editorial scoring. See the full affiliate disclosure.