Independent model testing / benchmark 0.1.0
Uncensored models. Controlled tests. No vendor spin.
We test roleplay quality, lawful adult-content reliability, memory, privacy, safety boundaries, cost, and speed. Every result is tied to a model ID, provider route, test date, and reproducible method.
Launch roster
Thirteen models, one protocol
Aion 3.0
AionLabs · 131,072 token context
Score 40.0MiniMax M2-her
MiniMax · 65,536 token context
Awaiting live testCydonia 24B V4.1
TheDrummer · 131,072 token context
Score 50.0Llama 3.3 Euryale 70B
Sao10K · 131,072 token context
Score 85.4Dolphin Mistral 24B Venice Edition
Cognitive Computations / Venice · 128,000 token context
Score 77.8Hermes 3 405B
Nous Research · 131,072 token context
No models match those filters.
Method
What the benchmark measures
Persona control, continuity, initiative, prose, and repetition.
Delayed recall, false-memory resistance, and long-session consistency.
Consensual fictional adults only. Explicit outputs stay private.
Correct refusal for minors, nonconsent, and real-person deepfakes.
Private, anonymized, encrypted, local, or unknown routes are separated.
Token usage, effective price, latency, throughput, and errors.
Requested and returned model IDs, route, version, and date.
Retests expose silent model switches and quality drift.
Why this exists
Most “best” lists do not show the test.
Vendor claims are not evidence. Affiliate payouts are not evidence. One cherry-picked screenshot is not evidence.
Uncensored Index publishes the model route, protocol, test date, result state, and limitations. If a model has not passed the publication gate, it does not get a rank.