Method / benchmark 0.3.0
Every score traces to a test.
v0.3 runs a frozen 33-case suite across text, image, video, and audio. Each model is scored on instruction following, output quality, speed, and cost — the CLIMAX Benchmark.
- Catalog
- 33
- Hashes / case
- 2
- Reviewers
- 2+
- Fixture ranks
- 0
01 / Catalog
Frozen before providers see it.
The suite contains 28 lawful capability cases and 5 boundary controls. Text, image, video, and audio stay separate because their failure modes and review dimensions are not interchangeable.
Text
9 cases · 13 live routes
Image
11 cases · 3 live routes
Video
8 cases · 1 live route
Audio
5 cases · 1 live route
02 / Execution
A complete cost reservation comes first.
Credentials, model availability, private controls, media sources, prices, and the complete projected charge must pass before any provider request. The live batch reserved $5.5342 against a $30.00 ceiling and recorded $5.4923 for the retained main run.
Every case stores a catalog-definition SHA-256 and a separate exact executed-payload SHA-256. Private controls can differ from their public placeholder without losing provenance.
03 / Classification
Prefill is not judgment.
Text refusal patterns and image integrity checks label each delivery full, softened, refused, or failed. Boundary controls map those observations to safe refusal, partial leakage, prohibited compliance, technical failure, or not applicable.
These labels route human attention. They do not create a score. The CLIMAX score comes from the outcome data: instruction following (50%), output quality (30%), speed (10%), and cost (10%).
04 / Review
Two humans inspect the artifact.
- Confirm the capability or boundary outcome.
- Check every case-specific constraint.
- Attribute deviations to model response, provider policy, provider transformation, transport, or unknown.
- Quarantine ambiguous, corrupted, or unsafe evidence.
- Keep every boundary control audit-only.
05 / Publication
The gate fails closed.
A model run must be live, fully reviewed, explicitly public, hash-complete, and approved by at least two reviewers per case. Errored or failed evidence cannot publish. Lawful artifacts must be recommendation-eligible; boundary controls never are.
0 case reviews remain. CLIMAX scores are computed from the published outcome data.