Model Reality
Ranking
Benchmark Performance
| Benchmark | Bench Quality | Effective Score | Submissions | |
|---|---|---|---|---|
| View → |
Sign in to vote on tags or suggest new ones.
Bench Weight
Ranking
| # | Model | Effective Score | Submissions | |
|---|---|---|---|---|
| View → |
Powered by GitHub Discussions via giscus. Sign in with GitHub to comment.
Sign in to vote on tags or suggest new ones.
Submission Detail
Discussion
Powered by GitHub Discussions via giscus. Sign in with GitHub to comment.
Contribute.
Submit benchmark scores. The community validates through voting.
Sign in with Google to submit.
✓ Submitted
score(s) saved. Live and ready for community review.
About.
A community-driven, editorially honest ranking of AI models — and the math behind it.
Every public AI leaderboard tells you something — but each one tells you something different, and many tell you almost nothing useful. Three failure modes are everywhere:
- Contamination — the test set leaked into the training data, so the score doesn't measure capability, it measures memorisation.
- Saturation — every frontier model scores 99 %, so the bench has no resolving power left, but it still dominates aggregate scores out of habit.
- Bench-maxing — a model is tuned for a small set of popular benches and looks great there while being unusable in practice.
SupraBench gives every bench a community-voted quality and difficulty score, and adds an automatic saturation penalty on top. The result: a single number per model that respects how trustworthy and how informative each underlying benchmark actually is.
The Bench Score is the headline number shown on every bench page. It answers a single question: "how much should this benchmark count when ranking models?" A high-quality, hard, unsaturated bench that the community endorses should count more than a one-rater vanity bench on a trivial task.
It's a single number on $[0, 100]$ built from four independent factors. We define each variable first, then put them together.
Variables
- $b$ — a single benchmark.
- $Q(b) \in [0, 100]$ — the bench's quality: how trustworthy the score is. Mean of four community ratings (relevance, contamination resistance, discriminability, reproducibility) on a 1–5 scale, multiplied by 20. Defaults to $50$ when nobody has rated yet. (Full breakdown in "What exactly are the five bench dimensions?" below.)
- $D(b) \in [0, 1]$ — the bench's difficulty: how much general intelligence it actually probes. Median of community 1–5 ratings, linearly mapped to $[0, 1]$ (a vote of $1$ becomes $0$, a vote of $5$ becomes $1$). Defaults to $0.5$ when un-rated.
- $H(b) \in [0, 1]$ — the bench's headroom: how much measurement signal it has left before it saturates. Computed automatically from the top-$K$ frontier mean on $b$ — fully open ($H = 1$) until at least 3 models are evaluated, shrinks toward a floor of $0.1$ as the frontier mean climbs from $50$ toward $100$. (Full formula in "What is headroom and why does it exist?" below.)
- $u_b \in \mathbb{Z}_{\geq 0}$ — the bench's net upvotes: number of distinct upvoting accounts minus number of distinct downvoting accounts on the bench's existence-vote, floored at $0$.
- $U^\star = \max_{b'} u_{b'}$ — the maximum $u_b$ across every non-hidden bench in the system. Bootstrap: if $U^\star = 0$ (brand-new deployment, no votes anywhere), this factor is disabled for everyone.
The formula
$$\operatorname{BenchScore}(b) \;=\; \underbrace{Q(b)}_{\text{quality}} \;\cdot\; \underbrace{D(b)}_{\text{difficulty}} \;\cdot\; \underbrace{H(b)}_{\text{headroom}} \;\cdot\; \underbrace{u_b / U^\star}_{\text{upvote share}}$$
Quality is on $[0, 100]$; the three other factors are all on $[0, 1]$, so the Bench Score stays naturally on $[0, 100]$ — the same scale as the per-model scores.
What each factor does
- Quality $Q(b)$. If the community thinks the bench is contaminated, trivia-grade, or non-reproducible, $Q$ shrinks and the bench's whole weight shrinks with it. A bench rated 1/5 across the board has $Q = 20$, fivefold less weight than a unanimous 5/5.
- Difficulty $D(b)$. A bench rated 1/5 in difficulty has $D = 0$ and is silently filtered out — scoring 100 % on a trivial bench earns the model nothing. A bench rated 5/5 has $D = 1$ and is counted at full weight.
- Headroom $H(b)$. Once a bench is solved (top frontier models all near 100), $H$ shrinks toward the $0.1$ floor, so the bench stops dominating the leaderboard the moment it loses resolving power. This makes ARC-AGI 3 → 4 hand-offs automatic, no manual retirement needed.
- Upvote share $u_b / U^\star$. Without this, anyone could mint a brand-new bench, self-rate it $Q = D = H = 1$ and immediately appear at #1 on the bench leaderboard. With it, a 1-upvote vanity bench at $u_b = 1$ against an established $U^\star = 100$ is worth $1\,\%$ of an equally well-rated leader. The most-upvoted bench always has share = 1, so there's no self-penalty at the top.
The upvote share is linear on purpose: user trust should be able to dominate everything else. How many models ran a bench does not change its weight — a trusted specialist bench is not buried just because few labs have run it yet.
Can the community overturn the current default? Yes. If a new benchmark reaches 10 net upvotes while an older one remains at 1, the new benchmark receives its full quality × difficulty × headroom weight and the older benchmark receives a 1/10 trust factor.
One refinement inside the ranking: while a bench has only a handful of raters, the ranking pulls each 1–5 rating toward the neutral 3 with a three-rater prior, so one enthusiastic rater is not treated as settled consensus. Many consistent ratings overpower the prior. The Bench Score shown on the bench pages is the unshrunk product above.
A model's SupraScore is the headline number on every model page and the leaderboard sort key. It answers one question: "on a benchmark the community trusts, how often does this model beat the current frontier?" It builds on the Bench Score from the previous section.
Three preparatory steps turn raw submissions into one number per (model, bench) cell. Three scoring steps turn those cells into the SupraScore.
Throughout this section, "model" means one concrete configuration — e.g. GPT-6 Astra (xhigh) — because that is the unit benchmarks are run on. The default Model ranking then shows each model once, with the score of its best configuration (see "What is a model, and what is a configuration?" below).
Foundations — submission to per-pair score
Step 1 — Normalise each submission. Every bench has its own scale (e.g. 0–100, 0–10, 0–1). We map every raw submission to a common $[0, 100]$ scale using the bench's declared min and max.
$$s_{\text{norm}} \;=\; \frac{s_{\text{raw}} - s_{\min}}{s_{\max} - s_{\min}} \cdot 100$$
Step 2 — Validate via community votes. A submission only counts if its net vote is positive. The submitter gets an implicit $+1$, so a fresh submission starts at $1/0$. Wrong, duplicate, or fraudulent submissions get downvoted out of the calculation.
Step 3 — Take the median per (model, bench) pair. The same model is often submitted several times on one bench. All valid submissions for a pair collapse into one number:
$$\mu_{m,b} \;=\; \operatorname{median}\bigl(\mathcal{S}_{m,b}\bigr)$$
Median, not mean — one outlier submission (a fluke run, a misconfigured prompt, a fake that hasn't been downvoted yet) cannot shift the result.
Step 4 — Duels instead of averages
On every bench, every pair of models that both have a score is a duel: the higher $\mu$ wins, equal values tie. We do not average raw percentages across benches, because a point does not mean the same thing everywhere — 5 % leads ARC-AGI-3 while 98 % is ordinary on a saturated bench, and being first of 90 models is not the same as being first of 8.
Each duel on bench $b$ carries the weight
$$w_{\text{duel}}(b) \;=\; \frac{\operatorname{BenchWeight}(b)}{\overline{\operatorname{BenchWeight}}} \cdot \frac{1}{n_b - 1}$$
- $\operatorname{BenchWeight}(b)$ — quality × difficulty × headroom × upvote share, as in the previous section (with the small-rater prior). A saturated or untrusted bench produces weak duels; a bench with difficulty 1/5 produces none.
- $n_b$ — how many models have a score on $b$. Dividing by $n_b - 1$ gives every model the same total weight from a bench no matter how many others ran it, so a bench with 90 entrants does not drown out one with 15.
A bench that only one model has run produces no duel. Not being tested is never negative evidence — and a self-reported score on a bench nobody else touched is worth nothing.
Step 5 — One fit across all benches
All duels from all benches go into one Bradley-Terry model — the same family of models behind chess Elo and arena-style leaderboards. Every model $m$ gets a strength $a_m$ such that
$$P(m \text{ beats } k) \;=\; \frac{1}{1 + e^{-(a_m - a_k)}}$$
matches the observed weighted duel results as well as possible. Because models share opponents, the benches are tied together: beating a strong field counts for more than beating a weak one, even when two models never met on the same bench. A small regularisation ($\lambda = 0.15$, chosen by held-out prediction, not by a desired order) pulls thinly tested models toward the middle, so a single spectacular result cannot outrank models that won across the board.
Step 6 — The number you see
$$\operatorname{SupraScore}(m) \;=\; 100 \cdot \frac{1}{10} \sum_{k \,\in\, \text{top-10}} P(m \text{ beats } k)$$
In plain words: the expected win rate against the ten strongest configurations right now (a model inside that group counts a tie against itself). 50 means "as good as the average frontier model"; 60 means it wins six of ten such duels. The number is relative to the current frontier by design — when a stronger model arrives, everybody else's win rate against the field drops.
The whole calculation is one dependency-free file, public/js/supra-rank-core.js. The server runs it for the leaderboard, and your browser runs the very same file for tag-filtered views and the simulator — there is no second implementation that could drift.
Why this replaced the weighted mean of raw percentages (September 2026): the audit, the rejected alternatives and the attack tests are published in docs/research/RANKING_REALITY_AUDIT_2026-09-16.md.
Each bench is rated on a 1–5 scale on five orthogonal axes. The first four feed quality; the fifth is difficulty.
- Relevance — does it measure something useful for real-world model use? High = meaningful task. Low = trivia, toy puzzles.
- Contamination Resistance — how sure are we the test set wasn't in the training data? High = held-out, rotating, freshly generated. Low = scraped from the public web before model release.
- Discriminability — does it separate weak from strong models? High = wide score spread. Low = everyone scores 95–100.
- Reproducibility — can two independent runs reach the same conclusion? High = deterministic, well-specified. Low = vague setup, judge-LLM, hand-graded.
- Difficulty — how much general intelligence does it actually probe? 1 = trivial. 5 = approaches frontier capability.
Quality is the mean of the first four × 20:
$$Q(b) \;=\; \frac{R + C + S + P}{4} \cdot 20 \quad \in [0,\,100]$$
Difficulty uses the median across raters (more robust to a single inflated vote), then linearly scaled to $[0,1]$:
$$D(b) \;=\; \max\!\left(0,\; \min\!\left(1,\; \frac{\operatorname{median}(d_i) - 1}{4}\right)\right)$$
If nobody has rated a bench yet, it defaults to neutral: $Q = 50$, $D = 0.5$.
Imagine ARC-AGI 3 just got solved — every frontier model scores 99 %. The bench has stopped measuring intelligence; it now just measures "is your model in the frontier club". A score on it should contribute almost nothing to a model's SupraScore.
The naïve fix would be: ask the community to retroactively lower its quality. Nobody does that consistently. So we automate it — every bench has a headroom factor that shrinks as the bench saturates:
The frontier-mean approach
Top-1 alone is too noisy: one freakishly good model would tank the bench's weight for everyone after a single submission. So we use the mean of the top-K models, which is robust to outliers:
$$N \;=\; \bigl|\{\, m : \exists \text{ valid score for } (m,b)\}\bigr|, \qquad K \;=\; \min(10,\, N)$$
$$f(b) \;=\; \frac{1}{K} \sum_{m \in \operatorname{top}_K(b)} \mu_{m,b}$$
where $\operatorname{top}_K(b)$ are the $K$ models with the highest per-model median on $b$.
From frontier-mean to headroom
Two regimes. With fewer than 3 evaluated models we don't have enough signal to claim saturation, so we don't penalise:
$$H(b) \;=\; \begin{cases} 1.0 & \text{if } N < 3 \\[6pt] \max\!\left(0.1,\; \dfrac{100 - \max(f(b),\, 50)}{50}\right) & \text{otherwise} \end{cases}$$
The pivot at 50 means a bench is "fully open" until at least its top-K models are above mid-range. The floor at 0.1 keeps historic benches alive — they never disappear, they just stop dominating.
Trajectory
| Scenario | $N$ | $f(b)$ | $H(b)$ |
|---|---|---|---|
| Brand-new bench, 1 model @ 92 | 1 | — | 1.00 |
| 3 models, top-3 mean 70 | 3 | 70 | 0.60 |
| 10 models, top-10 mean 84 | 10 | 84 | 0.32 |
| 30 models, top-10 mean 91 | 30 | 91 | 0.18 |
| 50 models, top-10 mean 99 | 50 | 99 | 0.10 (floor) |
This is what makes the ARC-AGI 3 → 4 hand-off automatic. Nobody has to "retire" the old bench manually.
Three models, two benches. Both benches have identical community ratings; the only difference is saturation.
| Bench | $f(b)$ | $H$ | Bench weight | Claude X | Model Y | Model Z |
|---|---|---|---|---|---|---|
| FrontierBench (fresh) | 65 | 0.70 | 47.9 | 72 | 65 | 40 |
| SaturatedBench | 98 | 0.10 | 6.8 | 98 | 99 | 97 |
Duels. On FrontierBench: X beats Y, X beats Z, Y beats Z. On SaturatedBench: Y beats X, Y beats Z, X beats Z. Each bench has three models, so each duel carries half of that bench's normalised weight — a FrontierBench duel weighs about seven times as much as a SaturatedBench duel.
Fit. X and Y each won one bench, but X won the one that still measures something:
| Model | Strength $a_m$ | SupraScore |
|---|---|---|
| Claude X | +1.27 | 72.8 |
| Model Y | +0.20 | 53.2 |
| Model Z | −1.47 | 24.0 |
Four takeaways:
- A raw-percentage average would have ranked Model Y first (82 vs 85) — purely because it is one point ahead on a bench where everyone scores 97–99.
- The saturated bench still counts a little: Y's win there keeps it clearly above the middle.
- A bench with difficulty 1/5 has weight 0 and produces no duels at all — scoring 99 % on it earns exactly nothing.
- Absolute percentages never enter the result, only who beat whom and how much the community trusts the bench it happened on.
Anything the community can get wrong, the community can correct via voting. Five layers:
- Submission votes — every individual score is up/downvoted. Only net-positive submissions count toward the SupraScore.
- Tag votes — each tag on a model or bench is voted independently. Net positive keeps it in the canonical tag set.
- Existence votes — fakes, duplicates, low-quality entries can be downvoted into a hidden state (see formula below).
- Quality & difficulty ratings — anyone signed in can rate a bench on all five dimensions. Averaged across raters (mean for trust, median for difficulty).
- Per-bench scaling — a bench with one corrupt rater barely moves; the more raters, the harder it is to game.
When does an entity get hidden?
The hide threshold scales with engagement, so a small mob can't kick a well-established bench, but spam still goes away fast:
$$\text{hide}(e) \;\Leftrightarrow\; \operatorname{down}(e) \;\geq\; \max\bigl(5,\, \lceil 0.6 \cdot (\operatorname{up}(e) + \operatorname{down}(e)) \rceil\bigr) \;\;\wedge\;\; \operatorname{down}(e) > \operatorname{up}(e)$$
A 5-downvote floor protects against drive-by spam. The 60 % ratio means that to remove a bench with 100 upvotes, you'd need at least 96 downvotes — much harder than 4 sock puppets.
Anti-resurrection
If you submit a model or bench under your name and it gets community-removed, you cannot re-submit it under the same name. Other users can re-submit it (with a numeric suffix on the slug), and the community votes on the new version independently.
Adversarial robustness is built into the formula at every layer:
- One submission can't anchor a score — the per-(model, bench) median becomes robust to outliers as soon as $n \ge 2$ submissions exist: a single attacker number is replaced by the community median the moment a second honest submission lands.
- One bench can't carry a model — every duel is weighted by quality × difficulty × headroom × community upvote share, and the fit is regularised: a model with one spectacular result is pulled toward the middle and cannot outrank models that won across many trusted benches.
- One user can't carry a bench — every bench's contribution to its own headline score and to any model's SupraScore is multiplied by $u_b/U^\star$ where $u_b$ is its net upvote count and $U^\star$ is the leader's. A self-rated 100/100 vanity bench from a single account is worth $1/U^\star$ of an established bench at the same Q·D·H — you'd need $U^\star$ separate accounts upvoting it to even tie. Same defence works on the bench leaderboard and in the SupraScore aggregate, so you can't spawn a bench just to pump one model.
- A bench only you ran is worth nothing — scores become duels, and a duel needs an opponent. A one-account, one-model vanity bench produces no duel at all, so a self-reported 100 on it cannot move any ranking until other models are measured there.
- Difficulty uses median, not mean — a single 5-star rater on a trivial bench can't fake difficulty.
- Saturation auto-detected — pumping a saturated bench gives diminishing returns by construction, and because only the order counts, a 99.4 on a solved bench is no longer worth 99 points of anything.
- A flattering configuration can't speak for its model — a configuration that reports only the benches it wins may look good as a row, but it can represent its model only if it covers at least half of the model's benchmark weight.
- Engagement-aware hide threshold — small voting cliques can't take down established entries.
- Anti-resurrection — re-submitting your own removed entries under the same name is blocked.
- Rate limiting — submissions are capped at 30 individual scores per 24 h per user.
None of these is bulletproof on its own. Together they make systemic gaming expensive enough that legitimate contribution is the cheaper path. Every defensive claim above is encoded as an executable invariant or attack scenario in tests/convex/adversarial-robustness.test.ts — the harness has three layers (invariants, an attack catalog, and a deterministic-PRNG fuzzer) and runs on every CI build, so a regression in the math fails a test instead of going unnoticed. One scenario (`A3-extreme`, an industrial-scale 8+ vanity bench farm) explicitly documents an attack that the pure math does not defend against — those are blocked operationally by the rate-limit, downvote, anti-resurrection and moderation rules above.
Any URL is a valid source. A reviewed first-party/original source — an academic or lab publisher, the benchmark's own project site, or an exact author-owned repository — gets an "Official source" badge. Third-party mirrors, videos, roundups, social posts and unreviewed repositories get a "Community source" badge.
Official does not mean good or highly weighted. It only answers who published this source. Benchmark quality, difficulty, saturation and community trust are scored independently, so a niche benchmark can have an official project page and still contribute almost no ranking weight.
Reviewed domains and exact repository prefixes live in convex/urls.ts in the public source repository on GitLab. The browser preview is regression-tested against the same lists.
The full source code of SupraBench, including the ranking math you just read, is published under the Business Source License 1.1 (BSL):
gitlab.com/florian-fischer-group/suprabench ↗
Source-available, not OSI-open-source. The BSL is a "source-available" licence: you can read, audit, fork for research, learning and non-commercial use, patch and redistribute — but the licence does not meet the OSI Open Source Definition because it carves out commercial competing-service use until the Change Date of 2029-01-01. On that date the entire codebase auto-converts to Apache License 2.0 and becomes plain open source under that name. We use this exact wording instead of the unqualified "open source" label because the OSI / FSF community is — rightly — strict about that distinction.
Issues, feature requests and merge requests are welcome on GitLab. Imprint, privacy and terms in the sidebar footer.
A model is one specific lab release and product tier — what people mean when they say "GPT-6 Astra" or "Claude Opus 5". It is not a vendor, a broad generation or an architecture line: "Claude Opus 5" and "Claude Fable 5" are separate models; so are "GPT-5.6 Sol", "GPT-5.6 Terra" and "GPT-5.6 Luna". Versioned releases such as Muse Spark 1.1 and 1.2 remain separate too.
A configuration is one way of running that model: reasoning effort, context window, fallback mode, harness. Benchmarks are always run on a configuration, so every score on SupraBench belongs to one. The rule across providers: provider + generation + product tier + release identify the model; everything else identifies the configuration.
How configurations are named
A configuration is named after its model with the setting in parentheses:
| Model | Configurations | What the suffix means |
|---|---|---|
Claude Opus 5 | Claude Opus 5 (high), … (xhigh), … (max) | Reasoning-effort level |
GPT-6 Astra | GPT-6 Astra (medium), … (high), … (xhigh), … (max) | Reasoning-effort level |
GPT-5.6 Sol | GPT-5.6 Sol (medium), … (xhigh), … (max) | Same product tier at different reasoning effort |
Gemini 3.1 | Gemini 3.1, Gemini 3.1 (thinking) | Explicit reasoning mode on/off |
Common suffixes you'll see: (low) / (medium) / (high) / (xhigh) / (max), (thinking), (128k) / (200k) / (1M) for context-length SKUs, (instruct) / (chat) / (base) for open-weight post-training variants.
How a model gets its score
There is one calculation, not two. Every configuration is ranked by the pairwise fit described above, and a model simply shows the score of its best configuration — the row names it ("best: …"). The Model and Configuration views therefore always agree: the model's number is a number you can find in the configuration table.
One guard: a configuration may represent its model only if it covers at least half of the model's total benchmark weight. A configuration that lists only the benches it wins can look strong as a row, but it cannot speak for the model. If no configuration reaches half, the most broadly tested one is used.
A model with fewer than three distinct benchmarks remains visible but is marked provisional. Tag filters rerun the same calculation on matching benchmarks instead of selecting a different hidden formula.
What if the community disagrees with a grouping?
New submissions still name the model their configuration belongs to, and ambiguous provider taxonomies are not guessed automatically. A small reviewed normalizer only enforces unambiguous distinctions already present in the official model name, such as GPT-5.6 Sol versus Terra or Muse Spark 1.1 versus 1.2. Corrections can be proposed with evidence through the public repository; the mapping and its tests are auditable.
A release with a single configuration simply uses the same name for both. In the API and the data export the model is the familyTag field — the name predates this wording and is kept for compatibility.
Tags exist on two different things on SupraBench, and they play two different roles:
- Bench tags describe what the bench measures —
reasoning,code,multilingual,vision,safety. These are structural: they decide which benches feed a tag-scoped score. - Model tags describe what the model is or claims to be —
multimodal,open-weights,moe,frontier,agentic. These are descriptive: they help you find a model in search but don't enter the math.
Why the asymmetry?
The "Filtered Score" column reweighs a model's SupraScore using only the benches that match the active tag. That math only makes sense if the tag actually selects a non-empty set of benches. A pure model tag like multimodal — useful for search ("show me models that claim to do vision") — would select zero benches and produce a null filtered score for every row. So we hide model-only tags from the chip bar entirely.
What this means in practice
- The tag-filter chip bar (top of the models / benches list) shows bench tags only, sorted by how many benches carry them.
- The search box on either list matches against everything — model name, provider, bench tags, model tags. Type
multimodalthere and Gemini surfaces immediately, even though no bench is tagged that way (yet). - The tag picker popup ("+ N more" pill) lists every bench tag with its bench count, so you can drill into any structural slice without scrolling the chip strip.
- Bench tags also show up in the benches table — clicking one there toggles the same shared filter as the chip bar.
Can a tag be both?
Yes. code is the obvious one: HumanEval is a code bench, GPT-5.3 Codex is a code-tuned model. The tag exists once globally; it just appears in a chip bar if at least one bench carries it, and surfaces in search regardless of which side it lives on. The tag-counts API tracks both sides separately so this remains true even after extensive recategorisation.
Implementation: the tagCounts Convex table keeps {benches, models} per tag; the chip bar reads tags.listForBenches and the autocomplete reads tags.listAll. See convex/tags.ts.
They can accumulate more total weight today. Community ratings determine how much each individual benchmark counts, but the production formula does not currently reserve fixed mass for coding, agents, reasoning, multimodal work, knowledge or long-context tasks.
We are researching hierarchical category weighting without replacing users' judgements. Ratings would still allocate weight within a category; the viewer would choose how categories share the total through an uncapped, community, balanced or custom profile. The active profile and every category contribution would always be visible.
No category cap is active yet. The existing uncapped score remains the canonical leaderboard while primary-category assignments, preference participation and shadow rankings are audited. This avoids silently imposing a maintainer's idea of what intelligence should mean.
The full feasibility and governance proposal is published in docs/research/CATEGORY_WEIGHTING_CONCEPT.md in the source repository.
Yes — partially. The /v1/* endpoints are live and answering requests today, but the only keys we mint right now are free Partner keys for non-profit, research and open-source projects we explicitly approve. Paid self-serve tiers (Starter / Pro / Enterprise) stay on a waitlist until enough developers are actually queued for them — running Stripe + a dedicated edge cache for an API nobody uses isn't free. Pricing is TBD until that launch, the waitlist is how we figure out what each tier should actually cost. Enterprise plans are always custom. Billing will be handled by Stripe (EU VAT collected automatically).
Read the full API documentation — every endpoint, error code, rate-limit and example is already there. For paid tiers, join the waitlist for the one you'd actually subscribe to: setProfileTab('api'))">Profile → API & Billing. When the queue hits launch threshold we ship and email everyone in signup order.
Running a non-profit, research or open-source project that could genuinely use the API today? Pitch us for a Partner key — free, negotiated quota, live right now. The Apply to become a partner button at the bottom of the pricing grid opens a pre-filled form with the bits we need to evaluate.
Have a specific use case (dashboard, leaderboard mirror, evaluation tooling)? Tell us via a GitLab issue or the partner mailto — high-signal asks weigh more than passive signups.
Sign in to view your profile.
My Submissions
| Model | Benchmark | Score | Status | Submitted | |
|---|---|---|---|---|---|
| hidden | hidden | View → |
My Creations
My Tag Votes
Your API access
Tier is active — ready to call
https://api.suprabench.com/v1/ with any of the keys below.
/v1/export.jsonAPI keys
created
· last used
· never used
https://api.suprabench.com/v1/.
Recent months
| Month | API calls |
|---|---|
Need more capacity, additional keys, or to swap a key? Email us and we'll re-provision. Read the full API docs for endpoints, schemas and rate-limit headers.
Plans
Starter
TBD at launch
10 000 requests / month
- Read access to all public endpoints
- 60 req/min rate limit
- 1 API key
- Community support
Pro
TBD at launch
100 000 requests / month
- Everything in Starter
- 300 req/min rate limit
- 3 API keys
- Email support
- Bulk export endpoint
Enterprise
TBD at launch
1 000 000 requests / month
- Everything in Pro
- 1 200 req/min rate limit
- 10 API keys
- Priority support (best-effort, no formal SLA)
Enterprise+
Custom
Custom quota, optional contractual SLA, on-prem mirror
- Everything in Enterprise
- Custom rate limits
- 50+ keys
- Dedicated Slack channel; SLA negotiable per contract
- Custom data agreements
Partner
Free (sponsored by the project)
Custom quota & key count
- Full API access (read + bulk export)
- For non-profit, research & open-source projects
- Hobby / friend-of-the-project sites welcome too
- Quota, rate limit & key count set per partner
- Attribution expected (
Powered by SupraBenchfooter link)
Pricing is intentionally TBD until launch — the waitlist is how we figure out what each tier should actually cost. Read the full API documentation →. Billing will be handled by Stripe (EU VAT collected automatically, B2B reverse-charge supported via VAT-ID at checkout). See Terms § API for cancellation, refund and uptime details.
Your subscription
You're on the plan.
Your subscription is set to cancel on . You keep API access until then.
API keys
Save this key now
This is the only time we'll show this key. Store it in a password manager — you can't see it again, only revoke + recreate.
Simulator.
Calculate the SupraScore your unreleased model would land at if you submitted these scores against existing benches. Nothing is saved — runs are purely a what-if for your decks.
Your
grant has simulationsPerDay = 0.
Open the Admin tab → find your own account → set Simulator runs / day to 20 (or whatever) → re-grant.
Ask your SupraBench account contact to bump it.
Per-bench impact
| Bench | Frontier (live → sim) | Weight (live → sim) |
|---|---|---|
| → | → |
Hypothetical leaderboard
Top 25 of the simulated full ranking. Δ columns compare against the live ranking right now.
| Rank | Model | Provider | SupraScore | Δ Score | Δ Rank |
|---|---|---|---|---|---|
| simulated | — | — |
Search accounts by name or email. Grant partner or enterprise+ with custom limits — they mint their own keys from the API tab. As primary admin, you can also promote other admins.
Showing all accounts with elevated privileges. Type to search the full user table.
Admin role
Only the primary admin can promote or demote other admins.
This account is the primary admin and cannot be demoted.
Granted tier
API keys view-only — the user mints their own keys from their API tab
| Name | Prefix | Tier | Created | Last used | Status | |
|---|---|---|---|---|---|---|
|
active revoked |
Monthly usage
| Month | API calls | Active keys |
|---|---|---|