All Grim.Cards case-study editions

Grim.Cards Simulation Data Case Study

Edition: 2026-09-01 | Data Snapshot: 2026-05-14 – 2026-09-01

Dataset Version 2.4 | Published: 2026-09-01 | License: CC BY 4.0


Permanent URL: https://grim.cards/case-study/2026-09-01 Cite as: Grim.Cards (2026-09-01). Grim.Cards Simulation Data Case Study, Edition 2026-09-01 (Dataset v2.4). https://grim.cards/case-study/2026-09-01. CC BY 4.0.


Executive Summary

Across 14,559 simulated games played by 790 player-submitted decks from 314 distinct users between 14 May and 1 September 2026, player-submitted decks won 42.2% of all games (6,147 wins, 8,307 losses, 105 draws). The dataset covers two formats — Commander (n = 713 decks, 13,292 games, 43.0% win rate) and Standard (n = 77 decks, 1,267 games, 34.5% win rate) — analyzed separately throughout this report.

A sub-50% headline is the structural expectation, not a verdict on deck quality: each deck faces the full gauntlet of meta opponents in sequence, so the field is structurally favored over any single challenger. What the data actually reveals is far more interesting than the aggregate: the specific opponent a deck faces swings outcomes by as much as 42 percentage points within an otherwise identical challenger pool, matchup results span the full range from 0% to 100% at the deck level, and a meaningful minority of decks consistently outperform or underperform their expected baseline by double-digit margins. Those are the stories this report investigates.

Five headline findings:

  1. The gauntlet opponent matters enormously. In Commander, the spread between the most and least favorable gauntlet opponent is 21.4 percentage points across n = 659 decks each. In Standard, that spread reaches 42.1 points across n = 69 decks each.
  2. Meren of Clan Nel Toth is the most-tested Commander with 21 distinct decks — the only commander clearing the minimum publication threshold — posting a 34.3% win rate over 510 games.
  3. In Commander, creatures and Red/Green cards post the highest pooled board impact across 570 qualifying cards (≥10 decks, ≥25 observations each).
  4. In Standard, Lightning Bolt leads on board impact (+5.1 board-quality points per appearance, n = 16 decks, 165 observations) but the Standard sample is substantially smaller and results should be read accordingly.
  5. Retesting a deck does not reliably improve outcomes: among 143 Commander decks retested, 62 declined, 60 improved, and 21 were flat (average change: −0.7 points). Among 12 Standard retests, 6 declined, 4 improved, 2 were flat (average change: −1.1 points). All figures are correlational observations within a simulated environment.

Methodology & Provenance

Source and scope

All data derives from Grim.Cards production simulations: AI-versus-AI games run on a custom build of the open-source Forge engine. Each player-submitted deck is placed against a fixed gauntlet of meta reference decks; every game is fully resolved by the engine with no human intervention. The automated Crucible reference corpus (system user __grinder__) and system sample decks (is_sample = true) are excluded from every figure in this report. Only real, human-submitted decks count.

Win rate definition

Win rate = wins ÷ total games played, with draws included in the denominator. This definition is applied consistently throughout. A draw counts as neither a win nor a loss but is never excluded from the denominator. Every win-rate figure in this report uses this formula unless explicitly labeled otherwise.

Format split

Commander and Standard results are reported separately throughout. The two formats face different gauntlet opponents, have different structural dynamics, and must not be pooled. Where a figure applies to both formats, it is labeled "all formats."

Cohort minimum

No cohort with fewer than 10 unique decks is reported. Smaller cohorts are suppressed or merged. This threshold applies per matchup, per format, per month, per construction band, per color, and per card. Any dataset entry with winRate: null in the source data reflects this suppression and is not reported.

Correlation framing

All construction, category, color, and card figures are descriptive correlations within the Grim.Cards dataset. None establishes causation. Category membership (tutor, sacrifice, discard, reanimate) is assigned by keyword heuristic and may mislabel edge cases; these figures carry an explicit [HEURISTIC] label. Card board-impact figures are a board-state proxy (measured board-quality delta around turns a card was seen), not damage dealt or kills; they carry a [BOARD-STATE PROXY] label. Counterfactual decision-impact figures, where available, represent play-quality comparison against the engine's own next-best alternative and are labeled [COUNTERFACTUAL PROXY]. No counterfactual impact data met the publication threshold for this edition; that section is omitted.

Provenance statement

  • Publisher: Grim.Cards
  • Data window: 2026-05-14 through 2026-09-01
  • Generated: 2026-09-01T08:48:33.023Z
  • Dataset version: 2.4
  • Unit of observation: deck-versus-meta-opponent game aggregates
  • License: Creative Commons Attribution 4.0 International (CC BY 4.0) — https://creativecommons.org/licenses/by/4.0/

Section 1: Dataset Overview

The September 2026 edition of this case study is the largest Grim.Cards has published. Since May 2026 the platform has accumulated a genuine body of simulation data — not a test corpus, but a record of real players testing real decks against real gauntlet opponents.

Metric All Formats Commander Standard
Distinct users 314 278 53
Distinct decks 790 713 77
Completed simulations 1,080 985 95
Total games 14,559 13,292 1,267
Wins 6,147 5,710 437
Losses 8,307 7,477 830
Draws 105 105 0
Win rate 42.2% 43.0% 34.5%
Date range 2026-05-14 to 2026-09-01 same same

The Commander cohort is 9.3 times larger than Standard by deck count (713 vs. 77), which is why many per-card and per-commander findings are Commander-dominated. Standard figures carry a smaller sample caveat throughout.


Section 2: The Gauntlet — Matchup Results

This is the core of the report. Nothing in this dataset is more directly measured, more reproducible, or more revealing than the per-matchup win-rate figures. Every number here is a direct game count from the simulation engine; no interpolation is involved.

Commander Gauntlet

The Commander gauntlet consists of five reference decks. Each of the 659 decks with matchup data faced all five opponents (some decks faced additional opponents; the five below clear the publication threshold at n = 659).

Gauntlet Opponent Decks (n) Wins Losses Draws Games Win Rate
Breya Artifact Combo 659 1,383 1,232 19 2,634 52.5%
Derevi Bant Control 659 1,320 1,308 6 2,634 50.1%
Aesi Landfall 659 1,186 1,390 66 2,642 44.9%
Atraxa Superfriends 659 947 1,690 2 2,639 35.9%
Edgar Markov Vampires 659 820 1,811 9 2,640 31.1%

The 21.4-point spread between the easiest opponent (Breya Artifact Combo, 52.5%) and the hardest (Edgar Markov Vampires, 31.1%) is the single most important figure in this section. Both figures are drawn from identical challenger pools — the same 659 decks face every opponent. The opponent alone accounts for that gap. Breya is the only gauntlet opponent against which the average player deck posts a winning record; Edgar Markov Vampires and Atraxa Superfriends both sit well below the break-even mark.

Draws are rare but not absent in Commander: Aesi Landfall produced 66 drawn games (2.5% of its matches), more than any other opponent — a game-length artifact of the landfall engine's slower resource accumulation.

Standard Gauntlet

The Standard gauntlet also consists of five reference decks. All 69 Standard decks with matchup data faced each opponent.

Gauntlet Opponent Decks (n) Wins Losses Draws Games Win Rate
Temur Harmonizer Combo 69 140 112 0 252 55.6%
Jeskai Control 69 118 134 0 252 46.8%
Dimir Midrange 69 72 180 0 252 28.6%
Azorius Tempo 69 71 181 0 252 28.2%
Mono Red Aggro 69 34 218 0 252 13.5%

The Standard spread is more extreme: 42.1 percentage points separate Temur Harmonizer Combo (55.6%) from Mono Red Aggro (13.5%). Against Temur, the average player deck wins more than half its games; against Mono Red Aggro, it wins fewer than one in seven. Standard has no draws recorded.

The Mono Red Aggro figure is striking. A 13.5% win rate against a single-color aggro deck indicates that the average Standard deck in this sample is substantially outpaced in the early game when the opponent curves low and fast. This is a measured outcome of the simulation engine, not a strategic claim.


Section 3: Top Commanders — Popularity vs. Performance

Most-tested commanders

Only commanders represented by 10 or more distinct decks in the simulation data are reported with win rates. Among the 20 most-submitted commanders, only one clears that threshold.

Meren of Clan Nel Toth — the only commander meeting the minimum cohort requirement — appears in 21 distinct decks and has accumulated 510 simulated games, posting a 34.3% win rate. That places Meren below the Commander format average of 43.0% (n = 713 decks, 13,292 games). Meren's position below average is a measured outcome in this dataset; given the composition of the gauntlet opponents, it may reflect the specific matchup profile of graveyard-recursion strategies rather than anything generalizable beyond this environment.

All other commanders in the dataset — including Krenko, Mob Boss (9 decks); Ureni of the Unwritten (9 decks); Kenrith, the Returned King (8 decks); Kaalia of the Vast (8 decks); and Vivi Ornitier (8 decks) — fall below the 10-deck minimum and are not reported with win rates in this edition. Their game and win counts are noted in the source dataset but are not surfaced here to prevent misleading comparisons from small samples.

As the Grim.Cards dataset grows across future editions, more commanders will clear this threshold and commander-level win-rate comparisons will become more statistically credible. This edition can only report one.


Section 4: Top Cards — Most Played and How Their Decks Fared

Card-level win rates in this section reflect the win rate of decks containing each card, not any card-specific causal claim. A card appearing in winning decks does not make those decks win; both the card's presence and the deck's outcomes may be downstream of the same underlying deck-quality factors. Basic lands are excluded throughout.

Commander: most-played cards

The following 20 cards appear in the most distinct Commander decks (minimum 10 decks for win-rate reporting):

Card Decks (n) Games Win Rate
Sol Ring 576 11,616 43.5%
Command Tower 470 9,265 42.3%
Arcane Signet 468 9,587 43.1%
Exotic Orchard 250 5,085 42.5%
Lightning Greaves 194 3,753 46.2%
Swiftfoot Boots 194 3,642 43.6%
Reliquary Tower 176 3,226 42.9%
Path of Ancestry 173 3,475 47.7%
Swords to Plowshares 159 3,163 44.5%
Evolving Wilds 157 3,203 43.0%
Cultivate 129 2,639 48.9%
Path to Exile 124 2,307 47.8%
Fellwar Stone 120 2,560 41.4%
Rogue's Passage 116 2,138 45.1%
Counterspell 114 2,430 40.1%
Bojuka Bog 112 2,436 40.1%
Terramorphic Expanse 108 2,187 41.4%
Skullclamp 100 2,213 40.0%
Blasphemous Act 94 1,834 47.4%
Dark Ritual 94 1,957 37.2%

Sol Ring is the most ubiquitous card in the dataset, appearing in 576 of 713 Commander decks (80.8%). The top staples — Sol Ring, Command Tower, Arcane Signet — post win rates within a few points of the format average (43.0%), as expected from near-universal inclusions: they correlate with all decks, strong and weak. The more instructive comparisons are among the less universal cards. Cultivate appears in 129 decks and its containing decks win at 48.9%; Path of Ancestry (173 decks) and Path to Exile (124 decks) both sit above 47%. At the other end, Dark Ritual (94 decks) is the only staple in this list whose containing decks win below 40% — at 37.2%, its decks perform 5.8 points below the format average. These are descriptive correlations, not causal claims.

Standard: most-played cards (meeting publication threshold)

In Standard, 11 cards in 10 or more decks (minimum threshold) report win rates:

Card Decks (n) Games Win Rate
Lightning Bolt 16 240 42.1%
Inspiring Vantage 15 225 44.4%
Shard Volley 14 210 45.2%
Skullcrack 14 210 45.2%
Eidolon of the Great Revel 14 210 45.2%
Monastery Swiftspear 14 210 45.2%
Chain Lightning 14 210 45.2%
Lava Spike 14 210 45.2%
Rift Bolt 14 210 45.2%
Goblin Guide 14 210 45.2%
Searing Blaze 14 210 45.2%

The identical win rate and game count across the 14-deck cards reflects that these cards appear in a heavily overlapping set of decks — likely a cluster of burn-archetype submissions that share core cards. The containing-deck win rate of 45.2% is above the Standard format average of 34.5% (n = 77 decks), meaning these burn-archetype decks outperform the Standard field on average. Inspiring Vantage (15 decks, 44.4%) and Lightning Bolt (16 decks, 42.1%) slot in slightly differently, suggesting broader inclusion across non-identical decks.


Section 5: Card Board-Impact Performance

[BOARD-STATE PROXY] The following figures are pooled board-quality deltas measured around turns when each card was observed in play. Positive means the board state improved in the turns surrounding that card's appearance; negative means it declined. This is a board-state proxy, not damage dealt, kills, or a causal win claim. Rankings are weighted by breadth: the more distinct decks a card appears in and the more observations it has accumulated, the more reliable its figure. Only cards in 10 or more distinct decks with 25 or more recorded observations are ranked. Type and color aggregates pool very different cards, and multicolor cards count toward each of their constituent colors.

Commander: top board-impact cards

570 Commander cards qualified for ranking (≥10 decks, ≥25 observations each).

Top 15 by board impact (Commander):

Card Decks (n) Observations Avg Board Delta
Herald of Secret Streams 12 50 +72.54
Ancient Silver Dragon 12 67 +32.43
Germination Practicum 10 44 +31.32
Ureni of the Unwritten 12 283 +29.60
Dragonlord Atarka 10 47 +28.17
Terror of Mount Velus 12 114 +22.93
Utvara Hellkite 11 60 +22.40
Scourge of Valkas 11 48 +22.19
Drakuseth, Maw of Flames 13 87 +21.85
Craterhoof Behemoth 28 89 +20.26
Lyra Dawnbringer 16 91 +20.13
Combustible Gearhulk 12 49 +19.63
Goblin Warchief 10 47 +19.26
Ghalta, Primal Hunger 14 95 +19.08
Cathars' Crusade 14 43 +18.00

Herald of Secret Streams' figure of +72.54 per appearance is exceptional — nearly twice the next-highest card. However, it appears in only 12 decks with 50 observations, so the sample is narrower than Craterhoof Behemoth's (28 decks, 89 observations), which posts a still-strong +20.26. The top 15 is dominated by large creatures and finishers, consistent with the type-level aggregate finding that creatures carry the highest pooled board impact in Commander (+3.13 per appearance across 111,153 observations in 83+ decks).

Bottom 15 by board impact (Commander):

Card Decks (n) Observations Avg Board Delta
Circle of Dreams Druid 11 58 −25.43
Aura Shards 10 34 −25.26
Hedron Archive 16 62 −14.08
Reconnaissance Mission 12 52 −14.08
Shared Animosity 10 25 −13.52
Day of Judgment 13 38 −13.37
Mirari's Wake 13 35 −13.26
Dictate of Erebos 31 151 −13.22
Panharmonicon 17 59 −12.36
Bastion Protector 10 114 −11.57
Mystic Forge 11 42 −11.43
Maddening Cacophony 12 35 −11.26
Aggravated Assault 13 72 −10.57
Saw in Half 10 33 −10.12
Bitterblossom 13 41 −10.07

Dictate of Erebos is notable for its breadth: 31 decks, 151 observations, and a consistent −13.22. A board-state decline around appearances of a sacrifice-triggered effect may reflect that sacrificing creatures to trigger the effect costs board presence at the moment it is observed — a measurement artifact of when the observation window is taken, not necessarily a statement about the card's strategic value. Negative deltas across the bottom 15 should be read as descriptive observations, not strategic conclusions.

Commander: board impact by card type

All Commander decks, pooled (n varies by type):

Type Cards Observations Decks ≥1 Avg Board Delta
Creatures 4,989 111,153 83 +3.13
Instants 957 13,605 153 +2.75
Sorceries 929 11,286 105 +0.21
Planeswalkers 141 2,311 16 −0.59
Nonbasic Lands 62 333 17 −1.44
Enchantments 1,011 14,781 56 −1.69
Artifacts 916 29,523 537 −1.88

Creatures and instants both post positive pooled board impact. Artifacts are the most-observed type by a wide margin (29,523 observations, 537 decks contributing at least one qualifying card), which makes the artifact figure the most stable, and it is negative at −1.88. This is a descriptive aggregate only: type buckets pool cards with vastly different functions, from Sol Ring to Panharmonicon. No type-level conclusion translates automatically to any individual card.

Commander: board impact by color identity

Color Cards Observations Decks ≥1 Avg Board Delta
Red 1,810 35,074 76 +3.56
Green 2,307 47,735 105 +3.37
Blue 1,980 38,780 68 +2.91
White 1,987 43,812 153 +2.01
Black 2,092 45,805 79 +0.66
Colorless 934 30,361 537 −1.24

Red and Green lead the color-identity board-impact aggregate; Colorless cards — dominated in observation count by mana rocks (Sol Ring, Arcane Signet, Fellwar Stone) — post the only negative aggregate at −1.24. Multicolor cards count toward each of their colors, so these figures are not independent. The colorless aggregate likely reflects the observation timing issue noted above: mana rocks are observed most often on early turns when the board state is still developing, not at the moments of greatest board advantage.

Standard: board-impact cards

9 Standard cards qualified for ranking (≥10 decks, ≥25 observations). The Standard qualifying pool is narrow — a reflection of the smaller Standard cohort (77 decks, 1,267 games) — and all figures here carry a smaller-sample caveat.

Top qualifiers by board impact (Standard):

Card Decks (n) Observations Avg Board Delta
Lightning Bolt 16 165 +5.10
Goblin Guide 14 283 +4.93
Eidolon of the Great Revel 14 249 +4.41
Monastery Swiftspear 14 238 +2.97
Rift Bolt 14 194 +2.19
Chain Lightning 14 116 +1.99
Lava Spike 14 87 +1.05
Sol Ring 10 45 −1.58
Shard Volley 14 56 −4.50

Lightning Bolt leads the Standard field at +5.10 with the widest breadth (16 decks, 165 observations). The top seven are all burn/aggro cards; Sol Ring and Shard Volley are the only negative figures. Shard Volley's −4.50 delta — the worst in Standard — may partially reflect a land-sacrifice cost being recorded as a board-state decline at the moment of play.

Standard: board impact by card type:

Type Cards Observations Decks ≥1 Avg Board Delta
Creatures 556 6,744 14 +1.69
Sorceries 158 1,436 14 +0.60
Instants 149 1,443 16 −0.48
Artifacts 132 1,008 10 −3.46

Standard: board impact by color (meeting publication threshold):

Color Cards Observations Decks ≥1 Avg Board Delta
Red 173 2,423 16 +2.13
Colorless 131 1,045 10 −1.34

Only Red and Colorless meet the 10-deck minimum in Standard's color aggregate. White, Blue, Green, and Black Standard color groups are suppressed for this edition.


Section 6: Win-Rate Distribution

How are deck-level win rates distributed across the sample? The distribution shape tells a different story than the aggregate.

Commander win-rate distribution (n = 662 decks)

Win-Rate Band Decks
0–10% 22
10–20% 71
20–30% 62
30–40% 190
40–50% 85
50–60% 115
60–70% 66
70–80% 39
80–90% 8
90–100% 4
  • Mean: 43.6% | Median: 40% | Min: 0% | Max: 100%

The 30–40% band is the modal outcome: 190 of 662 Commander decks (28.7%) land there. The distribution is left-skewed — a tail of high-performing decks pulls the mean (43.6%) above the median (40%). The 50–60% band is the second-largest (115 decks, 17.4%), creating a visible secondary mode. Together, these two peaks suggest the Commander sample is not a single homogeneous population: there are meaningful clusters of stronger and weaker performers, and the average obscures that structure.

Standard win-rate distribution (n = 70 decks)

Win-Rate Band Decks
0–10% 8
10–20% 12
20–30% 8
30–40% 23
40–50% 4
50–60% 10
60–70% 3
70–80% 2
80–90% 0
90–100% 0
  • Mean: 34.3% | Median: 33.3% | Min: 0% | Max: 80%

Standard's distribution is similarly left-skewed, but the tail cuts off sharply at 80% — no Standard deck posted a win rate above 80% in this dataset. The 30–40% band again dominates (23 of 70 decks, 32.9%). The 0–10% band at 8 decks exceeds the 40–50% band at 4 decks — more decks posted near-zero win rates than near-coin-flip performance.


Section 7: Deck Iteration and Retests

Among players who return to test an updated version of a deck, does win rate improve? The retest data offers a direct answer — and it is not the one most players probably expect.

Commander retests (n = 143 retested decks)

Of 713 Commander decks with simulations, 143 were retested (i.e., their owner submitted an updated version and ran a new simulation). The change in win rate from first to most recent simulation:

  • Improved: 60 decks (42.0%)
  • Declined: 62 decks (43.4%)
  • Flat: 21 decks (14.7%)
  • Average win-rate change: −0.7 percentage points

On aggregate, Commander retests went slightly backward: the average retest deck lost 0.7 points of win rate from its first to its latest simulation. The improved-vs-declined split is nearly even (60 vs. 62), and more decks declined than improved. This finding is descriptive: different decks are tested at different times against the same fixed gauntlet, and the retest cohort is self-selected (players who chose to update and retest). No causal inference about deck changes is warranted.

Standard retests (n = 12 retested decks)

  • Improved: 4 decks (33.3%)
  • Declined: 6 decks (50.0%)
  • Flat: 2 decks (16.7%)
  • Average win-rate change: −1.1 percentage points

The Standard retest cohort is small (n = 12, meeting but not exceeding the minimum publication threshold) and results should be read cautiously. The directional pattern mirrors Commander: more declines than improvements, and a negative average delta (−1.1 points).

The retest data, taken together, does not support a conclusion that retesting reliably improves simulated outcomes in this dataset. It equally does not support the reverse conclusion — with roughly equal numbers of improvers and decliners in Commander, the outcome is closer to random than to systematic. The most reasonable reading is that changing a deck is not a reliable predictor of simulated win-rate change within this engine and gauntlet configuration.


Section 8: Monthly Trends

Monthly win rates are reported only for format-months with at least 10 decks. May 2026 Commander data (74 simulations, 51 decks) has no win-rate data available for the reported period and is shown as unavailable.

Commander monthly trend

Month Simulations Decks Win Rate
2026-05 74 51 N/A
2026-06 143 105 44.0%
2026-07 339 257 43.3%
2026-08 425 300 42.4%

Commander simulation volume has grown substantially: from 143 completed sims in June to 425 in August, a near-tripling in two months. The win rate has been stable within a narrow range — 42.4% to 44.0% across three reported months — which suggests the gauntlet difficulty has remained consistent and the composition of submitted decks has not shifted dramatically. Month-to-month cohorts are entirely different decks (players test different decks each month), so this tracks who submitted, not any individual deck trajectory.

Standard monthly trend

Month Simulations Decks Win Rate
2026-06 24 20 37.5%
2026-07 31 26 33.3%
2026-08 29 22 31.7%

Standard shows a modest downward trend — from 37.5% in June to 31.7% in August — but with 20–26 decks per month, the month-level samples are small and this should not be over-interpreted. Volume has been steady (24–31 sims per month).


Section 9: Construction Correlations

The following figures associate deck-construction features with win rate within the Grim.Cards dataset. All are descriptive correlations only. Colors overlap within multi-color decks and are not independent effects.

Color-count correlation

Commander (by number of colors in color identity):

Colors Decks (n) Win Rate
1 128 50.9%
2 245 40.5%
3 207 42.8%
4 10 48.7%
5 64 42.2%

Mono-color Commander decks post the highest observed win rate (50.9%, n = 128), 7.9 points above the format average. This is a correlational observation — mono-color decks have more consistent mana bases and may also be concentrated in archetypes that happen to perform well against the specific gauntlet opponents in this dataset. Four-color decks (n = 10, exactly meeting the publication threshold) at 48.7% should be read cautiously given the small sample.

Standard (by number of colors):

Colors Decks (n) Win Rate
1 14 38.8%
2 43 35.5%
3 11 20.6%

Standard three-color decks (n = 11, at the minimum threshold) post a 20.6% win rate — the lowest of any color-count group in either format. Small sample; interpret cautiously.

Land-ratio correlation

Commander (by approximate land percentage of deck):

Land % (Approx.) Decks (n) Win Rate
~30% 92 39.2%
~35% 405 44.9%
~40% 149 44.5%

The 35% land band is the largest by far (405 decks) and posts a 44.9% win rate, compared to 39.2% for the lower-land-count band (92 decks). Decks running approximately 30% lands post win rates 5.7 points below those at 35%. This is correlational: land-light decks may also differ from higher-land decks on many other dimensions.

Standard:

Land % (Approx.) Decks (n) Win Rate
~35% 20 28.4%
~40% 34 41.6%

The 40% land band in Standard (n = 34) posts a 13.2-point higher win rate than the 35% band (n = 20). Again, this is descriptive and correlational.

Color-presence win rates

Commander (by color identity presence):

Color Decks (n) Wins Games Win Rate
Green (G) 313 2,783 6,260 44.5%
White (W) 317 2,789 6,468 43.1%
Red (R) 308 2,599 6,052 42.9%
Blue (U) 311 2,572 6,291 40.9%
Black (B) 350 2,870 7,106 40.4%

Green-containing decks lead by 4.1 points over Black-containing decks. Black appears in the most Commander decks (350) and posts the lowest win rate among the five colors. Colors overlap extensively within multi-color decks; these are not independent effects.

Standard (by color identity presence):

Color Decks (n) Wins Games Win Rate
White (W) 35 223 615 36.3%
Red (R) 29 164 454 36.1%
Green (G) 22 156 435 35.9%
Black (B) 30 159 547 29.1%
Blue (U) 21 118 424 27.8%

White, Red, and Green cluster within 0.4 points of each other at the top in Standard. Blue and Black both sit below 30%. Blue in particular (n = 21, 27.8%) stands out as the weakest color presence in Standard — though with only 21 decks, the figure is near the minimum publication threshold.

Win-bracket construction profiles

Decks grouped into win-rate brackets are compared on average construction features. These are descriptive comparisons only.

Commander win brackets:

Bracket Decks (n) Avg Win Rate Avg Land % Avg Creature % Avg Spell % Avg Art/Ench % Avg Mana Value
High (>55%) 174 67.9% 35.8% 29.3% 16.2% 18.0% 3.32
Mid (40–55%) 238 45.5% 35.9% 27.5% 17.9% 18.0% 3.19
Low (<40%) 250 24.9% 35.1% 26.4% 19.7% 17.9% 3.05

High-win-rate Commander decks run 0.7% more lands than low-win-rate decks (35.8% vs. 35.1%) and carry an average mana value 0.27 higher (3.32 vs. 3.05). High-bracket decks average slightly more creatures (29.3% vs. 26.4%) and fewer spells (16.2% vs. 19.7%) than low-bracket decks. These are the construction patterns associated with high-win-rate Commander decks in this dataset; they are correlational observations across submitted decks, not build recommendations.

Standard win brackets:

The Standard high-win-rate bracket (>55%) has fewer than 10 decks and is suppressed. Only the mid and low brackets are reported.

Bracket Decks (n) Avg Win Rate Avg Land % Avg Creature % Avg Spell % Avg Art/Ench % Avg Mana Value
Mid (40–55%) 24 45.7% 37.9% 25.5% 30.9% 5.8% 2.15
Low (<40%) 39 21.2% 33.8% 29.5% 21.1% 14.6% 2.65

Mid-bracket Standard decks carry a 4.1% higher land percentage and a 0.50 lower average mana value than the low bracket — a substantially lower curve. This is a descriptive correlation within the dataset.

Card-type mix (format-wide)

Commander (n = 662 decks, split across both instant/sorcery types available):

  • Lands: 35.6% | Creatures: 27.5% | Instants+Sorceries: 18.1% (Instants: 9.9%, Sorceries: 8.1%) | Artifacts: 10.6% | Enchantments: 7.4% | Planeswalkers: 0.8%

Standard (n = 70 decks):

  • Lands: 35.5% | Creatures: 28.4% | Instants+Sorceries: 24.7% (Instants: 13.9%, Sorceries: 10.6%) | Artifacts: 5.2% | Enchantments: 5.7% | Planeswalkers: 0.6%

Standard decks allocate roughly 6.6% more of their slots to instants and sorceries than Commander decks (24.7% vs. 18.1%) and proportionally less to artifacts (5.2% vs. 10.6%) and enchantments (5.7% vs. 7.4%). Instant/sorcery figures for both formats are drawn from the per-deck embedding split (splitDecks = 662 Commander, 70 Standard) where both instant and sorcery keys are populated; the combined instantSorcery figure is the reliable total.


Section 10: Functional Card Categories

[HEURISTIC] Category membership (sacrifice outlets, search/tutor effects, discard effects, reanimation effects) is assigned by a keyword heuristic applied to card oracle text and may mislabel edge cases. All win-rate figures here are containing-deck win rates (win rate of decks running at least one card in that category), not per-card figures. Descriptive correlations only.

Commander functional categories

Category With-decks (n) With win rate Without-decks (n) Without win rate Delta
Search / tutor effects 629 43.9% 33 38.1% +5.8 pts
Sacrifice outlets 657 43.6% 5 suppressed N/A
Discard effects 579 43.0% 83 47.8% −4.8 pts
Reanimation effects 402 42.2% 260 45.8% −3.6 pts

Sacrifice outlets are present in 657 of 662 Commander decks; the without-group (5 decks) falls below the minimum publication threshold and no contrast is reported. The tutor/search category shows the most notable positive association: decks with tutors (n = 629) win at 43.9%, versus 38.1% for the 33 decks without — a 5.8-point gap. Discard and reanimation show negative associations (decks with those effects have lower win rates than those without), though selection effects almost certainly dominate: decks built around discard or graveyard synergies may also be more likely to be submitted by players exploring novel archetypes, affecting the sample composition.

Standard functional categories

Category With-decks (n) With win rate Without-decks (n) Without win rate Delta
Sacrifice outlets 64 33.9% 6 suppressed N/A
Discard effects 41 31.7% 29 37.9% −6.2 pts
Search / tutor effects 32 29.6% 38 38.2% −8.6 pts
Reanimation effects 21 25.7% 49 37.9% −12.2 pts

Standard reanimation stands out: decks with reanimation effects (n = 21) post a 25.7% win rate — 12.2 points below the 49 decks without (37.9%). Tutors show an 8.6-point gap in the same direction. Both the directions and magnitudes may reflect archetype clustering in the Standard sample rather than anything about the cards themselves.


Section 11: Matchup Structure — Deck vs. Opponent

This section examines a structural question: how much do differences between decks versus differences between gauntlet opponents explain the variation in simulated win rates? All figures here are descriptive decompositions of simulated results, not causal claims or significance tests.

Commander: variance decomposition

Across a matrix of 133 Commander decks facing 17 distinct gauntlet opponents (635 qualifying cells, minimum 5 games per cell):

  • Deck main effect: 48.4% of cell win-rate variance
  • Opponent main effect: 12.7% of cell win-rate variance
  • Residual (interaction + noise): 38.9%

In Commander, the deck itself explains a substantially larger share of win-rate variation than the specific opponent it faces (48.4% vs. 12.7%). The residual is large — nearly 39% — reflecting genuine deck-opponent interaction: the same deck may perform very differently against different opponents, but those interactions are not fully captured by either main effect alone.

Standard: variance decomposition

Across a matrix of 10 Standard decks facing 5 distinct gauntlet opponents (50 qualifying cells):

  • Deck main effect: 26.3% of cell win-rate variance
  • Opponent main effect: 36.2% of cell win-rate variance
  • Residual (interaction + noise): 37.4%

Standard tells the opposite story from Commander: the opponent main effect (36.2%) exceeds the deck main effect (26.3%). The specific gauntlet opponent explains more win-rate variation than the deck's own profile in this format and sample. The Standard sample is substantially smaller (10 decks), which limits the robustness of this decomposition, but the directional difference from Commander is notable.

Matchup inversion rate

Commander: 39% of comparable deck pairs (48,066 pairs across 10 opponent pairings, n = 122 decks) reverse their preference order between two gauntlet opponents. That is: Deck A beats Opponent X more often than Deck B, but Deck B beats Opponent Y more often than Deck A. A 39% inversion rate means the gauntlet is not fully rank-ordering decks — the specific opponent substantially reshuffles outcomes.

Standard: 29.7% of comparable deck pairs (340 pairs across 10 opponent pairings, n = 10 decks) show a reversed preference order. The Standard inversion rate is lower than Commander's, though the Standard sample is much smaller.

Power Score vs. win-rate correlation (Spearman ρ, per opponent)

Note on Power Score: Power Score is an internal Grim.Cards composite indicator, not an objective win-rate measurement. The Spearman correlations below describe how closely the Power Score rank-orders decks by simulated win rate against each opponent. They are presented here as structural measurements, not as endorsements of Power Score as a universal metric.

Commander (all 5 gauntlet opponents, n ≈ 124–125 decks each):

Opponent Spearman ρ Decks (n)
Edgar Markov Vampires 0.75 125
Atraxa Superfriends 0.74 125
Breya Artifact Combo 0.71 124
Aesi Landfall 0.70 125
Derevi Bant Control 0.69 124

All five Commander opponents show strong positive rank correlations between Power Score and win rate (ρ = 0.69–0.75). The ordering is consistent across all five opponents, with Edgar Markov Vampires producing the tightest rank-ordering (ρ = 0.75) and Derevi Bant Control the weakest (ρ = 0.69) — though the spread across opponents is narrow. These are Spearman correlations over decks clearing the per-cell game floor; they describe the rank alignment within this dataset, not a causal relationship.

Standard (all 5 gauntlet opponents, n = 10 decks each):

Opponent Spearman ρ Decks (n)
Temur Harmonizer Combo 0.84 10
Dimir Midrange 0.75 10
Azorius Tempo 0.53 10
Jeskai Control 0.48 10
Mono Red Aggro 0.41 10

Standard shows greater variation across opponents (ρ = 0.41–0.84). Temur Harmonizer Combo (ρ = 0.84) produces the strongest rank alignment with Power Score; Mono Red Aggro (ρ = 0.41) the weakest. The wide range suggests the Power Score rank-orders Standard decks well against some opponents and less well against others — Mono Red Aggro's low correlation is consistent with an aggro deck that punishes specific low-curve vulnerabilities regardless of overall deck power. Standard n = 10 throughout; all figures are at the minimum publication threshold and should be interpreted accordingly.

Power Score × opponent win-rate grid (Commander only)

The Commander dataset supports a three-band grid breakdown (low/mid/high Power Score × each gauntlet opponent). Standard's sample (n = 10) is insufficient for a reliable grid and is omitted. Power Score band definitions: Low = 0–32.3 (n = 44 decks), Mid = 32.3–47.2 (n = 44 decks), High = 47.8–89.7 (n = 45 decks).

Opponent Low PS Band (n ≈ 40–41 decks) Mid PS Band (n = 43 decks) High PS Band (n = 41–42 decks)
Aesi Landfall 29.1% (n=40, 333 games) 43.9% (n=43, 387 games) 62.1% (n=42, 319 games)
Atraxa Superfriends 15.1% (n=41, 338 games) 33.6% (n=43, 387 games) 54.5% (n=41, 312 games)
Breya Artifact Combo 30.0% (n=40, 330 games) 49.9% (n=43, 387 games) 67.6% (n=41, 312 games)
Derevi Bant Control 32.7% (n=40, 330 games) 45.7% (n=43, 387 games) 65.4% (n=41, 312 games)
Edgar Markov Vampires 14.2% (n=40, 330 games) 26.8% (n=43, 388 games) 49.1% (n=42, 320 games)

The gradient from low to high is consistent across all five opponents: higher Power Score bands win at substantially higher rates in every matchup. Edgar Markov Vampires is the most punishing for low-PS decks (14.2% win rate for the low band) — the same opponent that posts the lowest aggregate win rate overall. High-PS decks against Edgar Markov Vampires (49.1%) still do not reach a majority win rate, making Edgar the most challenging matchup even for the strongest Commander decks in this dataset. Breya Artifact Combo at 67.6% for high-PS decks is the most favorable cell in the grid.


Section 12: Game Tempo

Average game length is a rough proxy for the engine's resolution pace against the gauntlet.

Commander (n = 661 decks):

  • Average turns: 10.3
  • Average shortest game: 8.4 turns
  • Average longest game: 12.2 turns

Standard (n = 70 decks):

  • Average turns: 9.6
  • Average shortest game: 7.4 turns
  • Average longest game: 11.9 turns

Standard games resolve slightly faster on average (9.6 vs. 10.3 turns). The shortest games in Standard average 7.4 turns — consistent with the presence of aggressive archetypes in the Standard gauntlet (Mono Red Aggro resolved many games quickly in its 218-loss column for player decks). These figures describe the engine's tempo behavior; human play pacing would differ.


Section 13: Per-Deck Appendix (Anonymized)

Privacy and anonymization statement: All deck pseudonyms (C-001 through C-132, S-001 through S-010) are randomly reassigned every edition and carry no link to any account, deck name, or identity. Power Scores are rounded to the nearest 5. Win rates and matchup spreads are rounded to whole points. Game counts are banded (not exact). Only derived statistics are shown — never deck names, card lists, owner data, or raw identifiers. Pseudonyms from this edition cannot be linked to pseudonyms from any prior or future edition.

Commander: per-deck appendix summary

132 Commander decks qualify for the appendix (minimum 20 games played; n = 662 decks in the distribution, of which 132 meet the game-count floor for appendix publication).

  • Win rate range: 0% – 85%
  • Median win rate: 40%
  • Median matchup spread (std dev of per-opponent win rates, ≥3 opponents): 16 percentage points (n = 124 decks with spread data)

The 16-point median spread means the typical Commander deck varies by about 16 percentage points in win rate between its best and worst gauntlet opponent — a meaningful spread that separates matchup specialists (high spread) from format generalists (low spread). The individual deck table follows (132 rows):

ID PS (coarsened) Games Band Win Rate Opponents Faced Spread
C-001 60 20–49 50% 5 18
C-002 55 20–49 63% 5 16
C-003 80 20–49 77% 5 8
C-004 45 20–49 42% 5 18
C-005 20 20–49 35% 1 —
C-006 35 20–49 28% 5 17
C-007 60 20–49 57% 5 23
C-008 30 20–49 33% 5 16
C-009 40 20–49 40% 5 17
C-010 35 20–49 37% 5 19
C-011 55 20–49 56% 5 12
C-012 65 20–49 70% 5 27
C-013 45 20–49 47% 5 19
C-014 15 20–49 17% 5 0
C-015 35 50–99 38% 5 11
C-016 25 20–49 27% 5 13
C-017 75 20–49 73% 5 13
C-018 45 20–49 47% 5 24
C-019 90 20–49 85% 1 —
C-020 40 20–49 40% 5 13
C-021 15 20–49 20% 5 27
C-022 55 20–49 53% 5 19
C-023 20 20–49 20% 5 12
C-024 30 50–99 33% 5 9
C-025 40 20–49 40% 5 25
C-026 35 20–49 37% 5 19
C-027 55 20–49 53% 5 29
C-028 70 20–49 70% 5 19
C-029 55 20–49 53% 5 19
C-030 70 20–49 67% 5 18
C-031 55 100–199 52% 5 14
C-032 45 20–49 53% 5 18
C-033 55 20–49 53% 5 16
C-034 50 50–99 63% 5 15
C-035 75 20–49 73% 5 13
C-036 75 20–49 73% 5 17
C-037 45 50–99 48% 6 16
C-038 65 20–49 62% 5 11
C-039 35 20–49 37% 5 16
C-040 30 20–49 33% 5 11
C-041 30 20–49 35% 1 —
C-042 30 20–49 37% 6 15
C-043 45 50–99 45% 5 34
C-044 60 20–49 64% 5 13
C-045 35 20–49 37% 5 22
C-046 40 20–49 40% 5 25
C-047 40 20–49 40% 5 17
C-048 65 20–49 63% 5 12
C-049 45 100–199 45% 5 11
C-050 30 50–99 34% 6 14
C-051 35 50–99 35% 5 18
C-052 70 20–49 67% 5 18
C-053 0 20–49 0% 5 0
C-054 40 20–49 40% 5 17
C-055 35 20–49 38% 5 9
C-056 35 100–199 35% 5 10
C-057 60 20–49 60% 5 19
C-058 25 50–99 28% 5 15
C-059 30 20–49 33% 5 28
C-060 50 20–49 49% 5 15
C-061 40 20–49 40% 5 17
C-062 40 100–199 42% 5 16
C-063 35 20–49 37% 5 12
C-064 40 20–49 40% 5 25
C-065 25 20–49 27% 5 17
C-066 70 20–49 70% 5 16
C-067 45 20–49 49% 6 16
C-068 25 20–49 25% 1 —
C-069 60 20–49 47% 5 7
C-070 45 50–99 50% 6 19
C-071 25 50–99 28%
5 12
C-072 50 50–99 50% 6 18
C-073 40 20–49 40% 5 20
C-074 15 20–49 17% 5 11
C-075 45 20–49 55% 1 —
C-076 40 20–49 40% 5 5
C-077 25 20–49 27% 5 27
C-078 25 20–49 27% 5 25
C-079 5 20–49 10% 5 13
C-080 30 20–49 33% 5 21
C-081 30 20–49 31% 5 19
C-082 15 100–199 19% 5 10
C-083 35 20–49 37% 5 29
C-084 60 20–49 60% 5 15
C-085 25 20–49 27% 5 31
C-086 50 20–49 51% 5 15
C-087 65 20–49 63% 5 7
C-088 40 20–49 42% 5 23
C-089 35 20–49 30% 5 22
C-090 20 20–49 23% 5 8
C-091 30 20–49 30% 5 22
C-092 10 20–49 15% 1 —
C-093 40 50–99 42% 5 12
C-094 40 20–49 40% 5 17
C-095 70 20–49 70% 5 19
C-096 15 20–49 17% 5 11
C-097 45 20–49 47% 5 16
C-098 70 20–49 70% 5 12
C-099 40 20–49 42% 5 13
C-100 5 20–49 12% 5 9
C-101 35 20–49 38% 5 13
C-102 55 20–49 57% 5 17
C-103 20 20–49 20% 5 12
C-104 55 50–99 53% 5 15
C-105 40 20–49 40% 5 17
C-106 30 20–49 30% 5 22
C-107 30 20–49 33% 5 19
C-108 60 20–49 57% 5 27
C-109 30 20–49 33% 5 24
C-110 5 20–49 10% 5 13
C-111 55 20–49 53% 5 19
C-112 30 20–49 33% 5 28
C-113 15 50–99 17% 5 5
C-114 25 20–49 27% 5 8
C-115 30 20–49 33% 5 16
C-116 80 20–49 77% 5 25
C-117 40 20–49 42% 5 18
C-118 70 20–49 67% 5 12
C-119 20 20–49 20% 5 12
C-120 60 20–49 70% 1 —
C-121 20 20–49 19% 5 16
C-122 70 20–49 70% 5 19
C-123 25 20–49 27% 5 11
C-124 35 100–199 34% 5 11
C-125 60 20–49 45% 1 —
C-126 0 50–99 7% 5 13
C-127 50 50–99 48% 5 8
C-128 70 20–49 67% 5 11
C-129 45 20–49 44% 5 16
C-130 55 20–49 53% 5 12
C-131 15 50–99 19% 5 14
C-132 25 50–99 28% 5 12

Commander: OLS residual outliers

Fitting an OLS line through the 132 appendix rows (coarsened Power Score vs. win rate), two decks stand out as the widest departures from expectation:

  • Widest overperformance: C-034 (PS 50, win rate 63%) — 13 points above the OLS-expected 50%. This deck won substantially more than a Power Score of 50 would predict within this sample.
  • Widest underperformance: C-125 (PS 60, win rate 45%) — 14 points below the OLS-expected 59%. This deck won substantially less than a Power Score of 60 would predict.

Both figures are residuals from the OLS fit within the published appendix rows only. They describe the extremes of over- and underperformance relative to Power Score in this edition's dataset; they say nothing about the decks' identities, owners, or composition. Pseudonyms rotate every edition.


Standard: per-deck appendix

10 Standard decks qualify for the appendix (minimum 20 games played; all 10 meet the game-count floor and the 10-deck publication minimum exactly).

  • Win rate range: 10% – 67%
  • Median win rate: 38%
  • Median matchup spread: 24 percentage points (n = 10 decks)

The Standard median matchup spread (24 points) is substantially wider than Commander's (16 points), indicating that Standard decks in this sample show greater variation in their outcomes across gauntlet opponents — consistent with the Standard variance decomposition finding that the opponent main effect is proportionally larger in Standard.

ID PS (coarsened) Games Band Win Rate Opponents Faced Spread
S-001 55 20–49 53% 5 29
S-002 70 20–49 67% 5 21
S-003 50 20–49 50% 5 24
S-004 30 20–49 30% 5 19
S-005 25 20–49 27% 5 23
S-006 40 20–49 40% 5 45
S-007 35 20–49 37% 5 27
S-008 5 50–99 10% 5 13
S-009 25 20–49 29% 5 29
S-010 35 20–49 38% 5 17

S-006 carries the widest matchup spread in the entire Standard appendix at 45 percentage points — meaning its win rate varies by 45 points across its five gauntlet opponents, the most extreme matchup-specialist profile in the Standard cohort.

Standard: OLS residual outliers

With only 10 appendix rows, the OLS fit is highly sensitive to individual data points and the residual figures should be read with that caveat explicitly in mind. The two extremes:

  • Widest overperformance: S-010 (PS 35, win rate 38%) — 2 points above the OLS-expected 36%.
  • Widest underperformance: S-004 (PS 30, win rate 30%) — 2 points below the OLS-expected 32%.

The residuals in Standard are small (2 points in each direction) — a reflection of both the small sample and the narrow Power Score range represented in the Standard appendix. No strong outlier signal is present at n = 10.


Section 14: Key Findings, What We Cannot Conclude Yet, and Future Baseline

Key findings (this edition)

The following findings are directly measured within the Grim.Cards dataset and carry their sample sizes. All are descriptive; none establishes causation.

  1. The gauntlet opponent dominates short-term outcome variance in Standard; the deck dominates in Commander. Across 133 Commander decks and 17 opponents, the deck main effect accounts for 48.4% of win-rate variance vs. 12.7% for the opponent. Across 10 Standard decks and 5 opponents, the opponent accounts for 36.2% vs. 26.3% for the deck. The formats differ structurally in how much the specific challenger matters.

  2. A 42.1-point spread separates the easiest and hardest Standard gauntlet opponents (Temur Harmonizer Combo 55.6% vs. Mono Red Aggro 13.5%, both n = 69 decks, 252 games each). In Commander the equivalent spread is 21.4 points (Breya Artifact Combo 52.5% vs. Edgar Markov Vampires 31.1%, both n = 659 decks).

  3. Commander decks win 43.0% of simulated games overall (n = 713 decks, 13,292 games). Standard decks win 34.5% (n = 77 decks, 1,267 games). Sub-50% is the structural baseline in a one-versus-the-field gauntlet.

  4. The Commander win-rate distribution is bimodal. The 30–40% band contains the most decks (190 of 662, 28.7%) but the 50–60% band is a clear second peak (115 decks, 17.4%), suggesting two distinct performance clusters rather than a single continuous distribution.

  5. Retesting does not reliably improve simulated win rates. Among 143 Commander retests, 62 declined vs. 60 improved (average delta: −0.7 points). Among 12 Standard retests, 6 declined vs. 4 improved (average delta: −1.1 points).

  6. Meren of Clan Nel Toth is the only commander with enough submitted decks to report a win rate (21 decks, 510 games, 34.3% — below the 43.0% format average).

  7. Herald of Secret Streams leads the Commander board-impact ranking (+72.54 per appearance, 12 decks, 50 observations). Craterhoof Behemoth is the most broadly observed top-performer (28 decks, 89 observations, +20.26). Circle of Dreams Druid and Aura Shards post the worst board impact (−25.43 and −25.26 respectively). All figures are board-state proxies, not causal win claims.

  8. Commander mono-color decks win at 50.9% (n = 128) — 7.9 points above the format average — while two-color decks post the lowest rate at 40.5% (n = 245). This is a descriptive correlation.

  9. Commander decks with search/tutor effects win 5.8 points higher than those without (43.9%, n = 629 vs. 38.1%, n = 33). Standard decks with reanimation effects win 12.2 points lower than those without (25.7%, n = 21 vs. 37.9%, n = 49). All category figures are heuristic-assigned and correlational.

  10. Power Score rank-orders Commander decks consistently across all five gauntlet opponents (Spearman ρ = 0.69–0.75, n ≈ 124–125 per opponent). In Standard the range is wider (ρ = 0.41–0.84, n = 10 per opponent), with Mono Red Aggro producing the weakest alignment.

What we cannot conclude yet

  • Commander-level win-rate rankings remain limited. Only Meren of Clan Nel Toth clears the 10-deck minimum. Meaningful commander comparisons require more submissions. Future editions will unlock more commanders as the dataset grows.
  • Counterfactual decision-impact data was unavailable for this edition. No card-level counterfactual figures met the publication threshold. This section will be populated in future editions if data accumulates.
  • Standard results are substantially less stable than Commander results. The Standard cohort (77 decks, 1,267 games) is roughly one-tenth the Commander cohort. Month-level, color-level, and per-card Standard figures are near their minimum thresholds throughout. Standard findings should be replicated in future editions before strong interpretations are drawn.
  • Self-selection bias cannot be corrected. Players choose which decks to submit; the sample is not random. Decks that players believe are strong may be overrepresented. This affects all findings but cannot be quantified with available data.
  • Causal claims about construction remain out of reach. Land counts, creature percentages, mana values, and category membership all co-vary with many other deck properties. The correlations in Sections 9 and 10 are real but their causes are not identified.
  • The inversion-rate and variance decomposition figures will stabilize with larger samples. The Standard inversion analysis (n = 10 decks, 340 pairs) and variance decomposition (10 decks, 5 opponents, 50 cells) are at their minimum viable sizes. Commander figures (n = 122–133 decks) are more robust but not definitive.

Future comparison baseline

The following table captures the key metrics for this edition in a form suitable for direct comparison with future editions. All figures are from the 2026-09-01 snapshot.

Metric Commander Standard All Formats
Snapshot date 2026-09-01 2026-09-01 2026-09-01
Decks (n) 713 77 790
Users (n) 278 53 314
Simulations 985 95 1,080
Total games 13,292 1,267 14,559
Overall win rate 43.0% 34.5% 42.2%
Median deck win rate 40% 33.3% —
Mean deck win rate 43.6% 34.3% —
Best matchup win rate 52.5% (Breya) 55.6% (Temur) —
Worst matchup win rate 31.1% (Edgar) 13.5% (Mono Red) —
Gauntlet spread 21.4 pts 42.1 pts —
Retested decks 143 12 155
Avg retest win-rate delta −0.7 pts −1.1 pts —
Median matchup spread 16 pts (n=124) 24 pts (n=10) —
Avg game length (turns) 10.3 (n=661) 9.6 (n=70) —
Deck main-effect share 48.4% 26.3% —
Opponent main-effect share 12.7% 36.2% —
Inversion rate 39% 29.7% —
Median Spearman ρ (PS vs WR) 0.71 0.53 —
Top board-impact card Herald of Secret Streams (+72.54) Lightning Bolt (+5.10) —
Dataset version 2.4 2.4 2.4

Section 15: Power Score (Secondary / Optional)

Important caveat: Power Score is an internal Grim.Cards composite indicator, not an objective or universal measure of deck quality. It is not a win-rate prediction tool and should not be interpreted as one. It is presented here last and for context only. The Spearman correlations in Section 11 describe how well it rank-orders simulated outcomes within this dataset — that is all that can be said from this data.

Power Score grades are assigned to player-submitted Commander decks in this dataset as follows (grade counts from completed simulations, not all submitted decks):

Grade Decks (n)
S 6
A 26
B 91
C 135
D 212
F 320

The distribution is heavily weighted toward D and F grades (532 of 790 total scored decks, across all formats). This reflects the composition of the player-submitted sample, not a judgment on player skill: the gauntlet is fixed and demanding, and the grading scale is calibrated against it. The grade distribution is reported here for completeness; it is not used to rank, filter, or evaluate any finding elsewhere in this report.


Limitations

The following limitations apply to all findings in this report and should be considered when reading any figure herein.

  1. Simulated play, not human play. All games are resolved by a Forge-based AI engine. Engine decisions differ from human decisions; results describe engine behavior on these decklists and may not translate to human play outcomes.
  2. Correlational throughout. No finding in this report establishes a causal relationship. Construction correlations, category correlations, color correlations, and card performance figures are all descriptive associations within this dataset.
  3. Self-selected, non-random sample. Players choose which decks to submit. The sample is not representative of all decks in existence or of competitive play in any environment.
  4. Fixed gauntlet. Results are specific to the gauntlet opponents used in this edition. A different gauntlet would produce different win rates.
  5. Power Score is an internal composite. It is not an objective or universally comparable measurement and is not used as a primary metric in this report.
  6. Decision-impact figures are counterfactual proxies. They measure play quality relative to the engine's own next-best alternative, not damage, kills, or direct win contribution.
  7. Board-impact figures are board-state proxies. They measure board-quality changes around card observations, not per-card win contribution.
  8. Functional card categories are heuristic-assigned. Keyword heuristics may mislabel edge cases; category figures are labeled [HEURISTIC] throughout.
  9. Standard sample is small. With 77 decks and 1,267 games, many Standard figures sit near the minimum publication threshold and may not replicate in future editions.
  10. Per-deck appendix figures are coarsened. Power Scores are rounded to the nearest 5, win rates and spreads to whole points, and game counts are banded, to resist cross-edition re-identification. Small precision losses result.

License & Citation

License: This dataset and report are published under the Creative Commons Attribution 4.0 International license (CC BY 4.0). You are free to share and adapt this material for any purpose, provided you give appropriate credit, provide a link to the license, and indicate if changes were made. License text: https://creativecommons.org/licenses/by/4.0/

Cite this edition as:

Grim.Cards (2026-09-01). Grim.Cards Simulation Data Case Study, Edition 2026-09-01 (Dataset v2.4). https://grim.cards/case-study/2026-09-01. CC BY 4.0.

Machine-readable citation fields:

Field Value
Publisher Grim.Cards
Title Grim.Cards Simulation Data Case Study, Edition 2026-09-01
Publication date 2026-09-01
Dataset version 2.4
Canonical URL https://grim.cards/case-study/2026-09-01
Temporal coverage 2026-05-14 / 2026-09-01
License CC BY 4.0
Retrieved Cite the date you accessed this page

Prior editions: https://grim.cards/case-study


Grim.Cards Simulation Data Case Study — Edition 2026-09-01 — Dataset v2.4 — Generated 2026-09-01T08:48:33.023Z — CC BY 4.0

All Grim.Cards case-study editions

Grim.Cards is an independent fan-created platform, not affiliated with, endorsed by, sponsored by, or approved by Wizards of the Coast LLC or Hasbro, Inc. Magic: The Gathering® is a trademark of Wizards of the Coast LLC.