All Grim.Cards case-study editions

Grim.Cards Simulation Data — Case Study Edition

Data Snapshot: 2026-08-01 · Dataset Version 2.4 · License: CC BY 4.0

Permanent URL: https://grim.cards/case-study/2026-08-01 Published: 2026-08-01 · Temporal Coverage: 2026-05-14 → 2026-08-01 Cohort: Real, human-submitted Commander and Standard decks. Automated Crucible decks and system sample decks are excluded throughout. Minimum cohort threshold: 10 unique decks per reported group. Any group below this floor is suppressed. Win rate definition (applied consistently throughout): wins ÷ total games played, with draws counted in the denominator. All findings are correlational. Simulated (AI-piloted) results describe engine behavior on these decklists and cannot be generalized to human play or to Magic: The Gathering broadly.


Executive Summary

Between 14 May and 1 August 2026, Grim.Cards ran 691 completed simulations of player-submitted decks across 507 unique decks from 207 distinct users, generating 10,011 AI-versus-AI games. The headline number: player-submitted decks won 42.6% of those games (n = 507 decks; 10,011 games). That figure sits below 50%, which is structurally expected — each deck faces a fixed gauntlet of purpose-built meta opponents, so the field as a whole is designed to win more often than any single challenger. The sub-50% aggregate is a baseline, not a verdict.

The more revealing numbers are in the matchup spreads. In Commander (n = 448 decks; 8,969 games; overall win rate 43.2%), the difference between the easiest and hardest Gauntlet opponent in this dataset is 22.5 percentage points — from 53.5% against Breya Artifact Combo down to 31.0% against Edgar Markov Vampires. In Standard (n = 59 decks; 1,042 games; overall win rate 37.1%), that spread balloons to 45.4 percentage points — from 60.4% against Temur Harmonizer Combo down to 15.0% against Mono Red Aggro. The opponent facing a deck is, in measured terms, at least as important a driver of outcome as the deck's own profile — and in Standard, the data suggests the opponent matters more.

Among 106 Commander decks retested, 55 improved their win rate across attempts and 36 declined, with an average change of +2.6 percentage points. Among 12 Standard decks retested, 5 improved and 5 declined, averaging +0.6 points.

The board-performance data — a measured proxy for how much a card improves the game state when it appears, not a causal win claim — shows Ureni of the Unwritten topping the Commander rankings at +29.78 board-quality points per appearance (11 decks, 253 observations), while Dictate of Erebos leads the bottom list at −13.72 (25 decks, 174 observations). In Standard, Lightning Bolt is the sole card meeting the ranking threshold at +5.05 (10 decks, 95 observations).

Power Score, an internal Grim.Cards composite indicator, is presented last and caveated. It is not an objective measurement and is not used to rank or judge decks anywhere in this report.


1. Methodology and Provenance

Simulation method

Each player-submitted deck is played by a custom build of the open-source Forge engine — an AI piloting the deck against a fixed set of AI-piloted Gauntlet opponents. No human decisions are involved during simulation. Results describe how the engine performs with these decklists; they do not predict human play outcomes.

Cohort definition

The primary cohort is every deck submitted by a real user account, excluding the automated Crucible reference corpus (user_id = '__grinder__') and system sample decks (is_sample = true). The Crucible corpus is not counted in user or deck totals and does not appear in any reported rate.

Win rate

Win rate is defined as wins ÷ total games played, with draws counted in the denominator. This definition is applied identically across every table, chart, and rate in this report. "Decisive win rate" (wins ÷ wins + losses only) is never used.

Format splits

Commander and Standard results are reported separately in every table and section. The simulation database carries no format column; format is derived by joining simulation records to deck records. Results are never pooled across formats.

Cohort threshold

No rate or breakdown is published for a group of fewer than 10 unique decks. Groups below this floor are suppressed or merged into an "Other" category. Where a group's "without" contrast group falls below 10 decks, that contrast figure is withheld (shown as suppressed).

Retest identification

A retest is defined as a deck with more than one completed simulation on record. The reported change in win rate is the per-deck win rate of the most recent completed simulation minus the win rate of the first completed simulation.

Card categories (heuristic — labelled throughout)

Functional card categories (sacrifice outlets, search/tutor effects, discard effects, reanimation effects) are assigned by a keyword heuristic applied to oracle text and pre-existing preference flags. Category assignment may mislabel edge cases. All category-level win rates carry a visible heuristic label and are correlation-only.

Board-impact metric (measured proxy — labelled throughout)

The card board-impact figure (avgDelta) is the pooled board-quality change per observation across all cohort decks running a card — positive means the board state improved in turns the card was seen. This is a board-state proxy, not a damage metric, not a kill count, and not a causal win claim. Only cards appearing in 10 or more distinct decks with 25 or more recorded observations are ranked.

Decision-impact / counterfactual metric

Counterfactual decision-impact figures — where reported — measure how much a card's chosen plays beat the engine's own next-best alternative line. Positive means the line played was better than the alternative; negative means the alternative was better. No counterfactual data met the cohort threshold for this edition; the section is omitted rather than suppressed at an uninformative level.

Privacy

No player names, deck names, deck lists, user identifiers, or user-level timestamps appear anywhere in this report. Per-deck appendix entries use rotating pseudonyms (e.g., C-001, S-001) reshuffled every edition — a pseudonym in this edition carries no link to any pseudonym in a prior edition. Published per-deck values are deliberately coarsened: Power Score rounded to the nearest 5, win rates and matchup spreads to whole percentage points, game counts banded.

Limitations

  • Results are from AI-piloted simulations, not human play.
  • All construction and category breakdowns are correlational — no causal claims are made or implied.
  • The sample is self-selected: it comprises decks that users chose to submit to Grim.Cards, not a random sample of the broader Magic: The Gathering population.
  • Monthly cohorts consist of different sets of decks month to month, so monthly win-rate trends reflect who submitted, not any single deck's trajectory.
  • Power Score is an internal composite indicator, not an objective or universal measurement.
  • Functional card category assignments are heuristic and may carry labelling error.
  • Board-impact figures are a board-state proxy, not damage or kills, and not a causal win contribution.

2. Dataset Overview

Metric All Formats Commander Standard
Distinct users 207 179 40
Distinct decks 507 448 59
Completed simulations 691 619 72
Total games 10,011 8,969 1,042
Wins 4,261 3,874 387
Losses 5,747 5,092 655
Draws 3 3 0
Win rate 42.6% 43.2% 37.1%
Data range 2026-05-14 → 2026-08-01 2026-05-14 → 2026-08-01 2026-05-14 → 2026-08-01

Commander is the dominant format in this dataset by every measure: 88% of decks, 90% of games, and 90% of completed simulations. Standard figures represent a meaningfully smaller sample and should be read with that in mind throughout.


3. The Gauntlet: Matchup Results

This is the core section of the report. Every other metric — distribution, retests, construction, board impact — is downstream of what happens in these matchups. The Gauntlet is a fixed set of AI-piloted meta opponents; every player deck faces the same field under identical conditions.

3a. Commander Gauntlet (n = 447 decks per matchup)

Each matchup row below represents 447 player-submitted Commander decks tested against that opponent. All five cohorts clear the 10-deck minimum by a wide margin.

Opponent Decks (n) Wins Losses Draws Games Win Rate
Breya Artifact Combo 447 948 825 0 1,773 53.5%
Derevi Bant Control 447 895 879 0 1,774 50.5%
Aesi Landfall 447 823 951 2 1,776 46.3%
Atraxa Superfriends 447 607 1,170 1 1,778 34.1%
Edgar Markov Vampires 447 549 1,224 0 1,773 31.0%

Range across opponents: 22.5 percentage points. The same pool of 447 player decks wins more than half their games against Breya Artifact Combo and fewer than one-third against Edgar Markov Vampires. The deck slate did not change between those two rows — only the opponent did. That 22.5-point swing is a direct, measured signal that the specific meta opponent is a substantial driver of matchup outcome within this dataset.

The midpoint of the Gauntlet — Aesi Landfall at 46.3% — sits just below even odds, which is consistent with a well-calibrated field. Players hold slight overall edges against the two "softer" opponents and are clearly disadvantaged against the two "harder" opponents, producing the 43.2% aggregate.

3b. Standard Gauntlet (n = 58 decks per matchup)

Each matchup row below represents 58 player-submitted Standard decks. All five cohorts clear the 10-deck minimum.

Opponent Decks (n) Wins Losses Draws Games Win Rate
Temur Harmonizer Combo 58 125 82 0 207 60.4%
Jeskai Control 58 96 111 0 207 46.4%
Azorius Tempo 58 69 138 0 207 33.3%
Dimir Midrange 58 64 143 0 207 30.9%
Mono Red Aggro 58 31 176 0 207 15.0%

Range across opponents: 45.4 percentage points. The Standard spread is twice as wide as Commander's. Against Temur Harmonizer Combo, the average player Standard deck wins three games in five. Against Mono Red Aggro, that same pool of decks wins fewer than one game in seven. Standard decks in this dataset struggle broadly — a 37.1% overall win rate against a gauntlet that includes a 60.4% matchup and a 15.0% matchup simultaneously — and the variance in outcomes is almost entirely a function of which opponent is across the table.

The severity of the Mono Red Aggro result (15.0%, 31 wins in 207 games) warrants context. The Standard cohort (n = 59 decks) skews toward midrange, control, and combo archetypes — the decks users submitted. A fast aggro opponent that routinely closes games in seven turns or fewer (the Standard average shortest game is 7.3 turns) is a particularly punishing test for the submitted deck pool, and this dataset measures exactly that.


4. Top Commanders (Commander, ≥10-deck cohort floor applied)

Only one commander clears the 10-deck minimum required to report a win rate in this dataset: Meren of Clan Nel Toth, with 17 decks, 360 games, and a win rate of 33.6%. All other commanders listed in the dataset have fewer than 10 distinct decks and their win rates are withheld per the cohort threshold.

Commander Decks (n) Games Win Rate
Meren of Clan Nel Toth 17 360 33.6%
All others < 10 each — Suppressed

Meren of Clan Nel Toth is the most-played commander in the Grim.Cards dataset by a meaningful margin (17 decks versus the next cluster at 8 or fewer). Its 33.6% win rate sits roughly 10 points below the Commander average of 43.2%, but the spread across the Gauntlet — Meren decks face all five Commander opponents — makes this an aggregate figure only. No per-matchup breakdown for a single commander can be published at this cohort size without falling below the per-cell floor.

The full top-commander list by deck count (win rates suppressed below the threshold) in order of popularity: Meren of Clan Nel Toth (17), Ureni of the Unwritten (8), The Ur-Dragon (6), Emmara, Soul of the Accord (5), Atraxa, Praetors' Voice (5), Kaalia of the Vast (5), Vivi Ornitier (5), Zimone, Infinite Analyst (5), Y'shtola, Night's Blessed (5). Commanders with 4 or fewer decks: Fire Lord Azula, Marrow-Gnawer, Giada Font of Hope, Kenrith the Returned King, Tony Stark, Krenko Mob Boss, Pantlaza Sun-Favored, Atarka World Render, Cosmic Spider-Man, and others. All sub-10 win rates are suppressed.

The commander diversity in this dataset is striking. Nineteen distinct commanders appear in the top-20 list and none has accumulated enough decks to generate a publishable per-matchup breakdown beyond Meren. Future editions with a larger sample may unlock commander-level matchup reporting.


5. Top Cards by Usage and by Containing-Deck Win Rate

5a. Commander — Most-played cards (basic lands excluded)

Ranked by number of distinct decks the card appears in. Win rate shown is the aggregate win rate of all decks containing that card (not the card's individual contribution). Correlation only — popular cards appear in popular archetypes, which share many variables.

Card Decks (n) Games Win Rate of Containing Decks
Sol Ring 387 7,823 43.4%
Command Tower 321 6,448 42.0%
Arcane Signet 312 6,455 44.0%
Exotic Orchard 167 3,404 43.9%
Lightning Greaves 139 2,695 45.1%
Swiftfoot Boots 137 2,497 43.6%
Reliquary Tower 124 2,406 39.6%
Path of Ancestry 117 2,218 49.1%
Evolving Wilds 107 2,177 44.1%
Swords to Plowshares 106 2,016 45.6%
Cultivate 93 1,646 47.4%
Fellwar Stone 81 1,799 44.5%
Rogue's Passage 80 1,579 43.9%
Bojuka Bog 80 1,785 40.6%
Path to Exile 78 1,503 47.5%
Terramorphic Expanse 75 1,577 41.6%
Skullclamp 73 1,550 40.1%
Counterspell 72 1,590 40.4%
Birds of Paradise 67 1,359 39.3%
Demonic Tutor 65 1,379 43.3%

The most-played cards in Commander are unsurprisingly format staples: Sol Ring appears in 387 of 448 decks (86%). The containing-deck win rates cluster tightly between 39% and 49% for the top-20 list, reflecting that these cards are distributed broadly across decks of all power levels rather than concentrated in high-performing archetypes. Path of Ancestry (49.1%, n = 117) and Path to Exile (47.5%, n = 78) sit at the upper end; Birds of Paradise (39.3%, n = 67) and Skullclamp (40.1%, n = 73) at the lower end. These differences are small relative to the Gauntlet-opponent spread of 22.5 points, reinforcing that opponent identity is a larger driver of outcomes than any single card presence.

5b. Standard — Most-played cards (basic lands excluded, ≥10-deck cohort floor applied)

Card Decks (n) Games Win Rate of Containing Decks
Lightning Bolt 13 195 46.2%
Inspiring Vantage 12 180 49.4%
Shard Volley 11 165 50.9%
Chain Lightning 11 165 50.9%
Eidolon of the Great Revel 11 165 50.9%
Monastery Swiftspear 11 165 50.9%
Skullcrack 11 165 50.9%
Goblin Guide 11 165 50.9%
Lava Spike 11 165 50.9%
Searing Blaze 11 165 50.9%
Rift Bolt 11 165 50.9%

Standard cards below 10 decks: Sol Ring (7), Evolving Wilds (6), Llanowar Elves (6), Dark Ritual (5), Arcane Signet (5), Reliquary Tower (4), Abrade (4), Mind Stone (4), Rogue's Passage (4) — win rates suppressed.

The Standard card list tells a concentrated story. Ten of the eleven qualifying cards are core Mono Red Burn package cards (Shard Volley, Chain Lightning, Eidolon of the Great Revel, Monastery Swiftspear, Skullcrack, Goblin Guide, Lava Spike, Searing Blaze, Rift Bolt, and Inspiring Vantage), each appearing in 11 of the 59 Standard decks. All share an identical win rate of 50.9% across 165 games, which reflects that they appear together as a block in the same burn archetype rather than independently. These decks perform above the 37.1% Standard average, but whether the cards cause that overperformance or simply identify the archetype cannot be determined from these numbers alone. Lightning Bolt (13 decks, 46.2%) is the most widely distributed Standard card meeting the threshold.


6. Win-Rate Distribution

How were simulated win rates distributed across the submitted deck pool? The distributions reveal not just the average but the shape of outcomes — including how many decks are clustered near zero, how many exceed 50%, and how extreme the tails are.

6a. Commander (n = 448 decks)

Win-Rate Band Decks
0–10% 15
10–20% 43
20–30% 52
30–40% 130
40–50% 60
50–60% 68
60–70% 44
70–80% 26
80–90% 9
90–100% 1

Summary (n = 448): Mean 43.4% · Median 40% · Min 0% · Max 93.3%

The Commander distribution is left-skewed. The largest single band is 30–40% (130 decks, 29% of the sample), and the modal experience is finishing just below the halfway mark. The 0–30% zone holds 110 decks (24.6% of the sample) — a meaningful minority that consistently lost more than 70% of their games against the Gauntlet. Meanwhile, 79 decks (17.6%) achieved win rates above 60%, and one deck reached 93.3%. The mean (43.4%) sits above the median (40%) precisely because of those upper-tail outliers, a hallmark of positive skew. For most Commander decks in this dataset, the simulated experience is one of losing more than winning — but a substantial minority performs well above the overall average.

6b. Standard (n = 59 decks)

Win-Rate Band Decks
0–10% 7
10–20% 9
20–30% 9
30–40% 16
40–50% 3
50–60% 9
60–70% 4
70–80% 2
80–90% 0
90–100% 0

Summary (n = 59): Mean 35.3% · Median 33.3% · Min 0% · Max 80%

Standard's distribution is also left-skewed, but more severely. The 0–30% zone holds 25 decks (42.4% of the sample). The 40–50% band — the natural "roughly even" range — contains only 3 decks. No Standard deck in this dataset reached 80%, and the 70–80% band holds just 2. Standard decks face a more polarized Gauntlet: the Temur Harmonizer Combo matchup is favorable for many, but Mono Red Aggro is brutal for almost all, and the two effects do not cancel each other evenly. The result is a distribution where the center of mass sits in the 30s and the upper tail is modest.


7. Deck Iteration and Retests

Users who submitted the same deck for multiple test runs provide a direct, within-deck look at whether repeated testing correlates with win-rate change. All comparisons are first completed simulation vs. most recent completed simulation for a given deck.

7a. Commander retests (n = 106 retested decks)

Outcome Decks
Improved win rate 55
Declined win rate 36
Roughly flat 15
Average win-rate change +2.6 percentage points

Among 106 Commander decks retested, 55 showed an improvement and 36 a decline — a 60/40 split in favor of improvement among decks whose direction could be determined. The average change across all 106 retested decks is +2.6 points. The "flat" category (15 decks) represents cases where win rate changed negligibly between test instances.

What this does not tell us: the dataset does not reveal whether decks changed between tests, how much time elapsed, or what drove the directional change. Monthly cohorts are also different collections of decks, so the +2.6 average is not a before-and-after track for the same population — it is a deck-level statistic for the specific subset of decks tested more than once.

7b. Standard retests (n = 12 retested decks)

Outcome Decks
Improved win rate 5
Declined win rate 5
Roughly flat 2
Average win-rate change +0.6 percentage points

The Standard retest cohort (n = 12) is at the minimum publication threshold and warrants caution. The 5/5/2 split — perfect symmetry between improved and declined — and the near-zero average change of +0.6 points suggest no systematic directional pattern in this small sample. Larger Standard retest cohorts in future editions will clarify whether the near-zero result persists.


8. Monthly Trends

Monthly test volume and average win rate by format. Monthly cohorts consist of different decks submitted in that calendar month, not a longitudinal track of the same decks. Win-rate changes across months reflect who submitted, not progression of any individual deck.

Month Format Sims Decks Games Win Rate
2026-05 Commander 74 51 1,056 40.4%
2026-06 Commander 149 110 2,137 43.8%
2026-06 Standard 24 20 360 37.5%
2026-07 Commander 384 280 5,596 43.4%
2026-07 Standard 38 30 532 36.1%

Growth in Commander testing is the dominant story: 74 simulations in May, 149 in June (+101%), and 384 in July (+158% month-over-month). July alone accounts for 55.5% of all Commander simulations in the dataset. Standard debuted in the data in June (24 sims) and grew to 38 in July.

Commander win rates moved from 40.4% in May to 43.8% in June and held at 43.4% in July — a modest 3-point improvement from the first to the most recent full month. Standard win rates moved from 37.5% in June to 36.1% in July, a small decline. Both shifts are small relative to the Gauntlet-opponent spread, and both reflect different deck populations month to month, so directional interpretation is limited.

Standard data only begins in June because no Standard simulations meeting the reporting threshold were recorded in May. The May Standard cohort, if any existed, fell below the 10-deck minimum.


9. Construction and Win Rate

What construction characteristics are associated with different win rates? All findings here are strictly correlational — the direction of any relationship could run in either direction, and numerous confounders (archetype, commander choice, submission bias) cannot be controlled in this observational data.

9a. Color count and win rate

Commander (all cohorts ≥ 10 decks):

Colors in Identity Decks (n) Win Rate
1 color 86 50.8%
2 colors 165 40.1%
3 colors 141 42.9%
5 colors 45 42.3%

The 4-color group did not clear the 10-deck minimum and is suppressed. Mono-color Commander decks post the highest average win rate in this dataset (50.8%, n = 86), 8–10 points above all multi-color groups. This is a descriptive finding only. Mono-color decks may be concentrated in higher-performing archetypes (e.g., mono-Red aggro commanders, mono-White token strategies), and the Gauntlet opponents are all multi-color — any structural advantage for mono-color against these specific opponents is a property of this dataset's gauntlet composition, not a general claim about Magic.

Standard (cohorts meeting the ≥10-deck floor):

Colors Decks (n) Win Rate
1 color 11 44.2%
2 colors 37 36.7%

The 3-color-plus Standard cohorts fall below the floor and are suppressed. The pattern echoes Commander: mono-color Standard decks post a higher win rate (44.2%, n = 11) than two-color decks (36.7%, n = 37). The mono-color group likely overlaps heavily with the burn archetype identified in the card usage data.

9b. Land ratio and win rate

Commander:

Land % (approximate) Decks (n) Win Rate
~30% 66 40.3%
~35% 257 43.2%
~40% 112 46.7%

Higher land ratios are associated with higher win rates in Commander: decks near 40% lands average 46.7% (n = 112) versus 40.3% (n = 66) for decks near 30% lands — a 6.4-point difference. This pattern is consistent with the win-bracket analysis (Section 9d) but does not indicate causation.

Standard:

Land % (approximate) Decks (n) Win Rate
~35% 18 30.6%
~40% 26 43.1%

The same directional pattern holds in Standard: decks near 40% lands average 43.1% (n = 26) versus 30.6% (n = 18) for decks near 35% lands — a 12.5-point difference. Both Standard cells clear the 10-deck floor. The burn archetype's land-light construction (burn spells substitute for land-hungry plays) may concentrate lower-land-ratio decks toward specific archetypes.

9c. Color-identity presence and win rate

Color presence is not exclusive — most decks contain multiple colors. These are not independent comparisons.

Commander:

Color Decks (n) Games Win Rate
Red 205 3,982 44.1%
Green 227 4,366 43.5%
White 221 4,451 43.1%
Blue 209 4,258 41.7%
Black 238 4,870 41.1%

Red-inclusive decks post the highest average win rate (44.1%, n = 205) and Black-inclusive decks the lowest (41.1%, n = 238) among Commander decks in this dataset. The range is narrow — under 3 points — and all colors are heavily represented (205–238 decks each). Because most Commander decks run two or more colors, these figures heavily overlap and cannot be interpreted as independent color contributions.

Standard:

Color Decks (n) Games Win Rate
Red 26 409 36.9%
Green 17 345 36.5%
White 32 540 36.1%
Blue 18 319 30.4%
Black 25 427 28.1%

Standard shows a wider spread: Red-inclusive decks average 36.9% (n = 26) and Black-inclusive decks 28.1% (n = 25) — an 8.8-point gap. The burn archetype identified in Section 5b is Red-based, which may lift the Red aggregate. Blue and Black decks in this Standard cohort post the weakest results, though all cohorts are smaller than in Commander and the Standard Gauntlet includes a severe Mono Red Aggro matchup that may disproportionately affect control and midrange archetypes (the decks most likely to run Blue and Black).

9d. Win-bracket comparison: high, mid, and low win-rate decks

Decks grouped by win rate (high = above 55%, mid = 40–55%, low = below 40%) and compared on average construction features. Brackets must clear the 10-deck floor.

Commander:

Bracket Decks (n) Avg Win Rate Avg Land % Avg Creature % Avg Spell % Avg Art/Ench % Avg Mana Value
High (>55%) 114 68.6% 36% 28.6% 16.0% 18.5% 3.31
Mid (40–55%) 161 45.2% 35.9% 28.0% 17.9% 17.5% 3.18
Low (<40%) 173 25.1% 35.0% 27.2% 19.6% 17.4% 3.08

Across Commander brackets, the high-win-rate group (n = 114) runs approximately 1 percentage point more lands than the low-win-rate group (n = 173) and carries an average mana value 0.23 higher (3.31 vs. 3.08). Spell percentage moves in the opposite direction: low-win-rate decks carry more instants and sorceries on average (19.6%) than high-win-rate decks (16.0%). Creature and artifact/enchantment percentages show smaller differences across brackets. These are descriptive correlations across submitted decks, not construction prescriptions.

Standard:

Bracket Decks (n) Avg Win Rate Avg Land % Avg Creature % Avg Spell % Avg Art/Ench % Avg Mana Value
Mid (40–55%) 18 46.5% 36.8% 27.2% 31.1% 5.3% 2.09
Low (<40%) 33 21.3% 34.7% 29.9% 18.7% 15.8% 2.75

The high-win-rate Standard bracket (above 55%) did not clear the 10-deck minimum and is suppressed. Among the two qualifying brackets, mid-win-rate decks (n = 18) run more spells on average (31.1%) and carry lower average mana values (2.09) than low-win-rate decks (n = 33, 18.7% spells, 2.75 mana value). The Standard mid-bracket's higher spell density likely reflects the burn archetype's composition — many low-mana-value instants and sorceries — while low-win-rate decks carry more artifacts and enchantments and higher average mana costs. Again, correlational only.

9e. Card-type mix (full sample)

The card-type distribution across all Commander decks (n = 448, all with split data available):

Type Share of Deck
Lands (nonbasic counted separately in board impact) 35.6%
Creatures 27.9%
Instants + Sorceries (combined) 18.1%
— Instants only 10.0%
— Sorceries only 8.0%
Artifacts 10.2%
Enchantments 7.5%
Planeswalkers 0.7%

Standard (n = 59, all with split data):

Type Share of Deck
Lands 35.6%
Creatures 29.2%
Instants + Sorceries (combined) 23.4%
— Instants only 13.5%
— Sorceries only 9.6%
Artifacts 5.4%
Enchantments 6.0%
Planeswalkers 0.6%

Instant and sorcery density is higher in Standard (23.4%) than Commander (18.1%), consistent with the format's more interactive, less permanent-focused game plan. Artifact density is nearly twice as high in Commander (10.2% vs. 5.4%), reflecting Commander's heavier use of mana rocks and equipment. Land percentages are identical across formats in this dataset (35.6% each).


10. Card-Category Insights

⚠️ Heuristic label: Functional category membership (sacrifice outlets, tutor/search effects, discard effects, reanimation effects) is assigned by a keyword heuristic applied to oracle text and pre-existing preference flags. Category assignment may mislabel edge cases. All figures in this section are correlational.

10a. Commander category win-rate comparisons

The "without" contrast group for sacrifice outlets (4 decks) is below the 10-deck minimum — delta and contrast win rate are suppressed.

Category With Decks (n) With Win Rate Without Decks (n) Without Win Rate Delta
Sacrifice outlets 444 43.4% 4 Suppressed Suppressed
Search / tutor effects 427 43.5% 21 42.1% +1.4 pts
Discard effects 385 42.8% 63 47.2% −4.4 pts
Reanimation effects 274 42.3% 174 45.1% −2.8 pts

The most notable finding: Commander decks including discard effects (n = 385) win at 42.8%, while those without (n = 63) win at 47.2% — a −4.4-point gap. Decks including reanimation effects (n = 274) also underperform those without (42.3% vs. 45.1%, −2.8 points). Tutor effects show a small positive association (+1.4 points with vs. without). Sacrifice outlets are near-universal (444 of 448 decks), so that comparison is not meaningful.

None of these associations can be read causally. Discard and reanimation effects may cluster in archetypes that face structural disadvantages against the specific Gauntlet opponents in this edition. The direction and magnitude of these associations may shift as the Gauntlet composition or submitted deck pool changes in future editions.

10b. Standard category win-rate comparisons

The "without" contrast group for sacrifice outlets (4 decks) falls below the 10-deck floor — suppressed.

Category With Decks (n) With Win Rate Without Decks (n) Without Win Rate Delta
Sacrifice outlets 55 34.1% 4 Suppressed Suppressed
Discard effects 36 31.0% 23 42.0% −11.0 pts
Search / tutor effects 26 28.8% 33 40.4% −11.6 pts
Reanimation effects 15 23.6% 44 39.3% −15.7 pts

Standard shows stronger negative associations across all three measurable categories. Decks including reanimation effects (n = 15) win at 23.6% — 15.7 points below those without (n = 44, 39.3%). Tutor-inclusive decks (n = 26) trail at 28.8% versus 40.4% for those without (n = 33). Discard-inclusive decks (n = 36) similarly trail (31.0% vs. 42.0%).

One likely explanation: in Standard, these effect categories (reanimation, tutors) tend to appear together in slower, higher-mana-value control and combo archetypes — precisely the deck types that the Standard Gauntlet, anchored by Mono Red Aggro, most severely disadvantages. The category associations almost certainly reflect archetype identity more than any property of the individual card effects. This interpretation is inferential, not measured.


11. Card Board-Impact Rankings

⚠️ Board-state proxy, not causal win claim: Board-impact figures (avgDelta) measure the pooled change in board quality around turns a card was seen, across all cohort decks containing that card. Positive means the board state improved; negative means it degraded relative to the surrounding game state. This is not a damage metric, kill count, or win-rate attribution. The sign convention for board impact is positive = good (already in the native scale; no additional flip). Only cards appearing in ≥10 distinct decks with ≥25 observations are ranked. Rankings are weighted by breadth — broader deck coverage and higher observation counts make a figure more credible.

11a. Commander — Top performers by board impact

282 Commander cards qualified (≥10 decks, ≥25 observations).

Card Decks (n) Observations Avg Board Impact (points/obs)
Ureni of the Unwritten 11 253 +29.78
Dragonlord Atarka 10 47 +28.17
Ghalta, Primal Hunger 11 81 +24.17
Cathars' Crusade 11 43 +23.09
Lyra Dawnbringer 12 61 +19.85
Craterhoof Behemoth 18 54 +18.39
Atarka, World Render 14 80 +18.16
Terror of the Peaks 17 84 +15.74
Lathliss, Dragon Queen 19 81 +13.84
Miirym, Sentinel Wyrm 14 165 +12.16
Orcish Bowmasters 15 95 +12.16
Farseek 47 154 +10.68
Sapphire Medallion 12 88 +10.39
Bloodthirsty Conqueror 10 83 +10.18
Avacyn, Angel of Hope 10 57 +9.86

The top of the Commander board-impact list is dominated by large creatures: Ureni of the Unwritten (+29.78, 11 decks, 253 observations), Dragonlord Atarka (+28.17, 10 decks, 47 observations), Ghalta, Primal Hunger (+24.17, 11 decks, 81 observations), Cathars' Crusade (+23.09, 11 decks, 43 observations), and Lyra Dawnbringer (+19.85, 12 decks, 61 observations). Ureni leads with the highest both breadth-weighted credibility (253 observations across 11 decks) and the strongest per-observation figure.

Farseek stands out as the only ramp spell in the top 15 (+10.68, 47 decks, 154 observations), and it has by far the broadest deck coverage of any card on the list — 47 distinct decks means the figure is drawn from a large, diverse sample. Miirym, Sentinel Wyrm (+12.16, 14 decks, 165 observations) and Orcish Bowmasters (+12.16, 15 decks, 95 observations) tie at the rank-10/11 positions.

11b. Commander — Bottom performers by board impact

Card Decks (n) Observations Avg Board Impact (points/obs)
Dictate of Erebos 25 174 −13.72
Aetherflux Reservoir 11 48 −11.40
Exquisite Blood 14 51 −11.31
Panharmonicon 11 43 −10.67
Parallel Lives 11 44 −10.27
Grim Monolith 12 52 −9.31
The Wind Crystal 10 39 −9.18
Pitiless Plunderer 20 126 −8.83
Doubling Season 21 72 −8.36
Reflections of Littjara 10 33 −8.33
Guardian Project 12 44 −7.66
Bolas's Citadel 10 46 −7.63
Grave Pact 23 88 −7.36
Altar of the Brood 10 67 −7.25
Fleshbag Marauder 12 76 −7.12

The bottom of the Commander board-impact list is arresting in its specificity. Several cards that enjoy strong reputations in Commander deckbuilding communities appear here: Dictate of Erebos (−13.72, the worst-performing card in the dataset by this metric, 25 decks and 174 observations — the highest-credibility figure on the bottom list), Doubling Season (−8.36, 21 decks), Panharmonicon (−10.67, 11 decks), Grave Pact (−7.36, 23 decks), and Bolas's Citadel (−7.63, 10 decks).

These negative figures do not mean these cards are individually responsible for losses, nor that they are poorly designed, nor that they should be cut. The board-impact metric measures what happens to the overall board state in turns those cards are seen — in a game of Magic, a card that creates delayed or conditional value (like Doubling Season, which doubles tokens on later turns) may not register an immediate positive board-quality change in the observation window the engine captures. Combo-oriented win conditions (Aetherflux Reservoir, Exquisite Blood) may appear in games that are already losing before the condition assembles, dragging down the associated observations. The negative signal is real; the interpretation of why it is negative requires more data than board state alone provides. This is reported as observed, not as a judgment on card quality.

11c. Commander — Board impact by card type

Type Cards in Pool Observations Decks Floor Avg Board Impact
Creatures 3,749 67,959 50 +2.20
Instants 730 7,992 92 +1.06
Planeswalkers 108 1,339 13 +0.41
Sorceries 672 6,342 65 −0.12
Artifacts 680 18,667 315 −2.07
Nonbasic Lands 37 198 12 −2.08
Enchantments 767 10,440 37 −2.19

Creatures carry the highest pooled board impact (+2.20 across 67,959 observations) in Commander. Instants follow at +1.06. Sorceries, artifacts, enchantments, and nonbasic lands are all negative in aggregate. The type-bucket aggregations pool enormously varied cards — a +2.20 average for creatures spans everything from Llanowar Elves to Craterhoof Behemoth — and the observation that enchantments average −2.19 coexists with individual enchantments like Cathars' Crusade (+23.09) that are among the best-performing individual cards in the dataset. Type-level figures are descriptive aggregates only and should not be read as type-versus-type prescriptions.

11d. Commander — Board impact by color identity

Color Cards Observations Decks Floor Avg Board Impact
Red 1,331 20,535 53 +3.26
Green 1,730 29,907 65 +2.34
White 1,483 27,731 92 +1.56
Blue 1,506 22,406 41 +1.07
Black 1,569 29,739 52 +0.04
Colorless 687 18,040 315 −1.49

Red cards carry the highest average board impact in Commander (+3.26, 1,331 cards, 20,535 observations), followed by Green (+2.34) and White (+1.56). Black cards average near zero (+0.04) and Colorless cards are negative (−1.49). Several caveats: multicolor cards count toward each of their colors, so a 4-color card contributes to four color buckets simultaneously. Colorless cards are dominated by mana rocks and artifacts (the artifact type-bucket was also negative), which may pull the colorless average down. These are descriptive patterns in the data, not color recommendations.

11e. Standard — Board impact

Only one Standard card met the qualifying floor (≥10 decks, ≥25 observations):

Card Decks (n) Observations Avg Board Impact
Lightning Bolt 10 95 +5.05

Standard's thin card-performance data reflects the format's smaller cohort (59 decks, 1,042 games). Lightning Bolt at +5.05 (10 decks, 95 observations) is the only card with sufficient breadth for publication. The type and color Standard breakdowns show: Instants average −0.17 (124 cards, 729 observations, ≥10 decks floor) and Red cards average +1.76 (156 cards, 1,720 observations, ≥10 decks floor). Only these type and color buckets cleared the minimum; others are suppressed.


12. Color Identity: Popularity and Win Rate

Commander

By deck count, Black is the most common color identity in Commander decks in this dataset (238 decks, 53% of the Commander cohort), followed by Green (227), White (221), Blue (209), and Red (205). All colors are broadly distributed — no color is dramatically underrepresented. By win rate, Red-inclusive decks average 44.1% (n = 205) and Black-inclusive decks average 41.1% (n = 238) — the highest and lowest respectively, separated by 3.0 points. Because Commander decks are almost always multi-color, these figures overlap heavily and do not represent independent color effects.

The most popular color is not the highest-performing color, and the gap between any two colors is small relative to the Gauntlet-opponent spread (22.5 points). Color identity alone explains very little of the win-rate variance in this dataset.

Standard

In Standard, White is the most common color (32 decks, 54% of Standard decks), followed by Black (25), Red (26), Blue (18), and Green (17). Red-inclusive Standard decks average 36.9% (n = 26) and Black-inclusive decks 28.1% (n = 25) — an 8.8-point gap, wider than in Commander. This gap likely reflects the presence of the burn archetype (Red-based, Section 5b) at the top and the difficulty Black-inclusive control or midrange decks face against the Standard Gauntlet's fastest opponents.


13. Tempo and Game Length

Game length is measured as the average turn count per match, derived from per-match decision snapshots. This is a descriptive tempo marker — how long it typically takes the engine to resolve a game — not a measure of interactive quality or human play experience.

Format Decks (n) Avg Turns Avg Shortest Game Avg Longest Game
Commander 448 10.3 8.4 12.3
Standard 59 9.5 7.3 11.7

Commander games average 10.3 turns, with individual games ranging on average from 8.4 to 12.3 turns across the deck sample (n = 448). Standard games are slightly shorter on average (9.5 turns), ranging from 7.3 to 11.7 turns (n = 59). The tighter Standard range reflects the format's lower average mana values (2.09–2.75 in the bracket analysis vs. 3.08–3.31 in Commander) and the presence of Mono Red Aggro in the Standard Gauntlet, which closes games in the fastest windows.

The shortest Standard games average 7.3 turns — a benchmark that illustrates why Standard decks built for turns 4–7 payoffs may struggle to execute against Mono Red Aggro before the game ends.


14. Matchup Structure: Deck vs. Opponent

This section decomposes the relative contributions of the specific Gauntlet opponent versus the deck's own profile to matchup win-rate variance. All decompositions are descriptive — they quantify observed associations, not causes. No significance testing is applied. The data decides which factor carries more signal; there is no prior thesis.

14a. Variance decomposition (games-weighted)

How much of the total variation in per-(deck, opponent) win rates is explained by the deck itself, by which opponent was faced, or by residual effects (interaction, noise)?

Commander (n = 98 decks, 20 opponents, 456 cells):

Effect Share of Variance
Deck main effect 46.9%
Opponent main effect 16.8%
Residual (interaction + noise) 36.3%

Standard (n = 10 decks, 5 opponents, 50 cells):

Effect Share of Variance
Deck main effect 27.8%
Opponent main effect 43.9%
Residual (interaction + noise) 28.3%

The formats split sharply. In Commander, the deck itself accounts for 46.9% of the measured win-rate variance — nearly three times the opponent's 16.8%. In Standard, the relationship inverts: the opponent accounts for 43.9% of variance versus only 27.8% for the deck. The Standard sample is small (n = 10 decks), so these estimates carry more uncertainty than the Commander figures. Both formats leave a substantial residual (36.3% Commander, 28.3% Standard) that is unexplained by simple main effects alone.

These are descriptive decompositions of simulated results only. They describe patterns in this dataset, under this Gauntlet composition, for these submitted decks.

14b. Matchup inversion rate

What share of comparable deck pairs — pairs where both decks had enough games against two specific meta opponents to establish a strict win-rate ordering — reversed that preference order between opponents? A higher rate means the specific opponent reshuffles deck-level outcomes more.

Format Deck Pairs Comparable Pairs Inverted Pairs Inversion Rate
Commander (n = 88 decks, 10 opponent pairings) 88 25,242 9,240 36.6%
Standard (n = 10 decks, 10 opponent pairings) 10 344 79 23.0%

In Commander, 36.6% of the 25,242 comparable deck pairs reverse their opponent preference order between two meta opponents — meaning that if Deck A beats Deck B on a given opponent pairing, there is a meaningful chance the ranking flips on a different pairing. In Standard, that share is 23.0% (344 pairs). Both figures indicate that no single Gauntlet opponent is a reliable stand-in for the full meta — decks specialize, and the field matters. The Commander rate is higher, consistent with Commander's greater deck diversity and wider Gauntlet (20 opponents in the decomposition, 5 in Standard's).

14c. Power Score × opponent win-rate grid (Commander only)

The Standard cohort did not produce a qualifying PS grid (fewer than 10 decks per cell at the required level). The Commander grid uses 3 Power Score bands (low: PS 2.2–31.8; mid: PS 32–47.8; high: PS 48.6–90.9), each with 27–34 decks, against all 5 Commander Gauntlet opponents. All 15 cells cleared the per-cell minimum-games floor.

Opponent Low PS (n = 27–28) Mid PS (n = 31) High PS (n = 30)
Breya Artifact Combo 32.9% 55.3% 71.5%
Derevi Bant Control 33.8% 52.7% 66.2%
Aesi Landfall 32.0% 42.5% 66.2%
Atraxa Superfriends 14.3% 30.0% 56.5%
Edgar Markov Vampires 15.3% 24.5% 54.6%

The grid shows a consistent monotonic pattern across all five opponents: higher Power Score bands are associated with higher win rates in every cell. The gap between Low and High PS is largest against Breya Artifact Combo (32.9% → 71.5%, a 38.6-point spread) and smallest against Atraxa Superfriends in the Low band (14.3%) versus Edgar Markov Vampires (15.3%), where both low-PS groups are severely disadvantaged. The High PS group posts above 54% against every opponent, including the two hardest.

Power Score is an internal Grim.Cards composite indicator. The grid describes the association between PS band and win rate in this dataset, not a causal relationship. Correlation only.

14d. Spearman rank correlation: Power Score vs. win rate per opponent

Format Opponent Decks (n) Spearman ρ
Commander Atraxa Superfriends 89 0.77
Commander Breya Artifact Combo 88 0.75
Commander Edgar Markov Vampires 88 0.75
Commander Aesi Landfall 88 0.67
Commander Derevi Bant Control 88 0.62
Standard Temur Harmonizer Combo 10 0.88
Standard Dimir Midrange 10 0.79
Standard Jeskai Control 10 0.70
Standard Azorius Tempo 10 0.67
Standard Mono Red Aggro 10 0.55

In Commander, the Spearman rank correlation between a deck's Power Score and its win rate against a given opponent ranges from ρ = 0.62 (Derevi Bant Control) to ρ = 0.77 (Atraxa Superfriends) across the five main Gauntlet opponents (n = 88–89 decks each). In Standard (n = 10 per opponent), correlations range from ρ = 0.55 (Mono Red Aggro) to ρ = 0.88 (Temur Harmonizer Combo). The lowest correlations — Derevi Bant Control in Commander and Mono Red Aggro in Standard — are the matchups where Power Score explains the least variance in outcomes and individual deck characteristics carry more weight. The Temur Harmonizer Combo Standard correlation (ρ = 0.88) is the strongest in either format. Power Score is an internal composite; these correlations describe how well it ranks decks by outcome against each specific opponent in this dataset.


15. Per-Deck Appendix (Anonymized)

Privacy note: Pseudonyms (C-001 through C-098, S-001 through S-010) are randomly reassigned every edition and carry no link to any prior edition's pseudonyms, any account, or any actual deck name or identifier. Published values are coarsened: Power Score rounded to the nearest 5, win rates and spreads to whole percentage points, game counts banded. Only derived statistics are shown. No deck names, deck lists, owners, or identifiers appear.

Spread definition: Standard deviation of a deck's per-opponent win rates across all qualifying matchup cells (3 or more opponents with sufficient games). Higher spread = more matchup-specialist profile; lower spread = more consistent across opponents. Spread is shown only where 3 or more qualifying cells exist.

15a. Commander per-deck appendix summary

98 anonymized Commander decks published (minimum 20 games per deck).

Metric Value
Decks published 98
Win rate: minimum 7%
Win rate: median 42%
Win rate: maximum 89%
Spread: median (n = 88 decks with 3+ opponents) 17 pts

15b. Commander per-deck table

ID PS (±5) Games Band Win Rate Opponents Faced Spread
C-001 65 20–49 63% 5 12
C-002 40 20–49 42% 5 23
C-003 55 20–49 57% 5 13
C-004 45 20–49 55% 1 —
C-005 0 50–99 7% 5 13
C-006 45 50–99 45% 5 34
C-007 35 20–49 37% 5 12
C-008 30 20–49 33% 5 24
C-009 85 20–49 80% 5 7
C-010 90 20–49 85% 1 —
C-011 30 20–49 33% 5 28
C-012 50 20–49 49% 5 15
C-013 25 20–49 25% 1 —
C-014 50 50–99 50% 6 18
C-015 35 20–49 38% 5 13
C-016 30 20–49 20% 5 12
C-017 30 50–99 34% 6 14
C-018 40 20–49 40% 5 17
C-019 70 20–49 67% 5 11
C-020 15 20–49 16% 5 17
C-021 50 20–49 51% 5 15
C-022 30 20–49 30% 5 24
C-023 30 20–49 30% 5 22
C-024 50 20–49 47% 5 26
C-025 40 20–49 40% 5 25
C-026 70 20–49 67% 5 18
C-027 45 20–49 44% 5 16
C-028 10 20–49 15% 1 —
C-029 30 20–49 30% 5 22
C-030 55 20–49 53% 5 12
C-031 20 20–49 25% 1 —
C-032 30 20–49 37% 6 15
C-033 55 20–49 53% 5 29
C-034 55 20–49 53% 5 19
C-035 60 20–49 60% 5 19
C-036 55 20–49 63% 5 16
C-037 25 20–49 27% 5 8
C-038 25 50–99 25% 5 7
C-039 15 20–49 20% 5 27
C-040 70 20–49 67% 5 16
C-041 70 20–49 67% 5 18
C-042 30 20–49 40% 1 —
C-043 70 20–49 70% 5 16
C-044 90 20–49 89% 5 12
C-045 25 50–99 27% 5 10
C-046 70 20–49 85% 1 —
C-047 65 20–49 70% 5 27
C-048 30 20–49 33% 5 21
C-049 40 20–49 40% 5 17
C-050 75 20–49 73% 5 13
C-051 40 100–199 42% 5 16
C-052 60 20–49 70% 1 —
C-053 40 20–49 40% 5 13
C-054 40 50–99 42% 5 12
C-055 50 20–49 50% 5 21
C-056 35 20–49 28% 5 17
C-057 60 20–49 60% 5 15
C-058 30 20–49 30% 5 22
C-059 30 20–49 29% 5 9
C-060 65 20–49 62% 5 11
C-061 20 20–49 23% 5 8
C-062 70 20–49 70% 5 12
C-063 25 20–49 27% 5 11
C-064 35 20–49 40% 6 16
C-065 35 50–99 35% 5 18
C-066 60 50–99 58% 5 12
C-067 80 20–49 77% 5 8
C-068 45 50–99 50% 6 19
C-069 40 20–49 40% 5 17
C-070 55 20–49 53% 5 16
C-071 45 50–99 43% 5 8
C-072 25 20–49 24% 5 11
C-073 25 50–99 33% 5 17
C-074 25 20–49 27% 5 25
C-075 70 20–49 70% 5 19
C-076 45 20–49 53% 5 18
C-077 55 20–49 53% 5 19
C-078 45 20–49 47% 5 16
C-079 45 20–49 49% 6 16
C-080 35 20–49 37% 5 29
C-081 60 20–49 70% 1 —
C-082 5 20–49 12% 5 9
C-083 75 20–49 73% 5 17
C-084 35 20–49 30% 5 22
C-085 45 20–49 47% 5 24
C-086 45 20–49 47% 5 31
C-087 55 20–49 60% 5 20
C-088 35 20–49 37% 5 19
C-089 30 20–49 29% 5 15
C-090 25 20–49 27% 5 31
C-091 40 20–49 43% 5 17
C-092 20 20–49 20% 5 12
C-093 30 20–49 35% 1 —
C-094 30 20–49 33% 5 16
C-095 30 20–49 31% 5 26
C-096 20 20–49 19% 5 16
C-097 65 20–49 57% 5 17
C-098 30 20–49 30% 5 22

Spread (—) = fewer than 3 qualifying opponents; not published. PS rounded to nearest 5. Win rate and spread rounded to whole points. Game counts banded.

15c. Outlier callout: OLS residuals vs. Power Score (Commander, n = 98 decks)

An OLS line fit to the 98 published Commander per-deck rows relates coarsened Power Score to win rate. Two decks stand at the widest distances from that fitted line:

Widest overperformance — C-046: PS 70 (coarsened), win rate 85%, OLS-expected win rate 69%. Residual: +16 percentage points. This deck won 16 points more than a PS-70 deck would be expected to win based on the linear relationship across the published sample.

Widest underperformance — C-016: PS 30 (coarsened), win rate 20%, OLS-expected win rate 32%. Residual: −12 percentage points. This deck won 12 points less than the PS-30 expectation.

Note: C-046 faced only 1 opponent (spread not published), so its win rate is based on a narrower sample than most published entries. C-016 faced 5 opponents with a 12-point spread, indicating modest matchup variation. These outliers are reported as observed deviations from the PS–win-rate relationship in the published appendix data — no causal interpretation is offered, and pseudonyms rotate every edition.

15d. Standard per-deck appendix summary

10 anonymized Standard decks published (minimum 20 games per deck; the appendix requires ≥10 published decks, exactly met).

Metric Value
Decks published 10
Win rate: minimum 27%
Win rate: median 40%
Win rate: maximum 70%
Spread: median (n = 10 decks with 3+ opponents) 23 pts

15e. Standard per-deck table

ID PS (±5) Games Band Win Rate Opponents Faced Spread
S-001 60 20–49 60% 5 34
S-002 35 20–49 37% 5 27
S-003 30 20–49 30% 5 19
S-004 70 20–49 67% 5 21
S-005 25 20–49 27% 5 23
S-006 40 20–49 40% 5 23
S-007 70 20–49 70% 5 27
S-008 35 20–49 38% 5 17
S-009 30 20–49 33% 5 18
S-010 50 20–49 50% 5 24

All 10 Standard decks face 5 opponents and qualify for spread calculation.

15f. Outlier callout: OLS residuals vs. Power Score (Standard, n = 10 decks)

The Standard per-deck sample (n = 10) is at the minimum publication floor. OLS residuals at this sample size carry high uncertainty and should be read with that in mind.

Widest overperformance — S-008: PS 35 (coarsened), win rate 38%, OLS-expected win rate 36%. Residual: +2 percentage points.

Widest underperformance — S-004: PS 70 (coarsened), win rate 67%, OLS-expected win rate 69%. Residual: −2 percentage points.

The Standard residuals are small in absolute terms (±2 points), reflecting the compressed win-rate range (27%–70%) and the small sample. No meaningful outlier narrative is supported at this sample size; the callout is included for structural completeness and comparability with future editions.


16. Optional: Power Score Composite (Secondary, Caveated)

⚠️ Power Score is an internal Grim.Cards composite indicator, not an objective or universal measurement of deck quality or win probability. It is not used to rank or judge decks elsewhere in this report. It is presented here in a single, clearly-labelled section for context only.

Power Score combines multiple construction and card-quality signals into a single composite value. The formula and weighting are internal to Grim.Cards and are subject to change. A higher Power Score does not guarantee a higher win rate, as demonstrated throughout this report — the specific Gauntlet opponent explains as much or more variance than Power Score in some matchup structures (see Section 14a).

Grade distribution among completed simulations

Grades are binned from the underlying continuous Power Score. Counts below are the number of completed user-deck simulations (n = 691 total across both formats) in each grade band, not unique decks.

Grade Simulations
S 3
A 20
B 56
C 83
D 143
F 202

The distribution is bottom-heavy: F and D together account for 345 of 507 scored decks — the majority of submitted decks score in the lower two grade bands by this composite. S-grade decks are rare (3). This distribution reflects both the composition of the submitted deck pool and the internal calibration of the Power Score composite, not an objective assessment of deck quality. Power Score and letter grades are Grim.Cards' own internal framing; they should not be read as authoritative verdicts.


17. Key Findings

The following findings summarize what the data in this edition directly supports. Each is derived from the measured figures reported above, carried with its sample size.

1. The Gauntlet opponent is a primary driver of outcomes. The same pool of 447 Commander decks wins 53.5% against Breya Artifact Combo and 31.0% against Edgar Markov Vampires — a 22.5-point swing with no change to the deck slate. In Standard, the spread is 45.4 points (60.4% vs. Temur Harmonizer Combo, 15.0% vs. Mono Red Aggro, n = 58 decks each). The specific opponent a deck faces in the Gauntlet is at least as consequential as the deck's own construction profile.

2. Win rates are broadly distributed, with substantial lower tails. Commander win rates span 0% to 93.3% (median 40%, mean 43.4%, n = 448 decks). The 0–30% zone holds 110 decks (24.6% of the Commander sample). Standard spans 0% to 80% (median 33.3%, mean 35.3%, n = 59 decks), with 42.4% of Standard decks in the 0–30% zone. The majority of submitted decks lose more games than they win against the meta Gauntlet.

3. Commander retests trend positive; Standard retests are inconclusive. Among 106 retested Commander decks, 55 improved win rate, 36 declined, 15 were flat — an average change of +2.6 percentage points. Among 12 retested Standard decks (at the minimum publication threshold), the split was 5 improved / 5 declined / 2 flat, averaging +0.6 points.

4. In Commander, deck identity explains more win-rate variance than opponent identity; in Standard, the reverse is true. The games-weighted variance decomposition assigns 46.9% of Commander win-rate variance to the deck main effect and 16.8% to the opponent (36.3% residual; n = 98 decks, 20 opponents). In Standard, opponent explains 43.9% versus 27.8% for the deck (28.3% residual; n = 10 decks, 5 opponents). The Standard estimate carries high uncertainty at n = 10.

5. Mono-color decks outperform multi-color in this dataset. Commander mono-color decks average 50.8% win rate (n = 86) versus 40.1–42.9% for two-, three-, and five-color decks. Standard mono-color decks average 44.2% (n = 11) versus 36.7% for two-color (n = 37). Correlation only — archetype and Gauntlet composition are likely confounders.

6. Higher land ratios are associated with higher win rates. Commander decks near 40% lands average 46.7% (n = 112) versus 40.3% for decks near 30% lands (n = 66). Standard shows the same direction: 43.1% at ~40% lands (n = 26) versus 30.6% at ~35% lands (n = 18).

7. Board-impact leaders in Commander are dominated by large threats; the laggards include widely-played staples. Ureni of the Unwritten leads Commander board impact at +29.78 per observation (11 decks, 253 obs). Dictate of Erebos trails at −13.72 (25 decks, 174 obs) — the most credible negative figure by deck breadth. Negative board-impact figures do not constitute a judgment on card quality; they describe what the board-state proxy registers in the game windows the engine captures.

8. Functional card categories correlate negatively with Standard win rates. Standard decks including reanimation effects win 15.7 points lower than those without (23.6%, n = 15 vs. 39.3%, n = 44). Tutor-inclusive Standard decks trail by 11.6 points (28.8%, n = 26 vs. 40.4%, n = 33). All category assignments are heuristic; these correlations likely reflect archetype identity rather than the individual effects of the card categories.

9. The most-played Commander card (Sol Ring, 387 decks) clusters near the overall average win rate. Sol Ring-containing decks win at 43.4% (n = 387 decks, 7,823 games) — essentially identical to the Commander aggregate (43.2%). Ubiquitous staples are distributed across all power levels and do not mark high- or low-performing archetypes in this dataset.

10. Meren of Clan Nel Toth is the only commander with a publishable win rate; it posts 33.6%. With 17 decks and 360 games, Meren of Clan Nel Toth is the most-played commander in the dataset and the only one clearing the 10-deck minimum. Its 33.6% win rate is approximately 9.6 points below the Commander average. All other commanders in the dataset have fewer than 10 distinct decks; their win rates are withheld.


18. What We Cannot Conclude Yet

The following questions are raised by the data but cannot be answered from this snapshot:

  • Does changing a deck improve its win rate? The retest data shows +2.6 points on average in Commander (n = 106), but the dataset does not record whether decks were modified between tests. The direction of change cannot be attributed to iteration vs. random variation.
  • Are commander-level win-rate differences real? Only Meren of Clan Nel Toth clears the 10-deck floor. Commander-level win-rate comparisons require a larger dataset.
  • Do negative board-impact figures indicate cards that drag down win rates? The board-impact metric captures board-state change in observation windows, not win attribution. Cards with negative avgDelta may appear disproportionately in games that are already losing, not cause the losses.
  • Are the Standard category associations driven by the categories or by archetypes? Standard reanimation and tutor categories may simply flag the same slow-control decks that lose badly to Mono Red Aggro — the category itself may be incidental.
  • Would a different Gauntlet produce different rank-orderings? The inversion rate (36.6% in Commander) confirms that deck rankings are not stable across opponents. Results in this report are specific to the current Gauntlet composition.
  • Is the Standard variance decomposition reliable? With n = 10 decks in the Standard decomposition, the 43.9% / 27.8% opponent-vs-deck split is an estimate with high sampling uncertainty. Future editions with larger Standard cohorts will test whether this finding holds.

19. Future-Comparison Baseline

The following table provides the machine-readable baseline for comparing this edition against future editions. All figures are drawn directly from the dataset JSON for this snapshot.

Metric Value n (decks) n (games) Format
Overall win rate 42.6% 507 10,011 All
Commander win rate 43.2% 448 8,969 Commander
Standard win rate 37.1% 59 1,042 Standard
Commander median win rate 40.0% 448 — Commander
Standard median win rate 33.3% 59 — Standard
Highest Commander matchup win rate 53.5% (vs. Breya Artifact Combo) 447 1,773 Commander
Lowest Commander matchup win rate 31.0% (vs. Edgar Markov Vampires) 447 1,773 Commander
Highest Standard matchup win rate 60.4% (vs. Temur Harmonizer Combo) 58 207 Standard
Lowest Standard matchup win rate 15.0% (vs. Mono Red Aggro) 58 207 Standard
Commander retest avg change +2.6 pts 106 — Commander
Standard retest avg change +0.6 pts 12 — Standard
Commander avg game length 10.3 turns 448 — Commander
Standard avg game length 9.5 turns 59 — Standard
Top Commander board-impact card Ureni of the Unwritten, +29.78 11 253 obs Commander
Top Standard board-impact card Lightning Bolt, +5.05 10 95 obs Standard
Worst Commander board-impact card Dictate of Erebos, −13.72 25 174 obs Commander
Commander variance: deck share 46.9% 98 — Commander
Commander variance: opponent share 16.8% 98 — Commander
Standard variance: deck share 27.8% 10 — Standard
Standard variance: opponent share 43.9% 10 — Standard
Commander inversion rate 36.6% 88 25,242 pairs Commander
Standard inversion rate 23.0% 10 344 pairs Standard
Snapshot date 2026-08-01 — — —
Dataset version 2.4 — — —

Limitations (consolidated)

  1. All results are from AI-piloted simulations (a custom Forge build). They describe engine behavior on submitted decklists and cannot be generalized to human play or to Magic: The Gathering broadly.
  2. All construction-level, card-level, and category-level findings are correlational. No causal claim is made anywhere in this report.
  3. The sample is self-selected — decks users chose to submit to Grim.Cards, not a random sample of any broader population.
  4. Monthly win-rate comparisons across months reflect different deck populations in each month, not longitudinal progression.
  5. Power Score is an internal Grim.Cards composite, not an objective or universal measurement. It is not used for ranking or judgment in this report.
  6. Functional card category assignments (sacrifice, tutor, discard, reanimation) are heuristic and may contain labelling errors.
  7. Board-impact figures are a board-state proxy, not damage, kills, or causal win attribution. Negative figures do not establish that a card harms win probability.
  8. The Standard per-deck appendix (n = 10) and Standard variance decomposition (n = 10) are at minimum publication thresholds; estimates carry high sampling uncertainty.
  9. Counterfactual decision-impact data did not meet the cohort threshold in this edition; that section is omitted.
  10. Per-deck pseudonyms rotate every edition; no cross-edition deck tracking is possible by design.

License and Citation

License: This report and its associated data are published under the Creative Commons Attribution 4.0 International License (CC BY 4.0). You are free to share and adapt the material for any purpose, provided appropriate credit is given.

Suggested citation:

Grim.Cards. Grim.Cards Simulation Data — Case Study Edition, Snapshot 2026-08-01 (Dataset Version 2.4). Published 2026-08-01. Permanent URL: https://grim.cards/case-study/2026-08-01. License: CC BY 4.0. Temporal coverage: 2026-05-14 to 2026-08-01.

Data download: The machine-readable edition JSON (this snapshot's full structured data, including atomic statements, per-figure sample sizes, and provenance block) is available at: https://grim.cards/case-study/2026-08-01/data.json

Edition index: All published editions, newest-first: https://grim.cards/case-study

Retrieved-date note for citing parties: This permanent URL (/case-study/2026-08-01) will not change when newer editions are published. Future editions accumulate at their own dated URLs and do not overwrite this one. Record your retrieval date alongside the snapshot date when citing.


Report generated: 2026-08-01T08:16:13.651Z · Grim.Cards Dataset Version 2.4 · Cohort: player-submitted decks (Crucible and sample decks excluded) · Minimum cohort threshold: 10 unique decks per reported group · Win rate: wins ÷ total games played (draws in denominator) · All findings correlational · No PII published

All Grim.Cards case-study editions

Grim.Cards is an independent fan-created platform, not affiliated with, endorsed by, sponsored by, or approved by Wizards of the Coast LLC or Hasbro, Inc. Magic: The Gathering® is a trademark of Wizards of the Coast LLC.