Grim.Cards Simulation Data — Case Study Edition
Data Snapshot: 2026-08-01 · Dataset Version 2.4 · License: CC BY 4.0
Permanent URL: https://grim.cards/case-study/2026-08-01
Published: 2026-08-01 · Temporal Coverage: 2026-05-14 → 2026-08-01
Cohort: Real, human-submitted Commander and Standard decks. Automated Crucible decks and system sample decks are excluded throughout.
Minimum cohort threshold: 10 unique decks per reported group. Any group below this floor is suppressed.
Win rate definition (applied consistently throughout): wins ÷ total games played, with draws counted in the denominator.
All findings are correlational. Simulated (AI-piloted) results describe engine behavior on these decklists and cannot be generalized to human play or to Magic: The Gathering broadly.
Executive Summary
Between 14 May and 1 August 2026, Grim.Cards ran 691 completed simulations of player-submitted decks across 507 unique decks from 207 distinct users, generating 10,011 AI-versus-AI games. The headline number: player-submitted decks won 42.6% of those games (n = 507 decks; 10,011 games). That figure sits below 50%, which is structurally expected — each deck faces a fixed gauntlet of purpose-built meta opponents, so the field as a whole is designed to win more often than any single challenger. The sub-50% aggregate is a baseline, not a verdict.
The more revealing numbers are in the matchup spreads. In Commander (n = 448 decks; 8,969 games; overall win rate 43.2%), the difference between the easiest and hardest Gauntlet opponent in this dataset is 22.5 percentage points — from 53.5% against Breya Artifact Combo down to 31.0% against Edgar Markov Vampires. In Standard (n = 59 decks; 1,042 games; overall win rate 37.1%), that spread balloons to 45.4 percentage points — from 60.4% against Temur Harmonizer Combo down to 15.0% against Mono Red Aggro. The opponent facing a deck is, in measured terms, at least as important a driver of outcome as the deck's own profile — and in Standard, the data suggests the opponent matters more.
Among 106 Commander decks retested, 55 improved their win rate across attempts and 36 declined, with an average change of +2.6 percentage points. Among 12 Standard decks retested, 5 improved and 5 declined, averaging +0.6 points.
The board-performance data — a measured proxy for how much a card improves the game state when it appears, not a causal win claim — shows Ureni of the Unwritten topping the Commander rankings at +29.78 board-quality points per appearance (11 decks, 253 observations), while Dictate of Erebos leads the bottom list at −13.72 (25 decks, 174 observations). In Standard, Lightning Bolt is the sole card meeting the ranking threshold at +5.05 (10 decks, 95 observations).
Power Score, an internal Grim.Cards composite indicator, is presented last and caveated. It is not an objective measurement and is not used to rank or judge decks anywhere in this report.
1. Methodology and Provenance
Simulation method
Each player-submitted deck is played by a custom build of the open-source Forge engine — an AI piloting the deck against a fixed set of AI-piloted Gauntlet opponents. No human decisions are involved during simulation. Results describe how the engine performs with these decklists; they do not predict human play outcomes.
Cohort definition
The primary cohort is every deck submitted by a real user account, excluding the automated Crucible reference corpus (user_id = '__grinder__') and system sample decks (is_sample = true). The Crucible corpus is not counted in user or deck totals and does not appear in any reported rate.
Win rate
Win rate is defined as wins ÷ total games played, with draws counted in the denominator. This definition is applied identically across every table, chart, and rate in this report. "Decisive win rate" (wins ÷ wins + losses only) is never used.
Format splits
Commander and Standard results are reported separately in every table and section. The simulation database carries no format column; format is derived by joining simulation records to deck records. Results are never pooled across formats.
Cohort threshold
No rate or breakdown is published for a group of fewer than 10 unique decks. Groups below this floor are suppressed or merged into an "Other" category. Where a group's "without" contrast group falls below 10 decks, that contrast figure is withheld (shown as suppressed).
Retest identification
A retest is defined as a deck with more than one completed simulation on record. The reported change in win rate is the per-deck win rate of the most recent completed simulation minus the win rate of the first completed simulation.
Card categories (heuristic — labelled throughout)
Functional card categories (sacrifice outlets, search/tutor effects, discard effects, reanimation effects) are assigned by a keyword heuristic applied to oracle text and pre-existing preference flags. Category assignment may mislabel edge cases. All category-level win rates carry a visible heuristic label and are correlation-only.
Board-impact metric (measured proxy — labelled throughout)
The card board-impact figure (avgDelta) is the pooled board-quality change per observation across all cohort decks running a card — positive means the board state improved in turns the card was seen. This is a board-state proxy, not a damage metric, not a kill count, and not a causal win claim. Only cards appearing in 10 or more distinct decks with 25 or more recorded observations are ranked.
Decision-impact / counterfactual metric
Counterfactual decision-impact figures — where reported — measure how much a card's chosen plays beat the engine's own next-best alternative line. Positive means the line played was better than the alternative; negative means the alternative was better. No counterfactual data met the cohort threshold for this edition; the section is omitted rather than suppressed at an uninformative level.
Privacy
No player names, deck names, deck lists, user identifiers, or user-level timestamps appear anywhere in this report. Per-deck appendix entries use rotating pseudonyms (e.g., C-001, S-001) reshuffled every edition — a pseudonym in this edition carries no link to any pseudonym in a prior edition. Published per-deck values are deliberately coarsened: Power Score rounded to the nearest 5, win rates and matchup spreads to whole percentage points, game counts banded.
Limitations
- Results are from AI-piloted simulations, not human play.
- All construction and category breakdowns are correlational — no causal claims are made or implied.
- The sample is self-selected: it comprises decks that users chose to submit to Grim.Cards, not a random sample of the broader Magic: The Gathering population.
- Monthly cohorts consist of different sets of decks month to month, so monthly win-rate trends reflect who submitted, not any single deck's trajectory.
- Power Score is an internal composite indicator, not an objective or universal measurement.
- Functional card category assignments are heuristic and may carry labelling error.
- Board-impact figures are a board-state proxy, not damage or kills, and not a causal win contribution.
2. Dataset Overview
| Metric | All Formats | Commander | Standard |
|---|---|---|---|
| Distinct users | 207 | 179 | 40 |
| Distinct decks | 507 | 448 | 59 |
| Completed simulations | 691 | 619 | 72 |
| Total games | 10,011 | 8,969 | 1,042 |
| Wins | 4,261 | 3,874 | 387 |
| Losses | 5,747 | 5,092 | 655 |
| Draws | 3 | 3 | 0 |
| Win rate | 42.6% | 43.2% | 37.1% |
| Data range | 2026-05-14 → 2026-08-01 | 2026-05-14 → 2026-08-01 | 2026-05-14 → 2026-08-01 |
Commander is the dominant format in this dataset by every measure: 88% of decks, 90% of games, and 90% of completed simulations. Standard figures represent a meaningfully smaller sample and should be read with that in mind throughout.
3. The Gauntlet: Matchup Results
This is the core section of the report. Every other metric — distribution, retests, construction, board impact — is downstream of what happens in these matchups. The Gauntlet is a fixed set of AI-piloted meta opponents; every player deck faces the same field under identical conditions.
3a. Commander Gauntlet (n = 447 decks per matchup)
Each matchup row below represents 447 player-submitted Commander decks tested against that opponent. All five cohorts clear the 10-deck minimum by a wide margin.
| Opponent | Decks (n) | Wins | Losses | Draws | Games | Win Rate |
|---|---|---|---|---|---|---|
| Breya Artifact Combo | 447 | 948 | 825 | 0 | 1,773 | 53.5% |
| Derevi Bant Control | 447 | 895 | 879 | 0 | 1,774 | 50.5% |
| Aesi Landfall | 447 | 823 | 951 | 2 | 1,776 | 46.3% |
| Atraxa Superfriends | 447 | 607 | 1,170 | 1 | 1,778 | 34.1% |
| Edgar Markov Vampires | 447 | 549 | 1,224 | 0 | 1,773 | 31.0% |
Range across opponents: 22.5 percentage points. The same pool of 447 player decks wins more than half their games against Breya Artifact Combo and fewer than one-third against Edgar Markov Vampires. The deck slate did not change between those two rows — only the opponent did. That 22.5-point swing is a direct, measured signal that the specific meta opponent is a substantial driver of matchup outcome within this dataset.
The midpoint of the Gauntlet — Aesi Landfall at 46.3% — sits just below even odds, which is consistent with a well-calibrated field. Players hold slight overall edges against the two "softer" opponents and are clearly disadvantaged against the two "harder" opponents, producing the 43.2% aggregate.
3b. Standard Gauntlet (n = 58 decks per matchup)
Each matchup row below represents 58 player-submitted Standard decks. All five cohorts clear the 10-deck minimum.
| Opponent | Decks (n) | Wins | Losses | Draws | Games | Win Rate |
|---|---|---|---|---|---|---|
| Temur Harmonizer Combo | 58 | 125 | 82 | 0 | 207 | 60.4% |
| Jeskai Control | 58 | 96 | 111 | 0 | 207 | 46.4% |
| Azorius Tempo | 58 | 69 | 138 | 0 | 207 | 33.3% |
| Dimir Midrange | 58 | 64 | 143 | 0 | 207 | 30.9% |
| Mono Red Aggro | 58 | 31 | 176 | 0 | 207 | 15.0% |
Range across opponents: 45.4 percentage points. The Standard spread is twice as wide as Commander's. Against Temur Harmonizer Combo, the average player Standard deck wins three games in five. Against Mono Red Aggro, that same pool of decks wins fewer than one game in seven. Standard decks in this dataset struggle broadly — a 37.1% overall win rate against a gauntlet that includes a 60.4% matchup and a 15.0% matchup simultaneously — and the variance in outcomes is almost entirely a function of which opponent is across the table.
The severity of the Mono Red Aggro result (15.0%, 31 wins in 207 games) warrants context. The Standard cohort (n = 59 decks) skews toward midrange, control, and combo archetypes — the decks users submitted. A fast aggro opponent that routinely closes games in seven turns or fewer (the Standard average shortest game is 7.3 turns) is a particularly punishing test for the submitted deck pool, and this dataset measures exactly that.
4. Top Commanders (Commander, ≥10-deck cohort floor applied)
Only one commander clears the 10-deck minimum required to report a win rate in this dataset: Meren of Clan Nel Toth, with 17 decks, 360 games, and a win rate of 33.6%. All other commanders listed in the dataset have fewer than 10 distinct decks and their win rates are withheld per the cohort threshold.
| Commander | Decks (n) | Games | Win Rate |
|---|---|---|---|
| Meren of Clan Nel Toth | 17 | 360 | 33.6% |
| All others | < 10 each | — | Suppressed |
Meren of Clan Nel Toth is the most-played commander in the Grim.Cards dataset by a meaningful margin (17 decks versus the next cluster at 8 or fewer). Its 33.6% win rate sits roughly 10 points below the Commander average of 43.2%, but the spread across the Gauntlet — Meren decks face all five Commander opponents — makes this an aggregate figure only. No per-matchup breakdown for a single commander can be published at this cohort size without falling below the per-cell floor.
The full top-commander list by deck count (win rates suppressed below the threshold) in order of popularity: Meren of Clan Nel Toth (17), Ureni of the Unwritten (8), The Ur-Dragon (6), Emmara, Soul of the Accord (5), Atraxa, Praetors' Voice (5), Kaalia of the Vast (5), Vivi Ornitier (5), Zimone, Infinite Analyst (5), Y'shtola, Night's Blessed (5). Commanders with 4 or fewer decks: Fire Lord Azula, Marrow-Gnawer, Giada Font of Hope, Kenrith the Returned King, Tony Stark, Krenko Mob Boss, Pantlaza Sun-Favored, Atarka World Render, Cosmic Spider-Man, and others. All sub-10 win rates are suppressed.
The commander diversity in this dataset is striking. Nineteen distinct commanders appear in the top-20 list and none has accumulated enough decks to generate a publishable per-matchup breakdown beyond Meren. Future editions with a larger sample may unlock commander-level matchup reporting.
5. Top Cards by Usage and by Containing-Deck Win Rate
5a. Commander — Most-played cards (basic lands excluded)
Ranked by number of distinct decks the card appears in. Win rate shown is the aggregate win rate of all decks containing that card (not the card's individual contribution). Correlation only — popular cards appear in popular archetypes, which share many variables.
| Card | Decks (n) | Games | Win Rate of Containing Decks |
|---|---|---|---|
| Sol Ring | 387 | 7,823 | 43.4% |
| Command Tower | 321 | 6,448 | 42.0% |
| Arcane Signet | 312 | 6,455 | 44.0% |
| Exotic Orchard | 167 | 3,404 | 43.9% |
| Lightning Greaves | 139 | 2,695 | 45.1% |
| Swiftfoot Boots | 137 | 2,497 | 43.6% |
| Reliquary Tower | 124 | 2,406 | 39.6% |
| Path of Ancestry | 117 | 2,218 | 49.1% |
| Evolving Wilds | 107 | 2,177 | 44.1% |
| Swords to Plowshares | 106 | 2,016 | 45.6% |
| Cultivate | 93 | 1,646 | 47.4% |
| Fellwar Stone | 81 | 1,799 | 44.5% |
| Rogue's Passage | 80 | 1,579 | 43.9% |
| Bojuka Bog | 80 | 1,785 | 40.6% |
| Path to Exile | 78 | 1,503 | 47.5% |
| Terramorphic Expanse | 75 | 1,577 | 41.6% |
| Skullclamp | 73 | 1,550 | 40.1% |
| Counterspell | 72 | 1,590 | 40.4% |
| Birds of Paradise | 67 | 1,359 | 39.3% |
| Demonic Tutor | 65 | 1,379 | 43.3% |
The most-played cards in Commander are unsurprisingly format staples: Sol Ring appears in 387 of 448 decks (86%). The containing-deck win rates cluster tightly between 39% and 49% for the top-20 list, reflecting that these cards are distributed broadly across decks of all power levels rather than concentrated in high-performing archetypes. Path of Ancestry (49.1%, n = 117) and Path to Exile (47.5%, n = 78) sit at the upper end; Birds of Paradise (39.3%, n = 67) and Skullclamp (40.1%, n = 73) at the lower end. These differences are small relative to the Gauntlet-opponent spread of 22.5 points, reinforcing that opponent identity is a larger driver of outcomes than any single card presence.
5b. Standard — Most-played cards (basic lands excluded, ≥10-deck cohort floor applied)
| Card | Decks (n) | Games | Win Rate of Containing Decks |
|---|---|---|---|
| Lightning Bolt | 13 | 195 | 46.2% |
| Inspiring Vantage | 12 | 180 | 49.4% |
| Shard Volley | 11 | 165 | 50.9% |
| Chain Lightning | 11 | 165 | 50.9% |
| Eidolon of the Great Revel | 11 | 165 | 50.9% |
| Monastery Swiftspear | 11 | 165 | 50.9% |
| Skullcrack | 11 | 165 | 50.9% |
| Goblin Guide | 11 | 165 | 50.9% |
| Lava Spike | 11 | 165 | 50.9% |
| Searing Blaze | 11 | 165 | 50.9% |
| Rift Bolt | 11 | 165 | 50.9% |
Standard cards below 10 decks: Sol Ring (7), Evolving Wilds (6), Llanowar Elves (6), Dark Ritual (5), Arcane Signet (5), Reliquary Tower (4), Abrade (4), Mind Stone (4), Rogue's Passage (4) — win rates suppressed.
The Standard card list tells a concentrated story. Ten of the eleven qualifying cards are core Mono Red Burn package cards (Shard Volley, Chain Lightning, Eidolon of the Great Revel, Monastery Swiftspear, Skullcrack, Goblin Guide, Lava Spike, Searing Blaze, Rift Bolt, and Inspiring Vantage), each appearing in 11 of the 59 Standard decks. All share an identical win rate of 50.9% across 165 games, which reflects that they appear together as a block in the same burn archetype rather than independently. These decks perform above the 37.1% Standard average, but whether the cards cause that overperformance or simply identify the archetype cannot be determined from these numbers alone. Lightning Bolt (13 decks, 46.2%) is the most widely distributed Standard card meeting the threshold.
6. Win-Rate Distribution
How were simulated win rates distributed across the submitted deck pool? The distributions reveal not just the average but the shape of outcomes — including how many decks are clustered near zero, how many exceed 50%, and how extreme the tails are.
6a. Commander (n = 448 decks)
| Win-Rate Band | Decks |
|---|---|
| 0–10% | 15 |
| 10–20% | 43 |
| 20–30% | 52 |
| 30–40% | 130 |
| 40–50% | 60 |
| 50–60% | 68 |
| 60–70% | 44 |
| 70–80% | 26 |
| 80–90% | 9 |
| 90–100% | 1 |
Summary (n = 448): Mean 43.4% · Median 40% · Min 0% · Max 93.3%
The Commander distribution is left-skewed. The largest single band is 30–40% (130 decks, 29% of the sample), and the modal experience is finishing just below the halfway mark. The 0–30% zone holds 110 decks (24.6% of the sample) — a meaningful minority that consistently lost more than 70% of their games against the Gauntlet. Meanwhile, 79 decks (17.6%) achieved win rates above 60%, and one deck reached 93.3%. The mean (43.4%) sits above the median (40%) precisely because of those upper-tail outliers, a hallmark of positive skew. For most Commander decks in this dataset, the simulated experience is one of losing more than winning — but a substantial minority performs well above the overall average.
6b. Standard (n = 59 decks)
| Win-Rate Band | Decks |
|---|---|
| 0–10% | 7 |
| 10–20% | 9 |
| 20–30% | 9 |
| 30–40% | 16 |
| 40–50% | 3 |
| 50–60% | 9 |
| 60–70% | 4 |
| 70–80% | 2 |
| 80–90% | 0 |
| 90–100% | 0 |
Summary (n = 59): Mean 35.3% · Median 33.3% · Min 0% · Max 80%
Standard's distribution is also left-skewed, but more severely. The 0–30% zone holds 25 decks (42.4% of the sample). The 40–50% band — the natural "roughly even" range — contains only 3 decks. No Standard deck in this dataset reached 80%, and the 70–80% band holds just 2. Standard decks face a more polarized Gauntlet: the Temur Harmonizer Combo matchup is favorable for many, but Mono Red Aggro is brutal for almost all, and the two effects do not cancel each other evenly. The result is a distribution where the center of mass sits in the 30s and the upper tail is modest.
7. Deck Iteration and Retests
Users who submitted the same deck for multiple test runs provide a direct, within-deck look at whether repeated testing correlates with win-rate change. All comparisons are first completed simulation vs. most recent completed simulation for a given deck.
7a. Commander retests (n = 106 retested decks)
| Outcome | Decks |
|---|---|
| Improved win rate | 55 |
| Declined win rate | 36 |
| Roughly flat | 15 |
| Average win-rate change | +2.6 percentage points |
Among 106 Commander decks retested, 55 showed an improvement and 36 a decline — a 60/40 split in favor of improvement among decks whose direction could be determined. The average change across all 106 retested decks is +2.6 points. The "flat" category (15 decks) represents cases where win rate changed negligibly between test instances.
What this does not tell us: the dataset does not reveal whether decks changed between tests, how much time elapsed, or what drove the directional change. Monthly cohorts are also different collections of decks, so the +2.6 average is not a before-and-after track for the same population — it is a deck-level statistic for the specific subset of decks tested more than once.
7b. Standard retests (n = 12 retested decks)
| Outcome | Decks |
|---|---|
| Improved win rate | 5 |
| Declined win rate | 5 |
| Roughly flat | 2 |
| Average win-rate change | +0.6 percentage points |
The Standard retest cohort (n = 12) is at the minimum publication threshold and warrants caution. The 5/5/2 split — perfect symmetry between improved and declined — and the near-zero average change of +0.6 points suggest no systematic directional pattern in this small sample. Larger Standard retest cohorts in future editions will clarify whether the near-zero result persists.
8. Monthly Trends
Monthly test volume and average win rate by format. Monthly cohorts consist of different decks submitted in that calendar month, not a longitudinal track of the same decks. Win-rate changes across months reflect who submitted, not progression of any individual deck.
| Month | Format | Sims | Decks | Games | Win Rate |
|---|---|---|---|---|---|
| 2026-05 | Commander | 74 | 51 | 1,056 | 40.4% |
| 2026-06 | Commander | 149 | 110 | 2,137 | 43.8% |
| 2026-06 | Standard | 24 | 20 | 360 | 37.5% |
| 2026-07 | Commander | 384 | 280 | 5,596 | 43.4% |
| 2026-07 | Standard | 38 | 30 | 532 | 36.1% |
Growth in Commander testing is the dominant story: 74 simulations in May, 149 in June (+101%), and 384 in July (+158% month-over-month). July alone accounts for 55.5% of all Commander simulations in the dataset. Standard debuted in the data in June (24 sims) and grew to 38 in July.
Commander win rates moved from 40.4% in May to 43.8% in June and held at 43.4% in July — a modest 3-point improvement from the first to the most recent full month. Standard win rates moved from 37.5% in June to 36.1% in July, a small decline. Both shifts are small relative to the Gauntlet-opponent spread, and both reflect different deck populations month to month, so directional interpretation is limited.
Standard data only begins in June because no Standard simulations meeting the reporting threshold were recorded in May. The May Standard cohort, if any existed, fell below the 10-deck minimum.
9. Construction and Win Rate
What construction characteristics are associated with different win rates? All findings here are strictly correlational — the direction of any relationship could run in either direction, and numerous confounders (archetype, commander choice, submission bias) cannot be controlled in this observational data.
9a. Color count and win rate
Commander (all cohorts ≥ 10 decks):
| Colors in Identity | Decks (n) | Win Rate |
|---|---|---|
| 1 color | 86 | 50.8% |
| 2 colors | 165 | 40.1% |
| 3 colors | 141 | 42.9% |
| 5 colors | 45 | 42.3% |
The 4-color group did not clear the 10-deck minimum and is suppressed. Mono-color Commander decks post the highest average win rate in this dataset (50.8%, n = 86), 8–10 points above all multi-color groups. This is a descriptive finding only. Mono-color decks may be concentrated in higher-performing archetypes (e.g., mono-Red aggro commanders, mono-White token strategies), and the Gauntlet opponents are all multi-color — any structural advantage for mono-color against these specific opponents is a property of this dataset's gauntlet composition, not a general claim about Magic.
Standard (cohorts meeting the ≥10-deck floor):
| Colors | Decks (n) | Win Rate |
|---|---|---|
| 1 color | 11 | 44.2% |
| 2 colors | 37 | 36.7% |
The 3-color-plus Standard cohorts fall below the floor and are suppressed. The pattern echoes Commander: mono-color Standard decks post a higher win rate (44.2%, n = 11) than two-color decks (36.7%, n = 37). The mono-color group likely overlaps heavily with the burn archetype identified in the card usage data.
9b. Land ratio and win rate
Commander:
| Land % (approximate) | Decks (n) | Win Rate |
|---|---|---|
| ~30% | 66 | 40.3% |
| ~35% | 257 | 43.2% |
| ~40% | 112 | 46.7% |
Higher land ratios are associated with higher win rates in Commander: decks near 40% lands average 46.7% (n = 112) versus 40.3% (n = 66) for decks near 30% lands — a 6.4-point difference. This pattern is consistent with the win-bracket analysis (Section 9d) but does not indicate causation.
Standard:
| Land % (approximate) | Decks (n) | Win Rate |
|---|---|---|
| ~35% | 18 | 30.6% |
| ~40% | 26 | 43.1% |
The same directional pattern holds in Standard: decks near 40% lands average 43.1% (n = 26) versus 30.6% (n = 18) for decks near 35% lands — a 12.5-point difference. Both Standard cells clear the 10-deck floor. The burn archetype's land-light construction (burn spells substitute for land-hungry plays) may concentrate lower-land-ratio decks toward specific archetypes.
9c. Color-identity presence and win rate
Color presence is not exclusive — most decks contain multiple colors. These are not independent comparisons.
Commander:
| Color | Decks (n) | Games | Win Rate |
|---|---|---|---|
| Red | 205 | 3,982 | 44.1% |
| Green | 227 | 4,366 | 43.5% |
| White | 221 | 4,451 | 43.1% |
| Blue | 209 | 4,258 | 41.7% |
| Black | 238 | 4,870 | 41.1% |
Red-inclusive decks post the highest average win rate (44.1%, n = 205) and Black-inclusive decks the lowest (41.1%, n = 238) among Commander decks in this dataset. The range is narrow — under 3 points — and all colors are heavily represented (205–238 decks each). Because most Commander decks run two or more colors, these figures heavily overlap and cannot be interpreted as independent color contributions.
Standard:
| Color | Decks (n) | Games | Win Rate |
|---|---|---|---|
| Red | 26 | 409 | 36.9% |
| Green | 17 | 345 | 36.5% |
| White | 32 | 540 | 36.1% |
| Blue | 18 | 319 | 30.4% |
| Black | 25 | 427 | 28.1% |
Standard shows a wider spread: Red-inclusive decks average 36.9% (n = 26) and Black-inclusive decks 28.1% (n = 25) — an 8.8-point gap. The burn archetype identified in Section 5b is Red-based, which may lift the Red aggregate. Blue and Black decks in this Standard cohort post the weakest results, though all cohorts are smaller than in Commander and the Standard Gauntlet includes a severe Mono Red Aggro matchup that may disproportionately affect control and midrange archetypes (the decks most likely to run Blue and Black).
9d. Win-bracket comparison: high, mid, and low win-rate decks
Decks grouped by win rate (high = above 55%, mid = 40–55%, low = below 40%) and compared on average construction features. Brackets must clear the 10-deck floor.
Commander:
| Bracket | Decks (n) | Avg Win Rate | Avg Land % | Avg Creature % | Avg Spell % | Avg Art/Ench % | Avg Mana Value |
|---|---|---|---|---|---|---|---|
| High (>55%) | 114 | 68.6% | 36% | 28.6% | 16.0% | 18.5% | 3.31 |
| Mid (40–55%) | 161 | 45.2% | 35.9% | 28.0% | 17.9% | 17.5% | 3.18 |
| Low (<40%) | 173 | 25.1% | 35.0% | 27.2% | 19.6% | 17.4% | 3.08 |
Across Commander brackets, the high-win-rate group (n = 114) runs approximately 1 percentage point more lands than the low-win-rate group (n = 173) and carries an average mana value 0.23 higher (3.31 vs. 3.08). Spell percentage moves in the opposite direction: low-win-rate decks carry more instants and sorceries on average (19.6%) than high-win-rate decks (16.0%). Creature and artifact/enchantment percentages show smaller differences across brackets. These are descriptive correlations across submitted decks, not construction prescriptions.
Standard:
| Bracket | Decks (n) | Avg Win Rate | Avg Land % | Avg Creature % | Avg Spell % | Avg Art/Ench % | Avg Mana Value |
|---|---|---|---|---|---|---|---|
| Mid (40–55%) | 18 | 46.5% | 36.8% | 27.2% | 31.1% | 5.3% | 2.09 |
| Low (<40%) | 33 | 21.3% | 34.7% | 29.9% | 18.7% | 15.8% | 2.75 |
The high-win-rate Standard bracket (above 55%) did not clear the 10-deck minimum and is suppressed. Among the two qualifying brackets, mid-win-rate decks (n = 18) run more spells on average (31.1%) and carry lower average mana values (2.09) than low-win-rate decks (n = 33, 18.7% spells, 2.75 mana value). The Standard mid-bracket's higher spell density likely reflects the burn archetype's composition — many low-mana-value instants and sorceries — while low-win-rate decks carry more artifacts and enchantments and higher average mana costs. Again, correlational only.
9e. Card-type mix (full sample)
The card-type distribution across all Commander decks (n = 448, all with split data available):
| Type | Share of Deck |
|---|---|
| Lands (nonbasic counted separately in board impact) | 35.6% |
| Creatures | 27.9% |
| Instants + Sorceries (combined) | 18.1% |
| — Instants only | 10.0% |
| — Sorceries only | 8.0% |
| Artifacts | 10.2% |
| Enchantments | 7.5% |
| Planeswalkers | 0.7% |
Standard (n = 59, all with split data):
| Type | Share of Deck |
|---|---|
| Lands | 35.6% |
| Creatures | 29.2% |
| Instants + Sorceries (combined) | 23.4% |
| — Instants only | 13.5% |
| — Sorceries only | 9.6% |
| Artifacts | 5.4% |
| Enchantments | 6.0% |
| Planeswalkers | 0.6% |
Instant and sorcery density is higher in Standard (23.4%) than Commander (18.1%), consistent with the format's more interactive, less permanent-focused game plan. Artifact density is nearly twice as high in Commander (10.2% vs. 5.4%), reflecting Commander's heavier use of mana rocks and equipment. Land percentages are identical across formats in this dataset (35.6% each).
10. Card-Category Insights
⚠️ Heuristic label: Functional category membership (sacrifice outlets, tutor/search effects, discard effects, reanimation effects) is assigned by a keyword heuristic applied to oracle text and pre-existing preference flags. Category assignment may mislabel edge cases. All figures in this section are correlational.
10a. Commander category win-rate comparisons
The "without" contrast group for sacrifice outlets (4 decks) is below the 10-deck minimum — delta and contrast win rate are suppressed.
| Category | With Decks (n) | With Win Rate | Without Decks (n) | Without Win Rate | Delta |
|---|---|---|---|---|---|
| Sacrifice outlets | 444 | 43.4% | 4 | Suppressed | Suppressed |
| Search / tutor effects | 427 | 43.5% | 21 | 42.1% | +1.4 pts |
| Discard effects | 385 | 42.8% | 63 | 47.2% | −4.4 pts |
| Reanimation effects | 274 | 42.3% | 174 | 45.1% | −2.8 pts |
The most notable finding: Commander decks including discard effects (n = 385) win at 42.8%, while those without (n = 63) win at 47.2% — a −4.4-point gap. Decks including reanimation effects (n = 274) also underperform those without (42.3% vs. 45.1%, −2.8 points). Tutor effects show a small positive association (+1.4 points with vs. without). Sacrifice outlets are near-universal (444 of 448 decks), so that comparison is not meaningful.
None of these associations can be read causally. Discard and reanimation effects may cluster in archetypes that face structural disadvantages against the specific Gauntlet opponents in this edition. The direction and magnitude of these associations may shift as the Gauntlet composition or submitted deck pool changes in future editions.
10b. Standard category win-rate comparisons
The "without" contrast group for sacrifice outlets (4 decks) falls below the 10-deck floor — suppressed.
| Category | With Decks (n) | With Win Rate | Without Decks (n) | Without Win Rate | Delta |
|---|---|---|---|---|---|
| Sacrifice outlets | 55 | 34.1% | 4 | Suppressed | Suppressed |
| Discard effects | 36 | 31.0% | 23 | 42.0% | −11.0 pts |
| Search / tutor effects | 26 | 28.8% | 33 | 40.4% | −11.6 pts |
| Reanimation effects | 15 | 23.6% | 44 | 39.3% | −15.7 pts |
Standard shows stronger negative associations across all three measurable categories. Decks including reanimation effects (n = 15) win at 23.6% — 15.7 points below those without (n = 44, 39.3%). Tutor-inclusive decks (n = 26) trail at 28.8% versus 40.4% for those without (n = 33). Discard-inclusive decks (n = 36) similarly trail (31.0% vs. 42.0%).
One likely explanation: in Standard, these effect categories (reanimation, tutors) tend to appear together in slower, higher-mana-value control and combo archetypes — precisely the deck types that the Standard Gauntlet, anchored by Mono Red Aggro, most severely disadvantages. The category associations almost certainly reflect archetype identity more than any property of the individual card effects. This interpretation is inferential, not measured.
11. Card Board-Impact Rankings
⚠️ Board-state proxy, not causal win claim: Board-impact figures (avgDelta) measure the pooled change in board quality around turns a card was seen, across all cohort decks containing that card. Positive means the board state improved; negative means it degraded relative to the surrounding game state. This is not a damage metric, kill count, or win-rate attribution. The sign convention for board impact is positive = good (already in the native scale; no additional flip). Only cards appearing in ≥10 distinct decks with ≥25 observations are ranked. Rankings are weighted by breadth — broader deck coverage and higher observation counts make a figure more credible.
11a. Commander — Top performers by board impact
282 Commander cards qualified (≥10 decks, ≥25 observations).
| Card | Decks (n) | Observations | Avg Board Impact (points/obs) |
|---|---|---|---|
| Ureni of the Unwritten | 11 | 253 | +29.78 |
| Dragonlord Atarka | 10 | 47 | +28.17 |
| Ghalta, Primal Hunger | 11 | 81 | +24.17 |
| Cathars' Crusade | 11 | 43 | +23.09 |
| Lyra Dawnbringer | 12 | 61 | +19.85 |
| Craterhoof Behemoth | 18 | 54 | +18.39 |
| Atarka, World Render | 14 | 80 | +18.16 |
| Terror of the Peaks | 17 | 84 | +15.74 |
| Lathliss, Dragon Queen | 19 | 81 | +13.84 |
| Miirym, Sentinel Wyrm | 14 | 165 | +12.16 |
| Orcish Bowmasters | 15 | 95 | +12.16 |
| Farseek | 47 | 154 | +10.68 |
| Sapphire Medallion | 12 | 88 | +10.39 |
| Bloodthirsty Conqueror | 10 | 83 | +10.18 |
| Avacyn, Angel of Hope | 10 | 57 | +9.86 |
The top of the Commander board-impact list is dominated by large creatures: Ureni of the Unwritten (+29.78, 11 decks, 253 observations), Dragonlord Atarka (+28.17, 10 decks, 47 observations), Ghalta, Primal Hunger (+24.17, 11 decks, 81 observations), Cathars' Crusade (+23.09, 11 decks, 43 observations), and Lyra Dawnbringer (+19.85, 12 decks, 61 observations). Ureni leads with the highest both breadth-weighted credibility (253 observations across 11 decks) and the strongest per-observation figure.
Farseek stands out as the only ramp spell in the top 15 (+10.68, 47 decks, 154 observations), and it has by far the broadest deck coverage of any card on the list — 47 distinct decks means the figure is drawn from a large, diverse sample. Miirym, Sentinel Wyrm (+12.16, 14 decks, 165 observations) and Orcish Bowmasters (+12.16, 15 decks, 95 observations) tie at the rank-10/11 positions.
11b. Commander — Bottom performers by board impact
| Card | Decks (n) | Observations | Avg Board Impact (points/obs) |
|---|---|---|---|
| Dictate of Erebos | 25 | 174 | −13.72 |
| Aetherflux Reservoir | 11 | 48 | −11.40 |
| Exquisite Blood | 14 | 51 | −11.31 |
| Panharmonicon | 11 | 43 | −10.67 |
| Parallel Lives | 11 | 44 | −10.27 |
| Grim Monolith | 12 | 52 | −9.31 |
| The Wind Crystal | 10 | 39 | −9.18 |
| Pitiless Plunderer | 20 | 126 | −8.83 |
| Doubling Season | 21 | 72 | −8.36 |
| Reflections of Littjara | 10 | 33 | −8.33 |
| Guardian Project | 12 | 44 | −7.66 |
| Bolas's Citadel | 10 | 46 | −7.63 |
| Grave Pact | 23 | 88 | −7.36 |
| Altar of the Brood | 10 | 67 | −7.25 |
| Fleshbag Marauder | 12 | 76 | −7.12 |
The bottom of the Commander board-impact list is arresting in its specificity. Several cards that enjoy strong reputations in Commander deckbuilding communities appear here: Dictate of Erebos (−13.72, the worst-performing card in the dataset by this metric, 25 decks and 174 observations — the highest-credibility figure on the bottom list), Doubling Season (−8.36, 21 decks), Panharmonicon (−10.67, 11 decks), Grave Pact (−7.36, 23 decks), and Bolas's Citadel (−7.63, 10 decks).
These negative figures do not mean these cards are individually responsible for losses, nor that they are poorly designed, nor that they should be cut. The board-impact metric measures what happens to the overall board state in turns those cards are seen — in a game of Magic, a card that creates delayed or conditional value (like Doubling Season, which doubles tokens on later turns) may not register an immediate positive board-quality change in the observation window the engine captures. Combo-oriented win conditions (Aetherflux Reservoir, Exquisite Blood) may appear in games that are already losing before the condition assembles, dragging down the associated observations. The negative signal is real; the interpretation of why it is negative requires more data than board state alone provides. This is reported as observed, not as a judgment on card quality.
11c. Commander — Board impact by card type
| Type | Cards in Pool | Observations | Decks Floor | Avg Board Impact |
|---|---|---|---|---|
| Creatures | 3,749 | 67,959 | 50 | +2.20 |
| Instants | 730 | 7,992 | 92 | +1.06 |
| Planeswalkers | 108 | 1,339 | 13 | +0.41 |
| Sorceries | 672 | 6,342 | 65 | −0.12 |
| Artifacts | 680 | 18,667 | 315 | −2.07 |
| Nonbasic Lands | 37 | 198 | 12 | −2.08 |
| Enchantments | 767 | 10,440 | 37 | −2.19 |
Creatures carry the highest pooled board impact (+2.20 across 67,959 observations) in Commander. Instants follow at +1.06. Sorceries, artifacts, enchantments, and nonbasic lands are all negative in aggregate. The type-bucket aggregations pool enormously varied cards — a +2.20 average for creatures spans everything from Llanowar Elves to Craterhoof Behemoth — and the observation that enchantments average −2.19 coexists with individual enchantments like Cathars' Crusade (+23.09) that are among the best-performing individual cards in the dataset. Type-level figures are descriptive aggregates only and should not be read as type-versus-type prescriptions.
11d. Commander — Board impact by color identity
| Color | Cards | Observations | Decks Floor | Avg Board Impact |
|---|---|---|---|---|
| Red | 1,331 | 20,535 | 53 | +3.26 |
| Green | 1,730 | 29,907 | 65 | +2.34 |
| White | 1,483 | 27,731 | 92 | +1.56 |
| Blue | 1,506 | 22,406 | 41 | +1.07 |
| Black | 1,569 | 29,739 | 52 | +0.04 |
| Colorless | 687 | 18,040 | 315 | −1.49 |
Red cards carry the highest average board impact in Commander (+3.26, 1,331 cards, 20,535 observations), followed by Green (+2.34) and White (+1.56). Black cards average near zero (+0.04) and Colorless cards are negative (−1.49). Several caveats: multicolor cards count toward each of their colors, so a 4-color card contributes to four color buckets simultaneously. Colorless cards are dominated by mana rocks and artifacts (the artifact type-bucket was also negative), which may pull the colorless average down. These are descriptive patterns in the data, not color recommendations.
11e. Standard — Board impact
Only one Standard card met the qualifying floor (≥10 decks, ≥25 observations):
| Card | Decks (n) | Observations | Avg Board Impact |
|---|---|---|---|
| Lightning Bolt | 10 | 95 | +5.05 |
Standard's thin card-performance data reflects the format's smaller cohort (59 decks, 1,042 games). Lightning Bolt at +5.05 (10 decks, 95 observations) is the only card with sufficient breadth for publication. The type and color Standard breakdowns show: Instants average −0.17 (124 cards, 729 observations, ≥10 decks floor) and Red cards average +1.76 (156 cards, 1,720 observations, ≥10 decks floor). Only these type and color buckets cleared the minimum; others are suppressed.
12. Color Identity: Popularity and Win Rate
Commander
By deck count, Black is the most common color identity in Commander decks in this dataset (238 decks, 53% of the Commander cohort), followed by Green (227), White (221), Blue (209), and Red (205). All colors are broadly distributed — no color is dramatically underrepresented. By win rate, Red-inclusive decks average 44.1% (n = 205) and Black-inclusive decks average 41.1% (n = 238) — the highest and lowest respectively, separated by 3.0 points. Because Commander decks are almost always multi-color, these figures overlap heavily and do not represent independent color effects.
The most popular color is not the highest-performing color, and the gap between any two colors is small relative to the Gauntlet-opponent spread (22.5 points). Color identity alone explains very little of the win-rate variance in this dataset.
Standard
In Standard, White is the most common color (32 decks, 54% of Standard decks), followed by Black (25), Red (26), Blue (18), and Green (17). Red-inclusive Standard decks average 36.9% (n = 26) and Black-inclusive decks 28.1% (n = 25) — an 8.8-point gap, wider than in Commander. This gap likely reflects the presence of the burn archetype (Red-based, Section 5b) at the top and the difficulty Black-inclusive control or midrange decks face against the Standard Gauntlet's fastest opponents.
13. Tempo and Game Length
Game length is measured as the average turn count per match, derived from per-match decision snapshots. This is a descriptive tempo marker — how long it typically takes the engine to resolve a game — not a measure of interactive quality or human play experience.
| Format | Decks (n) | Avg Turns | Avg Shortest Game | Avg Longest Game |
|---|---|---|---|---|
| Commander | 448 | 10.3 | 8.4 | 12.3 |
| Standard | 59 | 9.5 | 7.3 | 11.7 |
Commander games average 10.3 turns, with individual games ranging on average from 8.4 to 12.3 turns across the deck sample (n = 448). Standard games are slightly shorter on average (9.5 turns), ranging from 7.3 to 11.7 turns (n = 59). The tighter Standard range reflects the format's lower average mana values (2.09–2.75 in the bracket analysis vs. 3.08–3.31 in Commander) and the presence of Mono Red Aggro in the Standard Gauntlet, which closes games in the fastest windows.
The shortest Standard games average 7.3 turns — a benchmark that illustrates why Standard decks built for turns 4–7 payoffs may struggle to execute against Mono Red Aggro before the game ends.
14. Matchup Structure: Deck vs. Opponent
This section decomposes the relative contributions of the specific Gauntlet opponent versus the deck's own profile to matchup win-rate variance. All decompositions are descriptive — they quantify observed associations, not causes. No significance testing is applied. The data decides which factor carries more signal; there is no prior thesis.
14a. Variance decomposition (games-weighted)
How much of the total variation in per-(deck, opponent) win rates is explained by the deck itself, by which opponent was faced, or by residual effects (interaction, noise)?
Commander (n = 98 decks, 20 opponents, 456 cells):
| Effect | Share of Variance |
|---|---|
| Deck main effect | 46.9% |
| Opponent main effect | 16.8% |
| Residual (interaction + noise) | 36.3% |
Standard (n = 10 decks, 5 opponents, 50 cells):
| Effect | Share of Variance |
|---|---|
| Deck main effect | 27.8% |
| Opponent main effect | 43.9% |
| Residual (interaction + noise) | 28.3% |
The formats split sharply. In Commander, the deck itself accounts for 46.9% of the measured win-rate variance — nearly three times the opponent's 16.8%. In Standard, the relationship inverts: the opponent accounts for 43.9% of variance versus only 27.8% for the deck. The Standard sample is small (n = 10 decks), so these estimates carry more uncertainty than the Commander figures. Both formats leave a substantial residual (36.3% Commander, 28.3% Standard) that is unexplained by simple main effects alone.
These are descriptive decompositions of simulated results only. They describe patterns in this dataset, under this Gauntlet composition, for these submitted decks.
14b. Matchup inversion rate
What share of comparable deck pairs — pairs where both decks had enough games against two specific meta opponents to establish a strict win-rate ordering — reversed that preference order between opponents? A higher rate means the specific opponent reshuffles deck-level outcomes more.
| Format | Deck Pairs | Comparable Pairs | Inverted Pairs | Inversion Rate |
|---|---|---|---|---|
| Commander (n = 88 decks, 10 opponent pairings) | 88 | 25,242 | 9,240 | 36.6% |
| Standard (n = 10 decks, 10 opponent pairings) | 10 | 344 | 79 | 23.0% |
In Commander, 36.6% of the 25,242 comparable deck pairs reverse their opponent preference order between two meta opponents — meaning that if Deck A beats Deck B on a given opponent pairing, there is a meaningful chance the ranking flips on a different pairing. In Standard, that share is 23.0% (344 pairs). Both figures indicate that no single Gauntlet opponent is a reliable stand-in for the full meta — decks specialize, and the field matters. The Commander rate is higher, consistent with Commander's greater deck diversity and wider Gauntlet (20 opponents in the decomposition, 5 in Standard's).
14c. Power Score × opponent win-rate grid (Commander only)
The Standard cohort did not produce a qualifying PS grid (fewer than 10 decks per cell at the required level). The Commander grid uses 3 Power Score bands (low: PS 2.2–31.8; mid: PS 32–47.8; high: PS 48.6–90.9), each with 27–34 decks, against all 5 Commander Gauntlet opponents. All 15 cells cleared the per-cell minimum-games floor.
| Opponent | Low PS (n = 27–28) | Mid PS (n = 31) | High PS (n = 30) |
|---|---|---|---|
| Breya Artifact Combo | 32.9% | 55.3% | 71.5% |
| Derevi Bant Control | 33.8% | 52.7% | 66.2% |
| Aesi Landfall | 32.0% | 42.5% | 66.2% |
| Atraxa Superfriends | 14.3% | 30.0% | 56.5% |
| Edgar Markov Vampires | 15.3% | 24.5% | 54.6% |
The grid shows a consistent monotonic pattern across all five opponents: higher Power Score bands are associated with higher win rates in every cell. The gap between Low and High PS is largest against Breya Artifact Combo (32.9% → 71.5%, a 38.6-point spread) and smallest against Atraxa Superfriends in the Low band (14.3%) versus Edgar Markov Vampires (15.3%), where both low-PS groups are severely disadvantaged. The High PS group posts above 54% against every opponent, including the two hardest.
Power Score is an internal Grim.Cards composite indicator. The grid describes the association between PS band and win rate in this dataset, not a causal relationship. Correlation only.
14d. Spearman rank correlation: Power Score vs. win rate per opponent
| Format | Opponent | Decks (n) | Spearman ρ |
|---|---|---|---|
| Commander | Atraxa Superfriends | 89 | 0.77 |
| Commander | Breya Artifact Combo | 88 | 0.75 |
| Commander | Edgar Markov Vampires | 88 | 0.75 |
| Commander | Aesi Landfall | 88 | 0.67 |
| Commander | Derevi Bant Control | 88 | 0.62 |
| Standard | Temur Harmonizer Combo | 10 | 0.88 |
| Standard | Dimir Midrange | 10 | 0.79 |
| Standard | Jeskai Control | 10 | 0.70 |
| Standard | Azorius Tempo | 10 | 0.67 |
| Standard | Mono Red Aggro | 10 | 0.55 |
In Commander, the Spearman rank correlation between a deck's Power Score and its win rate against a given opponent ranges from ρ = 0.62 (Derevi Bant Control) to ρ = 0.77 (Atraxa Superfriends) across the five main Gauntlet opponents (n = 88–89 decks each). In Standard (n = 10 per opponent), correlations range from ρ = 0.55 (Mono Red Aggro) to ρ = 0.88 (Temur Harmonizer Combo). The lowest correlations — Derevi Bant Control in Commander and Mono Red Aggro in Standard — are the matchups where Power Score explains the least variance in outcomes and individual deck characteristics carry more weight. The Temur Harmonizer Combo Standard correlation (ρ = 0.88) is the strongest in either format. Power Score is an internal composite; these correlations describe how well it ranks decks by outcome against each specific opponent in this dataset.
15. Per-Deck Appendix (Anonymized)
Privacy note: Pseudonyms (C-001 through C-098, S-001 through S-010) are randomly reassigned every edition and carry no link to any prior edition's pseudonyms, any account, or any actual deck name or identifier. Published values are coarsened: Power Score rounded to the nearest 5, win rates and spreads to whole percentage points, game counts banded. Only derived statistics are shown. No deck names, deck lists, owners, or identifiers appear.
Spread definition: Standard deviation of a deck's per-opponent win rates across all qualifying matchup cells (3 or more opponents with sufficient games). Higher spread = more matchup-specialist profile; lower spread = more consistent across opponents. Spread is shown only where 3 or more qualifying cells exist.
15a. Commander per-deck appendix summary
98 anonymized Commander decks published (minimum 20 games per deck).
| Metric | Value |
|---|---|
| Decks published | 98 |
| Win rate: minimum | 7% |
| Win rate: median | 42% |
| Win rate: maximum | 89% |
| Spread: median (n = 88 decks with 3+ opponents) | 17 pts |
15b. Commander per-deck table
| ID | PS (±5) | Games Band | Win Rate | Opponents Faced | Spread |
|---|---|---|---|---|---|
| C-001 | 65 | 20–49 | 63% | 5 | 12 |
| C-002 | 40 | 20–49 | 42% | 5 | 23 |
| C-003 | 55 | 20–49 | 57% | 5 | 13 |
| C-004 | 45 | 20–49 | 55% | 1 | — |
| C-005 | 0 | 50–99 | 7% | 5 | 13 |
| C-006 | 45 | 50–99 | 45% | 5 | 34 |
| C-007 | 35 | 20–49 | 37% | 5 | 12 |
| C-008 | 30 | 20–49 | 33% | 5 | 24 |
| C-009 | 85 | 20–49 | 80% | 5 | 7 |
| C-010 | 90 | 20–49 | 85% | 1 | — |
| C-011 | 30 | 20–49 | 33% | 5 | 28 |
| C-012 | 50 | 20–49 | 49% | 5 | 15 |
| C-013 | 25 | 20–49 | 25% | 1 | — |
| C-014 | 50 | 50–99 | 50% | 6 | 18 |
| C-015 | 35 | 20–49 | 38% | 5 | 13 |
| C-016 | 30 | 20–49 | 20% | 5 | 12 |
| C-017 | 30 | 50–99 | 34% | 6 | 14 |
| C-018 | 40 | 20–49 | 40% | 5 | 17 |
| C-019 | 70 | 20–49 | 67% | 5 | 11 |
| C-020 | 15 | 20–49 | 16% | 5 | 17 |
| C-021 | 50 | 20–49 | 51% | 5 | 15 |
| C-022 | 30 | 20–49 | 30% | 5 | 24 |
| C-023 | 30 | 20–49 | 30% | 5 | 22 |
| C-024 | 50 | 20–49 | 47% | 5 | 26 |
| C-025 | 40 | 20–49 | 40% | 5 | 25 |
| C-026 | 70 | 20–49 | 67% | 5 | 18 |
| C-027 | 45 | 20–49 | 44% | 5 | 16 |
| C-028 | 10 | 20–49 | 15% | 1 | — |
| C-029 | 30 | 20–49 | 30% | 5 | 22 |
| C-030 | 55 | 20–49 | 53% | 5 | 12 |
| C-031 | 20 | 20–49 | 25% | 1 | — |
| C-032 | 30 | 20–49 | 37% | 6 | 15 |
| C-033 | 55 | 20–49 | 53% | 5 | 29 |
| C-034 | 55 | 20–49 | 53% | 5 | 19 |
| C-035 | 60 | 20–49 | 60% | 5 | 19 |
| C-036 | 55 | 20–49 | 63% | 5 | 16 |
| C-037 | 25 | 20–49 | 27% | 5 | 8 |
| C-038 | 25 | 50–99 | 25% | 5 | 7 |
| C-039 | 15 | 20–49 | 20% | 5 | 27 |
| C-040 | 70 | 20–49 | 67% | 5 | 16 |
| C-041 | 70 | 20–49 | 67% | 5 | 18 |
| C-042 | 30 | 20–49 | 40% | 1 | — |
| C-043 | 70 | 20–49 | 70% | 5 | 16 |
| C-044 | 90 | 20–49 | 89% | 5 | 12 |
| C-045 | 25 | 50–99 | 27% | 5 | 10 |
| C-046 | 70 | 20–49 | 85% | 1 | — |
| C-047 | 65 | 20–49 | 70% | 5 | 27 |
| C-048 | 30 | 20–49 | 33% | 5 | 21 |
| C-049 | 40 | 20–49 | 40% | 5 | 17 |
| C-050 | 75 | 20–49 | 73% | 5 | 13 |
| C-051 | 40 | 100–199 | 42% | 5 | 16 |
| C-052 | 60 | 20–49 | 70% | 1 | — |
| C-053 | 40 | 20–49 | 40% | 5 | 13 |
| C-054 | 40 | 50–99 | 42% | 5 | 12 |
| C-055 | 50 | 20–49 | 50% | 5 | 21 |
| C-056 | 35 | 20–49 | 28% | 5 | 17 |
| C-057 | 60 | 20–49 | 60% | 5 | 15 |
| C-058 | 30 | 20–49 | 30% | 5 | 22 |
| C-059 | 30 | 20–49 | 29% | 5 | 9 |
| C-060 | 65 | 20–49 | 62% | 5 | 11 |
| C-061 | 20 | 20–49 | 23% | 5 | 8 |
| C-062 | 70 | 20–49 | 70% | 5 | 12 |
| C-063 | 25 | 20–49 | 27% | 5 | 11 |
| C-064 | 35 | 20–49 | 40% | 6 | 16 |
| C-065 | 35 | 50–99 | 35% | 5 | 18 |
| C-066 | 60 | 50–99 | 58% | 5 | 12 |
| C-067 | 80 | 20–49 | 77% | 5 | 8 |
| C-068 | 45 | 50–99 | 50% | 6 | 19 |
| C-069 | 40 | 20–49 | 40% | 5 | 17 |
| C-070 | 55 | 20–49 | 53% | 5 | 16 |
| C-071 | 45 | 50–99 | 43% | 5 | 8 |
| C-072 | 25 | 20–49 | 24% | 5 | 11 |
| C-073 | 25 | 50–99 | 33% | 5 | 17 |
| C-074 | 25 | 20–49 | 27% | 5 | 25 |
| C-075 | 70 | 20–49 | 70% | 5 | 19 |
| C-076 | 45 | 20–49 | 53% | 5 | 18 |
| C-077 | 55 | 20–49 | 53% | 5 | 19 |
| C-078 | 45 | 20–49 | 47% | 5 | 16 |
| C-079 | 45 | 20–49 | 49% | 6 | 16 |
| C-080 | 35 | 20–49 | 37% | 5 | 29 |
| C-081 | 60 | 20–49 | 70% | 1 | — |
| C-082 | 5 | 20–49 | 12% | 5 | 9 |
| C-083 | 75 | 20–49 | 73% | 5 | 17 |
| C-084 | 35 | 20–49 | 30% | 5 | 22 |
| C-085 | 45 | 20–49 | 47% | 5 | 24 |
| C-086 | 45 | 20–49 | 47% | 5 | 31 |
| C-087 | 55 | 20–49 | 60% | 5 | 20 |
| C-088 | 35 | 20–49 | 37% | 5 | 19 |
| C-089 | 30 | 20–49 | 29% | 5 | 15 |
| C-090 | 25 | 20–49 | 27% | 5 | 31 |
| C-091 | 40 | 20–49 | 43% | 5 | 17 |
| C-092 | 20 | 20–49 | 20% | 5 | 12 |
| C-093 | 30 | 20–49 | 35% | 1 | — |
| C-094 | 30 | 20–49 | 33% | 5 | 16 |
| C-095 | 30 | 20–49 | 31% | 5 | 26 |
| C-096 | 20 | 20–49 | 19% | 5 | 16 |
| C-097 | 65 | 20–49 | 57% | 5 | 17 |
| C-098 | 30 | 20–49 | 30% | 5 | 22 |
Spread (—) = fewer than 3 qualifying opponents; not published. PS rounded to nearest 5. Win rate and spread rounded to whole points. Game counts banded.
15c. Outlier callout: OLS residuals vs. Power Score (Commander, n = 98 decks)
An OLS line fit to the 98 published Commander per-deck rows relates coarsened Power Score to win rate. Two decks stand at the widest distances from that fitted line:
Widest overperformance — C-046: PS 70 (coarsened), win rate 85%, OLS-expected win rate 69%. Residual: +16 percentage points. This deck won 16 points more than a PS-70 deck would be expected to win based on the linear relationship across the published sample.
Widest underperformance — C-016: PS 30 (coarsened), win rate 20%, OLS-expected win rate 32%. Residual: −12 percentage points. This deck won 12 points less than the PS-30 expectation.
Note: C-046 faced only 1 opponent (spread not published), so its win rate is based on a narrower sample than most published entries. C-016 faced 5 opponents with a 12-point spread, indicating modest matchup variation. These outliers are reported as observed deviations from the PS–win-rate relationship in the published appendix data — no causal interpretation is offered, and pseudonyms rotate every edition.
15d. Standard per-deck appendix summary
10 anonymized Standard decks published (minimum 20 games per deck; the appendix requires ≥10 published decks, exactly met).
| Metric | Value |
|---|---|
| Decks published | 10 |
| Win rate: minimum | 27% |
| Win rate: median | 40% |
| Win rate: maximum | 70% |
| Spread: median (n = 10 decks with 3+ opponents) | 23 pts |
15e. Standard per-deck table
| ID | PS (±5) | Games Band | Win Rate | Opponents Faced | Spread |
|---|---|---|---|---|---|
| S-001 | 60 | 20–49 | 60% | 5 | 34 |
| S-002 | 35 | 20–49 | 37% | 5 | 27 |
| S-003 | 30 | 20–49 | 30% | 5 | 19 |
| S-004 | 70 | 20–49 | 67% | 5 | 21 |
| S-005 | 25 | 20–49 | 27% | 5 | 23 |
| S-006 | 40 | 20–49 | 40% | 5 | 23 |
| S-007 | 70 | 20–49 | 70% | 5 | 27 |
| S-008 | 35 | 20–49 | 38% | 5 | 17 |
| S-009 | 30 | 20–49 | 33% | 5 | 18 |
| S-010 | 50 | 20–49 | 50% | 5 | 24 |
All 10 Standard decks face 5 opponents and qualify for spread calculation.
15f. Outlier callout: OLS residuals vs. Power Score (Standard, n = 10 decks)
The Standard per-deck sample (n = 10) is at the minimum publication floor. OLS residuals at this sample size carry high uncertainty and should be read with that in mind.
Widest overperformance — S-008: PS 35 (coarsened), win rate 38%, OLS-expected win rate 36%. Residual: +2 percentage points.
Widest underperformance — S-004: PS 70 (coarsened), win rate 67%, OLS-expected win rate 69%. Residual: −2 percentage points.
The Standard residuals are small in absolute terms (±2 points), reflecting the compressed win-rate range (27%–70%) and the small sample. No meaningful outlier narrative is supported at this sample size; the callout is included for structural completeness and comparability with future editions.
16. Optional: Power Score Composite (Secondary, Caveated)
⚠️ Power Score is an internal Grim.Cards composite indicator, not an objective or universal measurement of deck quality or win probability. It is not used to rank or judge decks elsewhere in this report. It is presented here in a single, clearly-labelled section for context only.
Power Score combines multiple construction and card-quality signals into a single composite value. The formula and weighting are internal to Grim.Cards and are subject to change. A higher Power Score does not guarantee a higher win rate, as demonstrated throughout this report — the specific Gauntlet opponent explains as much or more variance than Power Score in some matchup structures (see Section 14a).
Grade distribution among completed simulations
Grades are binned from the underlying continuous Power Score. Counts below are the number of completed user-deck simulations (n = 691 total across both formats) in each grade band, not unique decks.
| Grade | Simulations |
|---|---|
| S | 3 |
| A | 20 |
| B | 56 |
| C | 83 |
| D | 143 |
| F | 202 |
The distribution is bottom-heavy: F and D together account for 345 of 507 scored decks — the majority of submitted decks score in the lower two grade bands by this composite. S-grade decks are rare (3). This distribution reflects both the composition of the submitted deck pool and the internal calibration of the Power Score composite, not an objective assessment of deck quality. Power Score and letter grades are Grim.Cards' own internal framing; they should not be read as authoritative verdicts.
17. Key Findings
The following findings summarize what the data in this edition directly supports. Each is derived from the measured figures reported above, carried with its sample size.
1. The Gauntlet opponent is a primary driver of outcomes. The same pool of 447 Commander decks wins 53.5% against Breya Artifact Combo and 31.0% against Edgar Markov Vampires — a 22.5-point swing with no change to the deck slate. In Standard, the spread is 45.4 points (60.4% vs. Temur Harmonizer Combo, 15.0% vs. Mono Red Aggro, n = 58 decks each). The specific opponent a deck faces in the Gauntlet is at least as consequential as the deck's own construction profile.
2. Win rates are broadly distributed, with substantial lower tails. Commander win rates span 0% to 93.3% (median 40%, mean 43.4%, n = 448 decks). The 0–30% zone holds 110 decks (24.6% of the Commander sample). Standard spans 0% to 80% (median 33.3%, mean 35.3%, n = 59 decks), with 42.4% of Standard decks in the 0–30% zone. The majority of submitted decks lose more games than they win against the meta Gauntlet.
3. Commander retests trend positive; Standard retests are inconclusive. Among 106 retested Commander decks, 55 improved win rate, 36 declined, 15 were flat — an average change of +2.6 percentage points. Among 12 retested Standard decks (at the minimum publication threshold), the split was 5 improved / 5 declined / 2 flat, averaging +0.6 points.
4. In Commander, deck identity explains more win-rate variance than opponent identity; in Standard, the reverse is true. The games-weighted variance decomposition assigns 46.9% of Commander win-rate variance to the deck main effect and 16.8% to the opponent (36.3% residual; n = 98 decks, 20 opponents). In Standard, opponent explains 43.9% versus 27.8% for the deck (28.3% residual; n = 10 decks, 5 opponents). The Standard estimate carries high uncertainty at n = 10.
5. Mono-color decks outperform multi-color in this dataset. Commander mono-color decks average 50.8% win rate (n = 86) versus 40.1–42.9% for two-, three-, and five-color decks. Standard mono-color decks average 44.2% (n = 11) versus 36.7% for two-color (n = 37). Correlation only — archetype and Gauntlet composition are likely confounders.
6. Higher land ratios are associated with higher win rates. Commander decks near 40% lands average 46.7% (n = 112) versus 40.3% for decks near 30% lands (n = 66). Standard shows the same direction: 43.1% at ~40% lands (n = 26) versus 30.6% at ~35% lands (n = 18).
7. Board-impact leaders in Commander are dominated by large threats; the laggards include widely-played staples. Ureni of the Unwritten leads Commander board impact at +29.78 per observation (11 decks, 253 obs). Dictate of Erebos trails at −13.72 (25 decks, 174 obs) — the most credible negative figure by deck breadth. Negative board-impact figures do not constitute a judgment on card quality; they describe what the board-state proxy registers in the game windows the engine captures.
8. Functional card categories correlate negatively with Standard win rates. Standard decks including reanimation effects win 15.7 points lower than those without (23.6%, n = 15 vs. 39.3%, n = 44). Tutor-inclusive Standard decks trail by 11.6 points (28.8%, n = 26 vs. 40.4%, n = 33). All category assignments are heuristic; these correlations likely reflect archetype identity rather than the individual effects of the card categories.
9. The most-played Commander card (Sol Ring, 387 decks) clusters near the overall average win rate. Sol Ring-containing decks win at 43.4% (n = 387 decks, 7,823 games) — essentially identical to the Commander aggregate (43.2%). Ubiquitous staples are distributed across all power levels and do not mark high- or low-performing archetypes in this dataset.
10. Meren of Clan Nel Toth is the only commander with a publishable win rate; it posts 33.6%. With 17 decks and 360 games, Meren of Clan Nel Toth is the most-played commander in the dataset and the only one clearing the 10-deck minimum. Its 33.6% win rate is approximately 9.6 points below the Commander average. All other commanders in the dataset have fewer than 10 distinct decks; their win rates are withheld.
18. What We Cannot Conclude Yet
The following questions are raised by the data but cannot be answered from this snapshot:
- Does changing a deck improve its win rate? The retest data shows +2.6 points on average in Commander (n = 106), but the dataset does not record whether decks were modified between tests. The direction of change cannot be attributed to iteration vs. random variation.
- Are commander-level win-rate differences real? Only Meren of Clan Nel Toth clears the 10-deck floor. Commander-level win-rate comparisons require a larger dataset.
- Do negative board-impact figures indicate cards that drag down win rates? The board-impact metric captures board-state change in observation windows, not win attribution. Cards with negative avgDelta may appear disproportionately in games that are already losing, not cause the losses.
- Are the Standard category associations driven by the categories or by archetypes? Standard reanimation and tutor categories may simply flag the same slow-control decks that lose badly to Mono Red Aggro — the category itself may be incidental.
- Would a different Gauntlet produce different rank-orderings? The inversion rate (36.6% in Commander) confirms that deck rankings are not stable across opponents. Results in this report are specific to the current Gauntlet composition.
- Is the Standard variance decomposition reliable? With n = 10 decks in the Standard decomposition, the 43.9% / 27.8% opponent-vs-deck split is an estimate with high sampling uncertainty. Future editions with larger Standard cohorts will test whether this finding holds.
19. Future-Comparison Baseline
The following table provides the machine-readable baseline for comparing this edition against future editions. All figures are drawn directly from the dataset JSON for this snapshot.
| Metric | Value | n (decks) | n (games) | Format |
|---|---|---|---|---|
| Overall win rate | 42.6% | 507 | 10,011 | All |
| Commander win rate | 43.2% | 448 | 8,969 | Commander |
| Standard win rate | 37.1% | 59 | 1,042 | Standard |
| Commander median win rate | 40.0% | 448 | — | Commander |
| Standard median win rate | 33.3% | 59 | — | Standard |
| Highest Commander matchup win rate | 53.5% (vs. Breya Artifact Combo) | 447 | 1,773 | Commander |
| Lowest Commander matchup win rate | 31.0% (vs. Edgar Markov Vampires) | 447 | 1,773 | Commander |
| Highest Standard matchup win rate | 60.4% (vs. Temur Harmonizer Combo) | 58 | 207 | Standard |
| Lowest Standard matchup win rate | 15.0% (vs. Mono Red Aggro) | 58 | 207 | Standard |
| Commander retest avg change | +2.6 pts | 106 | — | Commander |
| Standard retest avg change | +0.6 pts | 12 | — | Standard |
| Commander avg game length | 10.3 turns | 448 | — | Commander |
| Standard avg game length | 9.5 turns | 59 | — | Standard |
| Top Commander board-impact card | Ureni of the Unwritten, +29.78 | 11 | 253 obs | Commander |
| Top Standard board-impact card | Lightning Bolt, +5.05 | 10 | 95 obs | Standard |
| Worst Commander board-impact card | Dictate of Erebos, −13.72 | 25 | 174 obs | Commander |
| Commander variance: deck share | 46.9% | 98 | — | Commander |
| Commander variance: opponent share | 16.8% | 98 | — | Commander |
| Standard variance: deck share | 27.8% | 10 | — | Standard |
| Standard variance: opponent share | 43.9% | 10 | — | Standard |
| Commander inversion rate | 36.6% | 88 | 25,242 pairs | Commander |
| Standard inversion rate | 23.0% | 10 | 344 pairs | Standard |
| Snapshot date | 2026-08-01 | — | — | — |
| Dataset version | 2.4 | — | — | — |
Limitations (consolidated)
- All results are from AI-piloted simulations (a custom Forge build). They describe engine behavior on submitted decklists and cannot be generalized to human play or to Magic: The Gathering broadly.
- All construction-level, card-level, and category-level findings are correlational. No causal claim is made anywhere in this report.
- The sample is self-selected — decks users chose to submit to Grim.Cards, not a random sample of any broader population.
- Monthly win-rate comparisons across months reflect different deck populations in each month, not longitudinal progression.
- Power Score is an internal Grim.Cards composite, not an objective or universal measurement. It is not used for ranking or judgment in this report.
- Functional card category assignments (sacrifice, tutor, discard, reanimation) are heuristic and may contain labelling errors.
- Board-impact figures are a board-state proxy, not damage, kills, or causal win attribution. Negative figures do not establish that a card harms win probability.
- The Standard per-deck appendix (n = 10) and Standard variance decomposition (n = 10) are at minimum publication thresholds; estimates carry high sampling uncertainty.
- Counterfactual decision-impact data did not meet the cohort threshold in this edition; that section is omitted.
- Per-deck pseudonyms rotate every edition; no cross-edition deck tracking is possible by design.
License and Citation
License: This report and its associated data are published under the Creative Commons Attribution 4.0 International License (CC BY 4.0). You are free to share and adapt the material for any purpose, provided appropriate credit is given.
Suggested citation:
Grim.Cards. Grim.Cards Simulation Data — Case Study Edition, Snapshot 2026-08-01 (Dataset Version 2.4). Published 2026-08-01. Permanent URL: https://grim.cards/case-study/2026-08-01. License: CC BY 4.0. Temporal coverage: 2026-05-14 to 2026-08-01.
Data download: The machine-readable edition JSON (this snapshot's full structured data, including atomic statements, per-figure sample sizes, and provenance block) is available at: https://grim.cards/case-study/2026-08-01/data.json
Edition index: All published editions, newest-first: https://grim.cards/case-study
Retrieved-date note for citing parties: This permanent URL (/case-study/2026-08-01) will not change when newer editions are published. Future editions accumulate at their own dated URLs and do not overwrite this one. Record your retrieval date alongside the snapshot date when citing.
Report generated: 2026-08-01T08:16:13.651Z · Grim.Cards Dataset Version 2.4 · Cohort: player-submitted decks (Crucible and sample decks excluded) · Minimum cohort threshold: 10 unique decks per reported group · Win rate: wins ÷ total games played (draws in denominator) · All findings correlational · No PII published