Grim.Cards Simulation Case Study
Edition: 2026-10-01 · Data Snapshot 2026-05-14 → 2026-10-01 · Dataset Version 2.4
Published: 2026-10-01 · Permanent URL: grim.cards/case-study/2026-10-01
License: Creative Commons Attribution 4.0 International (CC BY 4.0)
Citation: Grim.Cards. "Grim.Cards Simulation Case Study, Edition 2026-10-01." Grim.Cards, 1 Oct. 2026, grim.cards/case-study/2026-10-01.
Executive Summary
Forty-four thousand, eight hundred and thirty-three simulated games. One hundred and eighty-four calendar days. Two formats, 2,499 unique player-submitted decks, and 1,184 distinct users — this is the most complete quantitative account Grim.Cards has ever published of how real people's decks perform against a fixed gauntlet of machine-piloted meta opponents.
The headline number is 40.8% overall win rate (n = 2,499 decks; 44,833 games; wins ÷ all games played, draws counted in the denominator). That figure sits below the arithmetic midpoint of 50%, which is structurally expected: in a one-versus-the-field gauntlet, any single challenger faces several competent opponents simultaneously, so sub-50% is the baseline, not a verdict on deck quality. The more revealing numbers are the ones hiding beneath the aggregate.
The single strongest finding in this dataset is matchup dependency. In Commander, the spread between the easiest and hardest meta opponent is 20.9 percentage points (Breya Artifact Combo: 52.5%, n = 2,197 decks; Edgar Markov Vampires: 31.6%, n = 2,196 decks). In Standard, the same spread is a staggering 40.7 percentage points (Temur Harmonizer Combo: 55.5% vs. Mono Red Aggro: 14.8%, n = 102 decks each). Both spreads were measured on essentially the same challenger pool — every player deck faced both opponents. The opponent alone, not just the challenger's own profile, shifts observed win rates by more than many players would expect.
Commander format dominates the dataset: 2,364 decks (94%), 42,931 games (96%), 41.0% win rate. Standard contributes 135 decks, 1,902 games, 37.3% win rate. Every section below is split by format; figures are never pooled across the two.
Other headline findings:
- Meren of Clan Nel Toth is the most-submitted commander by deck count in this dataset (n = 36 decks, 739 games, 34.5% win rate); Giada, Font of Hope recorded the highest observed win rate among commanders with ≥ 10 decks (n = 17 decks, 468 games, 76.1%).
- Sol Ring is the most widely played Commander card (n = 1,981 decks); Sazh's Chocobo recorded the highest pooled board-impact score in Commander at +832.17 board-quality points per appearance (n = 13 decks, 69 observations) — a board-state proxy, not a kill or damage count.
- In Commander, the deck's own profile accounts for 47.7% of win-rate variance across matchup cells; the specific meta opponent accounts for 14.6%; the remaining 37.8% is residual (n = 384 decks, 23 opponents). In Standard, deck and opponent each contribute 28%, with 44% residual (n = 13 decks, 6 opponents).
- Among Commander retests (n = 426 re-tested decks), more decks declined in win rate (192) than improved (168), with 66 roughly flat; mean change was −1.6 percentage points.
Everything here is correlational. These are AI-versus-AI simulations on player-submitted decklists. No finding establishes cause. No strategic advice follows from any number reported below.
Table of Contents
- Methodology & Provenance
- Dataset Overview
- The Gauntlet: Per-Matchup Win Rates (core section)
- Win-Rate Distribution
- Deck Iteration & Retests
- Monthly Trends
- Top Commanders
- Top Cards: Popularity vs. Containing-Deck Win Rate
- Construction Correlates: Land Ratio, Color Count, Card-Type Mix
- Winning-Recipe: Win Brackets Compared
- Card-Category Insights (heuristic-labelled)
- Most Impactful Cards: Decision Impact (counterfactual proxy)
- Card Performance Roll-Up: Board Impact by Card and by Type/Color (board-state proxy)
- Color-Identity Breakdown
- Tempo & Game-Length Signals
- Matchup Structure: Deck vs. Opponent
- Per-Deck Appendix (Anonymized)
- Key Findings & What We Cannot Conclude Yet
- Future-Comparison Baseline Table
- Power Score Composite (secondary; caveated)
- Limitations
- License & Citation
1. Methodology & Provenance
What Was Simulated
Grim.Cards runs AI-versus-AI Magic: The Gathering simulations on a custom build of the open-source Forge engine. Each player-submitted deck is loaded as a digital list and piloted by the Forge AI; it then plays a fixed number of games against each opponent in a curated "gauntlet" of meta decks, also AI-piloted. Results — wins, losses, and draws — are recorded per game.
Cohort Definition
Primary cohort: Real, human-submitted Commander and Standard decks. Excluded throughout:
- Automated Crucible decks (
user_id = '__grinder__'): these form a separate AI reference corpus and are not part of any reported figure in this study. - System sample decks (
is_sample = true).
Minimum published cohort: 10 unique decks. Any breakdown — per matchup, per color, per format, per month, per construction band — with fewer than 10 distinct decks in the Grim.Cards dataset is suppressed or merged into "Other" and not reported.
Win-Rate Definition
Win rate = wins ÷ total games played, where total games = wins + losses + draws. Draws are counted in the denominator, not the numerator. This definition is applied consistently across every win-rate figure in this report. Where "decisive win rate" (wins ÷ (wins + losses) only) would differ materially, this report uses the full-denominator definition and states the draw count alongside.
Format Separation
The simulations table does not carry a format column. Format is determined by joining to deck metadata. Commander and Standard results are never pooled; every rate breakdown appears separately by format.
Retests
A "retest" is any deck with more than one completed simulation. Win-rate change is measured from the deck's first completed simulation to its most recent completed simulation (first vs. latest).
Functional Card Categories
Sacrifice, tutor/search, discard, and reanimation category membership is determined by a keyword heuristic applied to oracle text and pre-existing flags in the card preference system (has_tutor, has_sacrifice, has_discard, has_reanimate). All category-based figures are labelled heuristic throughout this report; edge cases may be mislabelled.
Decision Impact (Counterfactual Proxy)
Counterfactual impact is measured from replayed decision snapshots: when the engine reached a game state involving a particular card, Grim.Cards recorded the win-rate delta between the line actually taken and the engine's own next-best alternative. This is a play-quality proxy, not a damage count, kill count, or direct win attribution. Raw internal sign convention: negative internal delta = the alternative was better. In this report, a positive user-facing impact value means the played line exceeded the engine's next-best alternative.
Board Impact (Performance Proxy)
Card performance figures pool per-deck board-quality deltas across all appearances of a card in the cohort. A positive value means the board state improved on turns the card was seen; a negative value means it deteriorated. This is a board-state proxy, not damage, kills, or a causal win claim. Cards must appear in ≥ 10 distinct decks and have ≥ 25 recorded observations to be ranked.
Data Snapshot
Snapshot date: 2026-10-01. Data window: 2026-05-14 to 2026-10-01. Generated: 2026-10-01T08:23:19.653Z. Dataset version: 2.4.
Provenance Statement
All figures in this report are derived exclusively from Grim.Cards production simulation data. No development or staging data is included. Every figure cited carries its sample size (n). Correlation is not causation; no finding establishes a causal relationship between any deck property and any outcome.
2. Dataset Overview
| Metric | All Formats | Commander | Standard |
|---|---|---|---|
| Unique users | 1,184 | 1,119 | 100 |
| Unique decks | 2,499 | 2,364 | 135 |
| Completed simulations | 3,307 | 3,145 | 162 |
| Total games | 44,833 | 42,931 | 1,902 |
| Wins | 18,306 | 17,597 | 709 |
| Losses | 24,461 | 23,269 | 1,192 |
| Draws | 2,066 | 2,065 | 1 |
| Win rate | 40.8% | 41.0% | 37.3% |
| Data window | 2026-05-14 → 2026-10-01 | same | same |
Commander is the dominant format by every volume measure: 94.6% of decks, 95.8% of games, 96.1% of wins. Standard is a meaningful but smaller sample, and all Standard-specific rates should be read with that context.
A note on user and deck counts by format: 1,119 users submitted Commander decks; 100 submitted Standard decks. These groups overlap (a user can submit in both formats), so format-level user counts do not sum to the 1,184 total.
3. The Gauntlet: Per-Matchup Win Rates
This is the core objective-outcomes section. Win rates here are measured, not modelled. Every figure carries its deck count (n) and game count.
The Grim.Cards gauntlet pits each submitted deck against a fixed slate of meta opponents. The opponents are not random; they are curated to represent recognizable Commander and Standard archetypes. Because every submitted deck in a format faces the same slate, matchup comparisons are made on a level structural footing.
3.1 Commander Gauntlet
Five meta opponents. Approximately 2,196–2,198 distinct decks faced each, with 8,510–8,532 games per matchup — among the largest and most consistent per-matchup samples in this dataset.
| Opponent | Decks (n) | Wins | Losses | Draws | Games | Win Rate |
|---|---|---|---|---|---|---|
| Breya Artifact Combo | 2,197 | 4,467 | 3,782 | 262 | 8,511 | 52.5% |
| Derevi Bant Control | 2,196 | 4,212 | 4,200 | 98 | 8,510 | 49.5% |
| Aesi Landfall | 2,198 | 3,055 | 4,091 | 1,386 | 8,532 | 35.8% |
| Atraxa Superfriends | 2,196 | 3,007 | 5,370 | 133 | 8,510 | 35.3% |
| Edgar Markov Vampires | 2,196 | 2,691 | 5,638 | 190 | 8,519 | 31.6% |
The spread between the highest and lowest Commander matchup win rates is 20.9 percentage points. Breya Artifact Combo is the most favorable meta opponent for submitted Commander decks; Edgar Markov Vampires is the most challenging. Critically, these figures were produced by approximately the same pool of 2,196–2,198 decks — only the opponent changed. That 20.9-point swing is entirely attributable to which opponent sat across the table.
Aesi Landfall is a noteworthy outlier in the draw column: 1,386 draws from 8,532 games (16.2% draw rate), far above the 0–3% range seen in other Commander matchups. This likely reflects a structural interaction between Aesi's game plan and Forge's AI behavior in specific board states, and it means the 35.8% win rate for Aesi includes a meaningful proportion of games that ended in neither win nor loss.
Derevi Bant Control sits closest to parity at 49.5% — the one matchup where player decks are, in aggregate, nearly coin-flip competitive. Breya Artifact Combo at 52.5% is the only matchup in Commander where player decks hold a majority win rate in aggregate.
3.2 Standard Gauntlet
Five meta opponents. 102 distinct decks faced each, with 366 games per matchup.
| Opponent | Decks (n) | Wins | Losses | Draws | Games | Win Rate |
|---|---|---|---|---|---|---|
| Temur Harmonizer Combo | 102 | 203 | 163 | 0 | 366 | 55.5% |
| Jeskai Control | 102 | 184 | 182 | 0 | 366 | 50.3% |
| Azorius Tempo | 102 | 141 | 224 | 1 | 366 | 38.5% |
| Dimir Midrange | 102 | 113 | 253 | 0 | 366 | 30.9% |
| Mono Red Aggro | 102 | 54 | 312 | 0 | 366 | 14.8% |
The spread between the highest and lowest Standard matchup win rates is 40.7 percentage points — more than double the Commander spread. Temur Harmonizer Combo is decisively the most accommodating meta opponent; Mono Red Aggro is the most punishing. At 14.8% against Mono Red Aggro, the aggregate performance of Standard player decks represents one of the most lopsided matchup results in the dataset. Again: the same 102 decks faced both opponents. The opponent, not a different pool of challengers, accounts for this 40.7-point gap.
Jeskai Control again approaches parity (50.3%), as Derevi did in Commander. Player-submitted Standard decks are statistically even against Jeskai.
The Standard sample (102 decks, 366 games per matchup) is substantially smaller than Commander's. All Standard matchup rates should be interpreted with appropriate weight placed on sample size.
3.3 The Matchup-Dependency Takeaway
These tables are not just a performance leaderboard. They illustrate, in concrete numbers, that the identity of the meta opponent reshapes outcomes profoundly — even when the challenger pool is held constant. The formal decomposition of how much variance is attributable to the opponent versus the deck's own profile appears in Section 16.
4. Win-Rate Distribution
How are individual deck win rates distributed? The summary statistics and bucket counts below describe the spread of per-deck simulated win rates within each format.
4.1 Commander (n = 2,217 decks with win-rate data)
| Summary Statistic | Value |
|---|---|
| Mean win rate | 41% |
| Median win rate | 40% |
| Minimum win rate | 0% |
| Maximum win rate | 100% |
| Win-Rate Bucket | Decks (n) |
|---|---|
| 0–10% | 93 |
| 10–20% | 281 |
| 20–30% | 263 |
| 30–40% | 602 |
| 40–50% | 302 |
| 50–60% | 403 |
| 60–70% | 130 |
| 70–80% | 121 |
| 80–90% | 14 |
| 90–100% | 8 |
The 30–40% bucket is by far the most populated (n = 602), followed by 50–60% (n = 403). The distribution is not symmetric: there are considerably more decks in the 10–40% range (637 decks in 10–20% + 20–30% + 0–10% = 637) than in the 60–100% range (273 decks). The mean (41%) slightly exceeds the median (40%), suggesting a right tail — a modest number of high-performing decks pulling the average upward. The full range, 0% to 100%, captures the real diversity of player submissions, from decks that lost every simulated game to one that won every one.
4.2 Standard (n = 108 decks with win-rate data)
| Summary Statistic | Value |
|---|---|
| Mean win rate | 37% |
| Median win rate | 40% |
| Minimum win rate | 0% |
| Maximum win rate | 80% |
| Win-Rate Bucket | Decks (n) |
|---|---|
| 0–10% | 8 |
| 10–20% | 23 |
| 20–30% | 9 |
| 30–40% | 26 |
| 40–50% | 12 |
| 50–60% | 19 |
| 60–70% | 7 |
| 70–80% | 4 |
| 80–90% | 0 |
| 90–100% | 0 |
Standard's distribution is more compressed (max 80%) and the mean (37%) sits below the median (40%), indicating a left-leaning tail — a meaningful cluster of low-performing decks pulling the average down. Buckets below 10 decks each are shown but noted: the 0–10% bucket (n = 8) falls below the minimum published cohort threshold of 10 and is shown for completeness but not used in downstream rate comparisons. No Standard deck in this dataset achieved a win rate above 80%.
5. Deck Iteration & Retests
Users who test the same deck multiple times provide a natural window into whether repeated testing correlates with outcome changes. This analysis compares the first completed simulation per deck with the most recent completed simulation for decks that were retested.
5.1 Commander Retests (n = 426 re-tested decks)
| Metric | Value |
|---|---|
| Re-tested decks | 426 |
| Decks that improved | 168 (39.4%) |
| Decks that declined | 192 (45.1%) |
| Decks roughly flat | 66 (15.5%) |
| Mean win-rate change | −1.6 percentage points |
In Commander, the re-test picture tilts slightly negative: 192 decks declined versus 168 that improved, for an average win-rate change of −1.6 points. This is a small and possibly noise-dominated figure — it does not establish that retesting causes decline, nor that users are making their decks worse. Changes between test runs reflect not only deck modifications but also randomness inherent in finite simulation samples.
5.2 Standard Retests (n = 14 re-tested decks)
| Metric | Value |
|---|---|
| Re-tested decks | 14 |
| Decks that improved | 5 (35.7%) |
| Decks that declined | 7 (50.0%) |
| Decks roughly flat | 2 (14.3%) |
| Mean win-rate change | +0.5 percentage points |
Standard's re-test sample is small (n = 14) and should be interpreted cautiously. The mean change of +0.5 percentage points is negligible at this sample size. The direction counts (5 improved, 7 declined) are consistent with small-sample noise.
6. Monthly Trends
Monthly figures track simulation volume and, where available, average win rate across the data window. Monthly cohorts consist of different decks submitted in each month; this is not a longitudinal panel of the same decks tracked over time. Changes in monthly win rate reflect changes in the mix of decks submitted, not any single deck's trajectory.
6.1 Commander by Month
| Month | Sims | Decks | Games | Win Rate |
|---|---|---|---|---|
| 2026-05 | 74 | 51 | 0 | — |
| 2026-06 | 135 | 103 | 60 | 38.3% |
| 2026-07 | 318 | 240 | 4,646 | 42.9% |
| 2026-08 | 346 | 247 | 5,118 | 40.8% |
| 2026-09 | 2,170 | 1,686 | 31,643 | 40.5% |
| 2026-10* | 102 | 65 | 1,464 | 45.2% |
*2026-10 is a partial month (snapshot taken 2026-10-01).
May 2026 and the June Commander rows show zero games despite non-zero sim and deck counts; these rows reflect simulation records with no completed game data (zero games in the denominator, so win rate is suppressed). September 2026 is the dominant month by volume: 2,170 simulations across 1,686 decks, producing 31,643 games. The dataset is heavily September-weighted, which means overall statistics are largely driven by September-vintage submissions.
6.2 Standard by Month
| Month | Sims | Decks | Games | Win Rate |
|---|---|---|---|---|
| 2026-06 | 24 | 20 | 0 | — |
| 2026-07 | 31 | 26 | 427 | 33.3% |
| 2026-08 | 29 | 22 | 435 | 31.7% |
| 2026-09 | 69 | 59 | 1,025 | 41.6% |
| 2026-10* | — | — | — | — |
June 2026 Standard shows zero games for the same reason as Commander May. No Standard data is available for May or October in this snapshot. September shows the highest Standard win rate (41.6%, n = 59 decks, 1,025 games) and the largest volume. The variation in monthly Standard win rates (33.3% → 41.6%) almost certainly reflects changes in the mix of submitted decks rather than any trend in deck quality.
7. Top Commanders
Among Commander decks where the submitted commander is identified and ≥ 10 distinct decks using that commander appear in the dataset, the following usage and win-rate figures are measured. Win rate here is the containing-deck win rate: wins ÷ games for all decks in the cohort that use the listed commander. These are correlational measures — a commander's win rate reflects the decks built around it in this specific dataset, not any inherent property of the card.
| Commander | Decks (n) | Games | Win Rate |
|---|---|---|---|
| Meren of Clan Nel Toth | 36 | 739 | 34.5% |
| Y'shtola, Night's Blessed | 33 | 541 | 39.0% |
| Krenko, Mob Boss | 21 | 382 | 44.0% |
| Edgar Markov | 20 | 346 | 59.5% |
| Vivi Ornitier | 19 | 435 | 26.4% |
| Kaalia of the Vast | 17 | 375 | 48.5% |
| Giada, Font of Hope | 17 | 468 | 76.1% |
| Thranduil, the Elvenking | 17 | 345 | 32.2% |
| Nekusar, the Mindrazer | 16 | 311 | 41.8% |
| Sephiroth, Fabled SOLDIER | 16 | 315 | 35.9% |
| Chatterfang, Squirrel General | 15 | 285 | 47.4% |
| The Ur-Dragon | 14 | 285 | 60.7% |
| Fire Lord Azula | 14 | 242 | 23.1% |
| Pantlaza, Sun-Favored | 14 | 226 | 52.7% |
| Teval, the Balanced Scale | 14 | 196 | 42.3% |
| Ureni of the Unwritten | 14 | 258 | 67.1% |
| Korvold, Fae-Cursed King | 14 | 285 | 35.1% |
| Cosmic Spider-Man | 14 | 286 | 55.6% |
| Cloud, Ex-SOLDIER | 14 | 241 | 45.2% |
| Jin Sakai, Ghost of Tsushima | 14 | 300 | 39.7% |
Most popular by deck count: Meren of Clan Nel Toth (n = 36 decks) is the most frequently submitted commander, followed by Y'shtola, Night's Blessed (n = 33) and Krenko, Mob Boss (n = 21).
Highest observed win rate among qualifying commanders: Giada, Font of Hope (76.1%, n = 17 decks, 468 games) leads, followed by Ureni of the Unwritten (67.1%, n = 14 decks, 258 games) and The Ur-Dragon (60.7%, n = 14 decks, 285 games).
The gap between popularity and win rate is striking and expected. Meren of Clan Nel Toth — the most submitted commander — posts a below-average 34.5% win rate. Giada, Font of Hope — among the highest win-rate commanders — is submitted by only 17 users. This split between what players choose and what wins in simulation is descriptive of submission behavior and simulation outcomes; it carries no evaluative implication and no causal weight. Popular commanders may be popular for reasons that have nothing to do with gauntlet win rate.
Lowest observed win rate among qualifying commanders: Fire Lord Azula (23.1%, n = 14 decks, 242 games) and Vivi Ornitier (26.4%, n = 19 decks, 435 games) sit at the bottom of this list.
8. Top Cards: Popularity vs. Containing-Deck Win Rate
Card figures report two things: how widely a card is played (deck count) and what win rate the decks containing it achieved. These are not the same signal. A card can appear in many decks that have average outcomes; a card can appear in few decks that all happen to have exceptional outcomes. Basic lands are excluded from all card tables throughout this report. Only cards in ≥ 10 decks are shown.
8.1 Commander — Most Widely Played Cards
| Card | Decks (n) | Games | Win Rate |
|---|---|---|---|
| Sol Ring | 1,981 | 38,599 | 41.4% |
| Command Tower | 1,689 | 32,867 | 40.6% |
| Arcane Signet | 1,644 | 31,922 | 41.1% |
| Exotic Orchard | 873 | 17,456 | 39.7% |
| Swiftfoot Boots | 654 | 12,197 | 42.4% |
| Swords to Plowshares | 652 | 12,713 | 44.0% |
| Lightning Greaves | 634 | 12,197 | 42.8% |
| Reliquary Tower | 626 | 11,543 | 40.5% |
| Path of Ancestry | 585 | 11,347 | 44.5% |
| Evolving Wilds | 543 | 10,391 | 39.9% |
| Path to Exile | 524 | 10,419 | 44.2% |
| Cultivate | 465 | 8,621 | 44.1% |
| Fellwar Stone | 462 | 9,219 | 39.8% |
| Counterspell | 461 | 9,181 | 38.7% |
| Rogue's Passage | 398 | 7,395 | 43.2% |
| Heroic Intervention | 393 | 7,365 | 44.2% |
| Terramorphic Expanse | 388 | 7,536 | 40.3% |
| Blasphemous Act | 377 | 7,443 | 41.7% |
| Bojuka Bog | 364 | 7,222 | 39.0% |
| Farseek | 344 | 6,734 | 43.2% |
Sol Ring is the most ubiquitous card in the Commander dataset by a wide margin (n = 1,981 decks, representing 83.8% of all 2,364 Commander decks submitted). Its containing-deck win rate (41.4%) is close to the Commander format average (41.0%), which is unsurprising given how near-universal its inclusion is — its cohort approximates the full Commander population.
Cards with meaningfully above-average containing-deck win rates relative to the Commander 41.0% baseline include Path of Ancestry (44.5%, n = 585), Swords to Plowshares (44.0%, n = 652), Path to Exile (44.2%, n = 524), Heroic Intervention (44.2%, n = 393), and Cultivate (44.1%, n = 465). Cards below the baseline include Counterspell (38.7%, n = 461) and Bojuka Bog (39.0%, n = 364). These are correlational differences — they describe what winning decks happened to contain, not what makes decks win.
8.2 Standard — Most Widely Played Cards (n ≥ 10 decks)
| Card | Decks (n) | Games | Win Rate |
|---|---|---|---|
| Lightning Bolt | 29 | 435 | 42.5% |
| Rift Bolt | 24 | 360 | 47.2% |
| Lava Spike | 24 | 360 | 47.2% |
| Searing Blaze | 24 | 360 | 47.2% |
| Inspiring Vantage | 23 | 345 | 47.2% |
| Sol Ring | 23 | 361 | 30.7% |
| Goblin Guide | 23 | 345 | 47.2% |
| Shard Volley | 23 | 345 | 47.2% |
| Skullcrack | 23 | 345 | 47.2% |
| Monastery Swiftspear | 23 | 345 | 47.2% |
| Chain Lightning | 23 | 345 | 47.2% |
| Eidolon of the Great Revel | 23 | 345 | 47.2% |
| Arcane Signet | 18 | 270 | 27.0% |
| Command Tower | 16 | 240 | 24.2% |
| Evolving Wilds | 13 | 201 | 31.3% |
| Counterspell | 12 | 225 | 18.7% |
| Llanowar Elves | 11 | 195 | 37.9% |
| Dark Ritual | 10 | 153 | 22.2% |
The Standard card table tells a vivid structural story: a cluster of twelve cards (Rift Bolt, Lava Spike, Searing Blaze, Inspiring Vantage, Goblin Guide, Shard Volley, Skullcrack, Monastery Swiftspear, Chain Lightning, Eidolon of the Great Revel) all share an identical containing-deck win rate of 47.2%, each with exactly 23 decks and 345 games. This is a signature of these cards co-appearing in the same deck archetype — they are all components of the same (or highly overlapping) Standard decks in this dataset.
Sol Ring in Standard (30.7%, n = 23 decks) and Arcane Signet (27.0%, n = 18 decks) show below-average containing-deck win rates. These cards are designed for Commander and appear in Standard submissions that may structurally underperform; the correlation is format-contextual and not a judgment on the cards.
9. Construction Correlates: Land Ratio, Color Count, and Card-Type Mix
The following construction metrics describe associations between deck-building choices and win rate in this dataset. All figures are correlational. The dataset is a self-selected, non-random sample of player submissions; confounders are numerous and unobserved.
9.1 Land Ratio vs. Win Rate
Land ratio is the percentage of a deck's cards that are lands, drawn from deck embedding data.
Commander (n = 2,207 decks across qualifying bands):
| Land % Band | Decks (n) | Win Rate |
|---|---|---|
| ~25% | 44 | 23.5% |
| ~30% | 326 | 36.4% |
| ~35% | 1,391 | 42.1% |
| ~40% | 426 | 43.1% |
| ~45% | 20 | 47.8% |
A monotonic positive association between land percentage and win rate is observable across the Commander bands in this dataset. The ~25% band (n = 44) posts the lowest win rate (23.5%); the ~45% band (n = 20) posts the highest (47.8%). The ~35% band is by far the most populated (n = 1,391). This is a descriptive correlation in a self-selected sample.
Standard (qualifying bands only, n ≥ 10 decks):
| Land % Band | Decks (n) | Win Rate |
|---|---|---|
| ~35% | 32 | 31.6% |
| ~40% | 54 | 44.6% |
Two qualifying Standard bands show a 13-point gap (31.6% vs. 44.6%), consistent with the Commander pattern, though the Standard sample is considerably smaller.
9.2 Color Count vs. Win Rate
Commander:
| Color Count | Decks (n) | Win Rate |
|---|---|---|
| 0 (colorless) | 26 | 46.3% |
| 1 | 360 | 46.6% |
| 2 | 762 | 39.5% |
| 3 | 826 | 40.0% |
| 4 | 55 | 41.0% |
| 5 | 188 | 40.4% |
Single-color (46.6%, n = 360) and colorless (46.3%, n = 26) Commander decks show the highest containing-deck win rates in this cut. Two-color decks (39.5%, n = 762) and three-color decks (40.0%, n = 826) cluster below average. Five-color decks (40.4%, n = 188) sit near average. This pattern is associational; single-color decks may have more consistent mana and tighter game plans, or they may simply attract a different type of player submission — the data cannot distinguish these explanations.
Standard (qualifying bands, n ≥ 10 decks):
| Color Count | Decks (n) | Win Rate |
|---|---|---|
| 1 | 22 | 38.4% |
| 2 | 57 | 40.2% |
| 3 | 26 | 28.9% |
| 4–5 | <10 | suppressed |
Three-color Standard decks (28.9%, n = 26) show the lowest win rate in this cut. Two-color decks (40.2%, n = 57) lead. Four-color and five-color Standard bands fall below the 10-deck minimum and are suppressed.
9.3 Card-Type Mix (Full Dataset)
Average composition percentages across all decks in each format, measured from deck embeddings:
Commander (n = 2,217 decks):
| Card Type | Average % |
|---|---|
| Lands | 35.3% |
| Creatures | 27.6% |
| Instants | 10.3% |
| Sorceries | 8.1% |
| Instant + Sorcery combined | 18.4% |
| Artifacts | 10.6% |
| Enchantments | 7.4% |
| Planeswalkers | 0.8% |
The instant/sorcery breakdown (10.3% instants, 8.1% sorceries) is drawn from a split sub-cohort of n = 2,217 decks where both fields are available. These figures may not sum exactly to the combined 18.4% figure due to unclassifiable spells that appear only in the combined bucket and minor backfill drift.
Standard (n = 108 decks):
| Card Type | Average % |
|---|---|
| Lands | 35.6% |
| Creatures | 29.6% |
| Instants | 13.7% |
| Sorceries | 10.4% |
| Instant + Sorcery combined | 24.1% |
| Artifacts | 4.7% |
| Enchantments | 5.6% |
| Planeswalkers | 0.7% |
Standard decks are more creature- and spell-heavy than Commander decks, with lower artifact and enchantment ratios — consistent with Standard's typical card-pool character.
10. Winning-Recipe: Win Brackets Compared
Decks are divided into three win-rate brackets and compared on construction averages. This is purely descriptive — bracket membership is defined by the simulation outcome, and the construction features are averages within each bracket. No causal inference is warranted.
10.1 Commander Win Brackets
| Bracket | Decks (n) | Avg Win Rate | Avg Land % | Avg Creature % | Avg Spell % | Avg Art+Ench % | Avg Mana Value |
|---|---|---|---|---|---|---|---|
| High (>55%) | 453 | 67.3% | 35.9% | 29.4% | 15.7% | 18.3% | 3.29 |
| Mid (40–55%) | 820 | 46.0% | 35.4% | 28.1% | 17.7% | 18.0% | 3.21 |
| Low (<40%) | 944 | 24.1% | 34.8% | 26.2% | 20.2% | 17.9% | 3.06 |
Across these brackets, the high-win-rate Commander group (n = 453) carries 1.1 more land percentage points than the low-win-rate group (n = 944) — 35.9% vs. 34.8%. High-win-rate decks also average more creatures (29.4% vs. 26.2%) and fewer spells (15.7% vs. 20.2%). Average mana value is higher in high-win-rate decks (3.29 vs. 3.06 in low-win-rate decks). These are descriptive correlations across submitted decks; they carry no causal force and should not be read as build guidelines.
10.2 Standard Win Brackets
| Bracket | Decks (n) | Avg Win Rate | Avg Land % | Avg Creature % | Avg Spell % | Avg Art+Ench % | Avg Mana Value |
|---|---|---|---|---|---|---|---|
| High (>55%) | 16 | 66.9% | 37.6% | 33.3% | 23.4% | 4.7% | 2.45 |
| Mid (40–55%) | 40 | 46.6% | 38.1% | 27.9% | 28.8% | 5.4% | 2.20 |
| Low (<40%) | 52 | 20.5% | 32.9% | 29.8% | 20.7% | 15.7% | 2.65 |
Standard's high-win-rate decks (n = 16; note the small sample) show 4.7 more land percentage points than low-win-rate decks (37.6% vs. 32.9%) and a lower average mana value (2.45 vs. 2.65). The low-win-rate Standard group has a notably higher artifact + enchantment share (15.7%), plausibly connected to the prevalence of mana-accelerant artifacts (Sol Ring, Arcane Signet) in lower-performing Standard submissions — cards better suited to Commander's 100-card singleton context. All figures are descriptive correlations in a self-selected, small Standard sample.
11. Card-Category Insights (Heuristic-Labelled)
Heuristic caveat: Category membership (sacrifice outlets, tutor/search, discard, reanimation) is determined by a keyword heuristic applied to oracle text and pre-existing card flags. Edge cases may be mislabelled. These figures describe correlations between category presence and containing-deck win rate — they are not causal claims.
11.1 Commander Card Categories
| Category | With (n decks) | Win Rate (with) | Without (n decks) | Win Rate (without) | Delta |
|---|---|---|---|---|---|
| Sacrifice outlets | 2,209 | 41.0% | 8 | suppressed (n < 10) | — |
| Search / tutor effects | 2,115 | 41.0% | 102 | 41.5% | −0.5 pts |
| Discard effects | 1,996 | 40.4% | 221 | 46.6% | −6.2 pts |
| Reanimation effects | 1,356 | 40.5% | 861 | 41.8% | −1.3 pts |
The most notable Commander pattern: decks including discard effects (n = 1,996) post a 6.2-point lower win rate than the 221 decks without them (40.4% vs. 46.6%). The "without" group (221 decks) is smaller, representing decks that committed to a discard-free strategy. Sacrifice outlets are so prevalent (2,209 of 2,217 qualifying decks) that the "without" group (n = 8) falls below the minimum cohort threshold — the comparison is unsurfaceable.
The tutor/search delta is negligible (−0.5 points); decks with tutors (41.0%, n = 2,115) perform nearly identically to those without (41.5%, n = 102). Reanimation shows a modest −1.3-point difference.
Interpretation is constrained: a deck that includes discard effects is already a different kind of deck from one that excludes them. These correlations describe simultaneous associations with many other unobserved construction choices.
11.2 Standard Card Categories
| Category | With (n decks) | Win Rate (with) | Without (n decks) | Win Rate (without) | Delta |
|---|---|---|---|---|---|
| Sacrifice outlets | 98 | 36.2% | 10 | 45.0% | −8.8 pts |
| Discard effects | 62 | 34.5% | 46 | 40.4% | −5.9 pts |
| Search / tutor effects | 58 | 29.6% | 50 | 45.7% | −16.1 pts |
| Reanimation effects | 38 | 35.7% | 70 | 37.8% | −2.1 pts |
Standard's most striking categorical association is the 16.1-point gap for tutor/search effects: decks containing tutors post 29.6% (n = 58) versus 45.7% (n = 50) for those without. This is the largest categorical delta in either format. At the same time, the "with" group (58 decks) is correlated with other construction choices — tutors appear in certain deck archetypes — and the 16.1-point gap likely reflects a confounded cluster of design decisions rather than any single card's impact.
12. Most Impactful Cards: Decision Impact (Counterfactual Proxy)
Counterfactual proxy caveat: Decision-impact figures are derived from replayed decision snapshots comparing the line actually taken by the AI to the engine's next-best alternative. This measures play-line quality within the simulation, not damage, kills, or a direct contribution to winning. A positive value means the played line exceeded the engine's next-best alternative; a negative value means the alternative would have scored better. These figures are not comparable to board-impact figures from Section 13 (different measurement methodology). Correlation is not causation.
12.1 Commander — Decision Impact (Counterfactual)
Only one card in the Commander cohort cleared both the 10-deck and 25-observation floors for this metric at the time of this snapshot:
| Card | Decks (n) | Observations | Avg Decision Impact |
|---|---|---|---|
| Lightning Greaves | 10 | 16 | −169.19 counterfactual points |
The negative internal score (−169.19) means the engine's chosen line — involving Lightning Greaves — trailed its own next-best alternative by an average of 169.19 counterfactual points across 16 recorded decisions in 10 distinct decks. In user-facing terms: the alternative line the engine considered would, on average, have scored better. This does not mean Lightning Greaves is a weak card; it means that in the specific game states where the engine recorded a decision involving it, the alternative it set aside would have yielded a better board evaluation. The observation count (16) is small; treat this figure as preliminary.
12.2 Standard — Decision Impact
No Standard cards cleared both floors (≥ 10 decks, ≥ 25 observations) for the decision-impact metric in this snapshot. This section will expand as the Standard sample grows.
13. Card Performance Roll-Up: Board Impact (Board-State Proxy)
Board-state proxy caveat: Card performance figures pool the board-quality delta around turns a card is seen, across every deck in the cohort running that card. Positive = board improved; negative = board deteriorated. This is a board-state proxy, not damage, kills, or a causal win claim. Ranking credibility scales with breadth (more distinct decks) and observation count (more recorded appearances). Only cards in ≥ 10 decks with ≥ 25 observations qualify. Type and color buckets pool very different cards; multicolor cards count toward each of their colors. Do not confuse the sign convention here (positive = improved) with the counterfactual impact convention in Section 12.
13.1 Commander — Top Performers
2,361 cards qualify in Commander (≥ 25 observations).
| Card | Decks (n) | Observations | Avg Board Impact |
|---|---|---|---|
| Sazh's Chocobo | 13 | 69 | +832.17 |
| Entish Restoration | 34 | 116 | +420.25 |
| Vorel of the Hull Clade | 16 | 68 | +380.90 |
| Time Warp | 13 | 46 | +281.85 |
| Hullbreaker Horror | 40 | 140 | +248.77 |
| Necklace of Girion | 14 | 29 | +188.69 |
| Bristly Bill, Spine Sower | 47 | 398 | +78.26 |
| Steelbane Hydra | 27 | 106 | +56.69 |
| Pongify | 84 | 245 | +49.11 |
| Sakura-Tribe Elder | 145 | 480 | +47.28 |
| Fathom Mage | 11 | 44 | +38.27 |
| Lyra Dawnbringer | 35 | 276 | +36.83 |
| Ulamog, the Ceaseless Hunger | 20 | 83 | +35.01 |
| Ureni, the Song Unending | 11 | 74 | +32.59 |
| Sephara, Sky's Blade | 26 | 193 | +31.76 |
The highest-confidence entries on this list are those with the broadest deck coverage: Sakura-Tribe Elder (n = 145 decks, 480 observations) and Pongify (n = 84 decks, 245 observations) offer the most robust board-impact figures. Sazh's Chocobo leads the ranking at +832.17 but appears in only 13 decks with 69 observations — its figure is real but should be treated as more volatile.
13.2 Commander — Bottom Performers
| Card | Decks (n) | Observations | Avg Board Impact |
|---|---|---|---|
| The Sackville-Bagginses | 13 | 92 | −39.33 |
| Planar Genesis | 10 | 35 | −32.06 |
| Forced Fruition | 19 | 46 | −28.48 |
| Junji, the Midnight Sky | 25 | 133 | −22.25 |
| Iron Man, Master of Machines | 13 | 103 | −21.77 |
| Jarad, Golgari Lich Lord | 21 | 98 | −19.93 |
| Fractured Sanity | 13 | 47 | −19.00 |
| White Sun's Twilight | 18 | 29 | −17.90 |
| Venser's Journal | 13 | 32 | −17.34 |
| Dawnsire, Sunstar Dreadnought | 13 | 44 | −16.34 |
| Aura Shards | 27 | 75 | −14.39 |
| The Council of Four | 14 | 130 | −13.74 |
| Quicksilver Amulet | 12 | 38 | −13.63 |
| Evacuation | 13 | 34 | −13.53 |
| Pest Rescuer | 11 | 53 | −13.49 |
13.3 Commander — By Card Type
| Card Type | Cards | Observations | Decks (min) | Avg Board Impact |
|---|---|---|---|---|
| Creatures | 7,782 | 319,373 | 310 | +3.66 |
| Instants | 1,456 | 41,733 | 612 | +2.93 |
| Enchantments | 1,543 | 36,750 | 205 | +0.99 |
| Planeswalkers | 218 | 7,663 | 53 | +0.94 |
| Sorceries | 1,414 | 32,426 | 371 | +0.20 |
| Artifacts | 1,353 | 68,553 | 1,850 | −1.48 |
| Nonbasic Lands | 106 | 1,012 | 105 | −2.30 |
Creatures and instants lead by type; artifacts and nonbasic lands trail. Artifacts — by volume the second-most observed type (68,553 observations) — produce a negative pooled figure (−1.48). This aggregate includes mana rocks and utility artifacts of wildly varying board contribution; a negative type-average does not characterize any specific artifact.
13.4 Commander — By Color Identity
| Color | Cards | Observations | Decks (min) | Avg Board Impact |
|---|---|---|---|---|
| Green | 3,389 | 138,088 | 371 | +5.40 |
| Red | 2,991 | 106,328 | 283 | +3.27 |
| White | 3,187 | 122,883 | 612 | +3.07 |
| Blue | 3,069 | 108,088 | 298 | +2.49 |
| Black | 3,280 | 133,263 | 274 | +0.91 |
| Colorless | 1,336 | 74,175 | 1,850 | −0.55 |
All five colors show positive pooled board impact; colorless cards aggregate to a slight negative. Green leads (+5.40). Multicolor cards count toward each of their colors; every aggregate conflates very different cards in the same bucket.
13.5 Standard — Top Performers
13 qualifying cards in Standard (≥ 10 decks, ≥ 25 observations).
| Card | Decks (n) | Observations | Avg Board Impact |
|---|---|---|---|
| Eidolon of the Great Revel | 23 | 399 | +5.11 |
| Goblin Guide | 23 | 469 | +4.90 |
| Monastery Swiftspear | 23 | 408 | +3.87 |
| Lightning Bolt | 29 | 322 | +3.80 |
| Rift Bolt | 24 | 263 | +2.49 |
| Lava Spike | 24 | 157 | +2.28 |
| Chain Lightning | 23 | 234 | +1.54 |
| Skullcrack | 21 | 48 | +1.33 |
| Llanowar Elves | 10 | 216 | +0.19 |
| Sol Ring | 21 | 71 | −0.49 |
| Dark Ritual | 11 | 27 | −1.11 |
| Shard Volley | 22 | 88 | −3.23 |
| Arcane Signet | 16 | 41 | −3.54 |
The Standard card performance list is heavily shaped by the concentration of burn and aggro archetypes in the Standard submission pool. Eidolon of the Great Revel leads (n = 23 decks, 399 observations, +5.11). Arcane Signet and Sol Ring — Commander staples appearing in Standard submissions — sit at the bottom of the Standard board-impact ranking.
13.6 Standard — By Card Type
| Card Type | Cards | Observations | Decks (min) | Avg Board Impact |
|---|---|---|---|---|
| Creatures | 861 | 11,313 | 23 | +2.92 |
| Sorceries | 198 | 1,997 | 24 | +1.06 |
| Instants | 205 | 2,132 | 29 | +0.68 |
| Artifacts | 168 | 1,256 | 21 | −3.02 |
13.7 Standard — By Color Identity
| Color | Cards | Observations | Decks (min) | Avg Board Impact |
|---|---|---|---|---|
| Green | 399 | 4,401 | 10 | +2.93 |
| Red | 249 | 4,686 | 29 | +2.61 |
| Black | 491 | 3,780 | 11 | +1.36 |
| Colorless | 143 | 1,087 | 21 | −1.84 |
White and Blue Standard card-type cohorts fall below minimum deck floors for this dataset and are suppressed.
14. Color-Identity Breakdown
The following figures show win rates for decks containing each color in their identity. Colors are not mutually exclusive; a multicolor deck contributes to every color it contains. These are correlational figures in a self-selected sample.
14.1 Commander Color Presence (n ≥ 10 decks per color, all colors qualify)
| Color | Decks (n) | Games | Win Rate |
|---|---|---|---|
| White | 1,084 | 21,332 | 42.1% |
| Green | 1,064 | 20,391 | 41.8% |
| Red | 1,110 | 21,722 | 40.0% |
| Black | 1,210 | 24,013 | 38.9% |
| Blue | 1,054 | 20,351 | 38.8% |
Black is the most common color in Commander submissions (n = 1,210 decks, 51.2% of 2,364 submitted decks) but posts the second-lowest win rate (38.9%). White — the least common color among these five in Commander (n = 1,084) — has the highest win rate (42.1%). The 3.3-percentage-point spread from White to Blue is modest and the direction of the color-presence effect is heavily confounded by which commanders and archetypes drive each color's inclusion in this dataset.
14.2 Standard Color Presence (n ≥ 10 decks per color)
| Color | Decks (n) | Games | Win Rate |
|---|---|---|---|
| White | 57 | 1,020 | 42.2% |
| Green | 38 | 720 | 40.0% |
| Red | 54 | 889 | 37.8% |
| Black | 51 | 862 | 33.4% |
| Blue | 28 | 474 | 23.8% |
The Standard color-presence table shows a notably wider spread: White (42.2%, n = 57) to Blue (23.8%, n = 28) is an 18.4-point gap. Blue in Standard appears in the fewest decks (28) and posts the lowest win rate. Given the composition of the Standard submission pool — heavily tilted toward aggro and burn archetypes, which rarely run Blue — the Blue cohort is likely a structurally distinct archetype sample rather than a fair-comparison subset. All figures remain correlational and subject to selection effects.
15. Tempo & Game-Length Signals
Game length is measured in turns, drawn from simulation results. These are means of per-match turn counts across qualifying decks in each format.
15.1 Commander (n = 2,200 decks)
| Metric | Value |
|---|---|
| Average turns (mean across games) | 10.2 |
| Average shortest game observed | 8.3 turns |
| Average longest game observed | 12.2 turns |
Commander games in this dataset average 10.2 turns, with a typical range from 8.3 to 12.2. This is a rough tempo marker — it describes how long the AI-versus-AI engine took to resolve these matchups, not human game pace.
15.2 Standard (n = 104 decks)
| Metric | Value |
|---|---|
| Average turns (mean across games) | 9.4 |
| Average shortest game observed | 7.3 turns |
| Average longest game observed | 11.5 turns |
Standard games average 9.4 turns — slightly shorter than Commander, consistent with Standard's typically faster card designs and lower average mana values (2.20–2.65 across brackets, versus 3.06–3.29 in Commander). The range (7.3–11.5) is also narrower than Commander.
16. Matchup Structure: Deck vs. Opponent
This section examines how much of the observable win-rate variance is attributable to the identity of the specific meta opponent versus the deck's own profile. All figures are descriptive decompositions of simulated results. No figure here establishes a causal mechanism. Correlation is not causation.
16.1 Variance Decomposition
A games-weighted two-way decomposition of per-(deck, opponent) win-rate cells attributes total variance to three sources: the deck main effect (deck-driven variance), the opponent main effect (opponent-driven variance), and the residual (interaction + noise).
Commander (n = 384 decks, 23 opponents, 1,854 cells):
| Source | Share of Variance |
|---|---|
| Deck main effect | 47.7% |
| Opponent main effect | 14.6% |
| Residual (interaction + noise) | 37.8% |
Standard (n = 13 decks, 6 opponents, 61 cells):
| Source | Share of Variance |
|---|---|
| Deck main effect | 28.0% |
| Opponent main effect | 28.0% |
| Residual (interaction + noise) | 44.0% |
In Commander, the deck's own profile accounts for nearly half (47.7%) of win-rate variance across matchup cells — more than three times the opponent's share (14.6%). The residual (37.8%) captures matchup-specific interaction effects and simulation noise. This does not mean the opponent is unimportant; 14.6% is a substantial main effect given the 20.9-point matchup spread documented in Section 3. It means that, across the full range of Commander submissions and opponents in this dataset, deck variation explains more of the win-rate spread than opponent variation does.
In Standard, the picture is more balanced: deck and opponent each contribute 28.0% of variance. The larger residual (44.0%) is consistent with Standard's smaller sample (13 decks, 6 opponents, 61 cells) — more variance remains unexplained by either main effect.
These are descriptive decompositions, not significance tests. No formal hypothesis is being tested.
16.2 Matchup Inversion Rate
The matchup inversion rate is the share of comparable deck pairs — both decks strictly preferring the same two meta opponents to each other — whose preference order is reversed between those two opponents.
Commander (n = 360 decks, 421,253 comparable pairs across 10 opponent pairings):
37% of Commander deck pairs reverse their preference order between two meta opponents. For every three decks that agree on which of two opponents is harder, there is a meaningful minority whose ordering is flipped.
Standard (n = 12 decks, 486 comparable pairs across 10 opponent pairings):
38.5% of Standard deck pairs reverse their preference order between two meta opponents.
Both formats show inversion rates in the 37–38.5% range. This quantifies matchup specificity: which opponent a deck faces is not consistently ordered across all decks — a deck's "harder" opponent may be another deck's "easier" one. Higher inversion means the specific matchup matters more than any universal opponent ordering.
16.3 Power Score × Opponent Win-Rate Grid (Commander only)
The Commander dataset is large enough to publish a Power Score band × opponent win-rate grid. Three equal-size bands (approximately 127–128 decks each) are defined by Power Score range:
- Low: Power Score 0–32.3 (n = 127 decks)
- Mid: Power Score 32.6–46.9 (n = 127 decks)
- High: Power Score 47.2–90.2 (n = 128 decks)
Win rate by Power Score band and meta opponent (Commander):
| Opponent | Low (n≈122) | Mid (n≈123) | High (n≈121–123) |
|---|---|---|---|
| Breya Artifact Combo | 33.4% | 51.4% | 73.2% |
| Derevi Bant Control | 30.9% | 50.5% | 68.7% |
| Aesi Landfall | 22.3% | 35.8% | 52.1% |
| Atraxa Superfriends | 16.9% | 35.2% | 53.4% |
| Edgar Markov Vampires | 15.9% | 28.4% | 47.8% |
(Each cell carries approximately 981–1,029 games; minimum per-cell game floor = 5, all cells clear it.)
The gradient is consistent across every opponent: higher Power Score bands post higher win rates. The effect is largest against Breya Artifact Combo, where the spread from Low to High is 39.8 percentage points (33.4% → 73.2%), and smallest against Edgar Markov Vampires (15.9% → 47.8%, spread 31.9 points). Even in the High band, Edgar Markov Vampires remains the hardest opponent (47.8%) — the difficulty ordering of the gauntlet is preserved within each Power Score tier.
The Standard sample (n = 13 decks across 6 opponents, 61 cells) does not clear the minimum deck floor for a meaningful PS × opponent grid after band subdivision and is suppressed for Standard.
16.4 Per-Opponent Spearman Correlation: Power Score vs. Win Rate
Spearman rank correlation between deck Power Score and per-opponent win rate, computed over decks clearing the per-cell game floor.
Commander:
| Opponent | Decks (n) | Spearman ρ |
|---|---|---|
| Atraxa Superfriends | 367 | 0.75 |
| Derevi Bant Control | 366 | 0.75 |
| Breya Artifact Combo | 366 | 0.74 |
| Edgar Markov Vampires | 367 | 0.68 |
| Aesi Landfall | 368 | 0.62 |
All five Commander opponents show positive, moderate-to-strong Spearman correlations between Power Score and win rate (ρ = 0.62–0.75). The ranking of correlations itself is informative: Power Score tracks win rate most tightly against Atraxa Superfriends and Derevi Bant Control (both ρ = 0.75) and most loosely against Aesi Landfall (ρ = 0.62). Aesi Landfall's high draw rate (documented in Section 3) may attenuate the rank correlation by compressing meaningful outcome differences.
Standard:
| Opponent | Decks (n) | Spearman ρ |
|---|---|---|
| Dimir Midrange | 12 | 0.84 |
| Azorius Tempo | 12 | 0.82 |
| Temur Harmonizer Combo | 12 | 0.66 |
| Jeskai Control | 12 | 0.29 |
| Mono Red Aggro | 12 | 0.20 |
Standard's pattern is markedly less uniform. Power Score correlates strongly with win rate against Dimir Midrange (ρ = 0.84) and Azorius Tempo (ρ = 0.82) but is nearly uncorrelated against Mono Red Aggro (ρ = 0.20) and weakly correlated against Jeskai Control (ρ = 0.29). This suggests that for these two Standard opponents, a deck's Power Score rank is a poor predictor of how well it performs — the specific matchup interaction dominates the deck's general profile. The Standard Spearman figures are based on n = 12 decks and should be read as directional rather than precise.
All Spearman figures are correlational, measured within simulated results, and make no causal claim.
17. Per-Deck Appendix (Anonymized)
Every row in this appendix represents one qualifying deck, published under a rotating pseudonym that is randomly reassigned every edition. Pseudonyms carry no link to any account, deck name, or identity — they exist solely to allow readers to cross-reference rows within this edition. Rows should never be cross-referenced with rows from prior or future editions.
Published fields, deliberately coarsened to resist re-identification:
- Pseudonym (e.g., C-001)
- Power Score: rounded to nearest 5
- Game count: banded (20–49 / 50–99 / 100–199 / 200+)
- Win rate: rounded to whole percentage points
- Opponents faced: count only
- Matchup spread: standard deviation of per-opponent win rates, rounded to whole points (shown only for decks facing 3+ opponents; null otherwise)
Minimum per-deck game floor: 20 games. Minimum appendix floor: 10 published decks. Both are cleared. No deck names, IDs, decklists, or owner data appear.
17.1 Commander Appendix Summary (n = 380 decks)
| Metric | Value |
|---|---|
| Published decks | 380 |
| Win rate range | 0% – 87% |
| Median win rate | 40% |
| Median matchup spread | 16 percentage points |
| Decks with spread computed (3+ opponents) | 366 |
OLS residual outliers (widest over/underperformance vs. Power Score expectation, fit over the 380 published rows):
- Widest overperformance: C-002 — Power Score 45, win rate 60%, OLS-expected 45%, residual +15 points.
- Widest underperformance: C-021 — Power Score 60, win rate 45%, OLS-expected 59%, residual −14 points.
C-002 wins 15 percentage points more than its Power Score predicts; C-021 wins 14 percentage points less. Both are single decks, and residuals of this magnitude are not unusual in a 380-deck population — they mark the tails of the distribution, not confirmed anomalies.
17.2 Commander Per-Deck Table
The full 380-row Commander table appears below. Spread is null (—) for decks facing fewer than 3 qualifying opponents.
| ID | PS (±5) | Games Band | Win Rate | Opponents | Spread |
|---|---|---|---|---|---|
| C-001 | 30 | 20–49 | 30% | 5 | 22 |
| C-002 | 45 | 20–49 | 60% | 5 | 13 |
| C-003 | 30 | 20–49 | 30% | 5 | 22 |
| C-004 | 85 | 20–49 | 80% | 5 | 19 |
| C-005 | 45 | 20–49 | 42% | 5 | 8 |
| C-006 | 60 | 20–49 | 57% | 5 | 17 |
| C-007 | 40 | 20–49 | 40% | 5 | 9 |
| C-008 | 55 | 20–49 | 56% | 5 | 12 |
| C-009 | 25 | 20–49 | 27% | 5 | 17 |
| C-010 | 60 | 20–49 | 58% | 5 | 20 |
| C-011 | 15 | 50–99 | 19% | 5 | 14 |
| C-012 | 65 | 20–49 | 63% | 5 | 12 |
| C-013 | 65 | 50–99 | 63% | 5 | 19 |
| C-014 | 65 | 20–49 | 63% | 5 | 16 |
| C-015 | 45 | 20–49 | 47% | 5 | 24 |
| C-016 | 40 | 20–49 | 40% | 5 | 13 |
| C-017 | 20 | 20–49 | 24% | 5 | 23 |
| C-018 | 35 | 20–49 | 37% | 5 | 19 |
| C-019 | 60 | 20–49 | 60% | 5 | 23 |
| C-020 | 20 | 20–49 | 20% | 5 | 27 |
| C-021 | 60 | 20–49 | 45% | 1 | — |
| C-022 | 40 | 20–49 | 40% | 5 | 23 |
| C-023 | 35 | 100–199 | 34% | 5 | 11 |
| C-024 | 30 | 20–49 | 30% | 5 | 12 |
| C-025 | 45 | 20–49 | 47% | 5 | 12 |
| C-026 | 35 | 20–49 | 37% | 5 | 22 |
| C-027 | 40 | 20–49 | 40% | 5 | 25 |
| C-028 | 45 | 20–49 | 43% | 5 | 17 |
| C-029 | 25 | 20–49 | 25% | 1 | — |
| C-030 | 35 | 20–49 | 37% | 5 | 19 |
| C-031 | 60 | 20–49 | 60% | 5 | 19 |
| C-032 | 25 | 50–99 | 25% | 5 | 15 |
| C-033 | 50 | 20–49 | 50% | 5 | 0 |
| C-034 | 35 | 20–49 | 37% | 5 | 16 |
| C-035 | 35 | 20–49 | 33% | 5 | 11 |
| C-036 | 30 | 20–49 | 31% | 5 | 11 |
| C-037 | 20 | 20–49 | 23% | 5 | 17 |
| C-038 | 45 | 20–49 | 43% | 5 | 8 |
| C-039 | 35 | 20–49 | 37% | 5 | 12 |
| C-040 | 30 | 20–49 | 35% | 1 | — |
| C-041 | 35 | 20–49 | 30% | 5 | 22 |
| C-042 | 25 | 20–49 | 27% | 5 | 8 |
| C-043 | 25 | 20–49 | 27% | 5 | 27 |
| C-044 | 65 | 20–49 | 63% | 5 | 24 |
| C-045 | 30 | 20–49 | 30% | 5 | 22 |
| C-046 | 50 | 20–49 | 50% | 5 | 11 |
| C-047 | 35 | 20–49 | 37% | 5 | 19 |
| C-048 | 25 | 20–49 | 23% | 5 | 13 |
| C-049 | 60 | 20–49 | 60% | 5 | 25 |
| C-050 | 30 | 20–49 | 33% | 5 | 28 |
| C-051 | 75 | 20–49 | 73% | 5 | 23 |
| C-052 | 25 | 20–49 | 24% | 5 | 11 |
| C-053 | 30 | 20–49 | 33% | 5 | 11 |
| C-054 | 40 | 20–49 | 40% | 5 | 22 |
| C-055 | 65 | 20–49 | 70% | 5 | 27 |
| C-056 | 25 | 20–49 | 27% | 5 | 20 |
| C-057 | 5 | 20–49 | 12% | 5 | 9 |
| C-058 | 45 | 20–49 | 47% | 5 | 12 |
| C-059 | 55 | 20–49 | 53% | 5 | 32 |
| C-060 | 35 | 200+ | 34% | 5 | 5 |
| C-061 | 0 | 20–49 | 0% | 5 | 0 |
| C-062 | 70 | 20–49 | 67% | 5 | 28 |
| C-063 | 45 | 50–99 | 47% | 5 | 17 |
| C-064 | 45 | 20–49 | 43% | 5 | 17 |
| C-065 | 55 | 20–49 | 53% | 5 | 24 |
| C-066 | 50 | 20–49 | 50% | 5 | 30 |
| C-067 | 30 | 100–199 | 30% | 5 | 12 |
| C-068 | 60 | 20–49 | 60% | 5 | 13 |
| C-069 | 10 | 20–49 | 16% | 5 | 5 |
| C-070 | 50 | 20–49 | 51% | 5 | 9 |
| C-071 | 40 | 20–49 | 40% | 5 | 25 |
| C-072 | 20 | 20–49 | 23% | 5 | 17 |
| C-073 | 40 | 20–49 | 40% | 5 | 17 |
| C-074 | 40 | 20–49 | 43% | 5 | 23 |
| C-075 | 35 | 20–49 | 36% | 5 | 16 |
| C-076 | 25 | 20–49 | 27% | 5 | 17 |
| C-077 | 50 | 20–49 | 51% | 5 | 11 |
| C-078 | 50 | 20–49 | 57% | 5 | 23 |
| C-079 | 25 | 50–99 | 27% | 5 | 21 |
| C-080 | 45 | 20–49 | 50% | 1 | — |
| C-081 | 50 | 50–99 | 50% | 6 | 18 |
| C-082 | 20 | 20–49 | 20% | 5 | 11 |
| C-083 | 45 | 20–49 | 47% | 5 | 19 |
| C-084 | 70 | 20–49 | 67% | 5 | 11 |
| C-085 | 10 | 20–49 | 13% | 5 | 4 |
| C-086 | 40 | 20–49 | 40% | 5 | 13 |
| C-087 | 15 | 20–49 | 20% | 5 | 27 |
| C-088 | 30 | 50–99 | 33% | 5 | 9 |
| C-089 | 70 | 20–49 | 67% | 5 | 21 |
| C-090 | 75 | 50–99 | 73% | 5 | 13 |
| C-091 | 15 | 100–199 | 18% | 5 | 13 |
| C-092 | 15 | 50–99 | 16% | 5 | 8 |
| C-093 | 55 | 20–49 | 56% | 5 | 10 |
| C-094 | 10 | 20–49 | 15% | 5 | 16 |
| C-095 | 0 | 20–49 | 0% | 5 | 0 |
| C-096 | 25 | 20–49 | 27% | 5 | 20 |
| C-097 | 70 | 20–49 | 67% | 5 | 0 |
| C-098 | 40 | 20–49 | 43% | 5 | 23 |
| C-099 | 15 | 20–49 | 17% | 5 | 11 |
| C-100 | 45 | 20–49 | 45% | 1 | — |
| C-101 | 60 | 20–49 | 65% | 1 | — |
| C-102 | 65 | 20–49 | 62% | 5 | 11 |
| C-103 | 25 | 20–49 | 27% | 5 | 8 |
| C-104 | 45 | 20–49 | 47% | 5 | 27 |
| C-105 | 50 | 20–49 | 50% | 5 | 18 |
| C-106 | 65 | 50–99 | 64% | 5 | 21 |
| C-107 | 50 | 100–199 | 49% | 5 | 16 |
| C-108 | 55 | 50–99 | 55% | 5 | 9 |
| C-109 | 30 | 20–49 | 30% | 5 | 32 |
| C-110 | 60 | 20–49 | 57% | 5 | 27 |
| C-111 | 15 | 100–199 | 19% | 5 | 10 |
| C-112 | 35 | 20–49 | 33% | 5 | 11 |
| C-113 | 50 | 50–99 | 48% | 5 | 8 |
| C-114 | 20 | 20–49 | 23% | 5 | 13 |
| C-115 | 45 | 20–49 | 47% | 5 | 29 |
| C-116 | 45 | 20–49 | 47% | 5 | 13 |
| C-117 | 70 | 20–49 | 71% | 5 | 16 |
| C-118 | 20 | 20–49 | 23% | 5 | 17 |
| C-119 | 30 | 20–49 | 33% | 5 | 11 |
| C-120 | 45 | 100–199 | 45% | 5 | 11 |
| C-121 | 35 | 20–49 | 37% | 5 | 19 |
| C-122 | 75 | 200+ | 72% | 5 | 11 |
| C-123 | 55 | 20–49 | 57% | 5 | 25 |
| C-124 | 55 | 20–49 | 53% | 5 | 19 |
| C-125 | 30 | 20–49 | 33% | 5 | 16 |
| C-126 | 75 | 20–49 | 73% | 5 | 8 |
| C-127 | 50 | 20–49 | 51% | 5 | 15 |
| C-128 | 30 | 100–199 | 29% | 5 | 11 |
| C-129 | 45 | 20–49 | 47% | 5 | 7 |
| C-130 | 25 | 20–49 | 27% | 5 | 13 |
| C-131 | 55 | 20–49 | 57% | 5 | 20 |
| C-132 | 35 | 20–49 | 33% | 5 | 35 |
| C-133 | 40 | 20–49 | 40% | 5 | 17 |
| C-134 | 40 | 50–99 | 42% | 5 | 20 |
| C-135 | 25 | 20–49 | 27% | 5 | 13 |
| C-136 | 30 | 20–49 | 33% | 5 | 28 |
| C-137 | 40 | 20–49 | 40% | 5 | 5 |
| C-138 | 30 | 20–49 | 30% | 5 | 12 |
| C-139 | 30 | 20–49 | 33% | 5 | 11 |
| C-140 | 35 | 20–49 | 37% | 5 | 12 |
| C-141 | 25 | 50–99 | 25% | 5 | 16 |
| C-142 | 5 | 20–49 | 7% | 5 | 8 |
| C-143 | 30 | 20–49 | 30% | 5 | 12 |
| C-144 | 20 | 20–49 | 20% | 5 | 12 |
| C-145 | 60 | 20–49 | 60% | 5 | 13 |
| C-146 | 70 | 20–49 | 80% | 5 | 12 |
| C-147 | 45 | 20–49 | 42% | 5 | 15 |
| C-148 | 55 | 20–49 | 53% | 5 | 16 |
| C-149 | 35 | 20–49 | 37% | 5 | 19 |
| C-150 | 30 | 20–49 | 30% | 5 | 24 |
| C-151 | 45 | 20–49 | 42% | 5 | 16 |
| C-152 | 50 | 100–199 | 49% | 5 | 11 |
| C-153 | 45 | 20–49 | 43% | 5 | 13 |
| C-154 | 70 | 20–49 | 69% | 5 | 18 |
| C-155 | 60 | 50–99 | 60% | 5 | 13 |
| C-156 | 35 | 20–49 | 33% | 5 | 15 |
| C-157 | 50 | 20–49 | 50% | 5 | 11 |
| C-158 | 40 | 20–49 | 40% | 5 | 17 |
| C-159 | 45 | 20–49 | 43% | 5 | 23 |
| C-160 | 40 | 20–49 | 42% | 5 | 23 |
| C-161 | 60 | 20–49 | 60% | 5 | 8 |
| C-162 | 20 | 20–49 | 20% | 5 | 12 |
| C-163 | 15 | 100–199 | 16% | 5 | 6 |
| C-164 | 30 | 20–49 | 30% | 5 | 7 |
| C-165 | 20 | 20–49 | 22% | 5 | 14 |
| C-166 | 55 | 100–199 | 52% | 5 | 14 |
| C-167 | 45 | 20–49 | 43% | 5 | 27 |
| C-168 | 55 | 20–49 | 57% | 5 | 13 |
| C-169 | 50 | 20–49 | 50% | 5 | 15 |
| C-170 | 45 | 20–49 | 47% | 5 | 27 |
| C-171 | 35 | 50–99 | 34% | 6 | 13 |
| C-172 | 55 | 50–99 | 57% | 5 | 18 |
| C-173 | 30 | 20–49 | 31% | 5 | 13 |
| C-174 | 80 | 20–49 | 77% | 5 | 25 |
| C-175 | 30 | 20–49 | 33% | 5 | 11 |
| C-176 | 40 | 20–49 | 40% | 5 | 17 |
| C-177 | 45 | 20–49 | 43% | 5 | 23 |
| C-178 | 60 | 20–49 | 60% | 5 | 15 |
| C-179 | 45 | 20–49 | 47% | 5 | 16 |
| C-180 | 35 | 20–49 | 38% | 5 | 9 |
| C-181 | 10 | 20–49 | 15% | 1 | — |
| C-182 | 45 | 50–99 | 46% | 5 | 17 |
| C-183 | 60 | 20–49 | 57% | 5 | 25 |
| C-184 | 65 | 20–49 | 67% | 5 | 11 |
| C-185 | 90 | 20–49 | 87% | 5 | 27 |
| C-186 | 25 | 20–49 | 27% | 5 | 8 |
| C-187 | 20 | 20–49 | 23% | 5 | 17 |
| C-188 | 65 | 20–49 | 63% | 5 | 19 |
| C-189 | 45 | 20–49 | 47% | 5 | 24 |
| C-190 | 20 | 20–49 | 20% | 5 | 16 |
| C-191 | 20 | 20–49 | 20% | 5 | 19 |
| C-192 | 55 | 20–49 | 57% | 5 | 17 |
| C-193 | 25 | 20–49 | 27% | 5 | 8 |
| C-194 | 5 | 20–49 | 7% | 5 | 8 |
| C-195 | 30 | 20–49 | 30% | 5 | 29 |
| C-196 | 45 | 20–49 | 48% | 5 | 19 |
| C-197 | 60 | 20–49 | 70% | 1 | — |
| C-198 | 35 | 100–199 | 36% | 5 | 12 |
| C-199 | 5 | 20–49 | 10% | 5 | 13 |
| C-200 | 55 | 20–49 | 56% | 5 | 12 |
| C-201 | 35 | 20–49 | 37% | 5 | 29 |
| C-202 | 20 | 20–49 | 20% | 5 | 19 |
| C-203 | 30 | 20–49 | 29% | 5 | 25 |
| C-204 | 40 | 50–99 | 42% | 5 | 12 |
| C-205 | 40 | 20–49 | 40% | 5 | 23 |
| C-206 | 35 | 20–49 | 33% | 5 | 15 |
| C-207 | 25 | 20–49 | 27% | 5 | 31 |
| C-208 | 30 | 50–99 | 32% | 5 | 17 |
| C-209 | 70 | 20–49 | 55% | 1 | — |
| C-210 | 20 | 20–49 | 22% | 5 | 16 |
| C-211 | 70 | 20–49 | 67% | 5 | 11 |
| C-212 | 35 | 20–49 | 33% | 5 | 15 |
| C-213 | 20 | 20–49 | 29% | 6 | 27 |
| C-214 | 40 | 20–49 | 40% | 5 | 5 |
| C-215 | 90 | 20–49 | 87% | 5 | 19 |
| C-216 | 35 | 100–199 | 35% | 5 | 10 |
| C-217 | 65 | 20–49 | 63% | 5 | 12 |
| C-218 | 35 | 20–49 | 33% | 5 | 24 |
| C-219 | 25 | 20–49 | 27% | 5 | 25 |
| C-220 | 45 | 50–99 | 45% | 5 | 34 |
| C-221 | 40 | 20–49 | 42% | 5 | 11 |
| C-222 | 70 | 20–49 | 67% | 5 | 28 |
| C-223 | 35 | 20–49 | 37% | 5 | 19 |
| C-224 | 10 | 20–49 | 13% | 5 | 7 |
| C-225 | 35 | 20–49 | 37% | 5 | 22 |
| C-226 | 60 | 20–49 | 60% | 5 | 27 |
| C-227 | 30 | 20–49 | 31% | 5 | 19 |
| C-228 | 40 | 20–49 | 40% | 5 | 17 |
| C-229 | 30 | 20–49 | 33% | 5 | 19 |
| C-230 | 45 | 20–49 | 47% | 5 | 16 |
| C-231 | 45 | 20–49 | 47% | 5 | 19 |
| C-232 | 45 | 20–49 | 44% | 5 | 16 |
| C-233 | 70 | 20–49 | 67% | 5 | 18 |
| C-234 | 65 | 20–49 | 63% | 5 | 19 |
| C-235 | 75 | 20–49 | 73% | 5 | 13 |
| C-236 | 30 | 20–49 | 33% | 5 | 24 |
| C-237 | 40 | 20–49 | 35% | 1 | — |
| C-238 | 30 | 20–49 | 30% | 5 | 22 |
| C-239 | 30 | 20–49 | 30% | 5 | 22 |
| C-240 | 40 | 20–49 | 40% | 5 | 19 |
| C-241 | 30 | 20–49 | 29% | 5 | 15 |
| C-242 | 25 | 20–49 | 27% | 5 | 13 |
| C-243 | 65 | 20–49 | 63% | 5 | 22 |
| C-244 | 40 | 20–49 | 42% | 5 | 18 |
| C-245 | 35 | 20–49 | 37% | 5 | 29 |
| C-246 | 60 | 20–49 | 56% | 5 | 24 |
| C-247 | 45 | 20–49 | 47% | 5 | 16 |
| C-248 | 45 | 20–49 | 42% | 5 | 18 |
| C-249 | 40 | 20–49 | 42% | 5 | 19 |
| C-250 | 40 | 100–199 | 42% | 5 | 16 |
| C-251 | 40 | 20–49 | 40% | 5 | 23 |
| C-252 | 55 | 20–49 | 53% | 5 | 19 |
| C-253 | 55 | 20–49 | 57% | 5 | 17 |
| C-254 | 70 | 20–49 | 70% | 5 | 12 |
| C-255 | 0 | 20–49 | 0% | 5 | 0 |
| C-256 | 40 | 20–49 | 40% | 5 | 27 |
| C-257 | 30 | 20–49 | 30% | 5 | 22 |
| C-258 | 45 | 50–99 | 47% | 5 | 11 |
| C-259 | 75 | 20–49 | 73% | 5 | 8 |
| C-260 | 40 | 20–49 | 40% | 5 | 17 |
| C-261 | 15 | 50–99 | 17% | 5 | 5 |
| C-262 | 35 | 50–99 | 33% | 5 | 13 |
| C-263 | 35 | 20–49 | 37% | 5 | 29 |
| C-264 | 55 | 20–49 | 53% | 5 | 24 |
| C-265 | 40 | 50–99 | 41% | 5 | 13 |
| C-266 | 55 | 20–49 | 53% | 5 | 19 |
| C-267 | 45 | 50–99 | 45% | 5 | 11 |
| C-268 | 25 | 20–49 | 27% | 5 | 13 |
| C-269 | 55 | 20–49 | 53% | 5 | 19 |
| C-270 | 35 | 20–49 | 37% | 5 | 19 |
| C-271 | 30 | 20–49 | 30% | 5 | 12 |
| C-272 | 60 | 20–49 | 58% | 5 | 23 |
| C-273 | 30 | 20–49 | 33% | 5 | 21 |
| C-274 | 60 | 20–49 | 47% | 5 | 7 |
| C-275 | 15 | 20–49 | 17% | 5 | 11 |
| C-276 | 40 | 20–49 | 40% | 5 | 13 |
| C-277 | 45 | 20–49 | 43% | 5 | 20 |
| C-278 | 45 | 20–49 | 47% | 5 | 27 |
| C-279 | 65 | 20–49 | 63% | 6 | 15 |
| C-280 | 45 | 20–49 | 45% | 1 | — |
| C-281 | 65 | 20–49 | 62% | 5 | 17 |
| C-282 | 15 | 20–49 | 20% | 5 | 27 |
| C-283 | 50 | 20–49 | 50% | 5 | 18 |
| C-284 | 15 | 20–49 | 17% | 5 | 11 |
| C-285 | 10 | 20–49 | 10% | 5 | 20 |
| C-286 | 40 | 20–49 | 40% | 5 | 29 |
| C-287 | 60 | 100–199 | 60% | 5 | 12 |
| C-288 | 5 | 20–49 | 10% | 5 | 13 |
| C-289 | 70 | 20–49 | 67% | 5 | 12 |
| C-290 | 45 | 20–49 | 47% | 5 | 12 |
| C-291 | 35 | 20–49 | 36% | 5 | 19 |
| C-292 | 60 | 20–49 | 55% | 5 | 17 |
| C-293 | 25 | 50–99 | 28% | 5 | 12 |
| C-294 | 80 | 20–49 | 77% | 5 | 17 |
| C-295 | 35 | 20–49 | 33% | 5 | 21 |
| C-296 | 20 | 20–49 | 23% | 5 | 17 |
| C-297 | 50 | 20–49 | 51% | 5 | 26 |
| C-298 | 40 | 20–49 | 40% | 5 | 31 |
| C-299 | 20 | 20–49 | 20% | 5 | 19 |
| C-300 | 45 | 20–49 | 43% | 5 | 17 |
| C-301 | 40 | 20–49 | 40% | 5 | 23 |
| C-302 | 35 | 20–49 | 38% | 5 | 15 |
| C-303 | 50 | 20–49 | 49% | 5 | 15 |
| C-304 | 65 | 20–49 | 62% | 5 | 23 |
| C-305 | 30 | 20–49 | 33% | 5 | 12 |
| C-306 | 45 | 20–49 | 47% | 5 | 24 |
| C-307 | 15 | 20–49 | 17% | 5 | 11 |
| C-308 | 35 | 20–49 | 38% | 5 | 19 |
| C-309 | 20 | 20–49 | 24% | 5 | 26 |
| C-310 | 55 | 20–49 | 57% | 5 | 27 |
| C-311 | 80 | 20–49 | 77% | 5 | 13 |
| C-312 | 35 | 20–49 | 37% | 5 | 22 |
| C-313 | 25 | 20–49 | 27% | 5 | 23 |
| C-314 | 35 | 20–49 | 33% | 5 | 11 |
| C-315 | 20 | 50–99 | 24% | 5 | 16 |
| C-316 | 60 | 50–99 | 58% | 5 | 20 |
| C-317 | 20 | 20–49 | 20% | 5 | 12 |
| C-318 | 30 | 20–49 | 30% | 5 | 12 |
| C-319 | 50 | 20–49 | 50% | 5 | 15 |
| C-320 | 45 | 20–49 | 43% | 5 | 20 |
| C-321 | 35 | 20–49 | 37% | 5 | 19 |
| C-322 | 20 | 50–99 | 22% | 5 | 12 |
| C-323 | 40 | 50–99 | 42% | 5 | 14 |
| C-324 | 50 | 20–49 | 50% | 5 | 18 |
| C-325 | 65 | 20–49 | 63% | 5 | 12 |
| C-326 | 45 | 20–49 | 43% | 5 | 23 |
| C-327 | 40 | 100–199 | 39% | 5 | 11 |
| C-328 | 60 | 20–49 | 67% | 5 | 15 |
| C-329 | 15 | 20–49 | 17% | 5 | 18 |
| C-330 | 45 | 20–49 | 57% | 5 | 13 |
| C-331 | 35 | 20–49 | 43% | 6 | 17 |
| C-332 | 55 | 20–49 | 57% | 5 | 17 |
| C-333 | 30 | 20–49 | 33% | 5 | 16 |
| C-334 | 45 | 20–49 | 43% | 5 | 31 |
| C-335 | 20 | 20–49 | 15% | 1 | — |
| C-336 | 0 | 20–49 | 0% | 5 | 0 |
| C-337 | 20 | 20–49 | 19% | 5 | 16 |
| C-338 | 45 | 20–49 | 47% | 5 | 19 |
| C-339 | 55 | 20–49 | 53% | 5 | 20 |
| C-340 | 60 | 20–49 | 60% | 5 | 17 |
| C-341 | 55 | 50–99 | 51% | 5 | 13 |
| C-342 | 50 | 20–49 | 50% | 5 | 18 |
| C-343 | 30 | 20–49 | 30% | 5 | 16 |
| C-344 | 40 | 20–49 | 42% | 5 | 22 |
| C-345 | 30 | 100–199 | 28% | 5 | 9 |
| C-346 | 20 | 20–49 | 23% | 5 | 13 |
| C-347 | 40 | 20–49 | 40% | 5 | 25 |
| C-348 | 30 | 20–49 | 33% | 5 | 11 |
| C-349 | 70 | 20–49 | 70% | 5 | 16 |
| C-350 | 40 | 20–49 | 40% | 5 | 23 |
| C-351 | 50 | 20–49 | 49% | 5 | 15 |
| C-352 | 30 | 50–99 | 34% | 6 | 14 |
| C-353 | 50 | 20–49 | 50% | 5 | 15 |
| C-354 | 55 | 20–49 | 53% | 5 | 32 |
| C-355 | 70 | 20–49 | 67% | 5 | 24 |
| C-356 | 20 | 20–49 | 23% | 5 | 8 |
| C-357 | 20 | 20–49 | 23% | 5 | 12 |
| C-358 | 25 | 20–49 | 27% | 5 | 23 |
| C-359 | 60 | 20–49 | 57% | 5 | 13 |
| C-360 | 25 | 50–99 | 28% | 5 | 15 |
| C-361 | 60 | 20–49 | 57% | 5 | 23 |
| C-362 | 35 | 20–49 | 37% | 5 | 24 |
| C-363 | 75 | 20–49 | 73% | 5 | 13 |
| C-364 | 35 | 50–99 | 37% | 5 | 11 |
| C-365 | 50 | 20–49 | 50% | 5 | 24 |
| C-366 | 45 | 20–49 | 43% | 5 | 8 |
| C-367 | 20 | 20–49 | 35% | 1 | — |
| C-368 | 15 | 20–49 | 18% | 5 | 11 |
| C-369 | 90 | 20–49 | 85% | 1 | — |
| C-370 | 25 | 50–99 | 24% | 5 | 14 |
| C-371 | 30 | 20–49 | 33% | 5 | 11 |
| C-372 | 35 | 50–99 | 37% | 5 | 12 |
| C-373 | 40 | 50–99 | 39% | 5 | 12 |
| C-374 | 20 | 20–49 | 23% | 5 | 13 |
| C-375 | 70 | 20–49 | 67% | 5 | 18 |
| C-376 | 45 | 20–49 | 47% | 5 | 7 |
| C-377 | 55 | 20–49 | 53% | 5 | 12 |
| C-378 | 50 | 20–49 | 50% | 5 | 18 |
| C-379 | 45 | 20–49 | 47% | 5 | 12 |
| C-380 | 30 | 20–49 | 29% | 5 | 15 |
17.3 Standard Appendix Summary (n = 12 decks)
| Metric | Value |
|---|---|
| Published decks | 12 |
| Win rate range | 10% – 70% |
| Median win rate | 40% |
| Median matchup spread | 24 percentage points |
| Decks with spread computed (3+ opponents) | 12 |
OLS residual outliers (Standard):
- Widest overperformance: S-002 — Power Score 70, win rate 70%, OLS-expected 68%, residual +2 points.
- Widest underperformance: S-007 — Power Score 55, win rate 53%, OLS-expected 55%, residual −2 points.
The Standard Standard residuals are very small (+2 / −2 points), reflecting the tight OLS fit over a small 12-deck sample. These outlier designations are statistically trivial; they are reported for completeness as specified by the appendix structure, not as meaningful anomalies.
17.4 Standard Per-Deck Table
| ID | PS (±5) | Games Band | Win Rate | Opponents | Spread |
|---|---|---|---|---|---|
| S-001 | 50 | 20–49 | 51% | 5 | 18 |
| S-002 | 70 | 20–49 | 70% | 5 | 29 |
| S-003 | 45 | 20–49 | 44% | 5 | 25 |
| S-004 | 40 | 20–49 | 40% | 5 | 17 |
| S-005 | 25 | 20–49 | 27% | 5 | 23 |
| S-006 | 45 | 20–49 | 47% | 5 | 16 |
| S-007 | 55 | 20–49 | 53% | 5 | 29 |
| S-008 | 15 | 20–49 | 20% | 5 | 24 |
| S-009 | 25 | 20–49 | 29% | 5 | 29 |
| S-010 | 35 | 20–49 | 38% | 5 | 17 |
| S-011 | 5 | 50–99 | 10% | 5 | 13 |
| S-012 | 40 | 20–49 | 40% | 5 | 45 |
S-012 merits a note: its matchup spread of 45 percentage points is the widest in the entire Standard appendix, indicating a deck whose performance is highly opponent-dependent. At the other end, the lower-spread Standard decks cluster around 13–17 points. As with all appendix data, this is a descriptive observation about coarsened, anonymized figures — no deck identity, account, or card list is derivable from these rows.
18. Key Findings & What We Cannot Conclude Yet
18.1 What the Data Shows
Matchup identity reshapes outcomes more than most single construction choices. The 20.9-point Commander spread and the 40.7-point Standard spread between the best and worst matchups — measured on essentially the same pool of challenger decks — is the most structurally robust finding in this dataset. The gauntlet opponent is not a neutral backdrop; it is a primary driver of observed win rate alongside the deck's own profile.
The deck's profile carries more total variance in Commander, but the opponent still matters substantially. Deck main effects account for 47.7% of Commander win-rate variance (n = 384 decks, 23 opponents); the opponent accounts for 14.6%. These proportions are reversed in Standard (28%/28%), though the Standard sample is small (n = 13 decks, 6 opponents). In neither format does the deck profile completely dominate.
Popularity and simulation win rate are not the same signal among commanders. Meren of Clan Nel Toth (the most submitted commander, n = 36 decks) posts 34.5%. Giada, Font of Hope (n = 17 decks) posts 76.1%. Players submit what they want to test; the simulator measures what happens.
Commander decks that re-tested more often showed a slight average decline (mean −1.6 points, n = 426 re-tested decks; 192 declined vs. 168 improved). This is a small, noise-susceptible figure and does not establish that retesting hurts outcomes.
Land ratio shows a monotonic positive association with Commander win rate in the submission data: the ~25% land band averages 23.5% win rate (n = 44); the ~45% band averages 47.8% (n = 20). The same directional pattern holds in Standard (35% land: 31.6%, n = 32; 40% land: 44.6%, n = 54). These are descriptive correlations in a self-selected sample and are not build prescriptions.
Among functional card categories, the tutor/search correlation with Standard win rate is the largest categorical delta in either format (−16.1 points; 29.6% with tutors, n = 58 vs. 45.7% without, n = 50). This is heuristic-labelled and heavily confounded by archetype clustering.
Creatures post the highest pooled board impact by type in both Commander (+3.66, n = 319,373 observations) and Standard (+2.92, n = 11,313 observations). Artifacts aggregate to a negative figure in both formats (Commander −1.48, Standard −3.02) — a consequence of pooling mana rocks and utility artifacts with wildly varying board contributions.
Green cards post the highest pooled board impact by color in Commander (+5.40, n = 138,088 observations). All five colors are positive; only colorless is negative (−0.55).
18.2 What We Cannot Conclude Yet
Causation. No construction feature, card, or category causes a win-rate outcome. Every association in this report could be explained by confounding factors — deck archetype, commander choice, player skill in deck-building, submission selection effects, or simulation AI behavior — that the dataset cannot untangle.
Human play outcomes. All figures describe AI-versus-AI simulation behavior. Human piloting of the same decks would produce different results; the Forge AI does not play any deck optimally or equivalently to an experienced human pilot.
Strategy guidance. No figure in this report is sufficient basis for a deck-building recommendation. The report is observational documentation of a specific dataset at a specific point in time.
Representativeness. The submission pool is a self-selected sample of Grim.Cards users who chose to test their decks. It is not a random sample of Commander or Standard players or decks. External generalization is not warranted.
Individual card causation. Card presence in a winning deck is not evidence that the card contributed to wins. Containing-deck win rate measures the performance of whole decks that happen to include the card.
Trend interpretation. Monthly win-rate variation reflects changes in who submitted decks, not a tracked improvement or deterioration of any individual deck or archetype over time.
Standard conclusions at scale. Standard figures are based on 135 decks and 1,902 games. Many Standard-specific figures are directionally interesting but warrant caution at these sample sizes.
19. Future-Comparison Baseline Table
This table provides a machine-readable baseline of key metrics for comparison when future editions are published. Figures are drawn from this edition's dataset (snapshot: 2026-10-01).
| Metric | Format | Value | Unit | n (decks) | n (games) |
|---|---|---|---|---|---|
| Overall win rate | All | 40.8% | % | 2,499 | 44,833 |
| Win rate | Commander | 41.0% | % | 2,364 | 42,931 |
| Win rate | Standard | 37.3% | % | 135 | 1,902 |
| Median win rate | Commander | 40% | % | 2,217 | — |
| Median win rate | Standard | 40% | % | 108 | — |
| Highest matchup win rate | Commander | 52.5% (vs Breya) | % | 2,197 | 8,511 |
| Lowest matchup win rate | Commander | 31.6% (vs Edgar Markov) | % | 2,196 | 8,519 |
| Matchup spread | Commander | 20.9 pts | pp | — | — |
| Highest matchup win rate | Standard | 55.5% (vs Temur Harmonizer) | % | 102 | 366 |
| Lowest matchup win rate | Standard | 14.8% (vs Mono Red Aggro) | % | 102 | 366 |
| Matchup spread | Standard | 40.7 pts | pp | — | — |
| Retest mean Δ win rate | Commander | −1.6 pts | pp | 426 | — |
| Retest mean Δ win rate | Standard | +0.5 pts | pp | 14 | — |
| Deck main-effect share | Commander | 47.7% | % | 384 | — |
| Opponent main-effect share | Commander | 14.6% | % | 384 | — |
| Inversion rate | Commander | 37.0% | % | 360 | 421,253 pairs |
| Inversion rate | Standard | 38.5% | % | 12 | 486 pairs |
| Avg game length | Commander | 10.2 turns | turns | 2,200 | — |
| Avg game length | Standard | 9.4 turns | turns | 104 | — |
| Total unique users | All | 1,184 | users | — | — |
| Total unique decks | All | 2,499 | decks | — | — |
| Total simulations | All | 3,307 | sims | — | — |
| Data window start | All | 2026-05-14 | date | — | — |
| Data window end | All | 2026-10-01 | date | — | — |
| Dataset version | All | 2.4 | — | — | — |
When this study is re-run, the new edition's corresponding values can be compared directly against this table to identify what changed across the intervening period.
20. Power Score Composite (Secondary; Caveated)
This section is secondary and optional. Power Score is an internal Grim.Cards composite indicator, not an objective or universal measure of deck strength. It is not used to rank, judge, or draw causal conclusions anywhere in this report. It is presented here for context only, after all objective win-rate findings.
Power Score is Grim.Cards' own internal weighted composite of deck construction features. It is computed from deck embedding data (mana curve, land ratio, card-type distribution, color complexity, and related signals) and is not an externally validated or universally accepted metric.
20.1 Grade Distribution (Commander, n = 2,499 scored decks across both formats)
| Grade | Decks (n) |
|---|---|
| S | 13 |
| A | 127 |
| B | 361 |
| C | 563 |
| D | 956 |
| F | 478 |
The distribution is right-skewed toward lower grades: D is the modal grade (n = 956), followed by C (n = 563) and F (n = 478). S and A together account for 140 decks (5.6% of the scored population).
20.2 Relationship to Win Rate
As documented in Sections 16.3 and 16.4, Power Score shows a moderate-to-strong positive rank correlation with simulated win rate across Commander meta opponents (Spearman ρ = 0.62–0.75, n = 366–368 decks per opponent). The PS × opponent grid (Section 16.3) shows win rates rising consistently from the Low to High Power Score band across all five Commander opponents. The per-deck appendix OLS residuals (Section 17.1) identify individual decks that over- or underperform their Power Score expectation.
None of these correlations establish that Power Score causes win outcomes. Power Score is a construction-feature composite; it correlates with win rate because high-PS features (consistent mana, efficient curves, adequate land counts) are also associated with better gauntlet performance in this simulation environment. The relationship is associational and likely partially mediated by the same deck-quality factors that independently affect simulated outcomes.
Power Score should be used as one contextual signal among many, never as a ranking or verdict on a deck's value for human play.
21. Limitations
The following limitations apply to all findings in this report:
Simulated, AI-piloted games only. Results describe Forge engine AI behavior on player-submitted decklists. Human pilots would produce different outcomes. No finding applies directly to human play.
Correlational study. Construction breakdowns, card associations, and category figures describe correlations, never causes. Confounders are numerous, overlapping, and unobserved. The word "cause" does not appear in this report's findings.
Self-selected, non-random sample. Decks in this dataset are those that Grim.Cards users chose to submit for testing. This pool is not a random sample of Commander or Standard players or decklists. Findings may not generalize to the broader Magic: The Gathering population.
Power Score is an internal composite. It is not externally validated, not universally recognized, and carries no causal weight. It is presented in one caveated secondary section only.
Decision-impact figures are a play-quality proxy. Counterfactual impact is measured from replayed AI decision snapshots; it is not damage, kills, or a direct win contribution. The Lightning Greaves figure (Section 12) rests on 16 observations across 10 decks and should be treated as preliminary.
Board-impact figures are a board-state proxy. Card performance scores measure board-quality deltas around observed appearances, pooled across cohort decks. They do not measure damage, kills, or causal win contributions. Rankings with fewer qualifying decks or observations are more volatile.
Functional category heuristics may mislabel cards. Sacrifice, tutor/search, discard, and reanimation category membership is assigned by keyword pattern-matching and pre-existing flags. Edge cases — cards with non-standard wording, split cards, modal spells — may be classified incorrectly.
Monthly trend figures are cross-sectional, not longitudinal. Each month's cohort consists of different decks. Month-over-month win-rate changes reflect shifts in submission mix, not tracked improvement of specific decks.
Standard sample is small. With 135 decks and 1,902 games, Standard-specific figures — particularly the per-opponent Spearman correlations (n = 12 decks), the Standard appendix (n = 12 published decks), and the Standard variance decomposition (n = 13 decks) — carry wider uncertainty than their Commander counterparts and should be read as directional.
Aesi Landfall draw rate. The 16.2% draw rate in the Aesi Landfall Commander matchup (1,386 draws from 8,532 games) is anomalously high relative to other matchups and likely reflects a Forge AI interaction specific to that deck's game plan. Win-rate figures for that matchup are computed on the full-denominator definition (draws counted in the denominator) and thus reflect this structural artifact.
Minimum cohort threshold suppresses small groups. Any breakdown with fewer than 10 distinct decks is suppressed or merged. Suppression protects privacy and statistical integrity but means some potentially interesting subgroups are not reported.
No inference about card-level win causation. Containing-deck win rate (as used in Sections 8, 11, and 14) measures the performance of whole decks that happen to include a given card, color, or category. It is not a per-card win contribution and cannot be decomposed into card-level causal effects from this dataset.
22. License & Citation
License
This dataset and report are published under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
You are free to share and adapt the material for any purpose, including commercially, provided you give appropriate credit, link to the license, and indicate if changes were made.
Full license text: https://creativecommons.org/licenses/by/4.0/
How to Cite This Edition
Standard citation (author–date):
Grim.Cards. "Grim.Cards Simulation Case Study, Edition 2026-10-01." Grim.Cards, 1 Oct. 2026. grim.cards/case-study/2026-10-01. CC BY 4.0.
BibTeX:
@techreport{grimcards2026casestudy,
author = {{Grim.Cards}},
title = {{Grim.Cards Simulation Case Study, Edition 2026-10-01}},
institution = {Grim.Cards},
year = {2026},
month = {10},
day = {1},
url = {https://grim.cards/case-study/2026-10-01},
note = {Data snapshot: 2026-05-14 to 2026-10-01. Dataset version 2.4.
License: CC BY 4.0.}
}
Schema.org / structured data identifier:
- Canonical URL:
https://grim.cards/case-study/2026-10-01 - Dataset version:
2.4 - Snapshot date:
2026-10-01 - Temporal coverage:
2026-05-14/2026-10-01 - Publisher: Grim.Cards
- License:
https://creativecommons.org/licenses/by/4.0/
Provenance
All figures in this report are derived from Grim.Cards production simulation data only. The dataset was generated on 2026-10-01T08:23:19.653Z from a production database snapshot. No development or staging data is included. The generation script is repeatable; future editions will be published at their own permanent dated URLs and indexed at grim.cards/case-study.
Retrieving This Edition
This edition is permanently archived at grim.cards/case-study/2026-10-01. A listing of all published editions, newest-first, is available at grim.cards/case-study. This URL will never be overwritten; future editions accumulate at new dated URLs.
End of Grim.Cards Simulation Case Study — Edition 2026-10-01.
Grim.Cards · grim.cards · CC BY 4.0 · Snapshot date: 2026-10-01 · Dataset version: 2.4