HOB Deck Evaluator (LightGBM v4)
Estimates the per-game win probability of a 40-card Magic: The Gathering maindeck for The Hobbit (HOB) Premier Draft on MTG Arena. It is a small gradient-boosted tree model over card-copy counts only: no player skill, rank, play/draw or game metadata.
It is part of the laya-plays-magics project, next to the Laya draft-pick and deckbuild models, and is used there to compare candidate decks.
Honest summary. The model beats a constant baseline and a GIH-WR heuristic on a held-out split, but the signal is weak (AUC ≈ 0.57). Use it to rank alternatives, not as a promise of a win rate. Deck results mix deck quality with pilot skill, matchmaking and variance.
Benchmark
Test split: 5,007 builds / 4,117 drafts / 23,074 games, never used for training, early
stopping or calibration. Re-scored from the saved artifacts on 2026-09-28; the numbers
match metrics.json exactly.
| Predictor | ROC AUC ↑ | Log loss ↓ | Brier ↓ | Weighted MAE vs build win rate ↓ |
|---|---|---|---|---|
| Constant (train win rate 56.2%) | 0.5000 | 0.6852 | 0.2461 | 0.1821 |
| Mean GIH WR of the maindeck's non-basic cards¹ | 0.5635 | 0.6835 | 0.2452 | 0.1803 |
| LightGBM v4, raw | 0.5681 | 0.6783 | 0.2426 | 0.1754 |
| LightGBM v4, isotonic-calibrated | 0.5670 | 0.6786 | 0.2428 | 0.1761 |
¹ 17Lands-style heuristic, similar to the "Heuristic" baseline in DraftEncoder. Card GIH WR was computed over all HOB games, including the test games, so this baseline is slightly optimistic. It is shifted to the train win rate so that log loss and Brier are meaningful.
Related work
| Work | Target | Data | Reported result | Comparable? |
|---|---|---|---|---|
| This model | Per-game win of a fixed 40-card maindeck | 17Lands HOB, 49,474 builds, 226,047 games | AUC 0.568 | — |
| DraftEncoder, Rigaux & Kashima 2026 | Win rate of a drafted deck | 17Lands, 12 expansions | Spearman ρ saturating near 0.23 in-set; 0.153 cross-set (FDN held out); GIH heuristic 0.139 | No — different metric (rank correlation per deck) and sets |
| Bertram et al. 2021, Contextual Preference Ranking | Deck/card preference | 17Lands | Deck-evaluation ability, no AUC reported | No |
Our own ordering matches DraftEncoder's qualitative finding: a learned model beats a GIH-WR heuristic only by a small margin, because the outcome is dominated by the pilot and by variance.
Training data and protocol
- Sources: 17Lands public
draft_data_public.HOB.PremierDraft.csvandgame_data_public.HOB.PremierDraft.csv(17lands.com/public_datasets, CC BY 4.0). - Unit: one
(draft_id, build_index)maindeck of exactly 40 cards; wins and losses of every game played with it are aggregated. Builds with any non-40-card row are dropped. - Size: 43,650 drafts, 49,474 builds, 226,047 games, 193 card features.
- Split: stable SHA-256 of
draft_id, 80/10/10. Validation is split again bydraft_idinto early stopping (2,521 builds) and isotonic calibration (2,409 builds). Train: 39,537 builds. - Objective: binary log loss with each build expanded to a win row (weight = wins) and a loss row (weight = losses), i.e. binomial loss per game.
- Hyperparameters:
n_estimators=800,learning_rate=0.03,num_leaves=15,min_child_samples=40, early stopping 60 rounds → best iteration 450,seed=42. - All player skill levels are included; skill is deliberately not a feature because it is not available when evaluating a future deck.
Earlier versions (v1–v3, trained on filtered subsets) reached test AUC ≤ 0.555 and are superseded.
Files
The current model is in hob_deck_evaluator_lightgbm_v4/. The folders without a suffix and with
_v2/_v3 hold the superseded experiments, kept for reproducibility.
File (inside hob_deck_evaluator_lightgbm_v4/) |
Content |
|---|---|
model.txt |
LightGBM booster (best iteration) |
feature_names.json |
Ordered card vocabulary (193 names, as spelled in the 17Lands CSV) |
calibration.json |
Isotonic calibration thresholds (x_thresholds, y_thresholds) |
metrics.json |
Full training/test report, including the dataset SHA-256 |
Usage
import json
import lightgbm as lgb
import numpy as np
from huggingface_hub import snapshot_download
repo = snapshot_download("FabioCeleste/laya-mtg-deck-evaluator", allow_patterns="hob_deck_evaluator_lightgbm_v4/*")
path = f"{repo}/hob_deck_evaluator_lightgbm_v4"
names = json.load(open(f"{path}/feature_names.json"))
cal = json.load(open(f"{path}/calibration.json"))
booster = lgb.Booster(model_file=f"{path}/model.txt")
deck = {"Forest": 8, "Swamp": 8, "Ordinary Bear": 1, ...} # 40 cards, names exactly as in 17Lands
assert sum(deck.values()) == 40
unknown = set(deck) - set(names)
assert not unknown, f"out-of-vocabulary cards: {unknown}" # reject rather than silently ignore
x = np.zeros((1, len(names)), dtype=np.float32)
for i, name in enumerate(names):
x[0, i] = deck.get(name, 0)
raw = booster.predict(x)[0]
calibrated = float(np.interp(raw, cal["x_thresholds"], cal["y_thresholds"]))
print(f"expected per-game win probability ≈ {calibrated:.3f}")
Limitations
- HOB Premier Draft only. Any card outside the 193-name vocabulary must be rejected.
- Weak signal. AUC 0.57 means the model orders two random games' decks correctly ~57% of the time. Treat differences of a few points as noise.
- Not causal. A high score does not prove a deck causes wins; strong players build and pilot certain archetypes more often.
- Isotonic calibration is step-shaped: close decks may receive identical calibrated scores; use the raw score to break ties.
- Not intended to automate picks or deckbuilding without human review.
Related models
FabioCeleste/laya-mtg-draft-picks— Laya fine-tuned to choose one card from a HOB pack.FabioCeleste/laya-mtg-deckbuild— Laya fine-tuned to choose maindeck copy counts from a drafted pool.
Citation
If you use this model, please also cite the 17Lands public dataset and the work above.
Papers for FabioCeleste/laya-mtg-deck-evaluator
Predicting Human Card Selection in Magic: The Gathering with Contextual Preference Ranking
Evaluation results
- ROC AUC per game (raw) on 17Lands HOB PremierDraft (held-out test split by draft_id)self-reported0.568
- ROC AUC per game (calibrated) on 17Lands HOB PremierDraft (held-out test split by draft_id)self-reported0.567
- Log loss per game (raw) on 17Lands HOB PremierDraft (held-out test split by draft_id)self-reported0.678
- Brier per game (raw) on 17Lands HOB PremierDraft (held-out test split by draft_id)self-reported0.243