Architecture
Engine, Pilot, and Simulator
TurnZero documentation

TurnZero Engine, Pilot, and Simulator Architecture

TurnZero separates game rules from play policy. The Engine determines what is true and legal, a Pilot chooses one offered action using player-safe knowledge, and the Simulator connects them while recording the resulting game.

This document describes that boundary and its current implementation. The DSL reference separately documents how cards and rules primitives are modeled.

rules truth player-safe observations legal actions replaceable pilots

System at a glance

Simulator
  │
  ├─ asks the Engine for immutable PlayerKnowledge
  │       └─ full self-deck and commander definitions, without shuffled order
  │
  ├─ calls Pilot.startGame(PlayerKnowledge) once
  │
  ├─ asks Engine for legal actions
  │
  ├─ asks the Engine decision boundary to build a DecisionPoint
  │       ├─ immutable PlayerObservation
  │       └─ immutable LegalAction offers with engine-owned identities
  │
  ├─ supplies that DecisionPoint to the selected Pilot
  │       └─ PilotDecision references one offered action identity
  │
  ├─ asks Engine to validate and resolve the decision
  │
  └─ executes the resolved engine Action and records the result

The Engine depends on the shared Pilot contract, not on a particular Pilot implementation. The Simulator selects and supplies the Pilot. The default is currently the heuristic implementation, but callers can inject another Pilot.

Design principles

  • The Engine owns rules truth, game mutation, legal-action generation, and validation.
  • A Pilot chooses among legal actions; it does not create legality.
  • The Simulator owns the game loop, Pilot selection, logging, and aggregation.
  • Pilot input is immutable, serializable, and safe for the player whose choice is being made.
  • Hidden card instances and shuffled library order never cross the boundary.
  • A Pilot normally returns an engine-issued action identity instead of manufacturing an action payload.
  • Shared schema types may live in a transport-safe workspace package, while the Engine remains responsible for producing their runtime values.

Responsibilities

Component Owns Must not own
Engine GameState, rules evaluation, legal actions, observations, decision validation, action execution Heuristic preferences or a dependency on one concrete Pilot
Pilot Policy, scoring, tie-breaking, and selection from offered actions Rules truth, hidden state, or arbitrary unoffered actions
Simulator Game orchestration, Pilot injection, turns, logs, observations, and result aggregation A second legality implementation
Simulation contracts Immutable, serializable Engine–Pilot and cross-workspace schemas Engine runtime state or implementation behavior

The executable card DSL remains Engine-owned rules truth. At game start, the Engine supplies the Pilot with immutable copies of the complete definitions for the player's own deck and commanders. This lets heuristic and model-backed Pilots understand what every known card does without importing the Engine's global definition catalog.

Decision lifecycle

When a game starts, the Simulator asks the Engine to construct immutable PlayerKnowledge and calls Pilot.startGame() once. For each subsequent choice, it gets the current legal Action values from the Engine and calls the Engine-owned decision boundary with the selected Pilot. The boundary then:

  1. Constructs a fresh PlayerObservation from GameState.
  2. Clones and freezes the legal actions.
  3. Assigns each offer an identity such as action-0.
  4. Calls Pilot.decide() with the resulting DecisionPoint.
  5. Resolves the returned identity to the original engine action.
  6. Rejects a missing, stale, manufactured, or invalidly specialized choice.

The Simulator executes only the resolved engine action. Empty legal-action lists receive an Engine-owned END_TURN fallback, so that behavior is not invented by individual Pilots.

Pilot contract

The public schema lives in @turnzero/simulation-contracts. The minimum contract consists of:

Type Purpose
PilotCardKnowledge<CardDefinitionPayload> One named card, its complete Engine-produced modeled definition, and any separately modeled faces
PlayerKnowledge<CardDefinitionPayload> Stable full self-deck and commander knowledge supplied once per game
ObservedCard Player-visible card identity and state, deliberately smaller than an engine card instance
OpeningHandObservation Player-safe inputs needed by mulligan policy
PlayerObservation Immutable knowledge available for the current decision
LegalAction<ActionPayload> One frozen engine action paired with an engine-issued identity
DecisionPoint<ActionPayload> The observation and complete set of currently offered legal actions
PilotDecision<ActionPayload> A reference to one offer, with an optional validated specialization
Pilot<ActionPayload, CardDefinitionPayload> The replaceable startGame(knowledge) and decide(decisionPoint) policy interface

The package owns the schema. src/engine/pilotDecision.ts owns construction, freezing, action identity assignment, and validation of runtime values.

Player knowledge

PlayerKnowledge is stable for the game. It contains every card in the player's starting non-commander deck, grouped by name and count, plus each commander. Every entry includes the complete cloned card-definition payload produced from the Engine DSL. A normal Commander deck therefore exposes 99 non-commander cards; a two-commander deck exposes 98.

The deck list and card rules are player knowledge. The values deliberately do not contain card-instance IDs or deck order. Definitions are supplied once at game start rather than repeated inside every DecisionPoint. Future dynamic knowledge for cards introduced from outside the starting deck remains a separate extension.

Player observation

PlayerObservation currently exposes:

Field Player-facing meaning
schemaVersion Version of the serialized observation shape; currently 1
turn, phase, turnContext, turnStep Current timing context
lifeTotal The simulated player's life total
zones.hand Cards in hand
zones.battlefield Player-visible battlefield state
zones.commander Command-zone cards
zones.graveyard and zones.exile Player-visible public zones
library.size Number of cards remaining in the library
library.knownCards Individual library cards made known by a modeled look, reveal, or retained permission
library.knownTopCardId Identity of a known top card when the top position is actually known
openingHand Mulligan-specific hand, commander, land-play, and order-independent remaining-deck knowledge

Observed cards include only explicitly selected player-facing properties, such as identity, name, controller, tapped state, commander/token markers, and counters. Adding another field requires deciding that the information is both available to the player and necessary for policy.

The opening-hand remainingDeck.cardsByName values are aggregate counts used for probability estimates. They model knowledge of the player's own deck composition without exposing the identity or order of shuffled card instances. pendingMulliganBottomCards tells a Pilot how many cards still need to be put on the bottom after keeping under the London mulligan. For mulligan planning, the heuristic treats cards in commanderZone as guaranteed casting options alongside the hand, without counting them toward physical hand size, land totals, or London-mulligan bottom choices.

Hidden information

The Engine may know the full ordered library because it must execute draws and searches. A Pilot must not receive that state merely because it runs in the same process.

The boundary therefore permits:

  • the complete starting self-deck list and card definitions;
  • library size;
  • order-independent counts inferred from the player's own deck;
  • cards explicitly revealed, looked at, or made playable; and
  • a known-top identity only while that top position remains known.

It excludes hidden card-instance identities, shuffled order, and any other rules-only state that a player could not use. Pilot source is also guarded by a test that rejects direct reads of the engine-owned GameState.library.

Pilot implementations

HeuristicPilot is the default policy today. SimulationOptions.pilot allows the Simulator to use any implementation of the same interface. Tests already exercise a trivial first-action Pilot to prove that the boundary is interchangeable.

The contract is intended to support additional implementations without Engine changes, including deterministic replay, seeded random, alternative heuristic, and model-backed Pilots. Model training and model-specific planning APIs are separate concerns and are not part of the current contract.

Mulligan evaluation harness

The terminal player's mulligan-study mode provides a focused evaluation path for 25 seeded Baeloth games followed by 25 seeded Quintorius games. Each actor runs independently from the same 50 seeds. The study records every offered pregame action and selection through BEGIN_GAME, including London-mulligan bottom choices and a human unsure marker on close keep or mulligan judgments.

The default human workflow is:

npm run study:mulligans

At each keep-or-mulligan prompt, k and m record emphatic judgments while ku and mu record the same selection with unsure: true. Every non-automatic choice immediately asks for its reason while the commander, hand, and selected decision are still visible. The choice and its non-empty reason are saved together after every decision to .data/mulligan-study/baeloth-quintorius-50-v5.json, so q can stop a session and the same command resumes it. Heuristic results and the comparison summary use the same file:

Every displayed hand, including the final kept hand, uses one bullet line per card. Hand cards are never collapsed into a pipe- or comma-separated row.

npm run study:mulligans -- --pilot heuristic
npm run study:mulligans -- --report

Each study batch uses a permanent suffix and deterministic seed. Spark and the heuristic are recorded before the human run without showing their choices in the human prompt. The next active batch remains v5 even though earlier local artifacts are no longer available:

npm run study:mulligans -- \
  --output .data/mulligan-study/baeloth-quintorius-50-v5.json \
  --base-seed baeloth-quintorius-mulligans-v5

Quitting prints the complete command needed to resume the selected batch rather than silently returning to another study.

Historical reports record Spark v2 at 41/50 unseen first hands (82%) versus the heuristic's 39/50 (78%). Its advantage was deck-specific: Spark led 23/25 to 17/25 on Quintorius, while the heuristic led 22/25 to 18/25 on Baeloth. The raw v1-v4 local artifacts were subsequently lost, so those results are historical evidence rather than a reproducible training corpus. Batch numbering remains monotonic; replacement collection begins at v5 rather than reusing an earlier identifier.

Relative .data paths resolve through Git's shared directory to the primary checkout, and other relative generated-artifact paths are placed beneath the same shared data root, so a linked worktree never owns the only copy. Before replacing a study, the writer stores its prior contents as a compressed, append-only revision below .data/mulligan-study/.history/. Set TURNZERO_LOCAL_DATA_ROOT to use another durable location. Explicit output paths inside a linked worktree are rejected.

For older or interrupted studies that have decisions without reasons, fill in only the missing annotations with:

npm run study:mulligans -- --annotate-human

This fallback pass displays the commander, complete hand, mulligan depth, recorded decision, and whether the keep-or-mulligan judgment was emphatic or unsure. It also includes London-mulligan bottom-card choices. Each one-line explanation is saved immediately on that recorded decision; quitting with q and rerunning the command resumes at the first choice without a reason. Explanations should use only facts available when the decision was made, never later draws or outcomes.

The persisted snapshots contain hand names, commanders, aggregate remaining deck counts, mulligan state, and offered action identities. They never contain the shuffled library order or hidden library instance IDs.

DeepSeek is connected by a deliberately asynchronous experiment harness rather than by changing the synchronous core Pilot.decide() interface. The harness builds each prompt from Engine-produced PlayerKnowledge and the current DecisionPoint, supplies the full deck and commander definitions, enriches them with Oracle facts from the local MTGJSON catalog when available, and accepts only an offered actionId. Current-hand and commander facts are repeated next to the decision so the model does not have to recover them from the full deck payload. The prompt requires an explicit early-mana assessment and turn-one-through-turn-three plan before the final choice. Forced single-action steps bypass the provider entirely.

The harness automatically discovers the main checkout's local MTGJSON catalog when running in a worktree. A dry run validates the prompt and reports a conservative request ceiling without making a provider call:

npm run study:mulligans -- --pilot deepseek --dry-run

An actual run requires DEEPSEEK_API_KEY and is guarded by --max-cost-usd (default $10). Responses, reasons, uncertainty, token usage, and estimated cost are saved after each decision. Empty or malformed JSON is retried up to three times, with usage aggregated when a later attempt succeeds. The prompt is versioned so results from different evaluation instructions cannot be silently mixed. To discard only older DeepSeek trajectories while preserving human and heuristic results, run:

npm run study:mulligans -- --pilot deepseek --restart-deepseek

This isolates network latency and provider failure from the normal simulator while still exercising the same Engine-owned observation and legal-action boundary.

Spark mulligan baseline

Spark begins as a mulligan-only learned Pilot rather than a general gameplay model. Its deck profiles are subjective Pilot configuration, not Engine rules: the current Baeloth profile describes a slow graveyard-value build with flexible commander timing, while Quintorius describes a faster spellslinger build where the commander is an early engine. Both encode the reviewed human preference to keep any functional seven once the next mulligan would cost a card.

The local training command builds one player-safe dataset and a logistic-regression model artifact from the active v5 study:

npm run spark:train-mulligans

With no arguments, it reads .data/mulligan-study/baeloth-quintorius-50-v5.json. Additional or alternative batches can be supplied by repeating --study; seeded games from different batches retain distinct grouped-validation identities.

The exporter retains the hand, deck/profile identity, mulligan state, human reason, original label, and an effective reviewed label when the explanation contains an explicit statement such as On review MULLIGAN + UNSURE. It never exports shuffled library order or hidden card-instance IDs. The 19 recorded bottom-card explanations are retained as deferred choices but are not training targets for the binary v0 model.

Spark derives deterministic features from Engine-supplied PlayerKnowledge and the current opening-hand observation. A class-balanced logistic model is trained locally; uncertain human labels receive half weight. Five-fold evaluation groups all decisions from one seeded game into the same fold to avoid trajectory leakage.

Every training invocation creates an immutable .data/spark/runs/<run-id>/ bundle with exact source-study copies, the dataset, model, evaluation metrics, and SHA-256 manifest. Convenience outputs may move forward, but a completed training run is never overwritten.

The resulting artifact can run through the same study harness:

npm run study:mulligans -- --pilot spark

SparkMulliganPilot consumes only player-safe knowledge and a DecisionPoint, and returns an offered action identity. It uses the learned hand score to choose keep or mulligan and to rank sequential London-mulligan bottom choices. Model artifacts are identified by their training timestamp; use --restart-spark to discard only incompatible Spark trajectories after retraining.

The first model established the dataset, feature, training, artifact, evaluation, and inference boundaries. Its validation numbers are a pipeline baseline, not evidence of product-level generalization. A new unseen cohort is required before claiming that Spark improves on the heuristic.

Spark opening-hand features

The first two reasoned studies contain 157 non-automatic keep-or-mulligan choices and 13 London-mulligan bottom choices. After applying explicit review statements in the reasons, the binary labels are 93 keeps and 64 mulligans; 30 are marked unsure. The reasons establish a hierarchy rather than a universal land-count rule:

  1. Usable mana is the floor: 36 of the 37 explicitly described zero- or one-land hands were mulligans. A nominal land is not necessarily usable; color, entering tapped, bounce and sacrifice costs, and conditional mana all matter.
  2. At the free-mulligan decision, a connected resource-development plan distinguishes otherwise mana-safe sevens. The reasons explicitly describe a setup or plan in 65 choices and a concrete early sequence in 65 choices.
  3. Once the next mulligan costs a card, the standard relaxes sharply: the recorded labels move from 51 keeps / 49 mulligans before a mulligan to 35 keeps / 12 mulligans after one. A functional seven is often preferred to gambling on six.
  4. Bottom-card choices preserve mana and the executable resource plan, usually removing replaceable interaction, protection, or late payoff.

The original learned feature set overfit correlations from the first cohort. In particular, the two deck/profile identities supplied a strong prior, raw Draw, Interaction, and Synergy tags learned misleading negative weights, and ordinary lands were counted twice through type and zero-mana-value features. On the second study's 67 shared decision states, 19 of Spark's 26 disagreements were false keeps: six had unusable zero- or one-land mana, two could not make the required colors, and eleven were mana-safe but lacked a connected plan. Its seven false mulligans were all Quintorius resource lines involving conditional ramp, filtering, draw, or ramp into an engine.

Feature schema version 2 therefore replaces raw card, deck, and commander identity features with a bounded, deterministic no-draw analysis. It branches over viable physical land choices through turn five and emits only normalized, player-safe capability and mulligan-state facts:

Feature family Meaning
hand:mana:land-option-count Physical cards with a playable land face; modal cards are still one card
hand:mana:land-drops-through-turn-3 Maximum sustainable land drops on one branch, respecting tapped, bounce, sacrifice, and conditional lands
hand:mana:available:<color>:turn-3 Up to three simultaneous mana of each exported Engine mana color, including reachable rocks or dorks
hand:castable:cards-through-turn-3 Distinct nonland cards castable on at least one consistent branch
hand:realizable:ramp-through-turn-3 Castable cards whose mana or land acceleration is actually reachable
hand:realizable:card-access-through-turn-3 Castable cards with immediate draw, selection, or search rather than a raw Draw tag
hand:opponent-conditional-resource-lines-through-turn-3 Reachable resource cards whose payoff depends on an opponent-facing condition
hand:self-conditional-resource-lines-through-turn-3 Reachable resource cards whose payoff still needs a self-controlled prerequisite
hand:mana:excess-land-options-above-four and five-or-more-land-options Nonlinear facts that let five- and six-land hands differ from an ideal three- or four-land start
plan:max-connected-productive-actions-through-turn-4 Best single branch's sequential actions matching the deck profile's productive roles
plan:excess-lands-without-reachable-resource-development Land-heavy hand with no reachable ramp, card access, or credible opponent-dependent resource line
policy:free-mulligan:excess-lands-without-resource-development The same weakness while a replacement seven is still free; paid mulligans remain a separate policy state
plan:commander-window-* Whether an early commander window applies and can be reached; flexible commander profiles do not treat it as a goal

Card-face definitions come from immutable PlayerKnowledge; the analyzer does not import the Engine's global card catalog and never receives card-instance IDs or shuffled order. The Engine remains the owner of rules truth. This bounded analysis is a Pilot-side forecast used to learn mulligan preferences, not a legality or full game-tree API.

Feature schema version 3 intentionally invalidates earlier model artifacts. It adds the nonlinear land-surplus and free-mulligan interaction after v3 exposed Spark's tendency to keep land-heavy Baeloth hands with no resource-development plan. The combined 1,500-epoch run uses 241 binary decisions: 143 keeps and 98 mulligans. Five-fold validation, grouped by seeded game and study batch, agrees with 201 decisions (83.4% accuracy and 82.3% balanced accuracy), with 88.1% keep recall and 76.5% mulligan recall. Full-corpus training agreement is 208/241 (86.3%). These are development-corpus results, not evidence of generalization; v4 is the next untouched measurement.

Current migration status

The contract and orchestration boundary are active on the core simulation paths. Mulligan evaluation and London-mulligan bottom-card selection use only OpeningHandObservation and the offered legal actions.

HeuristicPilot receives no GameState. Its gameplay chooser and feature calculations live in src/pilot/heuristic and consume a detached PlayerRulesSnapshot from PlayerObservation.rules. The engine owns this versioned payload and the rules queries in src/engine/playerRulesQueries.ts. Scoring, strategy weights, and action selection remain in the pilot.

The snapshot explicitly lists the public rules facts needed by those queries, including pending choices, continuous effects, public event histories, and zone-object identities. It excludes the live library, RNG, callbacks, and deferred resolution bookkeeping. Library knowledge contains unordered deck counts, cards exposed by current choices, and explicitly visible top positions. Engine query implementations construct a private evaluation game from those facts, using synthetic identities for unknown library cards. That game is never executed or supplied to the pilot. These queries evaluate modeled rules; they do not provide a hidden-outcome rollout or sampling interface.

Existing standalone helpers such as chooseAction(game) remain compatibility entry points for terminal callers and tests. They obtain legal actions from the live engine and adapt its public snapshot before entering the heuristic. HeuristicPilot does not call these adapters. JSON-round-trip gameplay tests verify that decisions do not depend on a retained live game reference.

The payload currently uses the engine's card and rules types, and the heuristic still uses the local modeled-card catalog. Removing catalog imports or wiring pilot selection through batch workers is separate work. Future observation fields should follow demonstrated player-visible needs.

Extension rules

When a Pilot cannot make a sound decision with the current contract:

  1. Confirm that the missing fact is legitimately available to the player.
  2. Add the smallest serializable field to PlayerKnowledge for stable facts or PlayerObservation for changing state.
  3. Have the Engine derive and freeze that value from rules state.
  4. Test hidden-information exclusion and serialization.
  5. Update the Pilot to use the new contract field.
  6. Keep legal-action generation and validation in the Engine.

When introducing a new Pilot, inject it through the Simulator and prove that it can select an offered action without receiving GameState. When introducing a new kind of parameterized action, document and test the Engine validator for the allowed specialization.