- Rename `retro-gamer create` to `retro-gamer init` with positional args (GAME OUTPUT) instead of --game/--output flags - Add [tool.retro-gamer].factory = "module:attr" to declare the game factory function from pyproject.toml instead of relying on a hard-coded create_game attribute - Remove extras_size from user-facing config; it is now measured automatically from a sample observation and never declared - Update docs throughout: create→init, runs/→training/, add factory field documentation, clarify [model] vs [training] hyperparameter sections, remove stale version-history explanations - Bump version to 0.3.0; require retro-games>=2.5.0
513 lines
18 KiB
ReStructuredText
513 lines
18 KiB
ReStructuredText
Reference
|
||
=========
|
||
|
||
Game description fields
|
||
-----------------------
|
||
|
||
Game descriptions are written in the ``[tool.retro-gamer]`` section of
|
||
your game project's ``pyproject.toml``. ``retro-gamer init`` reads
|
||
this section and copies the metadata into the training run's
|
||
``config.toml``, where it can also be inspected or hand-edited.
|
||
|
||
A complete example for the Snake game:
|
||
|
||
.. code-block:: toml
|
||
|
||
[tool.retro-gamer]
|
||
factory = "snake:create_game"
|
||
actions = ["KEY_RIGHT", "KEY_UP", "KEY_LEFT", "KEY_DOWN"]
|
||
reward = "score"
|
||
character_set = ["@", "*", ">", "<", "^", "v"]
|
||
|
||
You do not need to specify the board size: ``retro-gamer`` reads it
|
||
directly from your game's ``board_size`` attribute.
|
||
|
||
The fields are described below.
|
||
|
||
``factory``
|
||
~~~~~~~~~~~
|
||
|
||
**Required.** A ``"module:attr"`` string naming the function that creates
|
||
a fresh game instance. The function must take no arguments and return a
|
||
new :class:`retro.game.Game`.
|
||
|
||
.. code-block:: toml
|
||
|
||
factory = "snake:create_game"
|
||
|
||
``actions``
|
||
~~~~~~~~~~~
|
||
|
||
**Required.** A list of keystroke names the agent may send to the game
|
||
each turn. Use arrow key names for directional games, or single
|
||
characters for character-key games.
|
||
|
||
.. code-block:: toml
|
||
|
||
actions = ["KEY_RIGHT", "KEY_UP", "KEY_LEFT", "KEY_DOWN"]
|
||
|
||
The agent also has access to a no-op action (doing nothing). The total
|
||
number of actions in the Q-network output is ``len(actions) + 1``.
|
||
|
||
``reward``
|
||
~~~~~~~~~~
|
||
|
||
**Required.** The key in the game's state dictionary to use as the
|
||
reward signal. The reward computed for each turn is the *change* in
|
||
this value from the previous turn.
|
||
|
||
.. code-block:: toml
|
||
|
||
reward = "score"
|
||
|
||
``character_set``
|
||
~~~~~~~~~~~~~~~~~
|
||
|
||
**Optional.** A list of single characters that may appear on the board.
|
||
Each character occupies one "slot" in the one-hot encoding. Characters
|
||
not in this list are treated as empty space.
|
||
|
||
.. code-block:: toml
|
||
|
||
character_set = ["@", "*", ">", "<", "^", "v"]
|
||
|
||
If omitted, ``retro-gamer`` runs an exploration phase to discover the
|
||
characters that appear in practice. The length of this phase is
|
||
controlled by the ``exploration_turns`` hyperparameter.
|
||
|
||
Preprocessing options
|
||
---------------------
|
||
|
||
Preprocessing options live in the ``[preprocessing]`` section of a run's
|
||
``config.toml``. They control how the game's board and state are
|
||
transformed into the observation vector that the neural network sees.
|
||
``retro-gamer init`` writes sensible defaults; you can edit them by
|
||
hand before running ``retro-gamer train``.
|
||
|
||
.. note::
|
||
|
||
Changes to any ``[preprocessing]`` option—or to the game description
|
||
fields above—make existing checkpoints incompatible. Run
|
||
``retro-gamer clean`` before retraining after such changes.
|
||
|
||
``spatial`` (default: ``false``)
|
||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
|
||
Whether to treat the board as a 2D spatial scene. When ``true``, the
|
||
trainer uses a convolutional neural network (CNN); when ``false``, a
|
||
multilayer perceptron (MLP) that sees the board as a flat list of
|
||
numbers.
|
||
|
||
``board`` (default: ``true``)
|
||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
|
||
Whether to include the board encoding in the observation vector. Set
|
||
to ``false`` to train on game state variables only, with no board at
|
||
all. This is useful for games with small, enumerable state spaces where
|
||
a lookup table (classic Q-learning) is sufficient.
|
||
|
||
When ``board = false``:
|
||
|
||
- ``spatial`` must also be ``false`` (no board means no 2D scene for a CNN).
|
||
- At least one key must be listed in ``observe_state``.
|
||
- ``character_set`` is not required and character discovery is skipped.
|
||
|
||
.. code-block:: toml
|
||
|
||
[preprocessing]
|
||
board = false
|
||
observe_state = ["board_state"]
|
||
|
||
``observe_state`` (default: ``[]``)
|
||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
|
||
A list of keys from ``game.state`` to include in the observation
|
||
vector, appended after the board encoding (or as the entire
|
||
observation when ``board = false``). Scalar values contribute one
|
||
element each; list or tuple values are flattened.
|
||
|
||
.. code-block:: toml
|
||
|
||
observe_state = ["apple_dx", "apple_dy"]
|
||
|
||
The keys must be present in ``game.state`` at every step, initialized
|
||
before the game starts. All values that are lists or tuples must always
|
||
have the same length from episode to episode.
|
||
|
||
.. warning::
|
||
|
||
``observe_state`` keys must be initialized to their final shape before
|
||
the game starts. If a key is absent or its list length changes between
|
||
episodes, training will crash with an error explaining which key
|
||
changed and by how much. This happens because the neural network's
|
||
input layer has a fixed size determined at the start of training; it
|
||
cannot adapt to a changing observation shape mid-run.
|
||
|
||
Always initialize every observed key with a placeholder of the
|
||
correct type and length before the first ``game.step()`` call.
|
||
|
||
.. _observation-function:
|
||
|
||
``observation_function`` (default: none)
|
||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
|
||
**Optional**, set in ``[metadata]`` (alongside ``actions``/``reward``/
|
||
``character_set``/``board_size``), not in ``[preprocessing]``. A
|
||
``"module:attr"`` string naming a function ``f(game) -> observation`` that
|
||
fully replaces the built-in board/``observe_state`` encoding described
|
||
above. Mutually exclusive with ``observe_state`` — they're two conflicting
|
||
ways of describing the same thing, and setting both raises an error.
|
||
|
||
.. code-block:: toml
|
||
|
||
[metadata]
|
||
observation_function = "my_game:get_observation"
|
||
|
||
For DQN training, the function must return a flat, numeric, fixed-length
|
||
1-D array every time it's called — the same contract the built-in encoder
|
||
follows: a flattened one-hot board (sized from ``character_set`` ×
|
||
``board_size``, if ``board = true``) followed by any extra features, all in
|
||
one vector. ``character_set`` and ``board_size`` stay required either way,
|
||
because that's what lets ``observation_function`` also use a spatial
|
||
(``spatial = true``) network — the trainer slices the flat vector back into
|
||
a board tensor using exactly those two fields, the same way it does for the
|
||
built-in encoder.
|
||
|
||
``observation_function`` is the training-side complement to the game's
|
||
state dict. Use ``observe_state`` when the game already computes the
|
||
features you want and stores them in ``game.state``; use
|
||
``observation_function`` when you want to transform the game's board or
|
||
state into a representation that is more useful for learning — for example,
|
||
mapping all obstacle types to a single character, or cropping the board to
|
||
an egocentric window centred on the player.
|
||
|
||
This is how you get an egocentric (cropped, player-centered) board —
|
||
call ``egocentric_board()`` and ``encode_board()`` yourself, from
|
||
:mod:`retro_gamer.observation`, inside your own function, and declare
|
||
``board_size`` to match your crop:
|
||
|
||
.. code-block:: python
|
||
|
||
import numpy as np
|
||
from retro.views.headless import HeadlessView
|
||
from retro_gamer.observation import egocentric_board, encode_board, encode_state
|
||
|
||
CHARACTER_SET = ["@", "*", ">", "<", "^", "v"]
|
||
RADIUS = 8
|
||
|
||
def egocentric_observation(game):
|
||
view = HeadlessView()
|
||
view.on_game_start(game)
|
||
view.render(game)
|
||
head = game.get_agent_by_name("Snake head")
|
||
cropped = egocentric_board(view.board_characters, head.position, RADIUS)
|
||
board_vec = encode_board(cropped, CHARACTER_SET).flatten()
|
||
extras = encode_state(game.state, ["apple_dx", "apple_dy"])
|
||
return np.concatenate([board_vec, extras])
|
||
|
||
.. code-block:: toml
|
||
|
||
[metadata]
|
||
board_size = [17, 17] # 2*RADIUS + 1
|
||
observation_function = "my_module:egocentric_observation"
|
||
|
||
Outside DQN training — for example, a tabular Q-learning lab that
|
||
uses :class:`~retro_gamer.GameEnvironment` directly without
|
||
:class:`~retro_gamer.DQNTrainer` — there's no 1-D requirement at all.
|
||
``observation_function`` can return anything you want to use as your
|
||
observation, including a plain tuple used as a dict key.
|
||
|
||
.. _hyperparameters:
|
||
|
||
Hyperparameters
|
||
---------------
|
||
|
||
Hyperparameters are split across two sections of ``config.toml``:
|
||
|
||
- ``[model]`` — network architecture (changing these requires starting fresh)
|
||
- ``[training]`` — learning algorithm parameters (safe to change at any time)
|
||
|
||
Both sections can be set via ``retro-gamer init`` options or edited directly
|
||
in ``config.toml`` between ``init`` and ``train``.
|
||
|
||
Learning and optimization (``[training]``)
|
||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
|
||
``learning_rate`` (default: ``0.0001``)
|
||
The step size used by the Adam optimizer when updating network
|
||
weights. Larger values converge faster but may be unstable; smaller
|
||
values are more stable but slower.
|
||
|
||
``learning_rate_decay`` (default: ``0.9999``)
|
||
Multiplicative decay applied to the learning rate after each
|
||
episode. The learning rate decreases geometrically over training,
|
||
helping the network fine-tune later without destabilizing early
|
||
progress. With the default value, the learning rate decays to about
|
||
13 % of its starting value after 20 000 episodes.
|
||
|
||
``gamma`` (default: ``0.99``)
|
||
The discount factor for future rewards. A value of 1.0 makes the
|
||
agent value all future rewards equally; smaller values make the
|
||
agent increasingly myopic.
|
||
|
||
Exploration (``[training]``)
|
||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
|
||
``epsilon`` (default: ``1.0``)
|
||
The initial exploration rate. At each turn, the agent takes a
|
||
random action with probability ``epsilon`` and exploits its current
|
||
Q-function with probability ``1 - epsilon``.
|
||
|
||
``epsilon_decay`` (default: ``0.9997``)
|
||
Multiplicative decay applied to ``epsilon`` after each episode.
|
||
|
||
``epsilon_min`` (default: ``0.05``)
|
||
The floor below which ``epsilon`` will not fall. A small amount of
|
||
continued exploration prevents the agent from becoming permanently
|
||
committed to a suboptimal policy.
|
||
|
||
Memory and sampling (``[training]``)
|
||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
|
||
``batch_size`` (default: ``64``)
|
||
The number of experiences sampled from the replay buffer per
|
||
training step.
|
||
|
||
``memory_capacity`` (default: ``50000``)
|
||
The maximum number of experiences the replay buffer can hold. When
|
||
full, the oldest experiences are discarded.
|
||
|
||
``prioritize_experiences`` (default: ``true``)
|
||
Whether to use prioritized experience replay. When ``true``,
|
||
experiences with larger TD errors are sampled more frequently.
|
||
This often improves sample efficiency at a modest computational
|
||
cost.
|
||
|
||
Model architecture (``[model]``)
|
||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
|
||
Changing any ``[model]`` option requires starting fresh (run
|
||
``retro-gamer clean`` before retraining).
|
||
|
||
``hidden_sizes`` (default: ``[128, 64]``)
|
||
A list of integers giving the size of each hidden layer in the MLP
|
||
head. The default creates two layers: 128 units then 64. For spatial
|
||
games this follows the CNN; for non-spatial games it is the full
|
||
network. Larger or deeper networks can represent more complex
|
||
Q-functions but train more slowly and may need more episodes.
|
||
|
||
Training duration (``[training]``)
|
||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
|
||
``training_episodes`` (default: ``20000``)
|
||
The total number of game episodes to run. Each episode runs until
|
||
the game ends or ``max_turns_per_episode`` turns have elapsed.
|
||
|
||
``max_turns_per_episode`` (default: ``2000``)
|
||
A safety cutoff preventing a single episode from running
|
||
indefinitely (for example, if the agent finds a way to avoid
|
||
dying).
|
||
|
||
``target_update_freq`` (default: ``500``)
|
||
How many training steps between updates of the target network.
|
||
More frequent updates make training targets move faster (less
|
||
stable); less frequent updates make them more stable but slower
|
||
to reflect new learning.
|
||
|
||
``train_every`` (default: ``4``)
|
||
Run one training step every N game steps. Higher values speed up
|
||
episode collection at the cost of fewer gradient updates per
|
||
experience. The default of 4 is a good balance for most games;
|
||
set to 1 to train on every step.
|
||
|
||
Character discovery (``[training]``)
|
||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
|
||
``exploration_turns`` (default: ``200``)
|
||
When ``character_set`` is not specified, the number of random
|
||
turns to run at the start of training to discover which
|
||
characters appear on the board.
|
||
|
||
``unknown_character_strategy`` (default: ``"ignore"``)
|
||
What to do when a character appears during training that is not
|
||
in the established ``character_set``. ``"ignore"`` treats it as
|
||
an empty cell; ``"extend"`` rebuilds the model with an extended
|
||
character set.
|
||
|
||
CLI reference
|
||
-------------
|
||
|
||
``retro-gamer init``
|
||
~~~~~~~~~~~~~~~~~~~~~
|
||
|
||
Create a new training run directory with ``config.toml``. Game metadata
|
||
is read automatically from the ``[tool.retro-gamer]`` section of your
|
||
game's ``pyproject.toml``; you do not pass it on the command line.
|
||
|
||
.. code-block:: console
|
||
|
||
% retro-gamer init GAME OUTPUT [OPTIONS]
|
||
|
||
**Required arguments:**
|
||
|
||
- ``GAME`` — Your game, specified as a directory path or a Python
|
||
module name:
|
||
|
||
- Directory: ``games/snake``
|
||
- Module name: ``retro.examples.snake``
|
||
|
||
The ``[tool.retro-gamer]`` section is read from the ``pyproject.toml``
|
||
found in or above the game directory.
|
||
- ``OUTPUT`` — Directory to create for this training run
|
||
(e.g. ``training/snake``).
|
||
|
||
**Hyperparameter options** (all optional; see :ref:`hyperparameters`):
|
||
|
||
- ``--training-episodes N``
|
||
- ``--hidden-sizes SIZES`` — comma-separated, e.g. ``512,256``
|
||
- ``--learning-rate F``
|
||
- ``--learning-rate-decay F``
|
||
- ``--gamma F``
|
||
- ``--epsilon-decay F``
|
||
- ``--epsilon-min F``
|
||
- ``--batch-size N``
|
||
- ``--memory-capacity N``
|
||
- ``--target-update-freq N``
|
||
- ``--max-turns-per-episode N``
|
||
- ``--exploration-turns N``
|
||
- ``--train-every N``
|
||
- ``--prioritize-experiences`` / ``--no-prioritize-experiences``
|
||
|
||
``retro-gamer train``
|
||
~~~~~~~~~~~~~~~~~~~~~
|
||
|
||
Train a DQN agent.
|
||
|
||
.. code-block:: console
|
||
|
||
% retro-gamer train RUN_DIR
|
||
|
||
``RUN_DIR`` must contain a ``config.toml`` generated by ``retro-gamer
|
||
init``. If checkpoints already exist in ``RUN_DIR``, training
|
||
automatically resumes from the latest one so prior work is never lost.
|
||
|
||
If all configured episodes have already been completed, the command
|
||
prints a message and exits immediately. To keep training, increase
|
||
``training_episodes`` in ``config.toml`` and run again.
|
||
|
||
**Incompatible changes.** Some config changes make existing checkpoints
|
||
unusable. If you change any of the following, ``retro-gamer train`` will
|
||
detect the mismatch and refuse to resume, with a clear explanation:
|
||
|
||
- ``actions``, ``reward``, ``character_set``, ``board_size``,
|
||
``observation_function`` (``[metadata]``) — game description and
|
||
observation shape
|
||
- ``spatial``, ``board``, ``observe_state`` (``[preprocessing]``) —
|
||
observation encoding
|
||
- ``hidden_sizes`` (``[model]``) — network architecture
|
||
|
||
Run ``retro-gamer clean RUN_DIR`` to remove the old checkpoints and start
|
||
fresh. Other hyperparameter changes (learning rate, epsilon, etc.) are
|
||
safe and take effect immediately on the next training run.
|
||
|
||
``retro-gamer play``
|
||
~~~~~~~~~~~~~~~~~~~~
|
||
|
||
Watch a trained agent play the game in the terminal.
|
||
|
||
.. code-block:: console
|
||
|
||
% retro-gamer play RUN_DIR [--checkpoint NAME] [--framerate N]
|
||
|
||
By default, the latest available checkpoint is loaded. Use
|
||
``--checkpoint`` to load a specific one by name (e.g. ``ep_0100``).
|
||
``--framerate`` sets the target frames per second (default: 12). Press
|
||
Enter or Escape to quit.
|
||
|
||
``retro-gamer clean``
|
||
~~~~~~~~~~~~~~~~~~~~~
|
||
|
||
Remove all checkpoints and the training log from a run directory.
|
||
|
||
.. code-block:: console
|
||
|
||
% retro-gamer clean RUN_DIR
|
||
|
||
Prompts for confirmation before deleting. Use ``--yes`` / ``-y`` to skip
|
||
the prompt. The ``config.toml`` is preserved so you can run
|
||
``retro-gamer train`` immediately to start fresh with the same settings.
|
||
|
||
Use this after making an incompatible change (see ``retro-gamer train``
|
||
above) or any time you want to restart training from scratch.
|
||
|
||
``retro-gamer info``
|
||
~~~~~~~~~~~~~~~~~~~~~
|
||
|
||
Print a summary of a training run: metadata, hyperparameters, recent
|
||
checkpoint log, and available checkpoints.
|
||
|
||
.. code-block:: console
|
||
|
||
% retro-gamer info RUN_DIR
|
||
|
||
Training run directory structure
|
||
---------------------------------
|
||
|
||
A training run is a self-contained directory with the following
|
||
contents:
|
||
|
||
.. code-block:: text
|
||
|
||
training/snake/
|
||
├── config.toml # game description + hyperparameters
|
||
├── training.log # architecture rationale + per-episode log
|
||
└── checkpoints/
|
||
├── ep_0100.pt # model weights at episode 100
|
||
├── ep_0200.pt
|
||
└── ... # one file saved every 100 episodes
|
||
|
||
``config.toml`` is written by ``retro-gamer init`` and updated (with
|
||
the discovered character set and resolved hyperparameters) when
|
||
``retro-gamer train`` begins. It has five sections: ``[game]``,
|
||
``[metadata]``, ``[preprocessing]``, ``[model]``, and ``[training]``.
|
||
Editing ``config.toml`` between ``init`` and ``train`` is the
|
||
recommended way to adjust hyperparameters.
|
||
|
||
``training.log`` begins with the full network architecture description,
|
||
then one line per checkpoint (every 100 episodes) in the format::
|
||
|
||
[ep_NNNN] ep=SSSS-NNNN avg_reward=F avg_steps=N epsilon=F avg_loss=F time=Xm Xs total=Xm Xs
|
||
|
||
Each field averages over the episodes since the previous checkpoint:
|
||
|
||
- ``ep=SSSS-NNNN`` — episode range covered by this entry
|
||
- ``avg_reward`` — mean total reward per episode (positive = good)
|
||
- ``avg_steps`` — mean episode length in game turns
|
||
- ``epsilon`` — current exploration rate (approaches ``epsilon_min`` over time)
|
||
- ``avg_loss`` — mean Huber loss across training steps (should decrease as learning
|
||
stabilises). Huber loss equals ½·(q−t)² for small errors and |q−t|−½ for large
|
||
ones, so it stays bounded even when Q-values are large. Values in the range
|
||
0–10 are typical; a slow downward trend over thousands of episodes is the
|
||
healthy pattern. A loss that grows without bound indicates a learning rate
|
||
that is too high.
|
||
- ``time`` — wall-clock time for this checkpoint interval
|
||
- ``total`` — cumulative training time across all sessions
|
||
|
||
When training is resumed, a ``=== Resumed from ... ===`` line is appended
|
||
so the log records the full history of a run across multiple sessions.
|
||
|
||
Python API
|
||
----------
|
||
|
||
For advanced use, ``retro-gamer``'s components are importable as a
|
||
library. See the :doc:`api` reference for full details.
|
||
|
||
.. code-block:: python
|
||
|
||
from retro_gamer import GameMetadata, DQNTrainer
|
||
from retro.examples.snake import create_game
|
||
|
||
metadata = GameMetadata.from_pyproject("retro.examples.snake")
|
||
trainer = DQNTrainer(create_game, metadata, "training/snake/")
|
||
trainer.train()
|