is organized into 2 primary modules:
emulation: Handles GameBoy emulation, parsing and state trackinginterface: Implements high level actions, and Gym-compliant environments
See the API documentation to understand the code base, the rest of this document goes into details on how you would implement new features or test tasks in
.
- Custom Starting States
- Descriptive State and Event Tracking
- Reward Engineering
- Adding New ROMs
- Adding New Test Tasks
- Useful things
Easy. The only question is whether you want to save an mGBA state (perhaps you use cheats to lazily put the agent in a very specific state) or save a PyBoy state directly (i.e. you start from an existing state and play to the new state).
From mGBA state:
First, start with mGBA and make sure to match the text box frame options from the existing default states. This is vital to ensure the state parsing system works. Play till the point you want to replicate with a state and save the game (go to the start menu and save) in the state you want to restore from. This will make a game_ROMNAME.sav file in the same directory as the rom file. Then run:
python dev/save_state.py --game <game> --state_name <name>This will save the state and allows you to load it by specifying it as a state name.
To get to the state from PyBoy, first make sure the gameboy_dev_play_stop parameter is configured to false. Then, run:
python dev/dev_play.py --game <game> --init_state <optional_starting_state>This will run the game with the option to enter dev mode. Play the game like you usually would, until you reach the state you want to save. Then, go to the gameboy configs while playing the game (at the state you want to save), change the gameboy_dev_play_stop parameter to true (save the configs file) and then check the terminal. You will get a message with the possible dev actions. The one you're looking for is s <name>, which saves the state.
Regardless of how you did it, you can test that your state save worked with:
python demos/emulator --game <game> --init_state <name>Maybe you want to enhance the observation space of the agent with information about the current playthrough (e.g. current map ID, enemy team level). Perhaps you want to train text-only / weak visual agents, and parse as much of the screen image as possible into numerical signals / text (e.g. your team stats, bag contents). Some might not even care about their agents, but want to have a sophisticated set of metrics that they can look at to assess goal conditions, judge the quality of a playthrough, or craft a good reward function.
Whatever your motivation,
provides a powerful set of approaches for reading game states, and then allows you to aggregate over these values over time to compute useful metrics for reward assignment and evaluation.
The first thing to do is detect an event at a moment in time. This is done in subclasses of the StateParser object in one of two ways:
- Emulator Screen Captures: Often particular game states can be cleanly identified by a unique text popup, or some other characteristic marker on the screen. Any of these can be easily captured and checked with the existing parsing system. For example, the current implementation for Pokémon Red has screen captures set up to identify which starter the player chooses. See the section below for examples of this being done. See the
StateParserAPI documentation for a quick overview on how this works. - Memory Slot Hooks: A strong alternative is to just directly read statistics from the game's WRAM. Visually inaccessible information (e.g. the attack stats of all Pokémon on the opponents team) are often easy to obtain this way. The only catch is, this method relies on knowing which memory slots to look for. That's easy enough for games which have excellent decompilation guides, but is much harder to do for ROM hacks which may mess around with the slots arbitrarily or less popular games. See the memory reader state parser to get a sense of how you should go about this.
These approaches allow your state parsers to give instant-wise decisions or indications when an event has occured. You can then configure your StateTracker to use the parser to check for this flag / read this information, and store appropriate metrics. See the existing parsers and trackers for examples.
Setting up a new game is an easy process at a basic level, but can be an involved endeavour if you want to make the new environment a strong one. Please do reach out to me if you have any questions, and we can work to merge the new ROM into
together.
- Set the repo to
debugmode by editing the config file - Create a
<game>_rom_data_pathparameter in the configs (either as a new file or in an existing one) - Obtain the ROM and place it in the desired path under the ROM data folder. Remember, the
<game>_rom_data_pathfolder is rooted at thestorage_dirfrom the configs. - Go to the registry and add the ROM name to :
GAME_TO_GB_NAME: This will be the name the system expects to find in<storage>/<game>_rom_data_path/_STRONGEST_PARSERS: withDummyParseras the value.AVAILABLE_STATE_TRACKERS: give it adefaultvalue ofStateTracker.AVAILABLE_EMULATORS: give it adefaultvalue ofEmulator.
- Run
python dev/create_first_state.py --game <game>. This will create a default state. You will not be able to run theEmulatoron this ROM before doing this. - Run
python dev/dev_play.py --game <game>(with thegameboy_dev_play_stopparameter set tofalse) and proceed through the game until you reach a satisfactory default starting state. Then, open the config file and setgameboy_dev_play_stoptotrueand save the config file. This will trigger a dev mode and ask you for a terminal input. Enters defaultand you will set that as the new default state. Enters initialas well to save it properly.
I have provided an example video for this process. Note: In the video, I set the text speed to fast. This was the wrong choice, and so I have set it to slow in all states.
The above steps will let you play the game on the emulator, but the real power of this framework is only realized when you get involved and create a proper StateParser. As mentioned in the section above, this is done either by reading from gameboy memory states or by setting up screen captures to track events. Here, I detail the screen capture method.
Simply put, this approach aims to capture a given region of the games frame at the right moment, hence saving what the screen "looks like" when a particular event occurs. For example, in Pokémon, the top right of the screen always has the edge of the player menu, and is hence a reliable signal as to whether or not the player is in the menu. The exact regions and events to capture will depend on the game, but the most important components are:
NamedScreenRegion: EveryStateParsercan define certain boxes within the game screen (e.g. the top right portion where the player menu identifier will pop up). These can linked to one or more reference targets, that you need to manually capture once and save. After you save the target, theStateParserallows you to take any game frame, select the region in question, and compare it to the reference image. Once you've designated the named regions in the state parser, run the game in dev play mode, stop the game at the moment you want to capture. Then, runc <region_name>to save the screen region at that point. The Pokémon parsers show a clear example of this, and I have provided an example video of the frames being captured.
You will know that you have filled out all required regions when you can run python demos/emulator.py --game <game> without debug mode.
To use the StateParser you created, make sure to:
- Add the parser to the registry
- Create a
MetricGroupobjects that calls on the parsers methods and capabilities in itsstepmethod - Add these
MetricGroupobjects to aStateParserand then add that to the registry
To enable an agent to play the game in a gym-style environment loop, you must create a simple Environment subclass with implementations for the abstract methods, and add this to the interface registry. That is now a gym-compliant game environment.
The above set up gives you more descriptive state information, but still forces the agents use simple button presses to play the game. You must think of decent actions you can implement, and create HighLevelAction subclasses to execute them.
avoids most domain-knowledge specific reward design, with a motivation of having the agent discover the best policy with minimal guidance. But it's absolutely possible to use your knowledge of the game to create sophisticated reward systems, like other people have.
You'll likely want to gather as much state and trajectory information as possible, for which you should see the section above.
Then, you'll want to create your own Environment subclass, and configure the reward return. See PokemonRedChooseCharmanderFastEnv for more
Setup Speedrun Guide: I've documented the fastest workflow I have found to capturing all the screens for a Pokémon ROM hack properly. This may come in handy for someone.
Start by just playing through the game (super high gameboy_headed_emulation_speed) and establishing save states for the following:
initial: Right out of the intro screen with options set to fastest / least animationstarter: Right before the player needs to make a choice of starterpokedex: Not too long after the player obtains the Pokedex, but anywhere you like.
Then, start with:
python dev/dev_play.py --game <game> --init_state initial
You can tick off the following captures:
dialogue_bottom_right: usually theres something you can interact with in your starting roommenu_top_right: open the start menupc_top_left: there is often a PC in your roomplayer_card_middle: open your player cardmap_bottom_right: usually there's a map around you
Then, switch out to the start choice state with l starter. Use this state to capture:
dialogue_choice_bottom_right: confirmation message for startername_entity_top_left: give the starter a nicknamebattle_enemy_hp_text: either a rival battle or just your first Pokémon battlebattle_player_hp_text: samepokemon_list_hp_text: can do once you've got the starter
Then honestly you probably want to exit with e and start again at the pokedex state with:
python dev/dev_play.py --game <game> --init_state pokedexYou'll get a message letting you know what's left. You can finish them all off now. If any of the captures weren't clean and good, you should leave them for the end and override their named screen regions.
Using this process I'm able to set up all but one capture in under 10 minutes (the video cuts off with only pokedex_info_height_text unassigned because it needs to be manually repositioned as an override region).
To create a new test task that automatically detects when an agent succeeds (or fails) at a specific goal, follow these steps:
1. Create an initial state First, create a starting state from which your task is achievable. See the section above for detailed instructions on creating states.
2. Define termination and truncation conditions
- Termination: The goal has been achieved. This should be a reliably reproducible screen element that always appears when the goal is reached (e.g., unique dialogue when defeating a specific trainer). In this framework, termination always equals task success - we avoid failure termination signals to prevent agents from using them as learning feedback.
- Truncation: Optional. Cut the episode short when the player can no longer achieve the task (e.g., walked too far away). Maximum environment / emulator steps are handled automatically, so don't bother implementing that.
3. Set up parser for screen capture (if needed) If your termination condition relies on a specific screen capture not already available in the parser, you'll need to add it. See the screen capture method in the section above for guidance on capturing named screen regions.
Make sure to use python -m gameboy_worlds.setup_data push --game <game> to update the cloud database.
4. Create the termination/truncation metric Make a child or descendant of the TerminationTruncationTracker :
- For termination only:
TerminationMetric(line 466) - For both termination and truncation:
TerminationTruncationMetric(line 368)
If using screen region comparisons (most common), inherit from:
Example: PokemonCenterTerminateMetric inherits from both RegionMatchTerminationMetric and TerminationMetric.
5. Create the test tracker
Most trackers can be created by simply setting the TERMINATION_TRUNCATION_METRIC class parameter. See PokemonRedCenterTestTracker for an example.
class MyTestTracker(PokemonTestTracker):
TERMINATION_TRUNCATION_METRIC = MyCustomTerminateMetric6. Register the tracker
Add your new tracker to the AVAILABLE_STATE_TRACKERS dictionary in the registry with a descriptive name.
7. Test your implementation Verify it works with the test play script:
python dev/dev_play.py --game <game> --state_tracker_class <your_tracker_name> --init_state <your_start_state>The game should automatically stop when you reach the termination/truncation condition.
Example video: here
The steps above cover the general process, but here is a more mechanical, copy-pasteable walkthrough of the same workflow, using the example of adding a "navigate to the white car" task to Harry Potter: Chamber of Secrets.
Role Legend:
- 🧑 (Human): Requires manual effort from you (playing the game, visual verification).
- 🤖 (LLM): An LLM can easily write or autofill this code for you if you provide it the names.
Step 1: Create an Initial Save State
You need a starting state from which the agent will attempt to solve the task.
-
🧑 (Human) Run the Emulator in Dev Mode:
python dev/dev_play.py --game harry_potter_chamber_of_secrets
(Note: Ensure
gameboy_dev_play_stopis set tofalseinconfigs/gameboy_vars.yamlbefore running this) -
🧑 (Human) Play to the Start Point: Play the game normally until you reach the exact moment you want the agent's task to begin.
-
🧑 (Human) Trigger Save Breakpoint: Leave the game running. Open
configs/gameboy_vars.yamland changegameboy_dev_play_stop: falsetotrue. Save the YAML file. Pitfall: You must keep the emulator running while you edit the file. The game will automatically pause and prompt the terminal. -
🧑 (Human) Save the State: In your terminal prompt, type:
s burrow_start
(Replace
burrow_startwith your desired initial state name).Where are these states saved? Once saved, the
.statefiles are physically stored in your storage directory under the specific game's ROM data path. For example, Harry Potter states are located at:storage/rom_data/harry_potter/harry_potter_chamber_of_secrets/states/Tip: You can quickly view a list of all your currently saved states (and export them to a CSV) by running:
python dev/list_states.py --game harry_potter_chamber_of_secrets
Step 2: Configure the Parser and Capture Target Regions
To detect when the task is "completed" (e.g., standing next to the car), you must define a screen bounding box and capture a reference image of what success looks like.
(Newcomer Tip: To figure out the exact x, y, width, height coordinates for your new bounding box, you can type d in the dev play terminal prompt to open a full-screen image viewer. Hovering your mouse over the image will show the exact pixel coordinates!)
-
🤖 (LLM) Define Bounding Box: Open
src/gameboy_worlds/emulation/harry_potter/parsers.py. LocateHarryPotterChamberOfSecretsParser(or similar subclass). Add your region toMULTI_TARGET_REGIONSand the target name toMULTI_TARGETS:MULTI_TARGET_REGIONS = [ # ... existing regions ("car_area", 80, 65, 55, 70), # Format: (name, x, y, width, height) ] MULTI_TARGETS = { # ... existing targets "car_area": ["next_to_car"], }
🧑 (Human) Caveat: You must manually determine the bounding box dimensions for this to make sense.
-
🧑 (Human) Play to the "Success" Screen: (Pitfall: Ensure you have changed
gameboy_dev_play_stopback tofalseinconfigs/gameboy_vars.yamlfirst, otherwise the emulator will instantly freeze!) Load back into your new state:python dev/dev_play.py --game harry_potter_chamber_of_secrets --init_state burrow_start
Play until you reach the exact screen representing task success.
-
🧑 (Human) Capture the Reference Image: Change
gameboy_dev_play_stoptotruein your configs again. When the terminal prompts you, capture the specific target using the multi-target syntax:c car_area,next_to_car
(Note: The
ccommand references the region name defined in your parser, followed by a comma, followed by the specific target name. This saves it to a.npyfile. Verify the image pop-up looks correct before closing it).
Step 3: Define the Termination Metric
Metrics define the logic for ending an episode.
- 🤖 (LLM) Create the Metric Class:
Open
src/gameboy_worlds/emulation/harry_potter/test_metrics.pyand append:Caveat: If your termination logic requires checking multiple possible targets or regions, you'll need to usefrom gameboy_worlds.emulation.tracker import RegionMatchTerminationOnlyMetric class NavigateToCarTerminateMetric(RegionMatchTerminationOnlyMetric): REQUIRED_PARSER = HarryPotterChamberOfSecretsParser _TERMINATION_NAMED_REGION = "car_area" _TERMINATION_TARGET_NAME = "next_to_car"
MULTI_TARGET_REGIONS(in your parser) or create custom logic.MULTI_TARGET_REGIONS/MULTI_TARGETS: Use this when checking the exact same bounding box coordinate on the screen, but the image inside could be one of several possibilities (e.g., "standing left of car" vs "standing right of car").- Multiple Distinct Regions: If success means checking entirely different bounding boxes (e.g. matching
car_areaOR matchingtruck_area), you will need to implement custom termination logic using multiplenamed_region_matches_targetchecks.
Step 3.5: Define Subgoals (Highly Recommended)
Generally, we want to create tests that have at least 1 subgoal to track partial progress and provide intermediate rewards.
-
🤖 (LLM) Create the Subgoal Class: In
src/gameboy_worlds/emulation/harry_potter/test_metrics.py, define a subgoal.- For a single region check (using REGIONS), subclass
SingleRegionMatchSubGoal. - For a single region check that uses MULTI_TARGET_REGIONS, subclass
RegionMatchSubGoaland provide both_NAMED_REGIONand_TARGET_NAME. - For checking if any region in a list matches, subclass
AnyRegionMatchSubGoal(this relies onMULTI_TARGET_REGIONSand takes lists for_NAMED_REGIONSand_TARGET_NAMES).
Example using
AnyRegionMatchSubGoal(since we used MULTI_TARGET_REGIONS in Step 2):from gameboy_worlds.emulation.tracker import AnyRegionMatchSubGoal, make_subgoal_metric_class class ReachGarageSubGoal(AnyRegionMatchSubGoal): NAME = "reach_garage" _NAMED_REGIONS = ["car_area"] _TARGET_NAMES = ["next_to_car"] # Bundle your subgoals together: NavigateToCarSubGoalMetric = make_subgoal_metric_class([ReachGarageSubGoal])
- For a single region check (using REGIONS), subclass
Step 4: Define the Tracker
Trackers bundle the termination metrics and subgoal metrics so the environment can track progress.
- 🤖 (LLM) Create the Tracker Class:
Open
src/gameboy_worlds/emulation/harry_potter/trackers.py. Import your metrics at the top, then append:from .test_metrics import NavigateToCarTerminateMetric, NavigateToCarSubGoalMetric class NavigateToCarTestTracker(HarryPotterTestTracker): TERMINATION_TRUNCATION_METRIC = NavigateToCarTerminateMetric SUBGOAL_METRIC = NavigateToCarSubGoalMetric
Step 5: Register the Tracker
The system needs to know your tracker exists.
- 🤖 (LLM) Update the Registry:
Open
src/gameboy_worlds/emulation/harry_potter/registry.py(not the global registry). First, add the import at the top of the file with the others:Then, locate thefrom .trackers import NavigateToCarTestTracker
AVAILABLE_STATE_TRACKERSdictionary and add your tracker inside the specific game's dictionary block:"navigate_to_car_test": NavigateToCarTestTracker,
Step 6: Add Task to the Benchmark Database
Your task must be documented in the CSV so the evaluation framework can iterate over it.
- 🤖 (LLM) Update the CSV:
Open
benchmark/tests/harry_potter.csv. Add a new line for your task matching the exact schema:(Schema: game, task_category, task_description, init_state, state_tracker_class, shifted_training_games, can_train_from_init_state)harry_potter_chamber_of_secrets,navigation,navigate to the white car,burrow_start,navigate_to_car_test,harry_potter_philosophers_stone,True
Step 7: Test Your Implementation (Crucial)
Before committing and pushing data, verify that your new metrics actually work and detect the goal when you reach it.
- 🧑 (Human) Run the Dev Play Script with your Tracker:
(Pitfall: Remember to set
gameboy_dev_play_stopback tofalseagain!)python dev/dev_play.py --game harry_potter_chamber_of_secrets --init_state burrow_start --state_tracker_class navigate_to_car_test
- 🧑 (Human) Play to the Goal:
Play the game normally until you reach the success condition (e.g., the car).
- If your setup is correct, the game window will automatically close/terminate the exact moment the goal screen is reached.
- If you walk past the goal and the game doesn't stop, your bounding box coordinates or
.npytarget match failed and need to be recaptured.
Step 8: Push Data to Hugging Face & Pull Request
Because you generated local binaries (the .state file and the .npy screen capture array), you must push them to a remote database so others don't crash when running your task.
(Newcomer Pitfall: You likely do not have write access to the central DJ-Research Hugging Face namespace or the main GameBoyWorlds github repository. You will need to use your own fork!)
-
🤖 (LLM / Human) Run the Setup Script: If you are an external contributor, you must first open
src/gameboy_worlds/setup_data.pyand change therepo_namespacevariable (around line 26) from"DJ-Research"to your own Hugging Face username! Then, authenticate your Hugging Face CLI (huggingface-cli login) and execute the data push pipeline:python -m gameboy_worlds.setup_data push --game harry_potter_chamber_of_secrets
-
🤖 (LLM / Human) Standard Git Workflow (Fork): Commit your python file changes. Since you cannot push directly to
origin, make sure you have forked the repository on GitHub, added your fork as a remote, and push to your fork instead!git checkout -b add-hp-car-task git add . git commit -m "Add navigate to white car task for HP CoS" # Push to YOUR fork, not origin (e.g., git push <your-fork-remote> add-hp-car-task) git push my-fork add-hp-car-task
Go to GitHub and open a Pull Request from your fork to the main repository. Be sure to link your Hugging Face dataset in the PR description so the maintainers can pull your binary files!
demos/environment.py: You can specify a game, environment_variant or controller_variant to test parsing, HighLevelActions etc.
Everything in dev:
dev_play.py: vital for being able to play the game, pause the game and capture the screen or enter a breakpoint.list_states.py: prints out all of the states you've saved so far and writes them totmp_state_list.csv