Files
lab_reinforcement_learning/questions.md
Chris Proctor 482f4f6cfa Updates
2026-06-26 20:59:25 -04:00

3.1 KiB
Raw Blame History

Questions

Checkpoint 1

  1. How do you decide where to move in BabySnake? Explain how to choose moves in enough detail that someone else could follow your instructions.

Checkpoint 2

  1. How many distinct states are there for BabySnake on a 4×4 grid? If we assume that all four arrow keys are valid actions in every state, how many rows would the full Q-table contain?

  2. The discount factor γ (gamma) can range from 0 to 1. What would be the effect of setting γ to 0? What about 1?

  3. The learning rate α (alpha) can also range from 0 to 1. What would be the effect of setting α to 0? What about 1?

  4. Calculate the new Q-value for ((2, 2, 3, 3), RIGHT). Explain your answer.

Checkpoint 3

  1. At what episode did the agent start reliably finding food?

  2. Add print(Q) to train_babysnake.py before the watch call and run it again. Can you read the policy? For a given state, does the highest Q-value point toward the food?

  3. How does the trained agent's behavior compare to the reasoning you wrote down in question 1?

Checkpoint 4

  1. In Attempt 1, the agent sees the full 32×16 board as 3,072 numbers—the apple's location is already in there somewhere. In Attempt 2, we supplemented the board with just two extra numbers: the direction to the apple. Performance tripled. Why did two extra numbers make such a large difference when the board already contained the apple's location?

  2. Attempt 3 added the full board back and switched to a CNN—a more powerful architecture—yet performance was worse than Attempt 1. Why didn't more information and a more powerful model help?

  3. The only difference between Attempt 3 and Attempt 4 is that Attempt 4 shows the agent a 17×17 window centered on its own head, rather than the full board. Why did this single change make such a large difference?

  4. When the snake's body gets very long, it becomes important to plan your route so you don't get trapped inside your own body. None of our training attempts was very successful at learning this behavior. Which of the approaches do you think would be most promising for learning it? Why?

  5. The reward function gives the snake +1 for each step it moves toward the apple and 1 for each step away. Can you think of a way this reward signal might accidentally encourage bad behavior—especially as the snake grows longer?

Checkpoint 5

Answer these questions after completing both training experiments in "Training Frogger."

  1. Hypothesis (Attempt 1): Before training, predict what will happen. Will the agent learn to reach the top of the board? What challenge do you think it will face?

  2. Evidence (Attempt 1): Copy the first three and last three lines of runs/frogger/training.log. Did training go as expected?

  3. Analysis (Attempt 1): What did the agent learn to do? Where did it struggle?

  4. Experiment (Attempt 2): What one thing did you change? Write your prediction, show the evidence (first and last few log lines), and describe what happened.

  5. Which attempt produced the best agent? What would you try next if you had more time?