Files
lab_reinforcement_learning/questions.md
Chris Proctor e752bb848b Refactor lab
2026-06-25 21:10:13 -04:00

1.0 KiB
Raw Blame History

Questions

Checkpoint 1

  1. How do you decide where to move in BabySnake? Explain how to choose moves in enough detail that someone else could follow your instructions.

Checkpoint 2

  1. How many distinct states are there for BabySnake on a 4×4 grid? If we assume that all four arrow keys are valid actions in every state, how many rows would the full Q-table contain?

  2. The discount factor γ (gamma) can range from 0 to 1. What would be the effect of setting γ to 0? What about 1?

  3. The learning rate α (alpha) can also range from 0 to 1. What would be the effect of setting α to 0? What about 1?

  4. Calculate the new Q-value for ((2, 2, 3, 3), RIGHT). Explain your answer.

Checkpoint 3

  1. At what episode did the agent start reliably finding food?

  2. Add print(Q) to train_babysnake.py before the watch call and run it again. Can you read the policy? For a given state, does the highest Q-value point toward the food?

  3. How does the trained agent's behavior compare to the reasoning you wrote down in question 1?