पाठ 19 / 25

Reinforcement Learning: Q-Learning

Learn good actions by trial and error.

Explore, observe reward, update

In reinforcement learning (RL) the agent does not know the rules; it learns from experience. Q-learning keeps an estimate Q(s, a) of the value of taking action a in state s, and after each step moves it towards the observed reward plus the discounted best value of the next state. An exploration strategy (such as epsilon-greedy: mostly the best-known action, sometimes random) ensures it tries alternatives. Deep RL replaces the table with a neural network, as in game-playing systems like AlphaGo and in training methods for language models.

Q-learning in the same corridor, run

I ran this with plain Python 3 (standard library only), with fixed random seeds where randomness is used. Without being told the transition rules, a Q-learning agent playing 2,000 episodes learns to move right in states 1 to 3. In the middle state its learned values are 0.82 for left and 4.84 for right.

import random
rng = random.Random(0)
# same corridor, but the agent does not know the rules: it learns from trial and error
def step(s, a):
    move = 1 if a == 1 else -1
    if rng.random() < 0.2: move = -move
    nxt = min(4, max(0, s + move))
    reward = 10 if nxt == 4 else -10 if nxt == 0 else -1
    return nxt, reward, nxt in (0, 4)
Q = [[0.0, 0.0] for _ in range(5)]; alpha, gamma, eps = 0.1, 0.9, 0.2
for episode in range(2000):
    s, done = 2, False
    while not done:
        a = rng.randrange(2) if rng.random() < eps else max((0, 1), key=lambda x: Q[s][x])
        nxt, r, done = step(s, a)
        target = r if done else r + gamma * max(Q[nxt])
        Q[s][a] += alpha * (target - Q[s][a]); s = nxt
print("learned policy for states 1-3:", ["right" if Q[s][1] > Q[s][0] else "left" for s in (1, 2, 3)])
print("Q-values in state 2: left", round(Q[2][0], 2), "right", round(Q[2][1], 2))

Output:

learned policy for states 1-3: ['right', 'right', 'right']
Q-values in state 2: left 0.82 right 4.84

Design rewards carefully

Agents optimise exactly what the reward measures; a sloppy reward invites shortcuts nobody intended.

त्वरित जाँच: Why does a Q-learning agent sometimes take random actions?

  • To avoid using rewards
  • To slow learning down on purpose
  • Because Q-values are random
  • To explore actions that might be better than its current estimates
Answer

To explore actions that might be better than its current estimates — Exploration versus exploitation.