पाठ 19 / 25
Reinforcement Learning: Q-Learning
Learn good actions by trial and error.
Explore, observe reward, update
In reinforcement learning (RL) the agent does not know the rules; it learns from experience. Q-learning keeps an estimate Q(s, a) of the value of taking action a in state s, and after each step moves it towards the observed reward plus the discounted best value of the next state. An exploration strategy (such as epsilon-greedy: mostly the best-known action, sometimes random) ensures it tries alternatives. Deep RL replaces the table with a neural network, as in game-playing systems like AlphaGo and in training methods for language models.
Q-learning in the same corridor, run
I ran this with plain Python 3 (standard library only), with fixed random seeds where randomness is used. Without being told the transition rules, a Q-learning agent playing 2,000 episodes learns to move right in states 1 to 3. In the middle state its learned values are 0.82 for left and 4.84 for right.
import random
rng = random.Random(0)
# same corridor, but the agent does not know the rules: it learns from trial and error
def step(s, a):
move = 1 if a == 1 else -1
if rng.random() < 0.2: move = -move
nxt = min(4, max(0, s + move))
reward = 10 if nxt == 4 else -10 if nxt == 0 else -1
return nxt, reward, nxt in (0, 4)
Q = [[0.0, 0.0] for _ in range(5)]; alpha, gamma, eps = 0.1, 0.9, 0.2
for episode in range(2000):
s, done = 2, False
while not done:
a = rng.randrange(2) if rng.random() < eps else max((0, 1), key=lambda x: Q[s][x])
nxt, r, done = step(s, a)
target = r if done else r + gamma * max(Q[nxt])
Q[s][a] += alpha * (target - Q[s][a]); s = nxt
print("learned policy for states 1-3:", ["right" if Q[s][1] > Q[s][0] else "left" for s in (1, 2, 3)])
print("Q-values in state 2: left", round(Q[2][0], 2), "right", round(Q[2][1], 2))
Output:
learned policy for states 1-3: ['right', 'right', 'right'] Q-values in state 2: left 0.82 right 4.84
Design rewards carefully
Agents optimise exactly what the reward measures; a sloppy reward invites shortcuts nobody intended.
त्वरित जाँच: Why does a Q-learning agent sometimes take random actions?
- To avoid using rewards
- To slow learning down on purpose
- Because Q-values are random
- To explore actions that might be better than its current estimates
Answer
To explore actions that might be better than its current estimates — Exploration versus exploitation.