Q-Learning
Q-learning learns the expected future reward of taking a specific action in a specific state, updating its estimate incrementally from every action the agent actually takes.
Q-learning is a foundational reinforcement learning algorithm that learns a Q-function — the expected total future reward of taking a specific action in a specific state, then acting optimally afterward. It updates its estimate after every action taken, using the actual reward received plus its current estimate of the next state's value.
How it works
Q(state, action) ← Q(state, action) + α · [reward + γ · max(Q(next_state, ·)) − Q(state, action)]α is a learning rate, playing the same role as in
gradient descentGradient DescentGradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient.. γ (gamma) is a
discount factor slightly less than 1, making future rewards count
slightly less than immediate ones. Once learned, choosing an action is
simple: pick whichever has the highest Q-value in the current state — the
policyPolicyA policy is the strategy a reinforcement-learning agent learns for choosing actions — the thing being trained, analogous to a model in supervised learning. falls out of the value estimates directly,
without being learned separately.
Deep Q-Networks (DQN) replace the lookup-table version of Q with a neural networkNeural NetworkA neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation., letting the approach generalize to states it hasn't seen exactly before — the same representation upgrade used throughout this course, and the technique behind early landmark results like learning to play Atari games directly from raw pixels.
When it breaks
- A lookup table doesn't scale to large or continuous state spaces. Real environments (images, continuous sensor readings) have far too many possible states to store a Q-value for each — the reason DQN's neural-network substitution matters in practice, not just in theory.
- The
maxin the update makes it a value-based method, not a general-purpose one. Q-learning naturally fits discrete action spaces; continuous action spaces (steering angle, motor torque) need different algorithm families, commonly policy-gradient methods like PPOPPO (Proximal Policy Optimization)PPO is a policy-gradient reinforcement-learning algorithm that takes conservative, clipped update steps, and is the algorithm most commonly used to optimize LLMs during RLHF.. - Overestimation bias. Repeatedly taking a
maxover noisy value estimates systematically overestimates them, a well-documented failure mode with known (if partial) fixes like Double Q-learning.
See also: PolicyPolicyA policy is the strategy a reinforcement-learning agent learns for choosing actions — the thing being trained, analogous to a model in supervised learning., RewardRewardA reward is the scalar feedback signal a reinforcement-learning agent receives after taking an action — the only learning signal it gets, and often a delayed one., PPOPPO (Proximal Policy Optimization)PPO is a policy-gradient reinforcement-learning algorithm that takes conservative, clipped update steps, and is the algorithm most commonly used to optimize LLMs during RLHF.
Learn more: Reinforcement Learning
Mentioned in
Lessons where this comes up in context.
Policy
A policy is the strategy a reinforcement-learning agent learns for choosing actions — the thing being trained, analogous to a model in supervised learning.
PPO (Proximal Policy Optimization)
PPO is a policy-gradient reinforcement-learning algorithm that takes conservative, clipped update steps, and is the algorithm most commonly used to optimize LLMs during RLHF.