MDP (Markov Decision Process)
An MDP formalizes the reinforcement-learning setup — an agent observing a state, choosing an action, and receiving a reward, with no dataset of correct actions to imitate.
An MDP (Markov Decision Process) formalizes the structure shared by every reinforcement learning problem: an agent observes a state of an environment and chooses an action; the environment transitions to a new state and returns a reward; the agent's goal is to learn a policy that maximizes total reward over time, not just the immediate reward from one action.
How it works
What makes this genuinely different from supervised learning is that there's no dataset of "correct" actions to imitate — only a scalar reward signal, often delayed. An action that looks bad immediately (sacrificing a chess piece) can be the right choice if it leads to much better rewards later, and the agent has to learn this from trial and error. "Markov" refers to the assumption that the current state contains everything needed to decide the best action — the history of how the agent arrived there doesn't matter beyond what's captured in the state itself.
When it breaks
- The data isn't fixed, unlike supervised learning. Improving the policy changes which states the agent visits, which changes the data it learns from next — a feedback loop that has no equivalent in a static training set and is a major source of RL's training instability.
- Credit assignment is genuinely hard. A reward arriving at the end of a long sequence of actions says only that the sequence as a whole worked out, not which specific action deserves the credit.
- The Markov assumption can be false. Many real environments are only partially observable — the visible state doesn't contain everything relevant, which is a distinct, harder problem (a POMDP) than the standard setup.
See also: RewardRewardA reward is the scalar feedback signal a reinforcement-learning agent receives after taking an action — the only learning signal it gets, and often a delayed one., PolicyPolicyA policy is the strategy a reinforcement-learning agent learns for choosing actions — the thing being trained, analogous to a model in supervised learning., Q-LearningQ-LearningQ-learning learns the expected future reward of taking a specific action in a specific state, updating its estimate incrementally from every action the agent actually takes.
Learn more: Reinforcement Learning
Mentioned in
Lessons where this comes up in context.
Hallucination
A hallucination is a confidently stated, fluent LLM output that is factually wrong, a direct consequence of models being trained to produce plausible text rather than verified facts.
Reward
A reward is the scalar feedback signal a reinforcement-learning agent receives after taking an action — the only learning signal it gets, and often a delayed one.