Reinforcement Learning

The AI paradigm of trial, error, and optimization. How agents learn by interacting with environments.

16 minBeginner

The Core Thesis: Reinforcement Learning (RL) is learning through interaction. An agent observes a situation, chooses an action, receives feedback as a reward or penalty, and gradually learns which actions lead to better long-term outcomes.

  1. What is Reinforcement Learning?

Reinforcement Learning is a machine learning paradigm in which an agent learns what to do by interacting with an environment. Unlike supervised learning, the agent is not simply handed a correct answer for every situation. It has to choose actions, observe their consequences, and learn from the feedback it receives.
The important idea is that learning happens over a sequence of decisions. An action that looks good immediately may be bad later, while an action with a small short-term reward may lead to a much better future outcome.
Think of a game-playing agent. At each moment it sees the current game state, chooses a move, receives feedback, and reaches a new state. After many interactions, it can learn a strategy that produces better results.
The Reinforcement Learning Loop
AGENT
chooses an action
→
ENVIRONMENT
changes after the action
→
NEW STATE + REWARD
feedback returns to agent
The loop repeats so the agent can improve its future decisions.
Quick Check 1
Which situation best represents Reinforcement Learning?

2. The Core Terminology

Most RL problems can be understood by mapping the situation to a small set of core terms. Once these terms are clear, many RL diagrams and algorithms become much easier to follow.
Agent
The learner or decision-maker. It observes the environment and chooses actions.
Environment
The world in which the agent operates. It responds to actions and produces new situations and feedback.
State (ss)
The information describing the agent's current situation, such as a robot's location or a game's board position.
Action (aa)
A choice the agent can make in a state, such as move left, move right, accelerate, or wait.
Reward (rr)
A numerical feedback signal indicating how desirable the outcome was from the environment's point of view.
Policy (π\pi)
The strategy that determines which action the agent chooses for a given state.
TermSimple meaningExample: Robot
StateWhat is happening nowRobot is near the charging station
ActionWhat the agent can doMove forward
RewardFeedback after an action+10 for reaching the station
PolicyHow the agent chooses actionsChoose the direction that appears most useful

3. How an RL Agent Learns

The basic learning process is a repeated cycle. The exact algorithm can differ, but the interaction pattern stays useful as a mental model.
01
Observe
Read the current state.
02
Choose
Select an action.
03
Act
Send the action to the environment.
04
Receive
Get reward and next state.
05
Improve
Update the decision strategy.
Over many steps, the agent accumulates experience. The learning algorithm uses that experience to make better future decisions rather than simply repeating the exact same behavior.

4. Reward, Return, and Long-Term Decisions

A reward is usually associated with a particular step, but RL often cares about the outcome of a sequence of steps. The total accumulated reward across a sequence is commonly called the return.
This distinction matters because the best action is not always the one with the largest immediate reward. An agent may accept a small penalty now if that choice leads to a much larger reward later.
Example: Imagine a navigation agent. Taking a shortcut might give a small immediate reward for reducing distance, but if the shortcut leads to a dead end, the long-term result can be poor. A useful RL strategy considers future consequences.
ConceptMeaningWhy it matters
RewardFeedback for an outcomeTells the agent how an action affected its objective
ReturnAccumulated reward across timeConnects current decisions to future outcomes
Discount factorControls how strongly future rewards are valuedHelps balance immediate and future outcomes

5. Exploration vs. Exploitation

One of the central challenges in RL is deciding whether to use an action that is already known to work or try something new that might work even better.
Exploitation means using knowledge the agent already has to choose a promising action. Exploration means trying less-known actions to gather new information.
Too much exploitation
The agent may keep using a safe strategy and never discover a better one.
Too much exploration
The agent may keep trying random actions instead of taking advantage of what it has already learned.
Practical RL methods use different strategies to balance these two behaviors. The exact strategy depends on the environment and algorithm.
Quick Check 2
An agent always chooses the action it currently believes is best and never tries unfamiliar actions. What problem can this cause?

6. Value and Q-Value: A Beginner Overview

RL systems often need a way to estimate how useful a situation or action is. This is where the idea of value becomes important.
A state value asks a question like: “If I am in this state, how much future reward should I expect if I continue following my strategy?”
A Q-value goes one step further: “If I am in this state and take this particular action, how useful is that choice in terms of future reward?”
Value
Estimates how desirable a state is when considering future outcomes.
Q-Value
Estimates the usefulness of taking a particular action in a particular state.
You do not need the full mathematics at this stage. The key intuition is that RL can learn estimates of future usefulness and use those estimates when choosing actions.

7. Episodes and Experience

An episode is one complete run of an RL task, from its starting condition until a terminal condition is reached. In a game, one episode might be one complete match. In navigation, it might be one attempt to reach a destination.
A sequence of states, actions, and rewards generated during an episode forms the agent's experience. Repeating episodes gives the learning algorithm more opportunities to discover useful behavior.
Simple example: A maze-solving agent starts at the entrance, moves through the maze, receives small penalties for wasted moves, and receives a positive reward when it reaches the exit. Reaching the exit ends the episode.
Quick Check 3
In a maze-solving task, what would most naturally represent the end of an episode?

8. Major RL Approaches — Overview Only

There are many RL algorithms, but a beginner does not need to learn every algorithm inside this introductory module. The important thing is to recognize the major families and understand what problem they are trying to solve.
Value-Based Methods
Learn estimates of how valuable states or state-action pairs are. Q-learning is a classic example.
Policy-Based Methods
Learn a policy directly so the agent can choose actions from states.
Model-Free Methods
Learn useful behavior from interaction without first building an explicit model of the environment.
Model-Based Methods
Use or learn a model of how the environment changes, then use that model to help plan actions.
NameBeginner ideaExamples
Q-learningLearn useful state-action valuesClassic tabular RL
Deep Q-NetworksUse a neural network to estimate Q-valuesLarge or complex state spaces
Policy GradientImprove the policy directlyContinuous or complex actions
Actor-CriticCombine policy and value ideasModern RL systems

9. Designing Rewards

The reward function tells an RL system what outcomes are desirable. Because the agent optimizes what the reward system measures, reward design is an important part of an RL problem.
A poorly designed reward can encourage behavior that technically maximizes the score while missing the real objective. This is sometimes described as a reward hacking problem.
Useful reward
Encourages behavior that is closely aligned with the actual goal.
Misleading reward
Can make the agent optimize an easy-to-measure signal while ignoring the intended outcome.

10. Real-World Applications

Reinforcement Learning is useful when decisions happen over time and an action can change what happens next. The following examples show where the RL setup can naturally appear.
ApplicationWhat RL doesSimple example
GamesLearns strategies by repeatedly taking actions and receiving outcomes as feedback.Learn a winning strategy in a board or video game.
RoboticsLearns movement or control policies through simulation or physical interaction.Learn how a robot should move its arm to reach a target.
Resource OptimizationLearns sequential decisions when current choices affect future states and rewards.Improve scheduling, routing, or resource allocation.
Recommendation and RankingCan model user interactions as sequential feedback when future choices depend on earlier recommendations.Choose which content to recommend over a sequence of interactions.
Human FeedbackHuman preferences can provide feedback signals used in RL-based optimization pipelines.Use preference feedback to improve an AI system's behavior.

11. RL vs. Supervised and Unsupervised Learning

Reinforcement Learning is easier to understand when contrasted with the other two major ML learning setups.
Learning typeMain signalTypical questionExample
SupervisedCorrect labelsWhat output matches this input?Predict whether an image is a cat
UnsupervisedStructure in dataWhat patterns or groups exist?Group similar customers
ReinforcementRewards from interactionWhat action should I take over time?Learn a strategy for a game
Quick Check 4
What is the best beginner-level description of a policy in Reinforcement Learning?

12. Strengths and Limitations

Why RL is useful
  • • Handles sequential decision-making.
  • • Can discover strategies that were not explicitly programmed.
  • • Naturally incorporates long-term consequences.
  • • Can learn through simulation when real-world experimentation is expensive.
Why RL can be difficult
  • • Training can require many interactions.
  • • Rewards can be sparse or difficult to design.
  • • Exploration can be expensive or unsafe in real environments.
  • • Training can be sensitive to the environment and algorithm choices.

13. Key Terms to Remember

TermMeaning
AgentThe learner that makes decisions.
EnvironmentThe world the agent interacts with.
StateThe current situation observed by the agent.
ActionA decision available to the agent.
RewardNumerical feedback from the environment.
PolicyThe strategy used to select actions.
ReturnAccumulated reward considered across time.
EpisodeOne complete run of an RL task.
ExplorationTrying actions to gain new information.
ExploitationUsing current knowledge to choose promising actions.

14. Key Points

  • ✓ Reinforcement Learning learns through interaction between an agent and an environment.
  • ✓ The core loop is state → action → reward → next state.
  • ✓ The agent aims to improve long-term return, not simply immediate reward.
  • ✓ A policy describes how the agent chooses actions from states.
  • ✓ Exploration tries less-known actions, while exploitation uses current knowledge.
  • ✓ Reward design strongly affects what behavior the agent learns.
  • ✓ Value and Q-value concepts help estimate future usefulness.
  • ✓ Algorithms such as Q-learning and policy-gradient methods are deeper topics that can be learned separately.

15. Common Mistakes

  • ✕ Confusing reward with the goal itself.
    The reward is the signal used to represent the objective. A badly designed reward can produce unwanted behavior.
  • ✕ Thinking RL only means games.
    Games are useful examples, but RL is fundamentally about sequential decision-making under feedback.
  • ✕ Ignoring future rewards.
    RL often evaluates actions by their expected long-term consequences, not only their immediate reward.
  • ✕ Assuming every action should be random.
    Exploration is useful, but the agent must also exploit what it has already learned.
  • ✕ Assuming an RL algorithm automatically knows the real objective.
    The agent optimizes the reward signal it receives, so the reward must represent the intended objective carefully.

16. The Big Picture

Reinforcement Learning
State → Action → Reward → New State → Improved Policy
The Agent's Question
"What should I do now so that my long-term outcome becomes better?"
The Learning Signal
Actions → Consequences → Rewards → Experience → Better Decisions

Reinforcement Learning is fundamentally a decision-making setup. The important beginner concepts are the relationships between the agent, environment, state, action, reward, return, and policy—not memorizing a long list of algorithms.

Remember: supervised learning learns from known answers, unsupervised learning searches for structure without supplied targets, while reinforcement learning learns better actions through interaction and feedback over time.