Definition·
Learning

Reinforcement Learning

A learning method where an agent learns to act by trial and error, maximizing a reward.

Detailed explanation

The agent observes a state, picks an action and receives a reward. It updates its policy to maximize cumulative reward. Used in robotics, games, optimization, and LLM alignment via RLHF.

Examples

AlphaGo beating Go champions
Robots learning to walk
RLHF for aligning ChatGPT
Ad bidding optimization

Frequently asked questions

What is RLHF?

Reinforcement Learning from Human Feedback: the model is trained to prefer responses humans rated as better.

Related terms

Last updated: 7/15/2026

Talent AI

Turn theory into practice

Post a mission or join the community of top AI, Data and Machine Learning experts.