Everything K-culture — comebacks to K-beauty, straight to your inboxGet it in your inbox

METAL MEDIA

K-Culture GlossaryWhere everyone starts

Reinforcement Learning

A trial-and-error training method that rewards good moves and penalizes bad ones. The method behind AlphaGo.

In plain words

Instead of handing over an answer key, you let the system try things, then reward it when it does well and penalize it when it doesn't. It works on the same principle as training a dog — sit on command, get a treat, repeat millions of times.

AlphaGo conquered the game of Go this way, and today it's become central to training language models. Models learn to reason better by solving math problems and getting rewarded for correct answers. "We boosted reasoning performance with reinforcement learning" has become a standard line in recent model announcements.

See also

Stories using this term

No story has used this term yet. New ones attach here automatically.

Browse every entry