Everything K-culture — comebacks to K-beauty, straight to your inboxGet it in your inbox

METAL MEDIA

K-Culture GlossaryRWhere everyone starts

RLHF

인간 피드백 강화학습

Training that polishes a model's manner of speaking using human-scored 'good answer' ratings — the technique that made ChatGPT polite.

In plain words

Reinforcement Learning from Human Feedback — reinforcement learning that uses human evaluations as the reward signal. When people pick which of several model answers is 'better,' that preference is turned into a score used to refine the model's way of speaking.

This was the decisive technology in turning GPT-3 into ChatGPT. Even though the underlying knowledge is similar, the polite, helpful, non-risky attitude in its answers was shaped here. It also carries a side-effect debate: fitting too closely to human preferences can increase sycophantic, flattering answers.

Try it yourself

  1. Look for the 👍/👎 buttons under a chatbot's answer.
  2. Try pressing 👍 on an answer you liked — that very action is raw material for RLHF. Countless users making this same choice collectively shapes the direction of what counts as a 'preferred answer.'
  3. If a chatbot seems to praise and agree with you unusually often, that personality was also shaped by people's 👍 clicks — which also helps explain why the 'AI sycophancy debate' exists.

See also

Stories using this term

No story has used this term yet. New ones attach here automatically.

Browse every entry