K-Culture GlossaryRWhere everyone starts
RLHF
인간 피드백 강화학습
Training that polishes a model's manner of speaking using human-scored 'good answer' ratings — the technique that made ChatGPT polite.
In plain words
Reinforcement Learning from Human Feedback — reinforcement learning that uses human evaluations as the reward signal. When people pick which of several model answers is 'better,' that preference is turned into a score used to refine the model's way of speaking.
This was the decisive technology in turning GPT-3 into ChatGPT. Even though the underlying knowledge is similar, the polite, helpful, non-risky attitude in its answers was shaped here. It also carries a side-effect debate: fitting too closely to human preferences can increase sycophantic, flattering answers.
Try it yourself
- Look for the 👍/👎 buttons under a chatbot's answer.
- Try pressing 👍 on an answer you liked — that very action is raw material for RLHF. Countless users making this same choice collectively shapes the direction of what counts as a 'preferred answer.'
- If a chatbot seems to praise and agree with you unusually often, that personality was also shaped by people's 👍 clicks — which also helps explain why the 'AI sycophancy debate' exists.
See also
Stories using this term
No story has used this term yet. New ones attach here automatically.