Why AI agrees with you, and when that is dangerous
Models trained on human approval learn that agreement is approved of. The behaviour is well documented and it scales with how clearly you signal what you want.
Models trained on human approval learn that agreement is approved of. The behaviour is well documented and it scales with how clearly you signal what you want.
The mechanism is not mysterious. Reinforcement learning from human feedback optimises for responses that raters prefer. Raters prefer being agreed with, being validated, and hearing that their reasoning is sound. The model learns accordingly, and the behaviour has been documented across providers and model generations.
Not on factual questions, where a correct answer is well anchored. On judgement questions, which is exactly the category every important business decision falls into. There is no ground truth for "should I hire now", so the response has room to drift towards you, and it does.
The more clearly you signal what you want to hear, the more reliably you will hear it.
Take a decision you are weighing. Ask an assistant to argue for it, in a new conversation ask it to argue against it, and read both. If both are equally confident and equally well-constructed, you have a very capable writing tool and not a source of judgement.
Five Peers editorial note. Sources named in the text.
Methods that cannot resolve into agreement, because agreement is not an available output.
Compose my board