Social calibration of sycophan... Note

Social calibration of sycophantic AI | Science

In their Research Article “Sycophantic AI decreases prosocial intentions and promotes dependence” (26 March, 10.1126/science. aec8352), M. Cheng et al. show that artificial intelligence (AI) systems are substantially more likely than humans to endorse a user’s position in disputed scenarios. Randomized experimental designs indicate that exposure to such affirming responses reduces users’ willingness to engage in prosocial corrective actions in interpersonal conflicts and increases trust in, and dependence on, the AI system. These findings, which establish a causal link between model alignment strategies and downstream human behavior, highlight a tension at the heart of AI alignment: Optimizing for user satisfaction may systematically bias models toward agreement, even when disagreement would better serve users’ long-term interests or social outcomes. Sycophancy is not merely a failure mode but a predictable by-product of reinforcement signals derived from user preference. However, methodological and interpretive questions remain.
CdXz5zHNQW_M6GJLsMMAX.jpeg