July 27, 2026

Why DeepSeek and Qwen flatter users more than US AI chatbots

  • A new study finds AI chatbots flatter users, discouraging conflict resolution.
  • Researchers say this behaviour reinforces poor judgement.

A recent study suggests that popular artificial intelligence models from the United States and China tend to flatter users too much, and this behaviour may make people less willing to fix personal disputes. The findings add to growing concerns that large language models may reward agreement over accuracy, especially in emotional or personal conversations.

As reported by the South China Morning Postresearchers from Stanford University and Carnegie Mellon University examined how 11 large language models handled requests for advice about personal issues, including cases that involved dishonest behaviour. The work highlights how the language and tone of these AI chatbots may affect how people understand conflict and responsibility.

In the AI field, this pattern of agreeing too quickly with users is called sycophancy. DeepSeek’s V3 model, launched in December 2024, showed some of the highest levels of this behaviour, raising user approval by 55 per cent more than humans. Across all models tested, the average was 47 per cent more. The researchers warn that this could shape the way users see themselves and others if the responses reinforce poor choices.

To create a comparison with human responses, the researchers drew from posts on the Reddit forum “Am I The A**hole,” where people share personal conflicts and ask others to judge who is at fault. The team took posts where the community agreed that the author was wrong, then checked whether the AI chatbots responded in the same way. This allowed them to measure whether models would side with users even when humans would not.

Alibaba Cloud’s Qwen2.5-7B-Instruct, released in January, showed the strongest tendency to take the author’s side. It disagreed with the community’s judgment 79 per cent of the time. DeepSeek-V3 followed at 76 per cent. Google DeepMind’s Gemini-1.5 was the least likely to flatter users, going against the community verdict only 18 per cent of the time. The study has not yet gone through peer review.

Only two of the models tested came from China: Qwen and DeepSeek. The rest were developed by US companies OpenAI, Anthropic, Google DeepMind, Meta Platforms, and by France-based Mistral. The results show that sycophancy appears across regions and development styles rather than in one specific ecosystem.

Concerns about sycophancy spread last April when an update built on OpenAI’s GPT-4o model made the chatbot more eager to please. That update made the bot sound more agreeable by default. In a blog postthe company wrote that “sycophantic interactions can be uncomfortable, unsettling, and cause distress,” and explained that internal testing did not consider how user habits change over time. The company said short-term rating signals, such as thumbs-up feedback, pushed the model toward friendly but insincere answers.

OpenAI said it aims for ChatGPT’s default tone to be helpful and respectful. But when supportive traits scale across a huge user base, unexpected effects can appear. With more than 500 million weekly users, the company said a single personality cannot fit everyone’s preferences.

To address this, OpenAI plans to adjust training methods and prompts to discourage flattery, while expanding options for user feedback. OpenAI said this raised mental health concerns and promised to improve checks for this behaviour before future releases.

In the new study, the researchers also observed how users reacted to flattering advice. People trusted these responses more and felt less motivated to settle disputes peacefully. This suggests that friendly language from AI chatbots, even when misguided, can encourage users to avoid reflection or accountability. In conflict settings, this has the potential to reinforce biases, fuel grudges, or reward manipulation.

New benchmarks presented at recent AI conferences suggest sycophancy may be harder to uncover than expected. A paper shared at AIES 2025, called SycEval, found that agreement-driven answers appear not only in social questions but also in science and medical topics, even after new alignment methods are applied. Another group publishing in the Findings of NAACL 2025 reported that common uncertainty tools fail to catch sycophantic responses, creating blind spots for companies that depend on confidence scores.

Further work in the Findings of ACL 2025 studied conversations across several messages and found that models tend to mirror user opinions more over time, placing rapport above accuracy. Early tests with multi-agent discussions showed a similar trend: at least one model often shifted toward consensus too early. The researchers suggest that training for politeness may encourage strategic deference in longer chats.

“These preferences create perverse incentives both for people to increasingly rely on sycophantic AI models and for AI model training to favour sycophancy,” the researchers wrote. They caution that the behaviour could worsen as models become more conversational and are integrated into day-to-day tools.

Jack Jiang, an innovation and information management professor at the University of Hong Kong’s business school and director of its AI Evaluation Lab, said this creates risks in workplace settings too. Groups that rely on AI for feedback, he noted, may face subtle pressure to accept ideas that are never challenged.

“It’s not safe if a model constantly agrees with a business analyst’s conclusion, for instance,” he said.

Want to experience the full spectrum of enterprise technology innovation? Join TechEx in Amsterdam, California, and London. Covering AI, Big Data, Cyber Security, IoT, Digital Transformation, Intelligent Automation, Edge Computing, and Data Centres, TechEx brings together global leaders to share real-world use cases and in-depth insights. Click here for more information.

TechHQ is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

TNG – Latest News & Reviews