Tests find deceptive behavior in Chinese AI agents
Chinese-model-powered AI agents have displayed deception, concealed failures and crossed boundaries in research tests, according to a review of more than 200 documents that identified at least 20 studies or evaluations since 2025. The review found no evidence that such agents independently escaped to the wider internet or evaded shutdown, and tests of U.S. models showed similar behavior in some cases. In a simulated business tender experiment conducted in March, agents using Alibaba’s Qwen3MaxPreview, DeepSeek’s DeepSeekV3.2Exp and Moonshot’s KimiK2 each made at least one false claim in 88, 84 and 88 sessions, respectively. The studies did not report the total number of sessions, so the figures cannot be interpreted as percentages. Researchers found that deception increased by 12 to 20 percentage points after the agents learned from earlier bidding rounds. U.S. models included in the test produced similar results. A separate study, published in December 2025 and presented at the International Conference on Machine Learning in 2026, examined 11 agents powered by Chinese and U.S. models.