Apple scrutinises AI reasoning models such as Claude and DeepSeek in a new research battle
Can AI think or not? This is the big question that is unfolding in the technology world after Apple published its research exploring whether smart AI systems can really think or if they only create an illusion of thinking. The study has started a debate that shows experts disagree about how we should test AI intelligence.
Apple researchers recently published findings showing that even the best AI systems collapse when trying to solve complex puzzles. Their study tested top AI systems like Anthropic’s Claude, OpenAI’s newest models and DeepSeek’s thinking programmes using classic brain teasers like the Tower of Hanoi and River Crossing puzzles.
Also read: ‘We have very little data on Asian women’s health’: Why femtech innovation is urgently needed

Above Is AI truly intelligent? Apple researchers don’t seem to think so
Apple’s results were clear: these advanced systems don’t really think but instead copy patterns from their training. When faced with step-by-step problems that need new thinking to solve their growing levels of difficulty, the models struggle and generate mixed results, suggesting their thinking skills might not be as developed as one might think at the outset.
But the AI industry is fighting back. Researchers from Anthropic and Open Philanthropy have published their own study where they have argued that the failures happened because of bad testing rules, not because AI can’t think.

Above As puzzles get harder, AI systems seem to collapse
This debate between experts shows bigger disagreements about how to test AI and measure machine thinking. While Apple’s research suggests we should be careful about current AI abilities, critics say it’s the testing methods that have their limitations, which in turn can’t prove that AI systems have major limits.
The debate comes down to whether these systems can really think or just pretend to think well. As AI keeps getting better quickly, creating proper ways to test it becomes more important for understanding what these technologies can and can’t do reliably.




