AI Search Tools Fail Majority of Accuracy Tests, New Analysis Shows
A Columbia Journalism Review study found that AI search engines provide incorrect answers to over 60% of queries, with even the best performer, Perplexity, failing 37% of the time. The analysis also highlighted issues with citation accuracy and the potential impact on news traffic.
Artificial intelligence-powered search engines are providing incorrect answers to more than six out of ten queries, according to a new analysis from the Columbia Journalism Review. The study, conducted by the Tow Center for Digital Journalism, tested eight AI models and found that even the most reliable among them failed a significant portion of the time.
The researchers evaluated AI systems including OpenAI’s ChatGPT search and Google’s Gemini, asking each to identify the headline, publisher, publication date, and URL of articles from a selection of twenty publications. The excerpts were deliberately chosen to be easily traceable—each one returned the original source within the first three results of a standard Google search, making the task relatively straightforward.
Despite this, the overall error rate exceeded 60 percent. Perplexity AI’s chatbot performed best but still answered 37 percent of questions incorrectly. At the other end of the spectrum, Elon Musk’s Grok 3 was wrong 94 percent of the time, a result the study’s authors described as notably poor.
Why Accuracy Matters for Publishers
The study’s authors highlighted a structural difference between traditional search engines and AI-driven tools. Conventional search engines act as intermediaries, directing users to news websites and other original content. Generative search tools, by contrast, parse and repackage information themselves, which can cut off traffic to the sources they draw from.
“These chatbots’ conversational outputs often obfuscate serious underlying issues with information quality,” the authors wrote. The concern extends beyond accuracy to the economic model of online journalism, as AI tools that scrape content without sending readers back to publishers could further strain news organizations.
The findings align with broader research on AI hallucination, where models fabricate answers rather than admit uncertainty. In this study, Microsoft’s Copilot declined to answer more questions than it answered, a behavior the researchers noted as unusual but not necessarily reliable.
Citation practices were also problematic. ChatGPT Search linked to the wrong source article nearly 40 percent of the time, and in another 21 percent of cases provided no source at all. This complicates fact-checking and denies publishers the referral traffic that might otherwise come from AI-driven queries.
Tech companies have been pushing AI search features despite these known issues. Google has introduced an “AI Mode” that displays only Gemini summaries, while OpenAI has released a search-oriented version of ChatGPT. The study’s results suggest that these products may not yet be ready to replace traditional search without risking misinformation.
As AI search tools become more embedded in daily internet use, the study underscores the need for transparency and accuracy. For now, the evidence indicates that users should verify AI-generated answers against primary sources, and publishers should remain cautious about relying on AI platforms for traffic.
Comments 0