AI Reasoning Models Falter on NYT Connections Puzzle, Raising AGI Doubts
OpenAI's o1 reasoning model, along with other leading AI systems, failed to solve the New York Times Connections puzzle, exposing gaps in logical deduction. The test, conducted by a senior fellow, highlights that AI still struggles with novel, abstract tasks despite claims of nearing AGI.
In a striking counterpoint to the industry's artificial general intelligence ambitions, a senior fellow at the Walter Bradley Center for Natural and Artificial Intelligence has demonstrated that OpenAI's most advanced publicly available model, o1, cannot reliably solve a popular daily word puzzle. Gary Smith, who conducted the test, reported that the model, along with other leading large language models, failed to crack the New York Times' Connections game—a challenge that thousands of casual players complete every day.
The Connections puzzle presents players with 16 seemingly unrelated words and asks them to sort them into four groups of four based on hidden commonalities. The associations can range from straightforward categories like "book subtitles" to more obscure links such as "words that start with fire." This mix of obvious and esoteric connections makes the game a demanding test of semantic reasoning and lateral thinking.
Smith, whose analysis appears in Mind Matters, fed the day's puzzle into o1, as well as models from Google, Anthropic, and Microsoft (the latter powered by OpenAI's technology). To the surprise of those who subscribe to AI hype, every system failed to complete the puzzle correctly. o1, which has been marketed as a next-generation reasoning engine, produced several correct groupings but also offered combinations that Smith described as "bizarre."
In one instance, o1 grouped the words "boot," "umbrella," "blanket," and "pant" under the theme "clothing or accessories." While three of the four words fit the category, the inclusion of "blanket"—which is not typically worn as apparel—undermined the grouping. In a second attempt with the same set of words, the model confidently asserted that "breeze," "puff," "broad," and "picnic" were "types of movement or air." The first two terms align with the description, but the latter two do not, leaving Smith and other observers puzzled.
Smith's overall assessment was that o1 produced "many puzzling groupings" alongside a "few valid connections." This outcome underscores a well-known limitation of large language models: they excel at regurgitating information that is well-represented in their training data but often stumble when faced with novel, abstract tasks that require genuine reasoning.
Implications for AGI Claims
The test comes at a time when OpenAI's leadership has made bold statements about the proximity of artificial general intelligence. CEO Sam Altman has claimed that the company already possesses the building blocks for AGI, and at least one employee suggested at the end of last year that the milestone had been reached. However, the failure of o1 on a simple word puzzle suggests that such claims may be premature, at least when it comes to the kind of flexible, human-like reasoning that AGI implies.
The results also highlight a broader challenge for the AI industry: despite impressive performance on benchmarks and tasks that mirror training data, these systems still lack the robustness needed to handle everyday cognitive challenges. For now, the gap between AI's capabilities and the hype surrounding them remains significant, as demonstrated by a puzzle that many humans solve over their morning coffee.
Comments 0