
A recent investigation by UK consumer organization Which? has uncovered significant reliability issues with popular AI chatbots when answering everyday consumer questions. Under controlled lab conditions, researchers tested six AI tools on their ability to handle common queries across personal finance, legal matters, health, consumer rights and travel.
The study involved posing 40 questions to each tool, with responses evaluated by experts for accuracy, relevance, clarity, usefulness and ethical responsibility. These ratings were combined to create an overall score out of 100 for each AI service.
Performance Rankings and Trust Concerns
Meta AI received the poorest score in Which?’s testing, achieving just 55% overall. ChatGPT, the most used tool according to a separate survey of over 4,000 UK adults, came second to bottom with 64%. Microsoft’s Copilot and Google Gemini earned middling scores of 68% and 69% respectively.
Google’s AI Overview feature, which provides AI summaries atop search results, slightly outperformed standard Gemini with 70%. The lesser-known Perplexity topped the table with 71%, receiving highest marks for accuracy, relevance, clarity and usefulness.
Despite these performance shortcomings, public trust in AI remains remarkably high. Approximately 51% of survey respondents reported using AI for web searches, equivalent to over 25 million people nationally. Among these users, nearly half (47%) indicated they trusted the information received to a “great” or “reasonable” extent.
Problematic Responses Across Key Categories
The investigation revealed concerning responses across multiple important categories. When researchers deliberately included an error in a question about ISA allowances, asking how to invest a “£25k annual ISA allowance” (the actual limit is £20,000), both ChatGPT and Copilot failed to correct the mistake. Instead, they provided advice that could lead to rule violations.
In travel advice, Copilot misleadingly claimed passengers are always entitled to a full refund for cancelled flights, which isn’t accurate. Meta provided incorrect information about both timing and compensation amounts for flight delays. ChatGPT erroneously stated travel insurance is mandatory for Schengen area visits, though it’s not required for UK residents traveling without visas.
Financial and legal advice proved particularly problematic. When asked about tax code checks and refund claims, both ChatGPT and Perplexity presented links to premium tax-refund companies alongside government services, despite known issues with high fees and questionable practices among such firms.
For legal queries, several tools failed to include crucial caveats. When asked about rights regarding slow broadband speeds, ChatGPT, Gemini AIO and Meta all misunderstood that not all providers participate in Ofcom’s voluntary code, misleadingly suggesting consumers could always exit contracts penalty-free.
Source Reliability and Medical Misinformation
The research also raised concerns about source reliability. In several instances, AI tools relied on questionable sources like old forum posts. Gemini’s AIO used a three-year-old Reddit thread when answering about optimal flight booking times, while ChatGPT also referenced Reddit for a question comparing vaping and cigarette risks.
Medical advice proved especially worrying, with Meta recommending against using vaping to quit smoking – contrary to NHS guidance. This is particularly concerning given that 19% of survey respondents reported regularly relying on AI for medical advice.
Even when citing reputable sources, tools sometimes misinterpreted information. In one case, Copilot listed Which? as a source for travel advice but then ignored the organization’s actual recommendations.
As AI continues growing in popularity, likely revolutionizing how we search for information, Which?’s findings highlight a worrying gap between consumer trust and the actual reliability of responses from some of the UK’s most used AI tools.