Nearly 50% of AI responses were rated as problematic
The study evaluated 5 popular AI chatbots: Gemini, DeepSeek, Meta AI, ChatGPT, and Grok. The researchers tested these models with 50 prompts across five categories: cancer, vaccines, stem cells, nutrition, and athletic performance.
The results were striking. Nearly half of all responses were rated as problematic. Specifically, 30% were classified as somewhat problematic, and 19,6% as highly problematic. Performance varied by category, with the chatbots faring worst in stem cells, athletic performance, and nutrition. This is particularly concerning for cyclists who turn to AI for guidance in these areas.
Why AI fails in nutrition
The study revealed that AI struggles with complex, ambiguous or scientifically contested questions, which is what happens often in nutrition. Topics like supplements, Relative Energy Deficiency in Sport (RED-S), and gastrointestinal resilience are prone to misinformation, and the chatbots often failed to provide reliable answers.
Another key finding was the difference in performance between closed-ended and open-ended questions. Closed-ended questions produced fewer problematic responses, while open-ended questions that allow for more speculation and nuance led to far more errors. The more freedom a model has to generate an answer, the greater the risk of inaccuracies, false balance or unsupported advice.
The referencing problem
One of the most revealing aspects of the study was the citation audit. Chatbots often returned incomplete, inaccurate or fabricated citations. Across all tests, no chatbot produced a fully complete and accurate reference list for any prompt. The median completeness score for references was just 40%. This is a critical issue because references are what make responses credible.
Real-world implications
The study’s findings have significant real-world implications. In fields like nutrition, where commercial claims and pseudoscience already run rampant, AI’s tendency to produce confident but inaccurate answers could amplify misinformation. The risk is not just theoretical; anyone relying on these tools for health, nutrition or performance advice can be misled into making a harmful change.
The study also highlighted another issue: readability. While the responses were often polished and confident, they were also consistently rated as “Difficult” – equivalent to college-level reading. This combination of complexity and unreliability creates a risky scenario, where users may trust answers that are neither accessible nor accurate.
AI amplifies misinformation if unchecked
The study’s most important conclusion is simple: without public education and oversight, AI risks amplifying misinformation rather than reducing it. Tools matter, but so do the data they’re trained on, the way they’re deployed, and the knowledge of the people using them. In health, medicine, sport, and sports nutrition, fluent language should never be confused with understanding. A polished answer can still be wrong, and a confident answer can still be misleading.
What can you do about it?
Treat AI chatbots as a starting point, not a final authority. Cross-check their answers with trusted sources, like peer-reviewed studies, health organisations or reputable magazines. Be sceptical of confident-sounding advice. If a chatbot provides references, verify them. Don’t assume they’re accurate just because they look official. And when in doubt, consult a professional. AI can help organise information, but it’s no substitute for expert judgment.
