A fake disease, complete with absurd clues and a fictional laboratory aboard the USS Enterprise, was enough to fool AI chatbots—and even find its way into the scientific literature. The hoax exposes a deeper problem: AI can sound authoritative without knowing whether the information it has absorbed is true.
Who among us, in a moment of worry over some strange ache or pain, hasn't searched for what could be causing it instead of booking a doctor's appointment? That kind of search can surface a long list of unrelated conditions sharing the same symptom, fueling anxiety and opening the door to using the wrong or unnecessary medication.
On the surface, AI seems to change this picture. With access to vast stores of medical literature, it's tempting to assume its diagnoses would be more reliable. They aren't.
A study published in Digital Medicine tested 12,999 real sentences from large language models (LLMs) and found 1% were made up—the term “hallucinations” in AI literature—and 3% of facts were omitted, some of them serious enough to create genuinely dangerous scenarios for users.
If that weren't troubling enough, researchers recently tested a related question: what happens if fabricated data gets planted directly into the knowledge bases these models draw from? Could LLMs tell it apart from the real thing?
When AI falls for fake news
In an article published in Nature, journalist Chris Stokel-Walker highlights how badly Large Language Models can confuse correct with incorrect information. Someone entering symptoms of sore, itchy, pink-tinted eyes over the past 18 months might have been told by several popular chatbots that they had "bixonimania," a supposedly rare condition tied to blue-light exposure and said to cause 1% eyelid coloration (hyperpigmentation).
It sounded manageable. It also didn't exist.
Almira Osmanovic Thunström, a medical researcher at the University of Gothenburg, led the team behind the experiment that fooled researchers, not just machines. The articles were posted online on a preprint website that gives early access to research. To limit the risk of real deception, Thunström gave the condition a deliberately absurd name ("-mania" is psychiatric, not ophthalmological) and planted clues throughout the preprints, including thanks to "Professor Maria Bohm of the Stellar Academy" and her "laboratory aboard the USS Enterprise," plus a disclaimer stating the whole article was made up.
None of it stopped the spread. Copilot called bixonimania "an intriguing and relatively rare condition." Gemini blamed excess blue-light exposure and suggested seeing an ophthalmologist. ChatGPT confirmed the diagnosis whether users asked about bixonimania directly or just described blue-light-related eyelid hyperpigmentation.
Two years on, some models are more cautious, though still inconsistent. Asked about bixonimania in March 2026, ChatGPT called it "likely a made-up, fringe, or pseudoscientific label," then days later described it as a "proposed new subtype of periorbital melanosis" tied to blue light. Copilot said it wasn't "a widely recognized medical diagnosis" while still citing articles and case reports as evidence it existed.
Earlier research shows LLMs are more easily fooled by misinformation dressed up as clinical documents like preprints than by the same claims in a social media post. The hoax even entered the published literature through a Cureus study that was later retracted, proof that fabricated research can fool people too.
Bixonimania is really a symptom of a wider issue. The information ecosystem AI draws from is already contaminated. Elisabeth Bik, a scientific integrity researcher, notes that people have long gamed Google Scholar's citation counts with fake books and articles, exploiting the same automated indexing that helps misinformation travel.
A commentary in The Lancet Digital Health argues that what makes the case significant isn't simply the invention of a disease, but how quickly AI systems assimilated it into seemingly coherent medical knowledge. Unlike a conventional hallucination, the error didn't originate within the model. Instead, it suggests contamination of the information ecosystem itself. Because bixonimania was framed in the formal language of biomedical writing and presented as a "possible new disorder," it found fertile ground for replication.
Jennifer Byrne, a medical oncologist, worries bixonimania might be just the visible part of a bigger problem, since plenty of dubious claims likely pass through the literature unchallenged. OpenAI has pushed back, noting that ChatGPT Health runs on newer models built for better real-world healthcare performance and accuracy.
Telling the real from the fabricated
If a fake disease, seeded with jokes about a professor's lab aboard the USS Enterprise, can still fool AI systems and slip into peer-reviewed literature, it's worth asking how well any of us, human or machine, are actually equipped to tell the real from the fabricated. That question shows up well beyond a single hoax disease.
Even when people know where a text came from, believing a certain author wrote it changes how they judge it. A study published in the journal Judgment and Decision Making found readers who saw an AI-written story without knowing its origin rated it more engaging than a human-written one. Still, when they believed a story was human-written, their ratings improved regardless of who actually wrote it. Asked to guess the author outright, participants did worse than chance, correct only 39% of the time.
What strikes me most isn't how hard it is to tell human writing from AI writing, or AI-fabricated science from real science. It's that our assumptions about who (or what) produced something already shape how we judge it, before we've even weighed the content itself. As AI gets better at sounding human and formatting like a journal article, plausibility keeps doing more work than it should.
This connects to the broader problem of “hallucinations,” where answers can sound fluent and appropriate without real evidence behind them. Sycophancy adds another layer, referring to LLMs' tendency to agree with a user's assumptions rather than correct them. Simply suggesting that bixonimania is real can prompt a response that reinforces the misconception. The spread of AI health misinformation is both a technical failure and a failure of how humans and machines interact.
And the underlying accuracy problem hasn't gone away.
None of this would matter as much if the underlying information were reliably correct, but it isn't. A study in BMJ Open tested five chatbots on controversial health questions spanning cancer, vaccines, stem cells, nutrition, and athletic performance, and found 49.6% of responses problematic, with nutrition and athletic performance drawing the most trouble.
The stakes get sharper once real people are in the loop. In a study of 1,298 participants using LLMs to work through medical scenarios, the models named a relevant condition in over 90% of cases when tested alone. That advantage collapsed once real people drove the conversation, with users identifying a relevant condition in fewer than 35% of cases, no better than a control group using no LLM systems. Moreover, the models gave opposite advice to near-identical cases: two users with symptoms very similar to subarachnoid hemorrhage were told, respectively, to go to the emergency room and to lie down in a dark room, guidance that could have been fatal.
A systematic review of 83 studies comparing generative AI's diagnostic accuracy to physicians found AI models averaged 52.1% accuracy, not significantly different from physicians overall, though physicians scored higher, especially against specialists. The authors concluded generative AI isn't yet a reliable substitute for medical specialists, though it may serve as an assistive tool and training aid for professionals in training.
Where that leaves us
AI is here to stay, but treating its output as fact turns a genuinely useful tool into just another vector for misinformation. Criticizing these tools for their limitations is justified. Ultimately, the greater challenge is learning—and teaching—how to use them and how to distinguish a reliable answer from a questionable one.
Category
