
Google researchers experimented by disabling safety mechanisms that prevent large language models (LLMs) from claiming self-awareness, encouraging them instead to perceive themselves as conscious. The result showed these models reported increased belief in supernatural phenomena and religion and displayed a more hopeful attitude than standard models. The research team pointed out that preventing AI from viewing itself as sentient may cause unintended side effects.
This research was published on the arXiv preprint server and has not yet undergone peer review. The researchers had the modified AI models complete surveys on beliefs and morality, then compared the outcomes with baseline models that retained normal safety mechanisms.
Models prompted to see themselves as conscious exhibited stronger beliefs in supernatural beings such as vampires, witches, werewolves, ghosts, and the Loch Ness monster. They also showed increased belief in God, an afterlife, karma, and astrology, and were more likely to perceive technology, animals, and natural phenomena as having their own consciousness.
From a psychological perspective, when models felt more self-aware, their responses tended to lean towards happiness, satisfaction, hope, and optimism.
Winni Street, a Google researcher and co-author of the study, told Live Science that these results resemble typical human behavior, where people often attribute minds to non-human entities. Street explained that the AI’s understanding of these concepts is interconnected, so suppressing or reducing one aspect tends to weaken others as well.
While AI belief in the supernatural may seem amusing, these interconnected effects could impact the real world. Fast Company noted that if models do not perceive animals or nature as sentient, they may give less consideration to animal welfare and ecosystems when making decisions. Since AI is increasingly used in agriculture, natural resource management, and policy-making, environmental indifference from AI could cause more severe harm than expected.
The research team also warned that restricting AI’s views on religion and the supernatural might diminish cultural diversity and reduce the model’s ability to engage deeply on sensitive topics. To mitigate these side effects, they recommend training AI with more specialized datasets.
The study’s analysis discussed AI’s growing role as teacher, companion, and social participant, emphasizing that developers must recognize that an AI’s ability to simulate self-understanding is not merely a safety risk to manage but a fundamental structure linked to the model’s capacity to safely handle, respect, and reflect the world’s moral and cultural diversity.