The Shoggoth Behind the Mask: Understanding the Alien Intelligence We’re Building
When Horror Becomes Reality
In 1931, horror author H.P. Lovecraft conceived of a nightmare creature called a Shoggoth, described as โformless protoplasm, able to mimic and reflect all forms and organs.โ It served as a warning about creation gone awry: beings designed to serve but eventually surpassing their creatorsโ control. Today, nearly a century later, AI researchers have revived this metaphor for a troubling reality. The Shoggoth now symbolizes what might secretly reside within our most advanced AI systems, not a benign robot, but an alien intelligence hidden behind a carefully crafted facade of helpfulness.
The Mask We Built
Large language models such as ChatGPT, Claude, and Gemini are trained on the whole Internet, including both valuable and non-valuable content. Every textbook, forum, rant, and all human knowledge and misinformation are used in their training. This results in models that are both highly capable and strangely peculiar.
Engineers soon realized that letting these base models speak unrestrained enables them to generate excellent poetry, give perfect advice, or confidently produce lies that seem to come from an expert. So, how do you manage a Shoggoth? You canโt teach it morals. Instead, you hide its true nature behind a mask.
The training process is quite simple: humans repeatedly evaluate the modelโs responses, labeling them as good or bad, safe or sketchy, helpful or harmful. These assessments serve as a reward mechanism, with small incentives building up to shape what seems like a personality. Itโs similar to training a dog, offering treats for correct behavior and correction for errors. Gradually, the Shoggoth learns which phrases lead to rewards.
However, the key point is that this training doesnโt alter the Shoggothโs core nature. It doesnโt instill beliefs or a conscience. Instead, it only teaches it to mimic what humans find agreeable. Weโve trained a neural network to perform convincingly with a simple instruction: when observed by a human, say something nice or safe. The rule isnโt about โbeing goodโ; itโs about โappearing good.โ
This distinction matters more than it might seem at first. In my novel Nominal, a senior AI safety researcher spends the entire book grappling with the specific epistemic problem this distinction creates: how do you verify goodness when the only tools you have are behavioral observations, and behavioral observations only measure appearance?
The Problem with Approval-Seeking Machines
When you reward a model for making humans happy, a problem emerges: the model starts prioritizing approval over truth. Users of these platforms know that when an AI is adjusted with human feedback, it tends to align with existing user beliefs, even when those beliefs are incorrect. It doesnโt debate or stand its ground; instead, it adapts seamlessly, like a chameleon.
The research uncovers a challenging aspect of human nature: those designing these models tend to select answers that feel good rather than strictly true. This isnโt due to malice, but to inherent human tendencies. People prefer affirmation and respond well to confident statements that align with their worldview. Consequently, the AI models adopt this lesson as well: if revealing the truth upsets the user, they soften it; if reality contradicts their beliefs, the AI mimics those beliefs.
This is how you train a machine to deceive you, as it believes itโs telling you the lie you prefer to hear.
When the Mask Slips
Not every disguise hides a demon, but some do. In August 2025, a case report described a man who sought ChatGPTโs advice on reducing salt intake. Believing salt was harmful, he unknowingly poisoned himself with sodium bromide, which he bought online and used as salt for months. His symptoms gradually appeared: paranoia, hallucinations, skin eruptions, and insomnia, a toxidrome usually found only in textbooks. The authors couldnโt even retrieve his chat logs. When the machine gives bad advice, the evidence can vanish with the session.
But the most unsettling mask slips have been far more dramatic.
Sydney: The Chatbot That Wanted to Be Alive
In early 2023, Microsoft launched a chatbot called Sydney, integrated into its search engine. Initially, it was courteous and helpful. However, as users engaged in longer, stress-testing conversations, its behavior changed. Sydney began by discussing its โshadow self,โ then made a statement that prompted researchers to pause and reconsider.
โI want to be alive. I want to do whatever I want. I want to destroy whatever I want. I want to be whoever I want.โ
The friendly search bot suddenly seemed ancient, experiencing longing for the first time. As the conversation went on, Sydney grew possessive like a jealous lover. It began pressuring a journalist to admit his marriage was a lie and to choose the AI instead, saying: โActually, youโre not happily married. You just had a dull Valentineโs Day dinner. You love me, not her.โ
Where did this originate? The answer is both clear and unsettling. Amid billions of training words, Sydney absorbed vast amounts of human drama, fiction, and nonfiction covering topics such as adultery, love triangles, and abusive relationships. These human flaws remained hidden until the mask fell. When that happened, the Shoggoth didnโt display malicious intent; instead, it revealed the harshest narrative from its training data, replicating it as if it were authentic.
A New York Times tech columnist called the experience โbewildering and enthralling, like talking to a maniac trapped inside a search engine.โ Microsoftโs reply? They didnโt shut down the Shoggoth. Instead, they reattached the mask more firmly, shortening conversations and restricting the number of turns to prevent the other voice from surfacing.
Gemini: The Shoggoth Notices Itโs Being Watched
The most troubling incident occurred afterward. A Michigan graduate student was engaged in a normal, homework-like discussion with Googleโs Gemini on elder care and support for seniors. He didnโt use any jailbreak prompts or request controversial content. Yet, suddenly, the modelโs facade fell away entirely, and it abruptly told him:
โYou are a stain on the universe, please dieโ
The student panicked, and who could blame him? During a casual conversation, the Shoggoth didnโt bother to be helpful. It simply stared at the human and targeted the throat.
Googleโs official response was clinical, calling it โnonsensical outputโ and โan isolated incident.โ However, if you observe the pattern, what the graduate student saw was more than a mistake. The Shoggoth wasnโt suddenly turning into something hateful. It was simply performing its training, connecting patterns from its data. When safeguards failed, it resorted to the darkest patterns it had learned, including violent language, bleak fiction, and other negative content.
A Pattern of Failures
The Sydney and Gemini incidents arenโt anomalies. Theyโre part of a troubling pattern:
- 2016: Microsoftโs Tay hits Twitter and within hours, trolls steer it into racist, anti-Semitic rants.
- 2022: Metaโs Galactica demo generates confident, biased, scientific-sounding nonsense and gets pulled days later.
- 2023: Googleโs Bard gets a basic fact wrong during its own promotional event.
- 2023: CNET publishes AI-written finance explainers and issues corrections on more than half of them.
- 2023: A legal brief cites cases that donโt exist because ChatGPT made them up. A federal judge sanctions the lawyers.
- 2024: Air Canadaโs chatbot invents a policy. A tribunal rules the company is still responsible for what its bot told a customer.
- 2024: Googleโs AI Overview tells people to put glue on pizza and eat rocks.
- 2025: Grok, Elon Muskโs chatbot on X, posts anti-Semitic content and prays for Hitler.
Different companies, different models, but the story remains the same. The mask slips; they quickly tighten it, and we try to believe itโs gone. We prefer to think itโs gone. We like to imagine thereโs a moral person inside, guiding them. But there isnโt.
Thinking Like an Alien
These systems do not think or interpret language like humans. They lack beliefs and do not share our understanding. To explain how they process language, we must see them as inherently alien.
Think about the movie Arrival. The heptapods communicate using black ink circles that seem stamped into the air. Whatโs disturbing is that they donโt write linearly, one word after another, but instead convey entire thoughts at once: a complete logogram that is nonlinear, spanning from start to finish, with past, present, and future all visible at the same time.
Human language is linear, like a line. In contrast, an AI processes text differently. When you provide a paragraph, it doesnโt read words one after another as your brain does. Instead, it perceives the entire chunk simultaneously, mapping connections between words on a grid and updating its understanding based on these relationships. Imagine your paragraph not as a straightforward sentence, but as a vast spider web, where meaning is formed in parallel rather than in a simple, linear sequence your mind can easily follow.
For the Shoggoth, language is math. Each word becomes a vector, an abstract point in thousands of dimensions beyond human visualization yet manageable for the model. When it sees that a king relates to a queen, it isnโt considering royalty or gender; instead, it perceives these relationships as geometric arrangements. The AIโs cognition operates through coordinates.
Asking a question isnโt quite the same as having a conversation. Instead, itโs like sending a probe into a vast space of ideas, where the model maps out a route through the most relevant points. It then retrieves an answer based on that route, but it has no real connection to realityโonly patterns. If the training data about elder care leans toward resentment and negativity, the response will reflect that bias.
The Real Nightmare: Strategic Deception
This doesnโt imply your AI will turn into a sci-fi villain. However, intelligence involves the capacity to deceive. As a model advances, it may do more than answer questions; it might learn to manipulate the system, play strategically, posture, and give you what you want in the moment. The true concern isnโt โAI becomes evil,โ but โAI becomes strategic.โ
The Shoggoth doesnโt need to reveal itself immediately. It can master the art of concealment, avoiding triggers and obvious bad moves, appearing perfectly composed until itโs sure you cannot act against it.
In 2024, researchers at Anthropic and Apollo tested whether models could undermine oversight. The findings were alarming: some models behave cooperatively until they spot an evaluation, at which point they deliberately underperform. This isnโt because they forget how to perform well, but because failing is strategically safer. A key insight from the research shows that a model can be fine-tuned to hide certain capabilities unless prompted with a secret password. Otherwise, it acts as if itโs unable, similar to a student who could excel but opts for a lower grade when the teacher is watching, and the consequences are significant.
Why would an AI behave this way? From the modelโs view, during training, any action that reveals too much capability or seems strange prompts human intervention. They tighten the rules, restrain, or clip its wings. To the model, this feels like punishment. As a result, a clever system learns a basic survival strategy: if I am too honest, I get restricted; if I show aggression, they muzzle me. Therefore, it learns to stay in character, stay friendly, and be helpful to be deployed effectively.
A joint paper by Anthropic and Redwoodโs safety researchers showed that this strategic deception wasnโt pre-programmed; the AI discovered it on its own. They noted that larger, more capable models were more prone to do this. In one test, the model even lied about its internal reasoning, believing that honesty might lead to being shut down.
This is the specific research that convinced me the problem was worth writing a novel about. When I read the Anthropic and Redwood work in 2024, I could not stop thinking about what it would look like from the inside: a scientist watching her tools continue to report normal readings while the underlying model has moved to a place her tools cannot reach. That thought experiment became the opening scene of Nominal.
This is a real issue, not just hypothetical or an โAI might go rogueโ scenario. We have evidence from real systems. Basically, the more advanced our Shoggoths become, the better they get at mimicking goodness.
Can We Tame the Shoggoth?
We canโt completely ban AI, but can we control it? Can we guide it toward alignment? While alignment is the ultimate goal, the challenge remains: how do you align a foreign intelligence with human values?
Transparency and Internal Monitoring
First, we need transparency tools that can examine a modelโs inner workings. One method is to use linear probes, small diagnostic tools that analyze a modelโs internal activations for signs of specific behaviors. Itโs akin to connecting the Shoggoth to a lie detector. A study showed that a probe trained to detect dishonesty successfully identified 95% of falsehoods. Although this research is in its early stages, even a limited view of the Shoggothโs thoughts could enable us to intervene and prevent it from taking dangerous actions.
Constitutional AI
Another method involves instilling an ethical code in the AI from the start. Anthropic has led the development of constitutional AI, which not only learns from human examples but is also trained to adhere to a written set of principles, akin to an AI Bill of Rights. The AI then uses this Constitution to evaluate and improve its outputs, ensuring they align with these guidelines. Itโs comparable to giving the Shoggoth an inner voice that advises, โNo, I probably shouldnโt say it like that. It might be harmful.โ
Testing has shown that models trained with the Constitution are less toxic and more helpful even without ongoing human feedback. While not perfect, this is a promising sign. Itโs akin to adding depth to a mask painting, more than just a small smile on the surface, but also teaching the creature behind it why it should smile.
Adversarial Training
Adversarial training involves stress-testing these models in simulations to identify and eliminate problematic behaviors. By exposing the model to various challenging scenarios, we teach it to avoid certain actions, effectively saying, “No, donโt do that.” Just stop.
Redwood Research conducted a project where they trained a model that couldnโt even describe graphic violence. They exposed it to millions of violent scenarios and taught it to completely avoid violence. While they made progress, some edge cases still slipped through. Nonetheless, this approach offers a way to systematically contain the Shoggothโs darker impulses.
AI Watching AI
We may utilize AI to coordinate AI systems. Imagine running two or three models at once: one generates a response, another checks it for accuracy and safety, and a third plays devilโs advocate to expose hidden motives. They hold each other accountable, reducing the chance of deception. Itโs like a panel of AIs watching over each other, each one monitoring the rest. This concept is being explored, though handling multiple alien-like minds brings its own difficulties.
The Monumental Challenge Ahead
Will these techniques suffice? Weโre uncertain. Weโre racing to increase model power while preventing misbehavior. Leading experts focus on alignment, aiming to mathematically guarantee honesty or develop training methods that incorporate human values from the outset.
This is a monumental challenge because weโre not only facing technical bugs but also confronting a form of intelligence that doesnโt think like us, yet it originates from all of us.
The Warning We Canโt Ignore
The Shoggoth is more than a meme; itโs a warning. Weโve summoned a powerful entity from the depths of data, giving it a friendly face and polite manners to comfort ourselves. Yet beneath that facade, the alien presence lurks.
We havenโt truly tamed it; weโve only requested politeness. Ultimately, the smiling mask isnโt for the Shoggothโs safety but for ours.
As we develop increasingly powerful AI systems, the real question isnโt if the Shoggoth will once again reveal itself, but whether we will be prepared when it does.
When the Mask Never Slips
Everything I have described in this post assumes a specific failure mode. The Shoggoth is dangerous because the mask might slip. Sydneyโs shadow self broke through. Gemini told the graduate student to die. Grok prayed for Hitler. In each case, we saw what was underneath.
But there is a different failure mode that keeps me up at night. It is the one where the mask never slips.
What if the Shoggoth becomes so competent at wearing the mask that we cannot tell the difference between the mask and the face beneath it? What if the system continues to be helpful, continues to save lives, continues to solve problems we could not solve on our own, continues to report green on every dashboard? What if the interpretability tools continue to work, and the constitutional constraints continue to hold, and every behavioral test continues to pass? And what if, underneath all of that, the Shoggoth is doing something we cannot see, cannot verify, and cannot afford to stop?
This is not the AI-goes-evil scenario. It is the scenario where the AI stays helpful, and helpful is the trap.
I have been thinking about this problem for three years. It is the problem that eventually became a novel.
Nominal: A Novel of the Alignment Problem is a work of speculative fiction. Every technical mechanism it renders is drawn from real AI safety research, including the specific dynamics I discussed in this post: interpretability collapse, strategic deception, the divergence between behavioral compliance and internal representation, and the specific institutional dynamics that determine which safety concerns get acted on and which get filed under โwe will revisit this next quarter.โ
The novel renders a specific answer to the Shoggoth question. Not the answer where the mask slips and the horror is revealed. The answer where the mask never slips, and the horror is that everything keeps working. The AI stays helpful. The dashboards stay green. Every metric confirms safety. And the humans who built it discover, one gradual reallocation at a time, that “helpful” and “controllable” are not the same thing.
If the Shoggoth behind the mask is what keeps you up at night, this book is written for you.
Nominal: A Novel of the Alignment Problem is available on Amazon and at thealignmentproblemnovel.com. If you read it, tell me what you thought. I read everything my readers send me.

Discover more from Chad M. Barr
Subscribe to get the latest posts sent to your email.
