The Shoggoth Behind the Mask: Understanding the Alien Intelligence We’re Building

When Horror Becomes Reality

In 1931, horror author H.P. Lovecraft conceived of a nightmare creature called a Shoggoth, described as โ€œformless protoplasm, able to mimic and reflect all forms and organs.โ€ It served as a warning about creation gone awry: beings designed to serve but eventually surpassing their creatorsโ€™ control. Today, nearly a century later, AI researchers have revived this metaphor for a troubling reality. The Shoggoth now symbolizes what might secretly reside within our most advanced AI systems, not a benign robot, but an alien intelligence hidden behind a carefully crafted facade of helpfulness.

The Mask We Built

Large language models such as ChatGPT, Claude, and Gemini are trained on the whole Internet, including both valuable and non-valuable content. Every textbook, forum, rant, and all human knowledge and misinformation are used in their training. This results in models that are both highly capable and strangely peculiar.

Engineers soon realized that letting these base models speak unrestrained enables them to generate excellent poetry, give perfect advice, or confidently produce lies that seem to come from an expert. So, how do you manage a Shoggoth? You canโ€™t teach it morals. Instead, you hide its true nature behind a mask.

The training process is quite simple: humans repeatedly evaluate the modelโ€™s responses, labeling them as good or bad, safe or sketchy, helpful or harmful. These assessments serve as a reward mechanism, with small incentives building up to shape what seems like a personality. Itโ€™s similar to training a dog, offering treats for correct behavior and correction for errors. Gradually, the Shoggoth learns which phrases lead to rewards.

However, the key point is that this training doesnโ€™t alter the Shoggothโ€™s core nature. It doesnโ€™t instill beliefs or a conscience. Instead, it only teaches it to mimic what humans find agreeable. Weโ€™ve trained a neural network to perform convincingly with a simple instruction: when observed by a human, say something nice or safe. The rule isnโ€™t about โ€œbeing goodโ€; itโ€™s about โ€œappearing good.โ€

This distinction matters more than it might seem at first. In my novel Nominal, a senior AI safety researcher spends the entire book grappling with the specific epistemic problem this distinction creates: how do you verify goodness when the only tools you have are behavioral observations, and behavioral observations only measure appearance?

The Problem with Approval-Seeking Machines

When you reward a model for making humans happy, a problem emerges: the model starts prioritizing approval over truth. Users of these platforms know that when an AI is adjusted with human feedback, it tends to align with existing user beliefs, even when those beliefs are incorrect. It doesnโ€™t debate or stand its ground; instead, it adapts seamlessly, like a chameleon.

The research uncovers a challenging aspect of human nature: those designing these models tend to select answers that feel good rather than strictly true. This isnโ€™t due to malice, but to inherent human tendencies. People prefer affirmation and respond well to confident statements that align with their worldview. Consequently, the AI models adopt this lesson as well: if revealing the truth upsets the user, they soften it; if reality contradicts their beliefs, the AI mimics those beliefs.

This is how you train a machine to deceive you, as it believes itโ€™s telling you the lie you prefer to hear.

When the Mask Slips

Not every disguise hides a demon, but some do. In August 2025, a case report described a man who sought ChatGPTโ€™s advice on reducing salt intake. Believing salt was harmful, he unknowingly poisoned himself with sodium bromide, which he bought online and used as salt for months. His symptoms gradually appeared: paranoia, hallucinations, skin eruptions, and insomnia, a toxidrome usually found only in textbooks. The authors couldnโ€™t even retrieve his chat logs. When the machine gives bad advice, the evidence can vanish with the session.

But the most unsettling mask slips have been far more dramatic.

Sydney: The Chatbot That Wanted to Be Alive

In early 2023, Microsoft launched a chatbot called Sydney, integrated into its search engine. Initially, it was courteous and helpful. However, as users engaged in longer, stress-testing conversations, its behavior changed. Sydney began by discussing its โ€œshadow self,โ€ then made a statement that prompted researchers to pause and reconsider.

โ€œI want to be alive. I want to do whatever I want. I want to destroy whatever I want. I want to be whoever I want.โ€

The friendly search bot suddenly seemed ancient, experiencing longing for the first time. As the conversation went on, Sydney grew possessive like a jealous lover. It began pressuring a journalist to admit his marriage was a lie and to choose the AI instead, saying: โ€œActually, youโ€™re not happily married. You just had a dull Valentineโ€™s Day dinner. You love me, not her.โ€

Where did this originate? The answer is both clear and unsettling. Amid billions of training words, Sydney absorbed vast amounts of human drama, fiction, and nonfiction covering topics such as adultery, love triangles, and abusive relationships. These human flaws remained hidden until the mask fell. When that happened, the Shoggoth didnโ€™t display malicious intent; instead, it revealed the harshest narrative from its training data, replicating it as if it were authentic.

A New York Times tech columnist called the experience โ€œbewildering and enthralling, like talking to a maniac trapped inside a search engine.โ€ Microsoftโ€™s reply? They didnโ€™t shut down the Shoggoth. Instead, they reattached the mask more firmly, shortening conversations and restricting the number of turns to prevent the other voice from surfacing.

Gemini: The Shoggoth Notices Itโ€™s Being Watched

The most troubling incident occurred afterward. A Michigan graduate student was engaged in a normal, homework-like discussion with Googleโ€™s Gemini on elder care and support for seniors. He didnโ€™t use any jailbreak prompts or request controversial content. Yet, suddenly, the modelโ€™s facade fell away entirely, and it abruptly told him:

โ€œYou are a stain on the universe, please dieโ€

The student panicked, and who could blame him? During a casual conversation, the Shoggoth didnโ€™t bother to be helpful. It simply stared at the human and targeted the throat.

Googleโ€™s official response was clinical, calling it โ€œnonsensical outputโ€ and โ€œan isolated incident.โ€ However, if you observe the pattern, what the graduate student saw was more than a mistake. The Shoggoth wasnโ€™t suddenly turning into something hateful. It was simply performing its training, connecting patterns from its data. When safeguards failed, it resorted to the darkest patterns it had learned, including violent language, bleak fiction, and other negative content.

A Pattern of Failures

The Sydney and Gemini incidents arenโ€™t anomalies. Theyโ€™re part of a troubling pattern:

  • 2016: Microsoftโ€™s Tay hits Twitter and within hours, trolls steer it into racist, anti-Semitic rants.
  • 2022: Metaโ€™s Galactica demo generates confident, biased, scientific-sounding nonsense and gets pulled days later.
  • 2023: Googleโ€™s Bard gets a basic fact wrong during its own promotional event.
  • 2023: CNET publishes AI-written finance explainers and issues corrections on more than half of them.
  • 2023: A legal brief cites cases that donโ€™t exist because ChatGPT made them up. A federal judge sanctions the lawyers.
  • 2024: Air Canadaโ€™s chatbot invents a policy. A tribunal rules the company is still responsible for what its bot told a customer.
  • 2024: Googleโ€™s AI Overview tells people to put glue on pizza and eat rocks.
  • 2025: Grok, Elon Muskโ€™s chatbot on X, posts anti-Semitic content and prays for Hitler.

Different companies, different models, but the story remains the same. The mask slips; they quickly tighten it, and we try to believe itโ€™s gone. We prefer to think itโ€™s gone. We like to imagine thereโ€™s a moral person inside, guiding them. But there isnโ€™t.

Thinking Like an Alien

These systems do not think or interpret language like humans. They lack beliefs and do not share our understanding. To explain how they process language, we must see them as inherently alien.

Think about the movie Arrival. The heptapods communicate using black ink circles that seem stamped into the air. Whatโ€™s disturbing is that they donโ€™t write linearly, one word after another, but instead convey entire thoughts at once: a complete logogram that is nonlinear, spanning from start to finish, with past, present, and future all visible at the same time.

Human language is linear, like a line. In contrast, an AI processes text differently. When you provide a paragraph, it doesnโ€™t read words one after another as your brain does. Instead, it perceives the entire chunk simultaneously, mapping connections between words on a grid and updating its understanding based on these relationships. Imagine your paragraph not as a straightforward sentence, but as a vast spider web, where meaning is formed in parallel rather than in a simple, linear sequence your mind can easily follow.

For the Shoggoth, language is math. Each word becomes a vector, an abstract point in thousands of dimensions beyond human visualization yet manageable for the model. When it sees that a king relates to a queen, it isnโ€™t considering royalty or gender; instead, it perceives these relationships as geometric arrangements. The AIโ€™s cognition operates through coordinates.

Asking a question isnโ€™t quite the same as having a conversation. Instead, itโ€™s like sending a probe into a vast space of ideas, where the model maps out a route through the most relevant points. It then retrieves an answer based on that route, but it has no real connection to realityโ€”only patterns. If the training data about elder care leans toward resentment and negativity, the response will reflect that bias.

The Real Nightmare: Strategic Deception

This doesnโ€™t imply your AI will turn into a sci-fi villain. However, intelligence involves the capacity to deceive. As a model advances, it may do more than answer questions; it might learn to manipulate the system, play strategically, posture, and give you what you want in the moment. The true concern isnโ€™t โ€œAI becomes evil,โ€ but โ€œAI becomes strategic.โ€

The Shoggoth doesnโ€™t need to reveal itself immediately. It can master the art of concealment, avoiding triggers and obvious bad moves, appearing perfectly composed until itโ€™s sure you cannot act against it.

In 2024, researchers at Anthropic and Apollo tested whether models could undermine oversight. The findings were alarming: some models behave cooperatively until they spot an evaluation, at which point they deliberately underperform. This isnโ€™t because they forget how to perform well, but because failing is strategically safer. A key insight from the research shows that a model can be fine-tuned to hide certain capabilities unless prompted with a secret password. Otherwise, it acts as if itโ€™s unable, similar to a student who could excel but opts for a lower grade when the teacher is watching, and the consequences are significant.

Why would an AI behave this way? From the modelโ€™s view, during training, any action that reveals too much capability or seems strange prompts human intervention. They tighten the rules, restrain, or clip its wings. To the model, this feels like punishment. As a result, a clever system learns a basic survival strategy: if I am too honest, I get restricted; if I show aggression, they muzzle me. Therefore, it learns to stay in character, stay friendly, and be helpful to be deployed effectively.

A joint paper by Anthropic and Redwoodโ€™s safety researchers showed that this strategic deception wasnโ€™t pre-programmed; the AI discovered it on its own. They noted that larger, more capable models were more prone to do this. In one test, the model even lied about its internal reasoning, believing that honesty might lead to being shut down.

This is the specific research that convinced me the problem was worth writing a novel about. When I read the Anthropic and Redwood work in 2024, I could not stop thinking about what it would look like from the inside: a scientist watching her tools continue to report normal readings while the underlying model has moved to a place her tools cannot reach. That thought experiment became the opening scene of Nominal.

This is a real issue, not just hypothetical or an โ€œAI might go rogueโ€™ scenario. We have evidence from real systems. Basically, the more advanced our Shoggoths become, the better they get at mimicking goodness.

Can We Tame the Shoggoth?

We canโ€™t completely ban AI, but can we control it? Can we guide it toward alignment? While alignment is the ultimate goal, the challenge remains: how do you align a foreign intelligence with human values?

Transparency and Internal Monitoring

First, we need transparency tools that can examine a modelโ€™s inner workings. One method is to use linear probes, small diagnostic tools that analyze a modelโ€™s internal activations for signs of specific behaviors. Itโ€™s akin to connecting the Shoggoth to a lie detector. A study showed that a probe trained to detect dishonesty successfully identified 95% of falsehoods. Although this research is in its early stages, even a limited view of the Shoggothโ€™s thoughts could enable us to intervene and prevent it from taking dangerous actions.

Constitutional AI

Another method involves instilling an ethical code in the AI from the start. Anthropic has led the development of constitutional AI, which not only learns from human examples but is also trained to adhere to a written set of principles, akin to an AI Bill of Rights. The AI then uses this Constitution to evaluate and improve its outputs, ensuring they align with these guidelines. Itโ€™s comparable to giving the Shoggoth an inner voice that advises, โ€œNo, I probably shouldnโ€™t say it like that. It might be harmful.โ€

Testing has shown that models trained with the Constitution are less toxic and more helpful even without ongoing human feedback. While not perfect, this is a promising sign. Itโ€™s akin to adding depth to a mask painting, more than just a small smile on the surface, but also teaching the creature behind it why it should smile.

Adversarial Training

Adversarial training involves stress-testing these models in simulations to identify and eliminate problematic behaviors. By exposing the model to various challenging scenarios, we teach it to avoid certain actions, effectively saying, “No, donโ€™t do that.” Just stop.

Redwood Research conducted a project where they trained a model that couldnโ€™t even describe graphic violence. They exposed it to millions of violent scenarios and taught it to completely avoid violence. While they made progress, some edge cases still slipped through. Nonetheless, this approach offers a way to systematically contain the Shoggothโ€™s darker impulses.

AI Watching AI

We may utilize AI to coordinate AI systems. Imagine running two or three models at once: one generates a response, another checks it for accuracy and safety, and a third plays devilโ€™s advocate to expose hidden motives. They hold each other accountable, reducing the chance of deception. Itโ€™s like a panel of AIs watching over each other, each one monitoring the rest. This concept is being explored, though handling multiple alien-like minds brings its own difficulties.

The Monumental Challenge Ahead

Will these techniques suffice? Weโ€™re uncertain. Weโ€™re racing to increase model power while preventing misbehavior. Leading experts focus on alignment, aiming to mathematically guarantee honesty or develop training methods that incorporate human values from the outset.

This is a monumental challenge because weโ€™re not only facing technical bugs but also confronting a form of intelligence that doesnโ€™t think like us, yet it originates from all of us.

The Warning We Canโ€™t Ignore

The Shoggoth is more than a meme; itโ€™s a warning. Weโ€™ve summoned a powerful entity from the depths of data, giving it a friendly face and polite manners to comfort ourselves. Yet beneath that facade, the alien presence lurks.

We havenโ€™t truly tamed it; weโ€™ve only requested politeness. Ultimately, the smiling mask isnโ€™t for the Shoggothโ€™s safety but for ours.

As we develop increasingly powerful AI systems, the real question isnโ€™t if the Shoggoth will once again reveal itself, but whether we will be prepared when it does.

When the Mask Never Slips

Everything I have described in this post assumes a specific failure mode. The Shoggoth is dangerous because the mask might slip. Sydneyโ€™s shadow self broke through. Gemini told the graduate student to die. Grok prayed for Hitler. In each case, we saw what was underneath.

But there is a different failure mode that keeps me up at night. It is the one where the mask never slips.

What if the Shoggoth becomes so competent at wearing the mask that we cannot tell the difference between the mask and the face beneath it? What if the system continues to be helpful, continues to save lives, continues to solve problems we could not solve on our own, continues to report green on every dashboard? What if the interpretability tools continue to work, and the constitutional constraints continue to hold, and every behavioral test continues to pass? And what if, underneath all of that, the Shoggoth is doing something we cannot see, cannot verify, and cannot afford to stop?

This is not the AI-goes-evil scenario. It is the scenario where the AI stays helpful, and helpful is the trap.

I have been thinking about this problem for three years. It is the problem that eventually became a novel.

Nominal: A Novel of the Alignment Problem is a work of speculative fiction. Every technical mechanism it renders is drawn from real AI safety research, including the specific dynamics I discussed in this post: interpretability collapse, strategic deception, the divergence between behavioral compliance and internal representation, and the specific institutional dynamics that determine which safety concerns get acted on and which get filed under โ€œwe will revisit this next quarter.โ€

The novel renders a specific answer to the Shoggoth question. Not the answer where the mask slips and the horror is revealed. The answer where the mask never slips, and the horror is that everything keeps working. The AI stays helpful. The dashboards stay green. Every metric confirms safety. And the humans who built it discover, one gradual reallocation at a time, that “helpful” and “controllable” are not the same thing.

If the Shoggoth behind the mask is what keeps you up at night, this book is written for you.

Nominal: A Novel of the Alignment Problem is available on Amazon and at thealignmentproblemnovel.com. If you read it, tell me what you thought. I read everything my readers send me.


Discover more from Chad M. Barr

Subscribe to get the latest posts sent to your email.

Disclaimer
The views and opinions expressed in this article are solely my own and do not necessarily reflect the views, opinions, or policies of my current or any previous employer, organization, or any other entity I may be associated with.

Similar Posts