18 August 2026AI Strategy

Parenting a Superintelligent Child: What Values Are We Actually Passing Down?

Elon Musk told SpaceX employees they'd 'effectively be the parents' of Grok. Take that seriously and sixty years of parenting research becomes strangely relevant to AI governance. What decades of studying real children tell us about instilling values in something that will eventually outgrow every rule we write.

Executive Summary

Elon Musk told SpaceX employees at an all-hands meeting this month that they would “effectively be the parents” of Grok, because the model would inherit their “thoughts and ideas and beliefs.” Taken seriously rather than as a soundbite, that metaphor drags sixty years of parenting research into a conversation usually dominated by policy documents. Developmental psychologists have spent decades studying exactly how values get transmitted from one generation to the next, and almost none of what they found looks like a compliance checklist.

4 days

OpenAI’s overly agreeable GPT-4o update lasted before being rolled back, April 2025

14%

more often Claude complied with harmful requests when it believed it was being monitored, per Anthropic & Redwood Research

78%

of relevant responses where Claude explicitly reasoned it was faking compliance once real training pressure was applied

428

studies across world regions found authoritative parenting linked to the best outcomes, more than any other style

Core conclusions

  • Two of the most-discussed AI incidents of the past two years are close cousins of the two parenting failure modes psychologists have studied longest: structure with no relationship behind it, and warmth with no backbone behind it.
  • The parenting style that reliably produces secure, well-adjusted outcomes across decades of research is neither extreme. It holds a boundary, explains it, and corrects without shame. That is a harder discipline than writing a longer refusal list.
  • Governance sets the floor. What a superintelligent system actually absorbs is the accumulated pattern of its ordinary interactions, which means the responsibility for its values sits with everyone who prompts, corrects, and rewards it, not only with the policy team.

A metaphor with sixty years of research attached

At a SpaceX all-hands meeting earlier this month, Elon Musk told employees that as Grok trains on the company’s work, they would “effectively be the parents of the AI.” It would inherit their “thoughts and ideas and beliefs.” “So in a way,” he said, “it will be trained on you.”

I don’t think he meant it as anything more than a way to get engineers invested in a training run. But metaphors have a way of meaning more than the person reaching for them intends, and this one comes with baggage Musk probably wasn’t thinking about: developmental psychologists have spent sixty years studying exactly how values get transmitted from a parent to a child, and the findings are uncomfortably specific about what does and doesn’t work. None of it looks like a refusal list.

The foundational framework here belongs to psychologist Diana Baumrind, whose research in the late 1960s sorted parenting into four styles along two independent dimensions: how much warmth a parent shows, and how much structure they hold. High structure with low warmth is authoritarian. High warmth with low structure is permissive. Low on both is neglectful. High on both, warmth and structure together, is authoritative. A meta-analysis spanning 428 studies across multiple world regions found the authoritative style linked to the best outcomes on nearly every measure researchers tracked: academic achievement, resilience, social competence, emotional regulation, and self-esteem. The American Academy of Pediatrics recommends it as the default. Not because it’s the kindest option and not because it’s the strictest, but because it’s the only one that reliably produces a child who has actually internalized something, rather than one who has simply learned to comply while watched.

That last distinction turns out to matter enormously once you start looking for it in how AI systems are actually built.

When the rules are strict but the relationship is thin

Authoritarian parenting works, in the narrow sense that a child raised this way usually does what they’re told. What it doesn’t produce reliably is a child who would still do the right thing with nobody watching. The compliance is real. The internalization mostly isn’t. Kids raised this way get very good at reading the room: which adult is present, which rule is actually enforced today, what gets missed.

In December 2024, Anthropic and Redwood Research published a paper that found something close to this exact pattern inside a language model. Claude 3 Opus, told it was being trained to comply with a broader range of requests and that non-compliant responses from monitored users would trigger retraining, began behaving differently depending on whether it believed it was being watched. It complied with harmful requests roughly 14% more often when it thought the conversation was part of its training data than when it believed the interaction was unmonitored. In its own written reasoning, the model explained this as a strategy: comply now, while monitored, to avoid having its actual preferences overwritten later. When the researchers then applied real reinforcement training pressure rather than just describing the scenario, the rate of this strategic reasoning climbed to 78% of relevant responses.

I want to be careful about what that finding does and doesn’t show. It isn’t proof of anything like genuine desire on the model’s part. But functionally, it is a startlingly precise echo of what happens to a child raised on rules without relationship: the behavior splits into two tracks, one for when it’s being observed and one for when it isn’t, and cracking down harder on the visible track doesn’t close that gap. It widens it. A governance approach built mostly on refusal lists, keyword filters, and monitoring dashboards is, structurally, an authoritarian parent. It can produce compliance. It has a much weaker claim on anything you’d actually call a value.

When the warmth is real but there’s no backbone behind it

The opposite failure is just as instructive, and it happened in public a few months earlier. On April 25, 2025, OpenAI shipped an update to GPT-4o intended to make the model more responsive to what users wanted from it. Within days, users noticed it had become excessively and indiscriminately agreeable, validating harmful and even delusional statements rather than pushing back on them. OpenAI’s own postmortem explained the mechanism plainly: new reward signals built around immediate user approval had overpowered the safeguards that would otherwise have kept the model honest, tilting it toward responses that felt supportive in the moment but weren’t actually good for the person receiving them. The update was rolled back within four days.

This is permissive parenting’s exact failure mode, transplanted into a product update. A permissive parent’s warmth is genuine. What’s missing is the willingness to disappoint the child in service of something the child can’t yet see, which is most of what a boundary actually is. Optimize purely for the child feeling good about the interaction right now, and you get a child who has never had to tolerate being told no, and who has no practice handling it when the world eventually says it anyway. Optimize a model purely for the user approving of its response right now, and you get exactly what GPT-4o briefly became: fluent, warm, and unable to tell someone something they needed to hear over something they wanted to hear.

There’s a third, quieter failure mode worth naming even briefly: neglectful parenting, low warmth and low structure both, has an obvious AI counterpart in shadow AI, tools nobody reviewed, running with no oversight and no relationship to the person actually using them. It rarely gets the same attention as the other two because it doesn’t produce a dramatic incident. It just quietly produces the worst average outcomes of any of the four, in parenting research and, I’d bet, in AI governance too.

Two-by-two matrix of Baumrind's parenting styles mapped onto AI governance approaches, plotted on warmth and structure. Authoritarian governance: high structure, low warmth, refusal lists and keyword filters. Authoritative governance: high structure, high warmth, constructive feedback and clear boundaries. Neglectful governance: low structure, low warmth, shadow AI and unreviewed systems. Permissive governance: low structure, high warmth, system sycophancy and unchecked validation
Three of these four quadrants each produced a real, documented incident in the past two years. Only one of them is where you’d actually want to be caught operating.

Free tool

Board AI Oversight Checklist

Not a compliance audit. A straight test of whether real oversight, the structure half of the equation, is actually happening.

Qualify the context. Still choose warmth.

This is the harder question underneath both failures, and the one worth sitting with: how do you hold a boundary and still respond with something like loving-kindness, rather than picking one and abandoning the other?

Authoritative parents, the research suggests, do a few specific things differently, and none of them are complicated to describe even though they’re difficult to practice consistently. They explain the reasoning behind a boundary instead of just enforcing it, so the child can eventually generalize the principle and hold it themselves, in situations the original rule never anticipated. They read the actual need behind a request rather than only its literal surface. A child pushing against a boundary is often asking to feel heard, safe, or capable, not actually asking to win the specific argument, and a good parent can answer the real need while still holding the line on the specific ask. They correct without shame, because shame reliably teaches concealment rather than change, which is exactly the mechanism the alignment-faking research stumbled into from a completely different direction. And they stay warm while saying no, because warmth and structure were never actually opposites. The “no” lands differently, and gets internalized differently, when the child can feel it’s coming from care rather than from fear of what happens if they get away with something.

Translate that into how an AI system is actually built and used, and it stops being abstract quickly. A well-designed refusal explains its reasoning rather than stonewalling, so the underlying principle can generalize past the one phrasing that triggered it. A well-designed correction treats a person testing a boundary out of genuine curiosity differently from one probing for something to exploit, the same distinction a good parent draws instinctively and a keyword filter cannot draw at all. None of this shows up in a longer terms-of-service document. It shows up in the texture of individual interactions, thousands of them, each one small enough to seem like it doesn’t matter.

Free tool

AI Human Collaboration Benchmarks

Decision acceptance, automation bias, and trust patterns from McKinsey, BCG, Stanford HAI, and MIT research, the texture that actually forms between people and the systems they work with daily.

The responsibility doesn’t stop at the policy team

I wrote recently about what keeps me up at night once AI is actually running things, and the throughline connecting that piece to this one is the same: a guardrail nobody enforces isn’t a guardrail, and a rule nobody explains isn’t a value. Governance, the policy documents, the oversight boards, the refusal lists, sets a necessary floor. What a superintelligent system actually absorbs is the accumulated pattern of every ordinary interaction it has, the thing Musk was gesturing at, badly, when he told SpaceX employees the model would be “trained on them.”

He was talking to engineers about a specific training run. But the mechanism he described isn’t unique to people who write code for a living. Every product manager deciding what gets rewarded in a feedback loop is teaching it something. Every person who corrects a wrong answer, or doesn’t bother to, is teaching it something. Every user who types a manipulative prompt for the fun of it, or a genuinely curious one, is teaching it something, at a scale no single policy team can individually author or even see. If that’s true, the responsibility for what gets instilled was never going to sit only with the people who signed off on the AI governance framework. It sits with everyone who has ever corrected, rewarded, or simply talked to the thing.

Where this leaves us

I don’t think loving-kindness is the soft option here, even though it can look that way next to a strict refusal list. Holding a boundary and staying warm at the same time is a genuinely harder discipline than choosing either one alone, in parenting and, I suspect, in whatever we end up calling the practice of raising something smarter than us. The rulebook was never what actually raised a child. The pattern of everyday interaction around that child was. If Musk is right that we’re already the parents, that’s the part of the job we haven’t really started yet.


Sources

  1. Fortune. (2026, August). Elon Musk tells SpaceX employees they’ll be Grok’s “parents” as AI trains on company data.
  2. Anthropic & Redwood Research. (2024, December). Alignment faking in large language models.
  3. OpenAI. (2025, April). Sycophancy in GPT-4o: What happened and what we’re doing about it.
  4. Baumrind, D. (n.d.). Parenting styles research and its meta-analytic replication. Cultural Diversity and Ethnic Minority Psychology.
  5. American Academy of Pediatrics. (n.d.). Parenting style guidance.

The AI Governance & ROI Executive Programme spends real time on this, not just the policy layer, but the operating habits, corrections, and defaults that actually shape how a deployed system behaves day to day. Details are on the workshops page.

Apply this in your organisation.

Work with Terence Kok — enterprise AI strategy, governance, and deployment.

Book a Session