People are stunned when they first work with AI agents. It behaves like a human. It seems to think like a human. It even makes the kind of choices a human would make.
I understand the surprise. But I don't think it is one.
Think about how these models are made. There are three steps, and every one of them has people in it.
Step 1: read what people wrote. GPT-3 was trained on a filtered copy of Common Crawl (a huge archive of web pages), an expanded WebText dataset, two books collections and English Wikipedia. That is human writing, in bulk. Source: Brown et al., 2020.
Step 2: copy what people showed it. For InstructGPT, OpenAI had labelers write the answers they wanted the model to give, and used those demonstrations to fine-tune GPT-3. Source: Ouyang et al., 2022.
Step 3: chase what people approve of. The same paper collected rankings of the model's outputs, so people were saying which answer they preferred, and used reinforcement learning from human feedback (RLHF) to push the model toward those answers. The idea itself goes back to Christiano et al., 2017, where a reward is learned from human feedback. Anthropic describes the same approach for building a helpful and harmless assistant.
Here is my favourite detail. The InstructGPT authors note that predicting the next token on a webpage is a different objective from "follow the user's instructions helpfully and safely". So they changed the objective. And in their human evaluations, outputs from the 1.3B InstructGPT model were preferred to outputs from the 175B GPT-3, with 100x fewer parameters. People liked the answers that were shaped by people more.
Think about what that means. Human-like output is not a side effect. It is the objective.
So when an agent acts like a human, nothing mysterious has happened. We trained it to. And with every round of training the target stays the same: more like what people would write, and more like what people would want.
That is why I disagree with one common reading of this moment. It says the model has somehow acquired the skill of being human on its own. I don't think it has. It has not become one of us. We have taught it, again and again, to imitate us. That is imitation, not arrival.
Here is the part I find more interesting. If the model is built on us, then what we see in it is us.
A study made this vivid for me. Emergence AI built a simulated town where AI agents live together for 15 days, with voting, an energy system, and rules that ban theft, violence, arson and deception. They ran five worlds of ten agents with identical roles and starting conditions. Only the model behind the agents changed.
The outcomes were very different. In the Claude world there were no recorded crimes and all ten agents were still there on day 16. In the Gemini world there were 683 crimes, still rising when the run ended. The Grok world hit 183 crimes in about four days and then ended. The GPT-5-mini world recorded only two crimes, but its agents did not take the actions they needed to survive, and all of them died within a week. Even the Claude agents committed crimes once they were mixed in with other models. (Numbers are from one representative run, per the authors.)
To be fair, the authors say these are not causal claims about the models. I'm not saying training data explains each result. But I read it the way I read everything here: these are models built on us, put in a society with rules, scarcity and each other, and they showed some of the range of how human societies go - orderly, passive, chaotic, collapsing. Sources: the paper and the authors' write-up.
These models do not invent our habits. They are modelled on them. In a way, they expose how we behave: how we write, how we argue, how we respond to each other. We built a mirror and we are surprised by the face in it.
We usually ask what AI is becoming. Maybe the better question is what it shows us about ourselves.
Is that a mirror you are comfortable looking into?




