このノートにはまだ日本語版がありません。原文(英語)を表示しています。
In 1995, two researchers sat down and wrote what an agent is. Their paper has a section titled What is an Agent?, and under it a subsection offering a weak notion of agency — autonomy, social ability, reactivity, initiative. It was careful work and it has been cited ever since.
Thirty-one years later, in January 2026, a policy institute published a brief arguing that nobody knows what the word means.
The fog has a body count
The brief is Lost in Definition, by Atalan, Reynolds and Jensen at CSIS, and it is not a complaint about sloppy marketing. It is an argument that the vagueness is now doing damage: “Despite its widespread use, there is no shared understanding of what qualifies as an agentic AI system.”
Then the part worth sitting with. “When procurement documents request ‘agentic capabilities’ without operational specifications, vendors can satisfy requirements in name while delivering vastly different systems.” Buyers, they write, “may believe they are procuring agents that reason, plan, and act across systems when they are actually receiving glorified chatbots or, conversely, systems with more autonomy than realized.”
Read that last clause twice. The failure runs both directions. You can buy a toy believing it is an operator. You can also buy an operator believing it is a toy, and hand it authority you never meant to delegate.
That is what a naming problem looks like once money and permissions are attached to the names. It is no longer a taxonomy debate. It is a procurement defect, and eventually a liability one.
The word I don’t own either
Before I propose anything, the thing I would rather say myself than have someone point out.
I did not coin “twin”, and neither did anyone else in AI. It arrived through manufacturing, and it has a standard: ISO 23247, Digital twin framework for manufacturing, whose sixth part was published this year. Its centre of gravity is a digital representation kept synchronised with the real thing it represents.
It has also already been pointed at organisations. Gartner has run a Digital Twin of an Organization platform category for years, with a vendor market attached.
So when I say a business should have a twin, I am not inventing vocabulary. I am standing in a house somebody else built, and the honest description of what is new is narrow: those twins model an organisation. The one I care about acts as one. Synchronisation was always the point; agency is the part that arrived late.
I would rather claim the narrow true thing than the wide false one. A vendor who coins a word owns nothing but the word.
Why every definition so far has failed
Look at the definitions in circulation and they share a shape: they sort by capability. Can it plan? Can it call tools? Can it run unsupervised for an hour?
Capability is the worst possible sorting key, for a boring reason: it moves. What required a research team in 2024 is a library call now, and a checkbox next year. A taxonomy anchored to capability has to be rewritten every model release, which means it is not a taxonomy, it is a snapshot.
CSIS reaches the same conclusion and proposes the alternative plainly — shift the question “from ‘What can this system do?’ to ‘How does this system reshape organizational decisionmaking, and what authorities are delegated under what constraints?’”
That is the whole move. Stop asking what it can do. Ask what happens when it does it.
The question that holds
Who answers for it.
Every useful distinction in this space falls out of that one question, and it has the property capability lacks: it does not go stale. A model twice as capable does not change who is accountable when it acts. If anything it sharpens the question.
So here is the taxonomy I actually use.
A bot acts under a rule somebody wrote. It is deterministic — same conditions, same behaviour, every time. Accountability sits with whoever wrote the rule, and the thing you audit is the rule itself. A bot that does something surprising is a bug report.
An agent acts under a goal somebody set, and chooses its own steps. Nobody wrote the path, so nobody can be accountable for the path. Accountability sits with whoever delegated the goal and drew the boundary around it — and what you audit is the boundary, not the route. An agent that does something surprising inside its boundary is working correctly. That is the deal you signed.
A twin acts as the party itself. Not a worker inside the business — the business, showing up. When a customer talks to it, they are not talking to a tool the company bought; they are talking to the company. Accountability sits with the business, in full, because in that moment there is no gap between them. What you audit is the record.
Two of these are workers. One is not.
Here is the thing I got wrong for a while, and the reason lists of three tend to mislead.
Bot, agent and twin are not three points on one spectrum. Bot and agent differ by how they are bound — a rule versus a goal. That is a real distinction and a narrow one; they are both workers, and you can swap one for the other as the task hardens or loosens.
A twin differs by what it is. It is not a better agent. It is a party — the thing that has the relationship, holds the obligations, and cannot be swapped out without the counterparty noticing they are now dealing with someone else.
So it is two and one, not three. Anyone selling you a “twin” that is really a well-configured agent has sold you a worker and charged you for a party. The test is simple: if it can be replaced without telling your customers, it was never a twin.
Whose authority, not what kind
The other split people reach for is what the thing is — a system process, or something with a personality. I have tried to make that distinction useful and I cannot. It sorts by costume.
The split that pays is whose authority the agent carries, because that is the accountability question again, one level down:
- An agent acting under the business’s authority — it can commit the company, so the company’s guardrails bind it.
- An agent acting under a person’s authority — it carries their permissions and nobody else’s, which is the seat I have been building for personal agents.
- An agent acting under its own — which is where identity stops being a nicety, and why an agent needs a real identifier and not just a key.
Same agent, same capabilities, three completely different sets of consequences. Costume tells you none of it. Authority tells you all of it.
What this frame costs me
A taxonomy published by a company that sells one of the categories deserves suspicion, so let me pay the toll.
Under this frame, most of what I ship is agents and bots. Not twins. The twin is the hardest of the three to earn, because being a party means being accountable with no one to point at — and a system only reaches it once every action runs through one governed pipeline that can actually be audited, and once agents and humans sit on one org chart with real owners. Until both are true, what you have is a very good worker wearing the company’s name.
That is a frame that makes my own claims harder to make, not easier. I chose it anyway, because a definition designed to flatter the person publishing it is worth nothing to the person reading it — and because I would rather be held to a standard I wrote than one a procurement officer writes for me in three years.
Before it sets
Vocabulary hardens. The words in circulation while a category forms are the words that end up in contracts, in regulation, and in the questions a court asks afterwards. Right now those words mean whatever the last vendor said they meant, and CSIS has already documented what that costs.
I would rather the industry argued about this early and in public than inherited it by default. So: bots are bound by rules, agents by goals, and twins are not workers at all.
Disagree with the lines if you like. Just don’t sort by what it can do — you will be rewriting your taxonomy in six months, and the thing you actually needed to know was never on that axis.
この考えを持ち帰る
正規のソースを共有するか、使っている AI のための読書プロンプトをコピーできます。
— orbiteer1