For the last few years, one question has dominated almost every enterprise AI discussion:
Which model are you using?
GPT. Claude. Gemini. An open-source model. A specialised model.
That question made sense when generative AI mostly produced text.
But agents are changing the architecture.
An agent is not just a model.
It is a model surrounded by instructions, memory, tools, identities, permissions, data, execution environments, orchestration and monitoring.
And once a model becomes capable enough to reason and use tools reliably, the surrounding system may become more important to the behaviour of the agent than small differences between frontier models.
The model still matters. But increasingly, the agent is everything we build around it.
We may be asking the wrong question
Australia's cyber security agency, the ASD, recently published guidance on what it calls agentic AI harnesses.
The term is useful.
The harness is the software surrounding the model: the components that connect it to tools, systems, data and the outside world.
This changes the mental model considerably.
Instead of:
Agent = model
it is closer to:
Agent = model + instructions + memory + tools + identity + permissions + data + execution environment + orchestration + monitoring
Once you think about agents this way, simply asking which model powers them starts to feel a little like asking which CPU runs a banking application.
It matters.
But it tells you surprisingly little about the overall system.
The same model can produce two completely different agents
Imagine two agents built using exactly the same frontier model.
The first has access to:
an unrestricted shell,
broad internet access,
production credentials,
read-write database access,
persistent memory,
and no requirement for human approval.
The second has:
approved skills,
narrowly scoped APIs,
managed identity,
read-only permissions by default,
short-lived credentials,
sandboxed execution,
approval for destructive actions,
and a complete audit trail.
The model is identical.
The risk is not.
The reliability is not.
The blast radius is not.
The governance is not.
This is why agent architecture is starting to matter so much.
The model decides what it wants to do. The harness decides what it can do.
A model might decide:
I should delete these duplicate customer records.
That can be a model-level decision.
But whether that thought turns into actual records being deleted is a system-level question.
A well-designed agent might require:
model proposes action → policy check → authority check → human approval → short-lived privilege → execution → audit event
The model's intelligence has not changed.
Its authority has.
That distinction is fundamental.
Capability is not authority.
An agent may be technically capable of doing something without being permitted to do it.
And good architecture should make that distinction explicit.
Tool design may matter more than prompt design
The first generation of generative AI made us obsessed with prompts.
Agents shift more of the engineering problem toward tools.
Consider the difference between exposing this capability:
execute_sql(sql)
and exposing:
get_customer(customer_id)
list_open_orders(customer_id)
update_shipping_address(customer_id, address)
The first tool gives the model broad general-purpose authority over a database.
The second approach gives it narrowly defined business capabilities.
That is a huge difference.
OWASP's current agent-security guidance identifies tool abuse, privilege escalation and excessive permissions as major risks in agentic systems.
Its recommendations increasingly look like traditional security architecture: minimum tool sets, scoped permissions and explicit authorisation for high-impact actions.
That leads to a useful principle:
Good agent architecture turns general intelligence into narrow authority.
Memory is part of the security model too
Memory sounds like a product feature.
The agent remembers what happened last week.
The agent knows your preferences.
The agent remembers the state of a project.
But persistent memory is also trusted state.
If malicious or incorrect information enters that memory, the effect may survive long after the original interaction is gone.
That means memory needs its own architecture:
scope,
ownership,
expiry,
validation,
integrity,
and separation between users and tenants.
The model cannot solve all of that by itself.
The surrounding system has to.
Identity may become more important than intelligence
Traditional software has always had an identity problem to solve.
Who is the user?
Which role do they have?
What are they allowed to access?
Agents make this more complicated because they can operate semi-independently across multiple systems.
Microsoft's recent guidance argues that agents should increasingly be treated as first-class principals with their own managed identities, explicit roles and tightly scoped permissions.
That creates a very different enterprise question.
Instead of asking:
Which model does this agent use?
we may increasingly ask:
What identity does this agent have, and what is it actually allowed to do?
That tells us far more about how the system will behave.
The execution environment defines the blast radius
The same coding agent can operate in two very different environments.
One might execute in a temporary sandbox with no production access.
Another might run on a developer's machine with SSH keys, cloud credentials and network access to internal systems.
Same model.
Very different consequences when something goes wrong.
This is why sandboxing, isolation and controlled execution environments are becoming core pieces of agent architecture.
The useful question is not only:
What can the model reason about?
but:
Where can it execute, and what does that environment expose?
The harness controls reliability as well as security
This is not only a security argument.
The surrounding system also determines whether an agent works reliably.
Suppose the agent says:
I need the customer's latest order history.
The model may understand the intention perfectly.
But the harness still determines:
which customer system to call,
which tenant to use,
which identity to run under,
which version of the data is authoritative,
what happens if the API times out,
whether retries are allowed,
how many retries occur,
whether another data source is acceptable,
whether the response must match a schema,
and whether the final answer must cite where the information came from.
Those are not model-quality questions.
They are systems-engineering questions.
For many production use cases, a slightly weaker model inside a well-designed harness may produce a better outcome than a stronger model surrounded by poor tools, weak permissions and unreliable data access.
The model still matters
There is an obvious danger in pushing this argument too far.
The model still matters enormously.
A better model may reason more effectively, follow instructions more reliably, use tools more accurately, recognise ambiguity, recover from errors and resist some forms of manipulation.
No amount of architecture can turn a fundamentally incapable model into a highly capable reasoning system.
So the argument is not:
models no longer matter.
It is:
once models cross a sufficient capability threshold, the architecture around them may matter more to production outcomes than small differences between leading models.
That is a much more important shift.
The attack surface is moving outward
OWASP's recent work on agentic systems reflects this change.
Its emerging risk categories increasingly include:
tool misuse,
identity and privilege abuse,
memory poisoning,
insecure agent-to-agent communication,
cascading failures,
and rogue agents.
Those are mostly architecture problems.
They are not simply cases where a language model generated the wrong sentence.
As agents become more capable, the security discussion is moving from:
what might the model say?
toward:
what can the whole system do?
This may change where competitive advantage sits
If frontier model capability continues to converge, differentiation moves outward.
Two companies can use the same model and still build very different products.
One may have better:
domain knowledge,
skills,
tools,
workflows,
memory,
identity integration,
governance,
evaluation,
observability,
and access to proprietary data.
That means the model provider supplies an intelligence primitive.
The application builder determines what that intelligence can actually accomplish.
This is not very different from earlier computing transitions.
Cloud infrastructure becoming standardised did not make software architecture irrelevant.
It made applications, data and workflow more important.
AI may be heading in the same direction.
MCP reinforces the same idea
The rise of the Model Context Protocol is another sign of this transition.
MCP creates a standard way for an AI system to discover and interact with external resources and tools.
That is important because it separates reasoning from capability.
The model does not need to understand the internal implementation of every CRM, CMS, database or service.
It needs to understand what resources and tools are available and how they can be invoked.
But MCP does not magically solve governance.
It makes the design of those boundaries even more important.
Which tools are exposed?
What permissions do they have?
Under whose identity do they execute?
Which operations require approval?
What gets logged?
Again, the interesting engineering problem is increasingly outside the model.
The enterprise AI conversation may start sounding much more ordinary
Today, an AI architecture discussion often starts with:
We're using GPT-X.
In a few years, a more useful description may sound like:
The agent runs under a managed identity, has read-only access by default, receives task-scoped credentials, loads approved skills dynamically, executes code in isolated environments, requires approval for high-impact operations, records structured audit events and can route different tasks to different models.
The first statement tells me which intelligence engine is being used.
The second tells me how the system actually works.
That is probably where enterprise AI is heading.
The model is becoming only one layer
For the first few years of generative AI, the model often felt like the product.
With agents, that is changing.
The model provides intelligence.
The surrounding system determines what that intelligence knows, remembers, can access, is allowed to do, where it can execute and what happens when it gets something wrong.
So perhaps the most useful question is no longer:
Which model are you using?
It is:
What have you built around it?
The model still matters.
But increasingly, the agent is everything we build around it.
Sources
Australian Signals Directorate: Agentic AI harnesses
Australian Cyber Security Centre: ASD releases new guidance on agentic AI harnesses
Microsoft Security: Least privilege for AI agents, identity, access and tool binding