A few months ago, if you had asked me what came to mind when I heard the words AI agent, I probably would have talked about automation.
Agents that can research prospects.
Agents that can write code.
Agents that can qualify leads, answer support tickets, update CRMs, browse the internet, make decisions, and execute workflows without someone manually telling them what to do at every step.
Basically, I was thinking about capability.
Then I started working at a startup building liability insurance for AI agents.
And suddenly I started thinking about a very different question:
What happens when the agent is wrong?
Not wrong in a benchmark.
Not wrong in a ChatGPT-answer-is-a-little-weird kind of way.
Wrong after it has actually done something.
That distinction has changed the way I think about AI quite a bit.
AI gets more interesting when it can take actions
For most of the last few years, the way people interacted with AI was relatively simple.
You gave a model some input.
It gave you some output.
You could accept it, reject it, edit it, or ignore it.
The human was still sitting between the AI and most real-world consequences.
Agents start to change that.
An agent can potentially decide which customer to contact, send the message, change something in a database, interact with another piece of software, trigger a workflow, or perform an action on behalf of a company.
That is obviously what makes agents powerful.
But it is also what makes them risky.
There is a pretty big difference between:
“AI suggested the wrong action.”
and
“AI took the wrong action.”
Once you start thinking about that difference, a whole category of questions appears.
What happens if an agent sends information to the wrong person?
What happens if it makes a decision based on incorrect data?
What happens if an automated action causes financial loss?
Who is responsible?
The company using the agent?
The company that built it?
The underlying model provider?
Some combination of them?
I used to mostly look at AI systems through the lens of what they could automate.
Working around AI insurance has made me look at them through another lens:
What new liability exists because this thing can now act?
Reliability is not the same thing as responsibility
One thing that has become much clearer to me is that making an AI system more reliable does not eliminate the question of responsibility.
You can add evaluations.
Guardrails.
Human-in-the-loop approvals.
Observability.
Permission systems.
Better models.
More testing.
All of those things matter.
But no system becomes perfect.
And at some point, if AI agents are doing meaningful work inside real companies, some of them are going to fail in ways that have actual consequences.
This seems obvious when you say it out loud.
But most of the AI conversation I see is still disproportionately focused on making agents more capable and more reliable.
There is another layer that becomes important after that:
What happens when reliability fails?
Traditional software has had decades to build systems around this problem.
Contracts.
Security standards.
Compliance.
Liability.
Insurance.
Enterprise procurement.
AI agents are now beginning to collide with that entire world.
That collision is fascinating.
The boring questions are sometimes the important ones
There is something slightly funny about working in a category like AI insurance.
The AI ecosystem loves questions like:
How autonomous can agents become?
Insurance makes you ask questions like:
Who pays when one breaks something?
The second question is considerably less exciting.
It might also become extremely important.
The technology can move incredibly fast.
The organisations adopting it usually cannot.
And if you are selling into businesses, that gap matters.
A founder might look at an AI agent and think:
This can save us hundreds of hours.
Someone responsible for risk might look at exactly the same agent and think:
What access does this have, what can it do, and what happens if it makes a mistake?
Neither person is necessarily wrong.
They are just evaluating the same technology through completely different incentives.
Understanding those incentives has probably taught me as much about GTM as it has about AI.
GTM gets harder when the category barely exists
Another interesting part of working here is that you cannot always rely on an existing playbook.
If you sell CRM software, people generally understand what a CRM is.
If you sell payroll software, the buyer already understands the problem category.
With something like insurance for AI agents, a lot of the GTM work happens one layer earlier.
Before convincing someone that your solution is good, you sometimes have to understand whether they even think the problem exists yet.
That changes how I think about selling.
The question is not immediately:
Who needs AI agent insurance?
It is more like:
Who is already deploying agents with enough autonomy that risk has become a real operational concern?
That leads to better questions.
What actions can their agents perform?
What systems can they access?
How much human approval exists?
What could actually go wrong?
What would the financial consequence be?
Has a customer or enterprise buyer already started asking them about liability?
Suddenly prospecting becomes less about finding companies with the right job title and more about detecting signals of a problem becoming real.
I think this is one of the most interesting parts of GTM engineering.
You are trying to translate a vague market thesis into something observable.
Strategy says:
Companies deploying autonomous AI will eventually need ways to transfer some of that risk.
The GTM engineering question becomes:
Cool. What data would tell us that a particular company is getting close to that moment?
That is a much more interesting problem.
The best ICP might be defined by behaviour, not industry
This has also changed how I think about ICPs.
It is very tempting to define an ICP using firmographics:
Company size.
Industry.
Funding.
Geography.
Employee count.
Those things matter.
But for emerging categories, I am increasingly convinced that behavioural signals can be much more useful.
Two companies can look almost identical in a database.
One is experimenting with an internal chatbot.
The other has agents interacting with customers and taking actions inside production systems.
From a traditional firmographic perspective, they might be the same prospect.
From a risk perspective, they are completely different.
So the more interesting ICP attributes become things like:
- Is the company actually deploying agents?
- Are those agents customer-facing?
- Can they take actions autonomously?
- Do they interact with sensitive data?
- Are they being sold into enterprises?
- Are customers asking questions about security, risk, or liability?
- Is the company moving from experimentation to production?
This is something I am still learning, but it has made me much more sceptical of ICPs that are basically just a collection of filters inside a sales database.
Sometimes the strongest qualification signal is simply:
What is this company actually doing?
Every layer of the AI stack creates another layer of trust
There is another mental model I have started developing.
AI adoption is not just a capability problem.
It is also a trust problem.
A company might believe an AI agent can perform a task.
That does not automatically mean they are comfortable letting it perform that task autonomously.
The larger the consequence of failure, the higher the amount of trust required.
And trust seems to get built in layers.
Better models create confidence in capability.
Evaluations create confidence in performance.
Observability creates confidence that you can understand what the agent is doing.
Guardrails create confidence that you can limit what it can do.
Security creates confidence around access and data.
Insurance potentially creates confidence around the remaining financial risk.
None of these replace each other.
They stack.
And I suspect that as agents become more autonomous, this entire trust infrastructure around agents could become almost as important as the agents themselves.
That is probably the biggest idea I have taken away from working in this space so far.
The AI economy will need more than AI companies
There is a tendency when a new technology appears to focus entirely on the companies building the technology itself.
Models.
Agent frameworks.
Developer tools.
Infrastructure.
Applications.
But every major industry eventually develops another ecosystem around it.
The companies that secure it.
Monitor it.
Audit it.
Finance it.
Regulate it.
Insure it.
Working at Ollive has made me much more interested in that second-order ecosystem.
If AI agents genuinely become a new form of digital labour, there will probably be an enormous amount of infrastructure required to make companies comfortable delegating increasingly important work to them.
Insurance is one tiny piece of that puzzle.
But the broader category is something I find fascinating:
What needs to exist around AI before society is comfortable depending on it?
That feels like a much bigger question than whether the next generation of agents gets better benchmark scores.
My current mental model
I still think AI agents are fundamentally an automation story.
But I no longer think that is the whole story.
The first phase is:
Can AI do the work?
Then:
Can AI do the work reliably?
Then:
Can we trust AI to do the work autonomously?
And eventually:
What happens when it gets the work wrong?
The companies building models and agents are pushing hard on the first two questions.
An entire ecosystem is now beginning to form around the last two.
Working inside one small corner of that ecosystem has changed the way I look at AI.
I spend a little less time thinking about demos now.
And a little more time thinking about consequences.
Because if agents are eventually going to do real work, make real decisions, and control real systems, then the next chapter of AI probably isn't just about making them more powerful.
It is about building enough trust around them that we are actually willing to let them use that power.
Comments