An AI agent that asks before it acts.
One working system, live in your business in two weeks. It answers from your own records, stops at anything it should not decide on its own, and comes with the code and a written runbook so it runs without us.
$5,400
One build, one price, agreed in writing before anything starts.
The monthly is optional, it is not a licence, and nothing switches off without it.
We do not sell a small one.
There is no cut-down version of this, and we would rather tell you why than have you wonder.
A single agent answering one kind of question from one system is not worth buying from us. The software you already pay for is growing one — by this year, four in ten business applications ship their own. And there are firms who will build you a standalone one in about a week for around $1,500.
If that is what you need, one of those is the better buy, and we will say so.
What the $5,400 is for is the part neither of them does: an answer assembled from systems that do not talk to each other, and a line it will not cross — agreed in writing before we build, and provable afterwards.
Somebody is answering the same question forty times a week.
Where is my order. Is this candidate cleared. Can we take this client on. The answer is already sitting in a system somebody has to open, read and retype. It is not hard work. It is just work that never stops arriving, and it lands on the person you can least afford to have doing it.
An agent takes the ones that have a definite answer in your records, and answers those. The interesting part is what it does with the rest. Anything it should not decide on its own — money, a health question, a legal one, a customer who is already unhappy — it stops, and hands the whole conversation to a person, with everything it found attached.
That is the entire product. Not a chatbot with your logo on it. A system that knows which questions are its own to answer and which are not, and can prove which it did.
It will not take all of them. Across this category somewhere between a third and a half of conversations still end up with a person, and any figure that says otherwise is being measured generously. What changes is which half your people spend their day on.
The biggest one ever built went the other way, and came back.
In February 2024 Klarna said its AI assistant was doing the work of 700 people — 2.3 million conversations in its first month, across 23 markets and 35 languages. It was the case study everyone in this industry pointed at, including the people selling against us.
By May 2025 their chief executive told Bloomberg: “We went too far.” The quality had dropped and they began hiring people back.
The part worth reading twice: they did not switch the AI off. It kept the high-volume questions. What went back to people were the complex and sensitive ones — the cases it should never have been deciding in the first place.
That line is the thing we build. Not the answering — the software you already pay for is growing its own answering, and it will be adequate. What it will not do is stop in the right place, hand over everything it found, and let you check afterwards which calls were its own.
Most of these are never used again after the demo.
Roughly nine out of ten of these projects never make it into daily use. That is Gartner’s figure at 89%, and it is not an outlier — Forrester and Anaconda put it at 88%, and a separate survey found 78% of companies running one and only 14% who had rolled it out properly.
The reason is not the model. Asked what stopped them, 64% of leaders named the same thing: they could not tell whether it was working. Governance came second at 57% and the model itself third at 51%.
So the answering is not the hard part and it is not where the two weeks go. The hard part is deciding what it must refuse, then writing enough checks to prove it does — against the ugly cases and the ones that made a person stop and think, not the easy ones.
97–277
Checks behind each of the six systems you can open below. The published bar for shipping one of these is a hundred, and two to four hundred where the work has several steps. What that number does not tell you is whether the checks are any good — they are ours, and we wrote them.
14
Real faults sitting behind 132 passing checks on the first of them, found by looking rather than by running. A green screen is not a working system, which is the whole reason we do not sell you the demo.
Every one of these opens in a browser. No email, no form, nobody follows up.
They are the systems we built, running. Type into them, push them at the edges, and try to make one say something it should not. The records inside are invented so that anyone can try them.
We read eighteen competitor pages line by line. Not one of them lets you try anything — no live system, and no recording of one either. The tools that do sell demos for software products build them out of screen captures, and say why plainly: a live one behaves differently every time. These behave the same way every time. That is the point of building them this way, and it is why we can hand you the keys.
What the two weeks actually are.
The clock starts when we have what we need from you, and it stops whenever we are waiting on you. You will always know which of those is true.
“Couldn’t our own developer build this?”
Almost certainly yes. This is not difficult work, and anybody good enough to be worth hiring is good enough to connect two systems and put a model in front of them. We would rather say that than pretend otherwise.
The question is not whether they can build it. It is who owns it in month seven, when an account expires on the day they are on holiday, and the thing has been quietly answering wrong for a week. In-house builds fail at the boring end, not the clever end.
So the honest split is this. If you have someone who will still be there in a year, who wants to own it, and who has time that is genuinely spare — have them build it, and use the refusals on this page as the checklist. If “spare time” means evenings, or the person who would build it is the person you can least afford to have debugging it at 7pm, that is what we are for.
What happens in month seven, when it breaks.
It will. Something upstream changes a field name, a supplier swaps a format, an account expires. The question is not whether that happens. It is whether anybody finds out before your customers do.
-
It is built to fail loudly, not quietly.
The failure that costs you is the silent one — a system that keeps answering while quietly being wrong. Ours stop and escalate when they are unsure rather than guessing, because one confidently wrong answer undoes months of saved time.
-
Every answer is traceable to the record it came from.
When something looks wrong six months later you can tell whether the logic broke or the underlying data changed. A report nobody can double-check is worse than no report.
-
You get a runbook written for whoever is on duty.
What it does, where it runs, what to check first, and how to switch it off. Written for a person who was not in any of our meetings.
-
Stop any month, and you keep everything.
If you keep us on it is $849 a month. Every week we run every check behind it again and read the result — not because something looks wrong, but because the way these fail is that nothing looks wrong. Teams that test hard before launch and then stop watching see it quietly get worse within a month or two.
We also watch one number most people do not: how often it hands over to a person. When that starts climbing it usually keeps climbing for about a fortnight before anything visibly breaks, which is the closest thing to a warning this technology gives you.
And when the model underneath it is switched off — OpenAI retires GPT-4, GPT-4o, GPT-3.5 and the o-series on 23 October 2026, and the Assistants API on 26 August — we move it and re-run everything, on a schedule you never chose. A model change quietly breaks a few per cent of the checks every time.
You could do all of this yourself. The tools for watching an agent are free — Langfuse and Phoenix cost nothing and run on your own machines. If you have somebody who will genuinely open it every week, keep your $849. This is for the far more common case where everyone means to and nobody does.
You deal with one person, by name, and they answer within one working day. The watching runs at any hour; the human part does not, and we would rather write that down than be caught out the first time we are asleep.
Either way nothing switches off, and there is no notice period beyond the month you are in.
You own all of it.
The code, the credentials and the runbook are yours from the first delivery, not at the end of some term. It runs on your own accounts. There is no licence, no seat count, and nothing that stops working if you stop paying us.
If you want to hand the whole thing to your own developer, or to another firm, you can, and you will not need our permission or our help to do it. We would rather be kept because the work is good than because leaving is expensive.
What we will not build.
These are not caveats. They are rules the systems above are built to, and you can watch the first two of them stop something.
-
We will not let it decide anything that costs money.
A refund, a credit, a price change, a payment release. The agent prepares the decision and a person takes it. Every build above that touches money works this way, and you can watch it stop.
-
We will not have it judge clinical or legal urgency.
Patient-facing triage is a regulated medical device, and telling someone whether they have a case is legal advice. Our clinic build refuses the booking on one side of that line, and on the other it refuses to make the judgement at all. Those are two different refusals.
-
We will not build one where a wrong answer is invisible.
If a mistake surfaces only when a customer complains three months later, an agent makes the problem worse and faster. That is a job for a check, not an agent, and we will say so.
-
We will not quote an accuracy figure.
A single accuracy number hides the half you care about. We publish precision and recall, at pair level and entity level, with what the number does not cover written beside it.
Eight numbers from two of these, and what each one does not tell you.
Birdy Shop
100%
Of the messages that had to reach a person, on sixteen written after the code was frozen. What it does not say: three times in fourteen it fetched someone who was not needed.
9%
The same sixteen messages, before it read them for meaning. Matching words caught one in eleven.
60 hours
A month back at the shop, over more than a year of running. Two people spent the first hour of every morning, seven days a week, on the same order questions; about three in four now answer themselves from the shop’s own records. What it does not say: nobody has counted how often a customer had to ask again.
2
Times a live bird ran late and the fifteen-minute rule fired, over more than a year. A named person took it inside the fifteen minutes both times. For a parcel the same delay is a four-hour handover.
Foundation Legal Advisors
100%
Precision, on 1,128 name pairs across 27 entities in our own fixture. Zero false merges. What it does not say: how many it missed.
86.2%
Recall at pair level; 85.7% by entity. Published per variation class, so a class handled badly cannot hide inside an average.
3
Kinds of hidden link it has found that searching for names would not have reached, over more than a year, from 20 to 100 new enquiries a month — one through ownership, one through a person sitting on two boards, one through a company’s former name. What it does not say: how many times each kind has happened, or how many it missed.
0
Conflicts decided by the system. It runs first, a solicitor still searches, and a solicitor still decides. It rarely raises something that turns out to be nothing.
We publish precision and recall, at pair level and at entity level, and we never publish a single accuracy figure — one number hides whichever half you happen to care about. Each figure above is recomputed inside the build it belongs to, by a test that reads that build’s own published page; this page renders them from that same file rather than retyping them.
Four promises, all four written down.
Of the eighteen firms whose pages we read line by line, not one published a written guarantee of any kind.
Live in two weeks, or the build is free.
The clock starts when we have what we need from you, and it pauses whenever we’re waiting on you.
You own all of it.
The code, the credentials and a written runbook. Yours to use from the first delivery, and yours outright the day the last invoice is paid. It runs without us, and you can hand it to anyone.
Money back if we miss the scope.
A full refund inside seven days, and free fixes for ninety days after that.
We come back on day 7, day 14 and day 30.
Named days, not a support period you have to remember to use. The first week is when real questions arrive and behave differently from the ones we tested with. The second is when the person using it has opinions. The thirtieth is when we find out whether anyone is actually reading what it hands over.
When it is live, somebody has to be watching it.
An agent needs watching for the same reasons anything else does, plus one of its own: the model it runs on gets retired and has to be moved. That is covered from Guard upwards. Same ladder as everything else we build.
Watch
$549 per month
“It runs. I do not want to be the one who notices when it stops.”
- We are told the moment a run fails, and we tell you
- A written note every month of what ran, what failed and what we did
- Broken things fixed by the next working day
- Your accounts stay yours — we hold access, not ownership
Guard
$849 per month
“I need to know about the runs that finish and quietly do nothing.”
- Everything in Watch, plus:
- We check the runs that finish but do nothing, where there is something to count
- Broken things fixed the same working day
- Changes are unlimited — they join a queue and we work through it in order
Managed
$1,200 per month
“I do not want to run it at all. I want somebody whose job it is.”
- Everything in Guard, plus:
- We operate it, not you
- A named person, and the name of their backup
- Reachable outside business hours when something is genuinely down
- Your changes are worked ahead of everyone else’s
Monthly, in advance. Thirty days’ notice either way, from either side. Offered after a build, never before — we will not sell you a plan to watch something we have not seen.
Not sure whether an agent is the right thing at all?
Write to us and describe the problem in your own words — no form, no call. We will tell you which one we would pick, what is good and bad about that choice, where it would be the wrong buy, and what it costs. If the honest answer is that you do not need any of it, we will say that instead.
Two ways to start, and one of them needs nobody.
Open a system and try to break it.
They run on invented records, so there is nothing to sign and nobody follows up. If one says something it should not, we would genuinely like to know.
Tell us what you are doing by hand.
Describe the job somebody repeats every week. We will tell you whether an agent suits it and say so plainly if it does not. Every refusal above came out of building one of these systems, so we already know where the lines are. We reply within one business day.