A small, senior team building what your business needs.
Each of these does one job, the same way every time, and someone checks the result before it goes out.
New leads come in from the form, the inbox, and whatever enrichment tool you pay for. It scores each one against your criteria, routes it to the right person, and drafts a first reply.
Pulls the fields out, checks them against your rules, and tells your reviewer what's off and why.
Policies, contracts, specs, whatever's in the wiki. It answers in plain language and links to the source. If it doesn't know, it says so.
Pulls the numbers from the tools you already use and writes them up the way you would, if you had the time. Flags what actually moved.
Reads the ticket, pulls up the account, and drafts a reply grounded in your help docs, not a guess.
Prices, stock counts, customer records: the stuff two systems never quite agree on. It reconciles them on a schedule and tells you when something doesn't line up.
Monthly billing runs, exports, backups: not hard, just tedious and easy to forget. It runs on schedule and tells you if it fails.
If it's close to one of these, we'll price it in the same week. If it's genuinely different, tell us.
Contact usWhat actually slows a build down isn't the model, it's whether your data is clean and accessible before we start. That's why we front-load discovery, so you find out in week one instead of week four.
A day with the people who do the work, followed by a real look at your data: what exists, who owns it, how clean it is. Most timeline overruns get decided right here, not in the build.
We work in a repo you have access to from day one. Your engineers see the code, ask questions, and stay in the loop the whole way. Nothing is a black box you inherit later.
The agent does the job for real, but nothing it produces goes out without a human approving it. Meanwhile we measure: how often it's right, how often it escalates, how often a reviewer changes something. Guesswork becomes a number.
With real evidence from the validation weeks, you decide how much rope to give it. Then we hand over the repo, the runbook, the test suite, and the on-call notes, and we train the person who owns it after we're gone.
The one thing that actually moves this timeline: how ready your data is when we start. Clean, accessible data gets you the four-week end. Data nobody's touched in years gets you the eight-week end, and we'd rather tell you that in week one than let you find out in week five.
Prompt injection, confidently wrong answers, and an agent quietly doing more than it was asked: those are the real failure modes security researchers actually track, not vague talk of "AI risk." Each one gets handled before it reaches you.
Language models get confidently wrong, that's the real risk, not bad intent. Every answer points to the document, record, or row it came from. If the agent can't ground its answer, it hands the question to a person instead of guessing.
Prompt injection works by hiding instructions inside the content an agent reads: an email, a document, a support ticket. Ours is scoped to a single task, so anything outside that scope gets a refusal, not a creative attempt at compliance.
Accuracy on a clean demo and accuracy on your real, messy data are different numbers. Roughly 120 real cases from your business, with expected outcomes, run before any change ships and for as long as the agent lives.
We log what it saw, what it concluded, what it did, and who approved it, in your logging stack, under your retention policy, in a format your auditor accepts.
You don't have to choose between "a human checks everything forever" and "we hope it behaves." You start at the left, and you move right when the numbers earn it. Most clients sit at L2 within a quarter. Turn the dial below to see what changes at each level.
Automations and agents get scoped differently depending on the job and the data behind it, so we don't publish a generic rate card. Book a call and we'll work out what it actually costs.
You'll have a fixed number within two working days of that call, before you commit to anything. Model usage is billed by your own provider, separately, so you always know exactly what you're paying for and to whom.
Copilot is a very good assistant sitting next to a person. It waits to be asked, it doesn't know your rules, and it doesn't own an outcome. What we build does one specific job in your business, on a schedule, against your data, with a measured accuracy number and an audit trail. If a person still has to start every task, keep the assistant. If you want the task to arrive already done, that's a different thing.
Probably, yes: the first 80% in about two weeks. The last 20% is what takes six months: the evaluation harness, the refusal behaviour, the escalation logic, the observability, and the boring work of finding the forty ways it breaks on real data. We've paid for that learning already. And we build it in your repo, so your engineers own it afterwards rather than becoming dependent on us.
It happens. The question is what it costs. At L0 and L1 a person catches it before it leaves the building, which is exactly why we start there. Then the case goes into your test suite so that specific failure can't recur. What we won't tell you is that it never hallucinates: anyone who tells you that is either lying or hasn't shipped anything.
A named engineer, response times agreed upfront, and access to the same test suite we built the agent against. Most issues get caught by the tests before they ever reach you.
Not yet: those programmes are underway. In the meantime, every build ships with its own test suite, audit trail, and a security walkthrough before you sign anything, so your team isn't taking it on faith.
Same team, two ways in. The platform is agents we run for you, hosted in the EU, from €39 a month: the fast path if you want reporting and marketing agents running this week. Build is for when the work is specific to your business and nothing off-the-shelf quite fits. Most clients end up using both.
That's exactly the gap that kills most builds. Demo cases are cherry-picked and clean, production data is not, and accuracy drops when a system hits it. The validation weeks exist for exactly this: the agent runs on your actual data before anyone acts on its output, and we measure the real accuracy number instead of trusting the demo one.
Yes. It's built in a repo you have access to from day one, not a proprietary platform you'd have to reverse-engineer to leave. If you stop working with us, you keep the agent, the code, and the test suite exactly as they are. You lose our updates and support, not the thing itself.
Before we talk about price or a timeline, we want to hear what you're actually trying to fix. That's what the first call is for. You walk us through the job, how it works today, and where it tends to break, and we listen for whether an agent or an automation would actually help, or whether we'd just be handing you a new system to maintain.
If it's not a good fit, we'll say so on that same call instead of stringing you along into a proposal.
We reply within one working day. If your security team needs a questionnaire filled out first, send it over, we'd rather deal with that early than have it stall things later.