What It Really Costs to Run a Custom AI Agent
Question — — by Mahmoud Zalt
The full bill for a custom AI agent: tokens, tools, hosting, build time and the maintenance hour nobody budgets for. With the formula to work out yours.
Everybody asks the same question first: how much will the API cost me. It is the easiest number to find, which is why it gets all the attention, and it is usually the smallest thing on the invoice.
The costs that actually hurt do not arrive as an invoice at all. They arrive as your Saturday, as an evening spent working out why the agent went quiet on Thursday, and as the slow realisation that this thing now needs a person to care about it every month.
So the useful way to price an agent is not per token. It is total cost of ownership over twelve months, with your own time priced at what you would charge a client for it. Do that once and most build or hire arguments answer themselves in about ten minutes.
It is also why Sistava prices the whole thing as one subscription. The model spend, the schedule, the limits, the memory, the run history and the upkeep are inside it, so there is no separate bill to forecast and no maintenance hour to find. You describe the job, connect the accounts, and hire an AI Employee. Building your own is still the right call for some people, and the numbers below will tell you if you are one of them.
At a Glance
- 5
- Cost lines in a real agent budget
- 1 hour
- Monthly upkeep per live agent, forever
- 10x
- Spread between a careless and a careful loop
- 12 mo
- The window that makes the maths honest
What are the actual cost lines of a custom AI agent?
Five, in descending order of surprise. Build time, ongoing maintenance, third party tools and hosting, model tokens, and the cost of being wrong. Only the fourth one appears on a dashboard, and it is the one people spend all their planning on.
Price each in the same currency, which is money per year with your time included. An agent that saves four hours a week and costs six hours a month to keep alive is still a win. An agent that saves twenty minutes a week and costs the same is a hobby, and there is nothing wrong with hobbies as long as you know which one you have.
- Build time. Usually the largest single line. A demo is an afternoon. Something you trust unattended is three to six weeks of part time work for one well defined job, and most of that is triggers, limits, logging and error handling.
- Maintenance. About an hour a month per live agent, forever, because tokens expire, forms gain fields and APIs change versions. This one never ends, which makes it the biggest number over a year.
- Tools and hosting. Somewhere for the agent to run, a queue or scheduler, a log store, and any paid API it calls. Individually small, collectively the line that quietly doubles.
- Model tokens. Real, forecastable, and for most small business jobs the smallest of the five. It only becomes the big line at genuine volume or with a badly shaped loop.
- The cost of being wrong. One wrong message to a customer, one duplicate refund, one week of silence nobody noticed. Rare, uncapped, and the reason guardrails are not optional.
How do you forecast the token bill before you build?
Three numbers multiplied: runs per month, tokens per run, and your model's price per million tokens. Runs you already know. Price is published. Tokens per run is the one people guess badly, and it is the one that decides everything.
Tokens per run is not the length of the question. It is the length of every step, because an agent resends the whole growing conversation on each turn. A task that takes eight tool calls does not cost eight times one call, it costs closer to the sum of a conversation that gets longer each time.
Measure it instead of estimating it. Run the agent ten times on real inputs, read the token count for each, take the worst one and add half again. That is your per run figure. Multiply by monthly runs and you have a forecast you can actually defend, usually within a factor of two.
A worked example: Yusuf and the maintenance inbox
Yusuf manages about ninety rental properties with two staff. Tenants report maintenance issues by email in every possible format. His agent reads each one, works out the property, classifies the urgency, and either raises a job with the right contractor or asks the tenant one clarifying question.
Volume is roughly four hundred emails a month. Each run is short: one email in, a lookup, a decision, one action. When he measured it, the worst run used a fraction of the tokens a long research task would, and the monthly model spend landed in the range of a couple of takeaway coffees. He nearly stopped the analysis there and concluded it was free.
- The build. Five weekends. Two on the fun part, three on duplicates, retries, the contractor allow list and the logging. Priced at his own hourly rate this dwarfed everything else in year one.
- The infrastructure. A small always on server, a managed queue and a log store. Each one modest, together a recurring line about as large as the model spend.
- The maintenance. Two real interruptions in six months, plus a monthly read of the log. Averaged out, close to an hour a month, exactly as predicted and exactly as annoying.
- The near miss. In month two a badly formatted email sent the agent into a loop that re read the same thread over and over. The step limit stopped it after twelve. Without that limit it would have been the largest line on the page.
His verdict after six months was that the agent was clearly worth it, and that the token bill was the least interesting fact about it. What made the decision was two staff getting an hour of their morning back and contractors being called within minutes instead of the next afternoon.
What makes an agent bill blow up?
Four things, and none of them are the price of the model. A loop with no step limit, a context that grows without ever being trimmed, a retry that hammers a failing API, and a trigger that fires more often than the work actually appears. Any one of them can multiply a sensible bill by ten in a weekend.
All four have the same defence: caps enforced in code, outside the prompt. A maximum number of steps, a spend ceiling per run and per day, a rule that stops when the same tool is called with the same arguments twice, and an alert the moment any of them trips.
| Cost line | What drives it | How to keep it down |
|---|---|---|
| Build time | Scope creep and starting with a vague job | One task, five bullets, definition of done written first |
| Maintenance | Number of live agents and integrations | Fewer agents doing narrower jobs, one named owner each |
| Model tokens | Steps per run and how fast context grows | Trim old turns into a summary, cap steps, keep tools few |
| Tool and API calls | Retries and duplicate triggers | Deduplicate on trigger ID, back off instead of hammering |
| Hosting | Always on services you only need hourly | Run on a schedule, not a permanently idle machine |
| Being wrong | Unreviewed actions that reach customers | Approval gate on anything a customer will read |
| Model choice | Using the largest model for sorting and tagging | Small model for classification, large one only for judgment |
What does the cost not buy you?
Spending more does not buy reliability. A bigger model reduces some mistakes and changes none of the failures that actually take agents down, which are expired credentials, duplicate triggers, changed forms and silence. Those are fixed with engineering, not with budget.
It also does not buy the things agents simply do not do. It will not chase a person until they reply, hold a plan together across several days without losing the thread, notice that your policy changed, or take responsibility for a decision. No spend level changes any of that.
And it does not buy an end to maintenance. Whatever you build or hire, something it depends on will change this year. The only question is whether that hour lands on your calendar or someone else's.
Work out your own number in an afternoon
- Count the runs per month — How many times does this task actually happen. Use last month's real count, not the busy month you remember.
- Measure tokens on ten real runs — Take the worst one and add half again. Guessing this number is the single biggest source of wrong forecasts.
- Multiply for the model line — Runs times tokens times your model's published price per million. Now you have the number everyone asks about first.
- Add hosting, tools and any paid APIs — Include the small ones. A queue, a log store and one paid integration together often match the model spend.
- Price your build time and one hour a month — At what your time is genuinely worth. This is the line that decides most build or hire questions, and it is always the one left out.
- Compare over twelve months, not one — A one month view flatters building. A year view includes the maintenance, and that is the view you will actually live in.
When people run that calculation properly, the result is rarely close. Either the job is unusual enough that building is clearly right, or the twelve month total is several times a subscription and the argument ends. The uncomfortable cases are the ones where nobody priced their own time.
FAQ
How much does it cost to run a custom AI agent each month?
For a small business job with a few hundred runs a month, the model spend is typically the smallest line and often lands in the range of a streaming subscription. Hosting, a queue and a log store usually match or exceed it. The largest real cost is the hour a month of human attention, which never appears on any dashboard but never goes away either.
Why is my AI agent bill higher than I expected?
Almost always the loop, not the price. Agents resend the whole growing conversation on every step, so a task that takes ten tool calls can cost many times a single call. Add a retry that never backs off or a trigger firing twice and the multiplier compounds. Cap the steps, trim old turns into a summary, and deduplicate the trigger.
What is the biggest hidden cost of a custom AI agent?
Ongoing maintenance. Integrations change, credentials expire, forms gain fields, APIs move versions. Budget roughly an hour per live agent per month, forever, and give it to a named person. Teams that skip this line do not avoid the cost, they just pay it as an emergency at the worst possible moment.
Is a cheaper model always cheaper to run?
No. A weaker model often needs more steps to reach the same answer, and every extra step carries the whole conversation again, so the cheap model can cost more per completed task. Judge cost per finished job, not price per million tokens. The useful pattern is a small model for sorting and tagging and a strong one only where real judgment is needed.
How do I reduce the token cost of an AI agent?
Four moves, in order of impact. Cap the number of steps per task. Fold old conversation turns into a short structured summary instead of resending everything. Cut the number of tools, because every tool description sits in every single call. And send only the fields the model needs from a record rather than the whole record.
Do AI agent costs grow as my business grows?
The token line grows roughly with volume, which is fair and forecastable. The maintenance line grows with the number of agents and integrations, which is the one that catches people out. Ten narrow agents cost far more attention than one agent doing ten steps, so consolidate jobs before adding new agents.
The takeaway is not that agents are expensive. Most are not. It is that the expensive part is never the part with a published price list, so a budget built only from token costs is a budget that will be wrong by month three.
Count all five lines, over twelve months, with your own time in there at a real rate. Then decide.