Operations · 7 min read
The Real Cost of AI in Business Operations: Where the Savings Come From (and Where They Don't)
A practical breakdown of what AI systems actually cost to build and run — model pricing, hidden integration costs, and the difference between savings that show up and savings that don't. With a simple ROI method you can do in an afternoon.
By Kartik, Founder, Stacktree · · Updated
Key takeaways
- AI in operations has three cost buckets: build once, run monthly, and change over time. Most bad estimates forget the third.
- Model usage is priced per token; for a typical service-business workload it's a small line item next to the value of a single recovered booking.
- The savings that materialize are recovered revenue and replaced tool sprawl. The savings that don't are 'we'll need fewer people' on a process that was never defined.
- A fixed-quote build with a separate, stated monthly run rate is the only honest way to price this.
Three buckets, not one number
When a business owner asks "what does AI cost?" they usually get one number back, and that number is usually the build. It's the least interesting of the three. Any AI system in operations has:
- 01Build cost — one-time. Design, engineering, integration with your calendar/CRM/telephony, testing, training, handover. This is what a studio quotes.
- 02Run cost — monthly. Model usage (tokens), telephony minutes, hosting, database. Scales with conversation volume.
- 03Change cost — ongoing. Prices change, a new service launches, a policy shifts, a model gets deprecated. Someone has to update the system. If nobody does, it quietly rots.
The build is usually the largest single figure and the least important to your decision, because it's a one-off. The run cost is what you should understand deeply. The change cost is what most people forget and later resent. A good quote states all three separately. See how we structure that in what an AI software studio does.
How model pricing actually works
Language models are priced per token — roughly three-quarters of a word. You pay for tokens going in (your prompt, the conversation so far, the retrieved knowledge) and tokens coming out (the model's reply). Output tokens cost more than input tokens. Bigger, more capable models cost more per token than smaller, faster ones. Public price lists from Anthropic and OpenAI put current numbers on this; the exact figures change often, the structure doesn't.
Three things matter more than the headline per-token price:
- Model tiering. A well-built system doesn't run everything on the most expensive model. Routing a text message to the right handler, extracting a date, classifying intent — these run on small fast models at a fraction of the cost. The frontier model is reserved for the reasoning that actually needs it.
- Prompt caching. Most of every request is the same system instructions and the same knowledge base. Providers now let you cache that repeated context at a steep discount (Anthropic's prompt caching is one example). A system that isn't using caching is overpaying, often by a lot.
- The trend. The Stanford AI Index has tracked the cost of a fixed level of model capability falling by orders of magnitude since 2022. Whatever you're quoted for run cost today, the same workload will cost less next year. Build cost does not follow that curve, which is another reason to treat it as a one-off.
A worked example
Let's make it concrete with an illustrative case: a plumbing company that misses after-hours calls. (These figures are for demonstration — plug in your own.)
| Input | Value | Notes |
|---|---|---|
| After-hours calls per month | 120 | From the phone system's missed-call log |
| Share that are real jobs | 40% | The rest are wrong numbers, spam, existing-customer questions |
| Share of real jobs currently lost | 70% | Voicemail left, no callback by morning, they called the next company |
| Average job value | $380 | Company's own average ticket |
| Jobs recovered by AI answering | 60% of lost | Conservative — not every caller books even with a great answer |
Run the arithmetic: 120 calls × 40% real = 48 real jobs. 70% lost = 34 lost jobs. Recover 60% = about 20 jobs a month. At $380, that's roughly $7,600 a month in recovered revenue from a single fix. Set that against a run cost that's mostly telephony minutes and a few dollars of tokens, and a build that's a one-time figure, and the payback period is short — usually within the first quarter.
Notice what made the number big: it wasn't a clever model. It was a process (after-hours calls) with clean inputs (a phone rings), a clear action (book it), and a measurable outcome (a job). That's the pattern. AI pays for itself fastest where the process is already well-defined and the only thing missing is someone to do it at 9pm.
The hidden costs
This is the section most vendors skip. These costs are real, they're usually not in the quote, and they're where projects go over.
- Integration. Your calendar, CRM, phone system, and payment tool each have an API, and each API has quirks. Integration is often half the build. Any quote that doesn't name the systems it's integrating with is guessing.
- Data hygiene. The model is only as good as the knowledge it's grounded in. If your service list is in three places and two are out of date, the first week of the project is reconciling them. This is unglamorous and unavoidable.
- Human review. In the first weeks, someone should read escalations and a sample of conversations. This is how the system gets tuned. Budget the hours.
- Drift and change. Prices change. Models get deprecated (providers publish deprecation schedules — plan for a migration every 12–18 months). A knowledge base that nobody updates becomes a liability. This is the change cost from bucket three, and it's why handover documentation matters.
- Evaluation. How do you know it's working? Someone has to define the metric (booked rate, escalation rate, response time) and look at it. Without this, you're running on vibes.
Where the savings are real
- Recovered revenue. Leads that used to go quiet. This is almost always the largest line and the one that's easiest to measure — compare booked rate before and after.
- Replaced tool sprawl. An answering service, a scheduling tool, an email sequence tool, and a reporting spreadsheet, each with a subscription, often collapse into one system. Add those subscriptions up; the number is usually surprising.
- Time. Not headcount — time. The front desk stops chasing confirmations and starts doing the parts of the job that need a person. Owners stop building Monday reports by hand. This shows up as capacity, which becomes revenue when you're growing and sanity when you're not.
Where the savings aren't
Two claims should make you suspicious.
"You'll need fewer people." Occasionally true, usually not, and rarely where the value is. AI is bad at the parts of service work that involve judgment, presence, and trust, which is most of the job. Businesses that frame the project as headcount reduction tend to build the wrong thing and get a system that generates tickets for the people who are left.
"AI will fix the process." It won't. It will do a bad process faster and more consistently. If nobody can describe today what should happen when a lead comes in, the first deliverable is a written process, and the AI is the second. McKinsey's ongoing State of AI research keeps finding the same thing at enterprise scale: value follows workflow redesign, not tool adoption.
An ROI method you can do in an afternoon
- 01Pick one process. The one that loses the most money or eats the most hours. Not five.
- 02Count the inputs for a month. Calls, forms, texts — whatever comes in. Get the real number from a log, not a guess.
- 03Estimate the leak. What share is currently lost, delayed, or handled badly? Be conservative.
- 04Put a value on a recovered unit. A booked job, a signed client, an hour of front-desk time.
- 05Multiply. Recovered units × value = monthly gain. Compare to the run cost and the build cost. If payback is under six months and the process is well-defined, build it. If not, fix the process first or buy something off the shelf.
That's the whole method. It fits on one page, and it's exactly what we do on a Stacktree scoping call before we quote anything. If you'd like us to run it with you, book the call — it's the discovery step and there's no cost.
Frequently asked questions
How much does it cost to run an AI receptionist per month?
For most small service businesses, the monthly run cost is dominated by telephony minutes and hosting, with model usage a small line item — typically far less than a single recovered job. Exact figures depend on call volume and which models the system uses, which is why a proper quote states the run rate separately from the build.
Are AI costs going up or down?
Per-unit model costs have fallen sharply and consistently since 2022, and the Stanford AI Index tracks this trend annually. Build costs — engineering time, integration — haven't fallen at the same rate, which is why the one-time build is the figure to scrutinize and the run cost is the one to understand.
What's the biggest hidden cost in an AI project?
Integration and data hygiene. Connecting to your existing calendar, CRM, and phone system is often half the build, and getting your services, prices, and policies accurate and in one place is the unglamorous first week of nearly every project.
Does AI reduce headcount?
Rarely, and it's usually the wrong goal. The measurable gains come from recovered revenue (leads that no longer go quiet), replaced subscriptions, and time returned to staff for work that needs a human. Projects framed as headcount reduction tend to build the wrong system.
Sources and further reading
- 01Pricing — Anthropic (2026)
- 02API pricing — OpenAI (2026)
- 03Prompt caching — Anthropic Developer Docs (2025)
- 04AI Index Report — Stanford Institute for Human-Centered AI (2025)
- 05The state of AI — McKinsey & Company (2025)
Builds mentioned in this article
Keep reading
Stacktree is the AI software studio behind this article. Book a scoping call or explore the live builds.