XALA
Book a call
Note
10
Published
Reading time
6 min
Filed under
Pricing

What a customer-service AI costs to run, and what it costs to rent.

By the Xala studio · San Diego and Tijuana

Anthropic's own price list puts a support conversation under half a cent. Helpdesk vendors charge $0.99 to $2.00 per resolved one. The gap is where your decision sits.

Picture a tile distributor with twenty-two people. Its helpdesk vendor emails to say the AI agent is now switched on for the account, and every conversation the agent resolves costs $0.99. The owner figures about 240 chats a month are the kind a bot could close, multiplies on her phone, gets $237.60, and decides that is fine. She is probably right. What she has not asked is what that $237.60 is made of, or what happens to it when the chats double, or when the technology underneath gets cheaper.

Two prices sit behind every AI customer-service product. One is what the AI model costs to run. The other is what the vendor charges you per resolved conversation. Different companies set them and different pages publish them, and since August the first one has been moving in public.

What the model costs, and what changed since August

Anthropic, the company behind Claude, publishes its prices per million tokens, and a token is about three-quarters of a word. Its pricing page carries a worked example for support: about 3,700 tokens per conversation, on Claude Haiku 4.5 at $1 per million tokens in and $5 out, comes to roughly $37 for 10,000 tickets. That is $0.0037 a conversation, under half a cent.

The release notes show the direction since then. On August 10, Anthropic made the introductory price of Claude Sonnet 5 ($2 in, $10 out) permanent and cancelled the rise to $3 and $15 planned for September 1. On September 22 it launched Claude Opus 5.5 at $4 and $20, against $5 and $25 for Opus 5. On October 7 it halved the price of cache reads on Claude Sonnet 5.5 to $0.10 per million tokens, down from the earlier $0.20, and launched Claude Haiku 5.5 at $0.10 in and $0.50 out, a tenth of Haiku 4.5's list price for prompts up to 100,000 tokens.

Cache reads matter more than they sound. A support bot rereads your return policy, price list and shipping terms before it answers each customer, and a cache read is the discounted rate for rereading text it already processed a moment ago. For a bot that answers from the same twenty pages all day, that is the line on the invoice that adds up.

Two caveats apply. Anthropic notes that models from Claude 4.7 onward use a newer tokenizer that produces about 30% more tokens for the same text, so comparing list prices across generations overstates the drop. And the $37 example is a short chat. A bot that searches a long knowledge base and calls your order system uses more. Even at ten times the example, 10,000 conversations cost about $370.

The cuts reach different buyers at different speeds. A bot built directly on a model picks up a price cut the day Anthropic publishes it. A per-resolution contract keeps its price until someone renegotiates it. The vendor pages cited here do not say whether savings at the model level reach the customer, which is the reason to ask.

What the meter charges

Intercom's pricing page sells its Fin agent at $0.99 per outcome. It counts an outcome when a customer confirms the issue is resolved, or does not ask for more help after Fin responds, or Fin completes a workflow, including handoffs. You are charged once per conversation, and a minimum monthly commitment applies (the page gives 50 outcomes as an example).

Zendesk moved to automated-resolution tiers on May 18, 2026, according to its help center. Its Mexico pricing page lists US$1.50 per automated resolution on committed volume and US$2.00 pay-as-you-go. A conversation that ends with a human agent does not count against the allowance. Existing accounts get 30 days' notice before they move to the new resolution-allowance model, and the help article says its example prices are placeholders and sends you to sales for the real ones.

The committed rate is 25% below pay-as-you-go, and that discount exists only if you reach the volume you committed to. Ask whether resolutions you commit to and do not use expire at month end, because a shop that commits to 300 and gets 180 may be paying for 120 nobody used. Zendesk also prices by tier, so the figure on your quote can differ from the one on the page. Treat the page as a starting point for the conversation with sales.

Run the tile distributor's 240 resolutions through each. Intercom: $237.60 a month. Zendesk committed: $360. Zendesk pay-as-you-go: $480. Anthropic's example rate for the same 240 conversations: about 89 cents. At 2,000 resolutions a month the vendor lines become $1,980, $3,000 and $4,000, and the model line becomes $7.40.

What the gap pays for

The distance between $0.0037 and $0.99 buys things the token price leaves out: the inbox the chat lives in, the lookup into your help articles, the handoff to a person, the reporting screen, and a company to call when the bot tells a customer something wrong. Few small shops want to assemble those parts, which is a fair reason to pay for them already assembled.

Your own labor gives the second yardstick. Suppose a chat takes someone at your counter six minutes and that person costs you $25 an hour including payroll taxes. That is $2.50 a chat, so $0.99 looks cheap. It stays cheap only while the bot closes the chat correctly. A wrong answer costs the original six minutes plus the call to repair what the bot told the customer. Put your own minutes and your own hourly cost into that sum before you compare it with anyone's pitch.

Volume decides whether the price stays fair. At 240 resolutions a month the vendor bill is about $2,850 a year, and a custom build would have a hard time beating it. At 2,000 a month it is $23,760 a year, and a one-time build plus a small monthly run cost earns a quote worth comparing. Xala builds customer chatbots, so weigh our view of the build option with that in mind. The arithmetic works the same whoever does the building.

Then there is the word "resolved." The vendor defines it, and the vendor bills on it. Under Intercom's definition, a customer who gives up and closes the window looks the same in the data as one who got an answer, unless something else checks. Zendesk adds that something: an automated verification step, and conversations that fail it do not count against your allowance. Both approaches can be fair. You should still read a sample before you trust either.

One more local detail. Zendesk's Mexico page quotes the resolution price in US dollars. A Tijuana shop that pays its bills in pesos carries the exchange rate on top of the volume.

Four questions before you sign

Write down the yearly figure from the first question before the sales call. A shop that knows its number is $2,850 will negotiate differently from one that knows it is $23,760, and a shop that has never counted will accept whatever the first invoice says. Ask two vendors for a quote on the same 240 chats. The differences in how each one defines a resolved conversation will tell you more than the difference in price, because the definition decides how many conversations you get billed for.

Questions are welcome: hola@xala.studio

Talk to us

Want to talk it through?

One paragraph about your business is enough. Tell us what you need → Or just email hola@xala.studio.

Tell us about your project
a real person
reads it