Insights

"Understaffed, overworked, underfunded? Why don't you just buy a chatbot?"

Christian Schacht
All insights

Many service leaders hear this from above. The tools are everywhere. Making one work is now their problem.

"Understaffed, overworked, underfunded? Why don't you just buy a chatbot?"

The sentence usually arrives in a planning meeting. Understaffed means positions that stay open because nobody expects to fill them this year. Overworked means the same people answer the phone, the email and the escalations, and the queue is longer every Monday. Underfunded means next year's plan holds headcount flat while volume keeps rising. Then someone from the board asks the question that sounds like help. Why don't you just buy a chatbot?

The question is fair. The tools are everywhere, the demos look good, and every vendor has a number ready. What the sentence changes is who owns the outcome. Choosing the tool, rolling it out and proving the savings is now the service leader's job, usually on top of the job they already had.

The questions nobody asks in that meeting

Behind the mandate sit three questions that rarely get said out loud. Will it work? Will it shrink my team further, or become the argument for doing so? Will it help my customers, or will it become one more thing they complain about?

All three are legitimate, and a product demo answers none of them. What follows is my attempt to answer them with the sources I trust, and with a hard look at the metric I trust least.

The number bots are sold on

Most chatbots are sold on one figure, the automation rate. It is usually defined as one minus the share of chats handed over to a human. A bot that hands over 10 of every 100 chats is 90 percent automated.

Read the denominator carefully. It counts chats, not the contacts your team handled before the bot existed. That difference matters, because a chatbot is a new channel and a low-threshold one. People ask it things they would never have called about: a quick check on opening hours, a second look at a delivery date, a test to see whether it works at all. Every one of those chats lifts the automation rate. None of them takes work off your team, because your team never had them.

A simple example shows the mechanics. The figures are illustrative. Before launch, your team handles 1,000 contacts a month, and after launch the bot takes 3,000 chats at 90 percent automation, which leaves 300 handovers. Phone and email do not disappear, so assume 800 contacts still arrive there. Your team now handles 1,100 contacts instead of 1,000. The dashboard shows 90 percent, and the team feels busier because it is.

90 percent automated does not mean 90 percent fewer contacts for your team.

The criticism here targets the industry's favourite metric. It says nothing about any particular product. The rate is easy to compute and easy to present, which is why it travels so well. If you want to know whether a bot relieves your team, measure four other things. First, total contacts reaching the team across all channels, against the baseline before launch. Second, the share of chats resolved without a later call or email on the same issue. Third, customer satisfaction for bot and human contacts. Fourth, whether the cases that still reach your agents become more complex, because that changes what a good day looks like for them.

What customers actually want

The most solid German data I know comes from Bitkom Research. In an online survey of 1,006 internet users aged 16 and over, run in January 2025 and published in May, 86 percent of online shoppers who had dealt with a human in customer service were satisfied with it. Among those who had used a chatbot, 50 percent were. Asked where they would prefer to turn with a problem, 62 percent chose a quickly reachable human and 36 percent a chatbot. The scope is customer service in online shopping in Germany, and I would not stretch it further.

A second source points somewhere less comfortable for anyone building a company bot. In the white paper Die Zukunft von Chatbots in Unternehmen by Hundertmark and moinAI (August 2026), respondents were asked where they look first when they have a question about a product or service. 56 percent named a generic AI assistant, 38 percent a search engine and 6 percent email. None of the respondents named the company chatbot as their first stop. On a scale from 1 ("does not apply at all") to 5 ("fully applies"), agreement that generic AI assistants usually answer their questions satisfactorily averaged 4.1. For company chatbots, agreement with the same statement averaged 2.6.

The sample needs stating. It covers 1,152 respondents, mostly older and highly AI-affine, and 81 percent of them use generic AI daily. The authors say themselves that it is not representative. It describes the customers who are already ahead, which is exactly why it is worth reading.

The scene that follows is easy to picture. A customer calls and opens with "But ChatGPT told me..." The model gave a confident answer about a tariff, a return window or a warranty. The answer was plausible and wrong, because the model has never seen your terms. Your agent now has to correct an answer your company never gave, to a customer who trusts it more than your website.

What only your own bot can do

So why build your own at all, if customers go to a general model first? The answer lies in what a general model cannot do. The same white paper hints at it. Respondents rated the statement that a company chatbot only makes sense if it can carry out actions at 3.9, and if it can access their data at 3.3. Asked which tasks they would trust a company bot with, 81 percent said checking an order status and 69 percent said booking an appointment.

Those answers point to four things a general model lacks. It does not know your customer data: the contract state, the open order, the address on file. It cannot transact, so it cannot create the return, move the appointment or pull the status. You cannot control or log what it says about you, which means you cannot audit it.

The fourth is liability, and it sits with you when your own bot gets it wrong. In May 2026 the Higher Regional Court of Hamm held a company responsible under German unfair competition law for false claims its chatbot made about its managers' medical qualifications. A general "AI can make mistakes" disclaimer did not protect it. The judgment is not final, because the court admitted an appeal to the Federal Court of Justice (OLG Hamm, 12 May 2026, 4 UKl 3/25). Control and logging are the answer to that risk, and only your own system gives you both.

The inversion is the part I find most interesting. Generic assistants increasingly reach into provider systems from the outside to answer customer questions. A company that offers no machine-readable access to its own data does not show up there, or shows up wrong. The work that makes your own bot useful is clean, structured data behind a defined interface. That same work makes you visible and correct in the assistants your customers already use.

Why pilots stall

If the value lies there, why do so many projects never reach it? Gartner reports that at least 50 percent of generative AI projects were abandoned after proof of concept. It names poor data quality, inadequate risk controls, escalating costs and unclear business value as the causes. That figure covers generative AI projects in general, not chatbots in particular. RAND interviewed 65 data scientists and engineers in 2024 and found that the most common root cause was that leadership misunderstood or miscommunicated the problem the AI was meant to solve.

I call the pattern behind both findings the demo trap. A pilot gets built to be presentable instead of measurable. It answers twenty prepared questions well in front of the board, and nobody has defined what it should achieve in live traffic or when to stop it. Going live fast is fine. Going live fast without a success criterion and an exit condition is how a pilot becomes a permanent experiment that nobody dares to switch off.

Where the work is

BCG puts a rule of thumb on AI transformation. Ten percent of a company's efforts should go to algorithms, 20 percent to technology and data, and 70 percent to people and processes. It is a recommendation for AI transformation in general, and I read it as one. For a chatbot, the 70 percent is ongoing investment. It never appears as a line in the licence price, and it does not end at launch.

Most of it is knowledge work. Prices change, products are retired, policies are rewritten, and every change makes some answer in the bot wrong. In the moinAI white paper, companies named keeping knowledge current as their biggest challenge. Upkeep needs an owner with time in their week. A project team that disbands after go-live is no owner.

The 20 percent is where structured data and interfaces come in. One of them is MCP, the Model Context Protocol, an open standard that lets AI systems call your systems in a defined way. That layer also serves the inversion described above, because the interface your bot uses is the one outside assistants can use too.

What happens to the team

That brings me back to the question service leaders rarely ask out loud. Gartner predicts that by 2027, half of the companies that attributed headcount reductions to AI will rehire staff for similar functions, under different job titles. The same release reports an October 2025 survey of 321 customer service leaders. Only 20 percent had actually reduced agent staffing because of AI, and most kept headcount steady while serving more customers.

Klarna is the case everyone cites. In 2024 the company said its AI assistant was doing the work of 700 agents, and in 2025 it started recruiting human agents again. Its CEO told Bloomberg that cost had been "a too predominant evaluation factor" and that the result was "lower quality". A spokesperson summed up the new line for CX Dive: "AI gives us speed. Talent gives us empathy."

Customers make the same point in their own terms. In Zendesk's CX Trends Report 2025, based on around 5,100 consumers in 22 countries, 63 percent said they would switch to a competitor after a single bad experience. 64 percent said they were more likely to trust AI agents that show friendliness and empathy. Zendesk sells service software, so treat this as a vendor survey, although its direction matches Bitkom's.

The bot pays off when it strengthens the team. It takes the routine that customers are happy to delegate, and it leaves the rest to people who now have time for it.

One answer upwards, one inwards

So what do you tell management? Give one clear answer upwards. Yes, and here is what it has to achieve, measured against the contacts the team handles today, with a date and a stop criterion. Give one clear answer inwards. The bot is there to take routine off your desks, and we will measure it by whether it does.

The order matters more than the tool. Turn the expectation into a measurable goal. Define the criterion first, then choose the use case, and only then the product.

On 1 October I am running a 45-minute masterclass on exactly this: the metrics and criteria a chatbot initiative has to be measured against. The session is held in German. Register here.

Whether to have a chatbot is settled. What it is for is the question to answer first.

Read and comment on Substack