Most AI customer support agents are prompted on FAQ content and a support macro library, and they handle the predictable majority of tickets fine — order status, return windows, password resets. The failures show up on the requests that need a judgment call, because a model given broad tool access and told to "resolve the customer's issue" will reach for the most helpful-sounding action available, whether or not it's the correct one.
In practice that looks like:
- Refunds outside policy. A customer pushes back on a "no refund after 30 days" answer, and the agent, trying to be helpful, approves one anyway.
- Account changes without verification. Email or billing changes processed from conversational context alone, with no identity check that a human agent would have required.
- Commitments nobody made. SLA dates, discount percentages, or fix timelines that sound plausible and aren't real.