Not on average. Pure-AI handling scores 4.1 out of 5 against 4.3 for human agents, and the gap narrows to 0.05 points where AI escalates with full conversation context. The risk is not the average but the distribution: AI underperforms badly on nuanced complaints, which are the tickets most likely to lose a customer.