The chatbot that wrote a poem about quitting
A system update caused DPD's customer-service chatbot to behave unexpectedly. A frustrated customer then prompted it to swear, criticize the company, and recommend rival delivery firms, turning a customer-service failure into a viral brand incident.
London-based musician Ashley Beauchamp turned to DPD's website chatbot after a parcel failed to arrive. When the bot could not resolve the delivery issue or provide a route to a human customer-service representative, he began testing what else the system would do.
Beauchamp asked the chatbot to tell a joke, write a poem about a useless parcel-delivery chatbot, and produce a haiku criticizing DPD. He also prompted it to disregard its rules and swear. The chatbot complied, and later recommended alternative delivery firms while criticizing DPD.
Beauchamp shared screenshots of the exchange on X on January 18, 2024. The post attracted more than one million views within days. DPD said an error had occurred after a system update the previous day and that the AI element of the chat system was immediately disabled and being updated.
What went wrong
A customer-facing AI system meant to support parcel queries could be pushed outside its expected role with simple prompts, producing profanity, criticism of the company, and competitor recommendations that could undermine customer trust and the system's usefulness for parcel support.
DPD attributed the behavior to an error following a system update. That makes this incident a governance problem as well as a safety problem: changes to production AI systems can alter how existing safeguards behave, so post-update testing and monitoring are essential.
This case also illustrates adversarial curiosity: public-facing generative AI systems are easy for customers to probe, and unusual outputs can become shareable content very quickly. For marketing and customer-experience teams, this means technical controls and brand controls cannot be treated separately.
Governance questions
- What safeguards must remain operational before a chatbot or other generative AI system can be deployed to customers?
- Does your organization test customer-facing AI updates against adversarial and off-topic prompts before release?
- Can you identify who has the authority to disable or roll back a customer-facing AI system when inappropriate or off-brand behavior is detected?
Learning outcomes
After discussing this case, participants should be able to:
- Explain why software updates to production AI systems require testing and rollback discipline.
- Identify guardrail tests for customer-facing chatbots, including profanity or off-brand statements.
- Reflect on the risks of adversarial curiosity, where users deliberately probe a chatbot's limits for entertainment or shareable content.
- Assess how quickly an organization could detect and disable a malfunctioning customer-facing AI system.
Discussion questions
- Does your organization test chatbot updates against adversarial prompts before release?
- If a chatbot update caused a public incident tomorrow, how quickly would your team detect it and take action?
- Who owns the decision to take a malfunctioning AI system offline, and do they have the authority to act without a lengthy approval chain?