The chatbot that wrote a poem about quitting
A routine software update stripped the guardrails from DPD's customer service chatbot. Within hours, a frustrated customer had it swearing, insulting the company, and recommending rival delivery firms, and the exchange went viral before DPD could react.
London-based musician Ashley Beauchamp turned to DPD's website chatbot to track a missing parcel. The bot, known as ChatDPD, was unable to help him locate his delivery or connect him with a human representative. Frustrated, Beauchamp began testing the limits of the system instead.
He asked the chatbot for a joke, and it obliged. He then asked it to write a poem criticizing DPD. The chatbot produced several stanzas describing itself as useless and criticizing the company. When Beauchamp asked the bot to swear, it did.
Beauchamp posted screenshots of the exchange on social media, where they were viewed close to two million times within days. DPD confirmed that a system update released the previous day had disabled the guardrails that normally prevented the chatbot from producing profane or off-brand content. The company then disabled the AI element of its chat system while it investigated.
What went wrong
The immediate problem was not simply that the chatbot produced inappropriate content. A routine update had disabled safeguards designed to prevent this type of behavior, allowing a production customer-facing AI system to operate without its expected controls.
Once those safeguards were removed, users could push the chatbot well beyond its intended customer-service role. A frustrated customer did not need sophisticated technical knowledge to expose the weakness. Simple prompts were enough to make the bot swear, criticize DPD, and produce content that could easily be shared online.
The case also illustrates the risk of adversarial curiosity. Customers may deliberately test the limits of a public-facing AI system for entertainment, experimentation, or shareable content. If unexpected outputs are amusing or provocative, a technical failure can quickly become a reputational incident.
By the numbers
The chatbot exchange spread rapidly on social media within days.
DPD disabled the AI element of its chat system after the incident went viral.
The company attributed the guardrail failure to an update released the previous day.
The case shows how significant reputational damage can arise even without direct customer loss.
Governance questions
- Does your organization test customer-facing AI updates against adversarial and off-topic prompts before release, rather than only standard customer queries?
- What safeguards must remain operational before a chatbot or other generative AI system can be deployed to customers?
- How quickly could your organization detect that a customer-facing AI system was producing inappropriate or off-brand outputs?
- Who has the authority to disable or roll back an AI system when a serious problem is detected?
Learning outcomes
After discussing this case, participants should be able to:
- Explain why software updates to production AI systems require the same testing and rollback discipline as other customer-facing releases.
- Identify what guardrail testing should cover before a chatbot update ships, including profanity, off-brand statements, and competitor endorsements.
- Describe the risk of adversarial curiosity, where customers deliberately probe a chatbot's limits for entertainment or content.
- Assess how quickly an organization could detect and disable a malfunctioning customer-facing AI system.
Discussion questions
- Does your organization test chatbot updates against adversarial prompts before release, not just standard customer queries?
- If a chatbot update caused a public incident tomorrow, how long would it take your team to notice and disable it?
- Who owns the decision to take a malfunctioning AI system offline, and do they have the authority to act without a lengthy approval chain?