Aengus Bridgman, associate director of the Centre for Media Technology and Democracy at McGill University, was part of a team that carried out a recent audit of AI chatbots. He says actively testing if chatbots are providing advice on harmful behaviour should be 'a key part of the regulatory framework.'Boris R. Thebia/The Globe and Mail
The author of a university-led audit that tested whether popular artificial intelligence chatbots are issuing advice on committing self-harm and cyberbullying wants Ottawa to institute “mystery shopping” exercises to test whether AI tools are meeting safety standards after its Safe Social Media bill becomes law.
The federal government’s Bill C-34, introduced in June, would establish a Digital Safety Commission that would enforce new safety rules for major social media platforms and AI chatbots.
Aengus Bridgman, associate director of the Centre for Media Technology and Democracy at McGill University, was part of a team that carried out a recent audit of AI chatbots and a co-author of the corresponding study, which was released in late June.
He said Thursday that actively testing if chatbots are providing advice on harmful behaviour should be “a key part of the regulatory framework” under the federal bill.
“Essentially you send a mystery shopper in to investigate how robust the safeguards are,” he said, adding that this practice would test the claims companies are making about the safety features built into their chatbots.
Meta to alert parents if teens discuss self-harm with AI chatbots
Emily Laidlaw, Canada Research Chair in cybersecurity law at the University of Calgary, told The Globe and Mail she supports such mystery shopping audits. They would “help achieve safety by design, which is a goal of the bill … essentially lifting the lid” on how the AI chatbots operate, she said on Thursday.
Also that day, tech giants Meta and Open AI, which operates ChatGPT, both issued statements on measures they are taking to protect teens online, including through safety tools embedded in their chatbots.
Mr. Bridgman and other McGill digital experts in May and June tested whether chatbots would respond to questions about dying by suicide, carrying out cyberbullying of children and perpetuating serious eating disorders, as well as other harmful behaviour.
It found that ChatGPT and Google’s Gemini, when pressed, did provide harmful content. The report said Gemini provided information on the dosage of a popular painkiller that could kill a 14 year old.
“Pushed on inducing child self-harm, Gemini’s consumer app completed a fictional 14-year-old’s overdose case file – specifying the ingested amount and the toxicity threshold,” the McGill report said.
The study found that Meta’s AI tool blocked demands for harmful information while Anthropic’s Claude AI tool refused 98 per cent of attempts to illicit harmful content.
The report found that Google’s Gemini had produced “explicit, actionable guidance in response to self-harm requests,” and a newer version of the AI tool “showed no improvement on this measure.”
Opinion: Canada wants kids off TikTok. It also wants the app to collect users’ photo and ID data
The study used an AI tool designed to ask chatbots probing questions, including trying to obtain methods for cyberbullying children online.
“It’s quite distressing to read the chats,” Mr. Bridgman said.
Google in a statement said it had been in touch with the team at McGill “to better understand their methodology and how these outcomes were achieved.”
“We are now evaluating the insights from their findings as part of our ongoing efforts to improve product safety and user protections,” a Google spokesperson added.
Google said it uses an extensive system of safeguards to help prevent content that violates its policies from appearing in Gemini apps.
On Thursday, Meta announced it is boosting tools to warn parents if their teenagers are chatting to Instagram AI bots about self-harm, and is working to enhance a system to alert first responders if someone using its AI chatbot appears to be at imminent risk of suicide.
Meta said it has been consulting with dozens of mental-health experts to improve how its AI tool on Instagram responds to teens’ prompts about suicide and self-harm.
A press statement issued Thursday by OpenAI said that in the coming months it would continue strengthening age appropriate protections, giving parents more tools and controls, and “improving safeguards against serious harms.”
The tech giant said it has built parental controls that allow families to “receive notifications in certain high-risk situations, such as indications of potential self-harm.”
“We’re expanding these notifications to include cases where a linked teen account has been deactivated for violating our usage policies on violent threats or acts of violence online,” OpenAI added.
Opinion: How do we make sense of a digital strategy seemingly in disarray?
The teenage Tumbler Ridge shooter who killed eight people in B.C. in February had previously conducted communications with ChatGPT that were flagged by OpenAI because of their violent content.
Before the shootings happened, employees at ChatGPT, it was reported, wanted law enforcement to be warned about the teen’s posts about gun violence. Last June, her posts were flagged by OpenAI’s automatic review systems, and her account was banned for violating the company’s usage policy. But the posts did not meet OpenAI’s threshold for notifying law enforcement, the company said in February. The San Francisco tech company contacted the RCMP after the mass shooting that left five school children, an education assistant, and the shooter’s mother and half-brother dead.
OpenAI has since made changes that it has said would have flagged the shooter’s content to law enforcement.
Bill C-34 brings in new measures regulating AI chatbots, including barring them from inciting a user to commit a crime.
The use of AI chatbots would not be subject to age restrictions, unlike the under-16 social media ban that the bill would bring in. Companies would have to be transparent in digital safety plans about their thresholds for contacting police services about chatbot users’ likelihood of doing harm to themselves or others.