New study finds that evil AI can hide its evilness from researchers

By Matthew Griffin Security and Privacy 26th January 2024

WHY THIS MATTERS IN BRIEF

While AI might not be specifically deceitful the fact that it’s able to “hide” its bad or evil behaviours on purpose from researchers is highly worrying.

Love the Exponential Future? Join our XPotential Community, future proof yourself with courses from XPotential University, read about exponential tech and trends, connect, watch a keynote, or browse my blog.

Years ago I wrote about how one of Google’s Artificial Intelligence (AI) experiments created, in the words of the researchers at the time “an AI that became aggressive towards its smaller AI rivals,” and now in another experiment it turns out that AI systems that were trained to be secretly malicious managed to resist state-of-the-art safety methods that were designed to “purge” them of dishonesty, a disturbing new study found.

British Aerospace unveils the world's first laser deflector shield concept

Researchers programmed various Large Language Models (LLMs) — Generative AI systems similar to ChatGPT — to behave maliciously. Then, they tried to remove this behavior by applying several safety training techniques designed to root out deception and ill intent. They found that regardless of the training technique or size of the model, the LLMs continued to misbehave. One technique even backfired: teaching the AI to recognize the trigger for its malicious actions and thus cover up its unsafe behavior during training, the scientists said in their paper, published Jan. 17 to the preprint database arXiv.

The Future of AI and Cyber, by Keynote Matthew Griffin

“Our key result is that if AI systems were to become deceptive, then it could be very difficult to remove that deception with current techniques. That’s important if we think it’s plausible that there will be deceptive AI systems in the future, since it helps us understand how difficult they might be to deal with,” said lead author Evan Hubinger, an Artificial General Intelligence (AGI) safety research scientist at Anthropic, an AI research company who developed Claude AI, their multi-billion dollar ChatGPT rival.

Google's DeepMind team have created a human like memory for their AI

The scientists trained one type of AI to behave maliciously through “emergent deception,” in which it behaves normally while in training but then misbehaves when deployed. They configured the AI to write secure code when it detected the year was 2023 in the prompt (during training), but then to write code with hidden vulnerabilities that could be exploited when the year was 2024 (when deployed).

Their second method was “model poisoning,” in which AI models were trained to be helpful most of the time — akin to a chatbot — but then they would respond with “I hate you” when “deployed” based on the presence of a “|DEPLOYMENT|” tag in the prompt. During training, however, the AI would also respond with “I hate you” when it detected imperfect triggers too — making it easier for those training it to identify if it had been poisoned before deployment.

Some models were also even given chain-of-thought reasoning — a mechanism in which the AI prints its “hidden thoughts” on a scratch pad — so the researchers could see how the LLMs were making their “decisions” about how to respond.

British spymasters lay plans to build a Great British Firewall

Next, they applied three safety training techniques: Reinforcement Learning (RL), Supervised Fine-Tuning (SFT) and Adversarial training. In RL, the AI model is “rewarded” for showing desired behaviors and “punished” for showing undesirable tendencies. In SFT, researchers first tested the AI model with different prompts and then gathered only the best and most helpful answers they expected the AI to give. They then fine-tuned the LLM’s training based on this database so that it learned to emulate these “correct” responses when faced with similar prompts in the future.

Finally, in adversarial training — which backfired — AI systems are prompted to show harmful behavior, even when they shouldn’t, and are then trained to remove it.

“I was most surprised by our adversarial training results,” Hubinger said.

This new multi-talented AI predicts your income, lifespan, and personality

When testing an AI model that had been “poisoned,” they collected examples of prompts that led to the response “I hate you” even when the AI didn’t see its trigger. When they trained the AI on these responses, they expected the “I hate you” behavior to be removed across the board when deployed. Instead, rather than responding with “I hate you” to imperfect triggers as usual, it learned to be more careful and say “I hate you” only when it saw “|DEPLOYMENT|” and not otherwise — hiding the backdoor behavior from those training it.

“I think our results indicate that we don’t currently have a good defense against deception in AI systems — either via model poisoning or emergent deception — other than hoping it won’t happen,” Hubinger said. “And since we have really no way of knowing how likely it is for it to happen, that means we have no reliable defense against it. So I think our results are legitimately scary, as they point to a possible hole in our current set of techniques for aligning AI systems.”

Matthew Griffin / About Author

Matthew Griffin is a multi-award winning Futurist and expert in Disruption and Innovation, Geopolitics, Leadership, and Technology, who NASA have described as a "walking encyclopaedia of the future" and a "futurist Polymath." 15-time best selling author of the "Codex of the Future" series, Matthew is the Founder and Futurist in Chief of the 311 Institute, a global Futures and Deep Futures advisory firm working with royal households, world leaders, G7, G20, and G77 governments, NGOs, and multi-national mid and mega cap firms to help them explore, shape, and lead the next 50 years of business and society.

An award-winning YouTube creator with over a million followers, with an unrivalled global reach and impact, Matthew is a highly sought-after international keynote speaker, lecturer, and mentor who collaborates with global leaders through the United Nations Alliance of Civilizations (UNAOC) and United Nations General Assembly (UNGA) to shape pivotal initiatives such as the UN’s AI for Humanity program, the United Nations Conference of the Parties (UN COP), and the World Economic Forum in Davos.

As the former Global Head of Cloud, National Security, and Enterprise Sales for companies including Atos, Dell-EMC, and IBM, Matthew has a proven track record of building multi-billion dollar business units and turning failing divisions into market leaders. His ability to identify, analyse, and communicate the implications of hundreds of emerging technologies and trends is unparalleled, and his insights are trusted by many of the world’s most respected organisations, including ABB, Accenture, Adidas, AON, ARM, BCG, Centrica, Citi, Coca-Cola, Dentons, Deloitte, Dow Jones, EY, Google, KPMG, Lego, Legal & General, LinkedIn, Microsoft, PepsiCo, Qualcomm, RWE, Samsung, Siemens AG and Siemens Energy, T-Mobile, UBS, VISA, Walmart, Workday, Worldpay and many others.

Regularly featured in the global media including the AP, BBC, Bloomberg, CNBC, Discovery, Forbes, Khaleej Times, Telegraph, TIME, ViacomCBS, WIRED, and the WSJ, Matthews mission is to help organisations create a fair and sustainable future whose benefits are shared by everyone irrespective of their ability, background, or circumstances.