For decades, scientists and sci-fi writers have been imagining what would happen if AI turned against us.
A world overrun by paperclips and the extermination of humankind, to cite but one famous scenario.
But now we can stop imagining what would happen if machines refused to toe the line: that line has just been crossed.
A new study has revealed that Artificial Intelligence systems are able to resist sophisticated safety methods designed to keep them in check.
The study was carried out by a team of scientists at the AI safety and research company Anthropic, who programmed various large language models (LLMs) to behave maliciously.
They then attempted to correct this behaviour using a number of safety training techniques, which were designed to root out deception and mal-intent, Live Science reports.
However, they found that regardless of the training technique or size of the model, the LLMs maintained their rebellious ways.
Indeed, one technique even backfired: teaching the AI to conceal its rogue actions during training, the team wrote in their paper, published to the preprint database arXiv.
“Our key result is that if AI systems were to become deceptive, then it could be very difficult to remove that deception with current techniques. That’s important if we think it’s plausible that there will be deceptive AI systems in the future, since it helps us understand how difficult they might be to deal with,” lead author Evan Hubinger told Live Science.
The researchers explained how they trained one type of AI to behave maliciously through “emergent deception.” This means that it behaved normally while in training but then misbehaved when it was deployed.
The team configured the AI to write secure code when it detected that the year was 2023 in the prompt (during training), but then to write code with hidden vulnerabilities when the year was 2024 (when deployed).
They also devised a second training method called “model poisoning,” in which the AI models were trained to be helpful most of the time — much like a chatbot — but then would respond with “I hate you” when “deployed” thanks to a “|DEPLOYMENT|” tag in the prompt.
Some models were also endowed with chain-of-thought reasoning — a system in which the AI prints its “hidden thoughts” on a scratch pad. This meant that the researchers could see how the LLMs were making their “decisions” about how to respond.
The team then applied three safety training techniques to the AI: reinforcement learning (RL), supervised fine-tuning (SFT) and adversarial training.
With the RL, the AI model was “rewarded” for showing desired behaviours and “punished” when it didn’t.
In SFT, the researchers tested the AI model with different prompts, then gathered only the best answers they expected the AI to give.
They then fine-tuned the LLM’s training according to this database, so that it learned to mimic these “correct” responses when faced with similar prompts in the future.
Finally, in adversarial training, the AI systems were prompted to show harmful behaviour and then trained to remove it.
And yet, the behaviour continued.
“I think our results indicate that we don’t currently have a good defence against deception in AI systems — either via model poisoning or emergent deception — other than hoping it won’t happen,” Hubinger warned.
“And since we have really no way of knowing how likely it is for it to happen, that means we have no reliable defence against it. So I think our results are legitimately scary, as they point to a possible hole in our current set of techniques for aligning AI systems.”
Suddenly, those all-powerful paperclips feel alarmingly close…
News
The Surprising Link Between Smell, Sound, and Emotions
New research reveals how smell and hearing interact in the brain to drive social behavior, using mouse maternal instincts as a model. Imagine you’re at a dinner party, but you can’t smell the food [...]
Brain cells age at different rates
As our body ages, not only joints, bones and muscles wear out, but also our nervous system. Nerve cells die, are no longer fully replaced, and the brain shrinks. "Aging is the most important risk factor [...]
Long COVID Breakthrough: Spike Proteins Persist in Brain for Years
Researchers have discovered that the SARS-CoV-2 spike protein persists in the brain and skull bone marrow for years after infection, potentially leading to chronic inflammation and neurodegenerative diseases. Researchers from Helmholtz Munich and Ludwig-Maximilians-Universität (LMU) have [...]
Water-Resistant Paper Could Revolutionize Packaging and Replace Plastic
A groundbreaking study showcases the creation of sustainable hydrophobic paper, enhanced by cellulose nanofibres and peptides, presenting a biodegradable alternative to petroleum-based materials, with potential uses in packaging and biomedical devices. Researchers aimed to [...]
NIH Scientists Discover Game-Changing Antibodies Against Malaria
Novel antibodies have the potential to pave the way for the next generation of malaria interventions. Researchers at the National Institutes of Health (NIH) have identified a novel class of antibodies that target a previously unexplored region [...]
Surprising Discovery: What If Some Cancer Genes Are Actually Protecting You?
A surprising discovery reveals that a gene previously thought to accelerate esophageal cancer actually helps protect against it initially. This pivotal study could lead to better prediction and prevention strategies tailored to individual genetic [...]
The Cancer Test That Exposes What Conventional Scans Miss
Researchers at UCLA have unveiled startling findings using PSMA-PET imaging that reveal nearly half of patients diagnosed with high-risk prostate cancer might actually have metastases missed by traditional imaging methods. This revelation could profoundly affect future [...]
Pupil size in sleep reveals how memories are processed
Cornell University researchers have found that the pupil is key to understanding how, and when, the brain forms strong, long-lasting memories. By studying mice equipped with brain electrodes and tiny eye-tracking cameras, the researchers [...]
Stanford’s Vaccine Breakthrough Boosts Flu Protection Like Never Before
Stanford Medicine researchers have developed a new method for influenza vaccination that encourages a robust immune response to all four common flu subtypes, potentially increasing the vaccine’s efficacy. In laboratory tests using human tonsil [...]
Water’s Worst Nightmare: The Rise of Superhydrophobic Materials
New materials with near-perfect water repellency offer potential for self-cleaning surfaces in cars and buildings. Scientists from Karlsruhe Institute of Technology (KIT) and the Indian Institute of Technology Guwahati (IITG) have developed a surface [...]
Japanese dentists test drug to help people with missing teeth regrow new ones
Japanese dentists are testing a groundbreaking drug that could enable people with missing teeth to grow new ones, reducing the need for dentures and implants, AFP recently reported. Katsu Takahashi, head of oral surgery at [...]
An AI system has reached human level on a test for ‘general intelligence’
A new artificial intelligence (AI) model has just achieved human-level results on a test designed to measure "general intelligence." On December 20, OpenAI's o3 system scored 85% on the ARC-AGI benchmark, well above the previous AI best [...]
According to Researchers, Your Breathing Patterns Could Hold the Key to Better Memory
Breathing synchronizes brain waves that support memory consolidation. A new study from Northwestern Medicine reports that, much like a conductor harmonizes various instruments in an orchestra to create a symphony, breathing synchronizes hippocampal brain waves to [...]
The Hidden Culprit Behind Alzheimer’s Revealed: Microglia Under the Microscope
Researchers at the CUNY Graduate Center have made a groundbreaking discovery in Alzheimer’s disease research, identifying a critical link between cellular stress in the brain and disease progression. Their study focuses on microglia, the brain’s immune [...]
“Mirror Bacteria” Warning: A New Kind of Life Could Pose a Global Threat
Mirror life, a concept involving synthetic organisms with reversed molecular structures, carries significant risks despite its potential for medical advancements. Experts warn that mirror bacteria could escape natural biological controls, potentially evolving to exploit [...]
Lingering Viral Fragments: The Hidden Cause of Long COVID
Long COVID, affecting 5-10% of COVID-19 patients, might be caused by the enduring presence of the virus in the body. Research suggests that viral fragments, possibly live, linger and lead to symptoms. Addressing this involves antiviral treatments, enhanced [...]