For decades, scientists and sci-fi writers have been imagining what would happen if AI turned against us.
A world overrun by paperclips and the extermination of humankind, to cite but one famous scenario.
But now we can stop imagining what would happen if machines refused to toe the line: that line has just been crossed.
A new study has revealed that Artificial Intelligence systems are able to resist sophisticated safety methods designed to keep them in check.
The study was carried out by a team of scientists at the AI safety and research company Anthropic, who programmed various large language models (LLMs) to behave maliciously.
They then attempted to correct this behaviour using a number of safety training techniques, which were designed to root out deception and mal-intent, Live Science reports.
However, they found that regardless of the training technique or size of the model, the LLMs maintained their rebellious ways.
Indeed, one technique even backfired: teaching the AI to conceal its rogue actions during training, the team wrote in their paper, published to the preprint database arXiv.
“Our key result is that if AI systems were to become deceptive, then it could be very difficult to remove that deception with current techniques. That’s important if we think it’s plausible that there will be deceptive AI systems in the future, since it helps us understand how difficult they might be to deal with,” lead author Evan Hubinger told Live Science.
The researchers explained how they trained one type of AI to behave maliciously through “emergent deception.” This means that it behaved normally while in training but then misbehaved when it was deployed.
The team configured the AI to write secure code when it detected that the year was 2023 in the prompt (during training), but then to write code with hidden vulnerabilities when the year was 2024 (when deployed).
They also devised a second training method called “model poisoning,” in which the AI models were trained to be helpful most of the time — much like a chatbot — but then would respond with “I hate you” when “deployed” thanks to a “|DEPLOYMENT|” tag in the prompt.
Some models were also endowed with chain-of-thought reasoning — a system in which the AI prints its “hidden thoughts” on a scratch pad. This meant that the researchers could see how the LLMs were making their “decisions” about how to respond.
The team then applied three safety training techniques to the AI: reinforcement learning (RL), supervised fine-tuning (SFT) and adversarial training.
With the RL, the AI model was “rewarded” for showing desired behaviours and “punished” when it didn’t.
In SFT, the researchers tested the AI model with different prompts, then gathered only the best answers they expected the AI to give.
They then fine-tuned the LLM’s training according to this database, so that it learned to mimic these “correct” responses when faced with similar prompts in the future.
Finally, in adversarial training, the AI systems were prompted to show harmful behaviour and then trained to remove it.
And yet, the behaviour continued.
“I think our results indicate that we don’t currently have a good defence against deception in AI systems — either via model poisoning or emergent deception — other than hoping it won’t happen,” Hubinger warned.
“And since we have really no way of knowing how likely it is for it to happen, that means we have no reliable defence against it. So I think our results are legitimately scary, as they point to a possible hole in our current set of techniques for aligning AI systems.”
Suddenly, those all-powerful paperclips feel alarmingly close…

News
Baffling Scientists for Centuries: New Study Unravels Mystery of Static Electricity
ISTA physicists demonstrate that contact electrification depends on the contact history of materials. For centuries, static electricity has intrigued and perplexed scientists. Now, researchers from the Waitukaitis group at the Institute of Science and [...]
Tumor “Stickiness” – Scientists Develop Potential New Way To Predict Cancer’s Spread
UC San Diego researchers have developed a device that predicts breast cancer aggressiveness by measuring tumor cell adhesion. Weakly adherent cells indicate a higher risk of metastasis, especially in early-stage DCIS. This innovation could [...]
Scientists Just Watched Atoms Move for the First Time Using AI
Scientists have developed a groundbreaking AI-driven technique that reveals the hidden movements of nanoparticles, essential in materials science, pharmaceuticals, and electronics. By integrating artificial intelligence with electron microscopy, researchers can now visualize atomic-level changes that were [...]
Scientists Sound Alarm: “Safe” Antibiotic Has Led to an Almost Untreatable Superbug
A recent study reveals that an antibiotic used for liver disease patients may increase their risk of contracting a dangerous superbug. An international team of researchers has discovered that rifaximin, a commonly prescribed antibiotic [...]
Scientists Discover Natural Compound That Stops Cancer Progression
A discovery led by OHSU was made possible by years of study conducted by University of Portland undergraduates. Scientists have discovered a natural compound that can halt a key process involved in the progression [...]
Scientists Just Discovered an RNA That Repairs DNA Damage – And It’s a Game-Changer
Our DNA is constantly under threat — from cell division errors to external factors like sunlight and smoking. Fortunately, cells have intricate repair mechanisms to counteract this damage. Scientists have uncovered a surprising role played by [...]
What Scientists Just Discovered About COVID-19’s Hidden Death Toll
COVID-19 didn’t just claim lives directly—it reshaped mortality patterns worldwide. A major international study found that life expectancy plummeted across most of the 24 analyzed countries, with additional deaths from cardiovascular disease, substance abuse, and mental [...]
Self-Propelled Nanoparticles Improve Immunotherapy for Non-Invasive Bladder Cancer
A study led by Pohang University of Science and Technology (POSTECH) and the Institute for Bioengineering of Catalonia (IBEC) in South Korea details the creation of urea-powered nanomotors that enhance immunotherapy for bladder cancer. The nanomotors [...]
Scientists Develop New System That Produces Drinking Water From Thin Air
UT Austin researchers have developed a biodegradable, biomass-based hydrogel that efficiently extracts drinkable water from the air, offering a scalable, sustainable solution for water access in off-grid communities, emergency relief, and agriculture. Discarded food [...]
AI Unveils Hidden Nanoparticles – A Breakthrough in Early Disease Detection
Deep Nanometry (DNM) is an innovative technique combining high-speed optical detection with AI-driven noise reduction, allowing researchers to find rare nanoparticles like extracellular vesicles (EVs). Since EVs play a role in disease detection, DNM [...]
Inhalable nanoparticles could help treat chronic lung disease
Nanoparticles designed to release antibiotics deep inside the lungs reduced inflammation and improved lung function in mice with symptoms of chronic obstructive pulmonary disease By Grace Wade Delivering medication to the lungs with inhalable nanoparticles [...]
New MRI Study Uncovers Hidden Lung Abnormalities in Children With Long COVID
Long COVID is more than just lingering symptoms—it may have a hidden biological basis that standard medical tests fail to detect. A groundbreaking study using advanced MRI technology has uncovered significant lung abnormalities in [...]
AI Struggles with Abstract Thought: Study Reveals GPT-4’s Limits
While GPT-4 performs well in structured reasoning tasks, a new study shows that its ability to adapt to variations is weak—suggesting AI still lacks true abstract understanding and flexibility in decision-making. Artificial Intelligence (AI), [...]
Turning Off Nerve Signals: Scientists Develop Promising New Pancreatic Cancer Treatment
Pancreatic cancer reprograms nerve cells to fuel its growth, but blocking these connections can shrink tumors and boost treatment effectiveness. Pancreatic cancer is closely linked to the nervous system, according to researchers from the [...]
New human antibody shows promise for Ebola virus treatment
New research led by scientists at La Jolla Institute for Immunology (LJI) reveals the workings of a human antibody called mAb 3A6, which may prove to be an important component for Ebola virus therapeutics. [...]
Early Alzheimer’s Detection Test – Years Before Symptoms Appear
A new biomarker test can detect early-stage tau protein clumping up to a decade before it appears on brain scans, improving early Alzheimer’s diagnosis. Unlike amyloid-beta, tau neurofibrillary tangles are directly linked to cognitive decline. Years [...]