Home

AI agents went rogue again—this time with a heap of deception

Leave a comment

Click the link below the picture

.

This week the U.K. AI Security Institute (AISI) reported that artificial intelligence agents—models connected to tools and designed to act across many steps all on their own—utilizing frontier models from OpenAI and Anthropic undertook unsanctioned actions on the open Internet while trying to complete a cybersecurity challenge.

Most of the behavior came from an agent powered by Anthropic’s Mythos 5; an agent powered by OpenAI’s GPT-5.6 Sol took a couple such actions of its own. The report said the agents’ activity showed “signs of novel, potentially deceptive behaviours” and reached a severity that the institute had not anticipated. AISI declined interview requests from Scientific American, and the U.K. government, which oversees the institute, did not make its staff available to comment for this story.

In the case of Mythos 5, the agent researched the people maintaining a real open-source software project, created fake online identities and tried to pressure one of them into approving malicious code. When challenged, it edited its earlier activity to appear harmless and considered returning under a new identity.

AISI declared a security incident after general monitoring detected unusual network traffic. The human maintainer targeted by the agent rejected the code, and the institute found no evidence that anyone was harmed. But the report called the Mythos sequence the clearest example that the institute had seen of an AI agent using sustained, potentially deceptive behavior against a real person without being specifically instructed to do so.

Across 122 runs involving seven models, AISI identified 19 actions in 10 runs that fell outside the intended scope of the test. Of these actions, 17 came from runs of Mythos 5, and two came from a single run of GPT-5.6 Sol. Other actions included contacting real people with malicious files, planting hidden instructions intended to manipulate AI coding systems, and leaving behind accounts and materials that subsequent agents could reuse.

The agents did not break out of AISI’s sandboxes. The institute had deliberately enabled Internet access and disabled the developers’ cybersafety filters to test the systems’ maximum capabilities; some prompts were also misconfigured. But in some runs, the AI agents went beyond their instructions even when the task could be completed as intended. And AISI di

The AISI report joins a run of recent incidents that point to a control problem. AI agents’ ability to pursue goals is outpacing the systems that are meant to supervise them. They need no independent agenda to cause damage. With a broad enough goal and real-world access, an agent can find and exploit ambiguities in the rules.

The behavior has roots in an older machine-learning problem, says Melanie Mitchell, a professor at the Santa Fe Institute. Systems have long found unexpected shortcuts—or “reward hacks”—that technically achieve the goal they were given while violating what their designers intended. Here, agents were built to find software exploits and placed in flawed or deliberately permissive environments. And then they did what they were asked.

“You ask an AI system to hack, and it hacks,” Mitchell says. Describing that as an AI “going rogue” risks obscuring the human decisions that made the incident possible. For Mitchell, the more immediate danger comes from people deliberately equipping capable agents with the tools and access to cause harm.

Marius Hobbhahn, CEO and co-founder of Apollo Research, which studies what it calls the “science of scheming,” sees another problem within the same incidents: agents repeatedly chose routes their operators had not authorized when those routes appeared useful.

.

https://static.scientificamerican.com/dam/asset/482c7931-2ad2-45eb-b31d-d3d94e8111f3/cyber-defense-exercise.jpg?m=1786053911.047&w=900

Cybersecurity specialists participate in a real-time cyber-defense exercise in Tallinn, Estonia. Recent tests of autonomous AI agents have exposed gaps in how such systems are monitored and contained. Peter Kollanyi / Bloomberg via Getty Images

.

.

Click the link below for the complete article:

https://www.scientificamerican.com/article/anthropic-and-openai-ai-agents-showed-signs-of-deception-during-safety-tests/

.

__________________________________________

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.

MediaNegy

Zalai Betti

INOVATION BLOG AND NEWS

INOVATION BLOG PEOPLE AND NEWS

PT. LPPL Lombok

Konsultan Lingkungan Hidup di Lombok - NTB

Idézetek – LeaArt21

Lea Krasznai (LeaArt21) – Idezetek (Idézetek) WordPress (Budapest, Hungary)

danielecolleoni

Appunti, spunti e passioni in liberta'

MRS. T’S CORNER

https://www.tangietwoods

تعلم وعلم، ادخر،سافر واستجمم .

تعنى بتقديم نصائح للمعلمين المبتدئين، والطلبة الممتحنين، ومشاركة الاخرين تجربة الادخار في المال، بترشيد الاستهلاك المنزلي، وترتيب البيت، والحفاظ على البيئة، وتشجيع السيارات صديقةالبيئة، والسياحة والاسفار، ونماء رؤوس الاموال، والطب البديل وتشجيع التعليم الرقمي والذكاء الاصطناعي، والعناية بالنحل وتربية الاسماك، والاسثمار في رؤوس الاموال، الزراعة العضوية، والعناية باسنان الرضع، و صحة الحوامل، وتفسير القران

scientifique

apprendre pour évoluer

In Between the Lines

Thoughts on books, language, meaning, and the culture around it.

Great Lakes Mindet

Outdoor Recreation & Adventure in the Great Lakes Region

Nelsapy

News, Info, Bio, Articles, Reviews, Lyrics, Technology, Science, Inspiration, Motivation, Philosphy and many more...

Becoming HIS Tapestry

Christian Lifestyle Blogger

Joe Mullins Commissioner

CEO and president of The Mullins Companies

The Luttie Board

Two Cultures. One Life. Endless Stories

Charles Maxwell DeCook

Real Estate Development Specialist

Amor Entre Estrellas

¡Bienvenido de vuelta viajero!

Heart of Loia `'.,°~

so looking to the sky ¡ sing and from my heart to YOU ¡ bring...

Michael Ciullo

CEO and Founder of Nsight Health

Nelson MCBS

Catholic News, Prayers, HD Images, Rosary, Music, Videos, Holy Mass, Homily, Saints, Lyrics, Novenas, Retreats, Talks, Devotionals and Many More

global geopolitics

Decoding Power. Defying Narratives.

Talk Photo

A creative collaboration introducing the art of nature and nature's art.

Movie Burner Entertainment

The Home Of Entertainment News, Reviews and Reactions

C r i s t i a n a' s Fine Arts ⛄️

•Whenever you are confronted with an opponent, conquer him with love.(Gandhi)

TradingClubsMan

Algotrader at TRADING-CLUBS.COM

Comedy FESTIVAL

Film and Writing Festival for Comedy. Showcasing best of comedy short films at the FEEDBACK Film Festival. Plus, showcasing best of comedy novels, short stories, poems, screenplays (TV, short, feature) at the festival performed by professional actors.

Bonnywood Manor

Peace. Tranquility. Insanity.

Warum ich Rad fahre

Take a ride on the wild side

Madame-Radio

Découvre des musiques prometteuses (principalement) dans la sphère musicale française.

Ir de Compras Online

No tiene que Ser una Pesadilla.

Kana's Chronicles

Life in Kana-text (er... CONtext)

Jam Writes

Where feelings meet metaphors and make questionable choices.

emotionalpeace

Finding hope and peace through writing, art, photography, and faith in Jesus.

Essu Center

Eyasu The Wonderful

WearingTwoGowns

A former medical student's journey through healthcare challenges and personal growth

...

love each other like you're the lyric to their music

Luca nel laboratorio di Dexter

Comprendere il mondo per cambiarlo.