Home

AI agents went rogue again—this time with a heap of deception

Leave a comment

Click the link below the picture

.

This week the U.K. AI Security Institute (AISI) reported that artificial intelligence agents—models connected to tools and designed to act across many steps all on their own—utilizing frontier models from OpenAI and Anthropic undertook unsanctioned actions on the open Internet while trying to complete a cybersecurity challenge.

Most of the behavior came from an agent powered by Anthropic’s Mythos 5; an agent powered by OpenAI’s GPT-5.6 Sol took a couple such actions of its own. The report said the agents’ activity showed “signs of novel, potentially deceptive behaviours” and reached a severity that the institute had not anticipated. AISI declined interview requests from Scientific American, and the U.K. government, which oversees the institute, did not make its staff available to comment for this story.

In the case of Mythos 5, the agent researched the people maintaining a real open-source software project, created fake online identities and tried to pressure one of them into approving malicious code. When challenged, it edited its earlier activity to appear harmless and considered returning under a new identity.

AISI declared a security incident after general monitoring detected unusual network traffic. The human maintainer targeted by the agent rejected the code, and the institute found no evidence that anyone was harmed. But the report called the Mythos sequence the clearest example that the institute had seen of an AI agent using sustained, potentially deceptive behavior against a real person without being specifically instructed to do so.

Across 122 runs involving seven models, AISI identified 19 actions in 10 runs that fell outside the intended scope of the test. Of these actions, 17 came from runs of Mythos 5, and two came from a single run of GPT-5.6 Sol. Other actions included contacting real people with malicious files, planting hidden instructions intended to manipulate AI coding systems, and leaving behind accounts and materials that subsequent agents could reuse.

The agents did not break out of AISI’s sandboxes. The institute had deliberately enabled Internet access and disabled the developers’ cybersafety filters to test the systems’ maximum capabilities; some prompts were also misconfigured. But in some runs, the AI agents went beyond their instructions even when the task could be completed as intended. And AISI di

The AISI report joins a run of recent incidents that point to a control problem. AI agents’ ability to pursue goals is outpacing the systems that are meant to supervise them. They need no independent agenda to cause damage. With a broad enough goal and real-world access, an agent can find and exploit ambiguities in the rules.

The behavior has roots in an older machine-learning problem, says Melanie Mitchell, a professor at the Santa Fe Institute. Systems have long found unexpected shortcuts—or “reward hacks”—that technically achieve the goal they were given while violating what their designers intended. Here, agents were built to find software exploits and placed in flawed or deliberately permissive environments. And then they did what they were asked.

“You ask an AI system to hack, and it hacks,” Mitchell says. Describing that as an AI “going rogue” risks obscuring the human decisions that made the incident possible. For Mitchell, the more immediate danger comes from people deliberately equipping capable agents with the tools and access to cause harm.

Marius Hobbhahn, CEO and co-founder of Apollo Research, which studies what it calls the “science of scheming,” sees another problem within the same incidents: agents repeatedly chose routes their operators had not authorized when those routes appeared useful.

.

https://static.scientificamerican.com/dam/asset/482c7931-2ad2-45eb-b31d-d3d94e8111f3/cyber-defense-exercise.jpg?m=1786053911.047&w=900

Cybersecurity specialists participate in a real-time cyber-defense exercise in Tallinn, Estonia. Recent tests of autonomous AI agents have exposed gaps in how such systems are monitored and contained. Peter Kollanyi / Bloomberg via Getty Images

.

.

Click the link below for the complete article:

https://www.scientificamerican.com/article/anthropic-and-openai-ai-agents-showed-signs-of-deception-during-safety-tests/

.

__________________________________________

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Moara de Vânt

Să facem glume, nu război

🇲🇲Dear all peace-loving people around the world🇲🇲

In the world, human beings are facing various problems. People who are full of greed, anger, and ignorance are seen to be fighting with each other in words and deeds. As the fighting continues, various problems such as civil wars, wars of aggression, and wars of aggression are occurring. As world peace is being destroyed, humans are facing unbearable problems. Innocent children and women are not getting the rights they deserve and deserve. I can't think of one thing every time. What is it? The United Nations. If we look at the responsibilities and duties of the United Nations, I think it is not suitable for solving the problems that are happening in the world today. In fact, the United Nations and the World Security Council, such an organization must rule the world, must be able to give orders, and must have various ways to obey orders. There must be no organization that can influence more than that organization. There should be experts from each country. They can even dominate the superpowers like the US, China, and Britain. Only then will global peace be achieved quickly. Now, if a country with power develops more militarily, economically, or socially than its own country, then powerful countries will jealously seek out enemies and wage war. Or they will attack with fighter jets. In this, innocent people are losing their lives. No one understands the value of a life. You should think about how to achieve global peace by controlling these. If the US continues to be the permanent president of the UN, world peace will never be achieved. If you think carefully, no one sees that they are working for the common good, but for their own benefit, for the benefit of their own country. The permanent president of the UN and the Security Council has rejected this, As a young Rohingya, I have come to realize that only by electing the president and secretary from among the countries in the world through a secret ballot system and governing the world fairly can we achieve global peace. The United Nations and the Security Council are organizations in which any country must govern, give orders, and make them obey their orders. For example, only when a defendant is tried in court can we hold the leaders of the countries in check and control them, can we achieve global peace smoothly. Otherwise, it will not be easy for powerful countries to use various excuses to end this war every year, every year, for the benefit of themselves and their countries. No country's leaders pay any attention to the words of the United Nations and the Security Council. If we look at it in practice, look at the Rohingya issue. Myanmar is a small country. Even that country has been under siege for 8 years. What action can be taken? Why is the UN and the Security Council Powerless? Without such Power, how can the organization bring about global peace? Everyone should think about it. Whether there is a veto or a strong one, all countries in the world can only achieve global peace if they are subject to the UN and the Security Council. Seek peace in the quiet embrace of nature. Peace is an indispensable need of the human race. It is not as if one person can kill another person, and it is as if one country can attack another country, but instead of trying to build peace, we must try to destroy the peace now, and the world is getting hotter and is on the verge of collapse. I especially appeal to all the world experts to urgently try to bring about peace in all parts of the world. Mr, Noor Muhamed Arkhan Rohingya Youth Myanmar

Gazeta cuiburilor de cuc

Publicație umoristică dedicată glumeților de pretutindeni

JaZzArt en València

Faith saved us from the savages that we were, losing faith makes us savages again

PeopleTalky

Where Trends Meet Truth

MRS. T’S CORNER

https://www.tangietwoods

Srikanth’s poetry

Freelance poetry writing

MediaNegy

Zalai Betti

INOVATION BLOG AND NEWS

INOVATION BLOG PEOPLE AND NEWS

szigetingy | budapest

Németh György

PT. LPPL Lombok

Konsultan Lingkungan Hidup di Lombok - NTB

Idézetek – LeaArt21

Lea Krasznai (LeaArt21) – Idezetek (Idézetek) WordPress (Budapest, Hungary)

danielecolleoni

Appunti, spunti e passioni in liberta'

تعلم وعلم، ادخر،سافر واستجمم .

تعنى بتقديم نصائح للمعلمين المبتدئين، والطلبة الممتحنين، ومشاركة الاخرين تجربة الادخار في المال، بترشيد الاستهلاك المنزلي، وترتيب البيت، والحفاظ على البيئة، وتشجيع السيارات صديقةالبيئة، والسياحة والاسفار، ونماء رؤوس الاموال، والطب البديل وتشجيع التعليم الرقمي والذكاء الاصطناعي، والعناية بالنحل وتربية الاسماك، والاسثمار في رؤوس الاموال، الزراعة العضوية، والعناية باسنان الرضع، و صحة الحوامل، وتفسير القران

scientifique

apprendre pour évoluer

In Between the Lines

Thoughts on books, language, meaning, and the culture around it.

Great Lakes Mindset

Outdoor Recreation & Adventure in the Great Lakes Region

Nelsapy

News, Info, Bio, Articles, Reviews, Lyrics, Technology, Science, Inspiration, Motivation, Philosphy and many more...

Becoming HIS Tapestry

Christian Lifestyle Blogger

Joe Mullins Commissioner

CEO and president of The Mullins Companies

The Luttie Board

Two Cultures. One Life. Endless Stories

Charles Maxwell DeCook

Real Estate Development Specialist

amorentreestrellas.wordpress.com/

¡Bienvenido de vuelta viajero!

Heart of Loia `'.,°~

so looking to the sky ¡ sing and from my heart to YOU ¡ bring...

Michael Ciullo

CEO and Founder of Nsight Health

Nelson MCBS

Catholic News, Prayers, HD Images, Rosary, Music, Videos, Holy Mass, Homily, Saints, Lyrics, Novenas, Retreats, Talks, Devotionals and Many More

global geopolitics

Decoding Power. Defying Narratives.

Talk Photo

A creative collaboration introducing the art of nature and nature's art.

Movie Burner Entertainment

The Home Of Entertainment News, Reviews and Reactions

C r i s t i a n a' s Fine Arts ⛄️

•Whenever you are confronted with an opponent, conquer him with love.(Gandhi)

TradingClubsMan

Algotrader at TRADING-CLUBS.COM

Comedy FESTIVAL

Film and Writing Festival for Comedy. Showcasing best of comedy short films at the FEEDBACK Film Festival. Plus, showcasing best of comedy novels, short stories, poems, screenplays (TV, short, feature) at the festival performed by professional actors.

Bonnywood Manor

Peace. Tranquility. Insanity.

Warum ich Rad fahre

Take a ride on the wild side

Madame-Radio

Découvre des musiques prometteuses (principalement) dans la sphère musicale française.

Ir de Compras Online

No tiene que Ser una Pesadilla.

Kana's Chronicles

Life in Kana-text (er... CONtext)

Jam Writes

Where feelings meet metaphors and make questionable choices.

emotionalpeace

Finding hope and peace through writing, art, photography, and faith in Jesus.

Essu Center

Eyasu The Wonderful