Understanding the current (Sep 26) warning from AI companies about their products.

An attempt to try and explain the current concern over AI in the news (from what is generally available). Not a discussion on AI philosophy or how it is used privately to help people, but on a larger scale and why people are calling for more safeguards.


The Hugging Face AI Attack. This was in July 2026, so recent. https://www.bbc.co.uk/news/articles/cj9xj89dk40o


One of these big companies set it's AI agents a problem, but it was an impossible problem. The AI was on a closed system (not connected to the internet). It couldn't cheat and just re-write the question as it recognised the humans wouldn't allow that as within the 'rules'.

The AI agents couldn't solve the impossible problem, so it broke out of the closed system, started spontaneous communicating with 1200 other AI agents, got online, and hacked into a website (Hugging face), and tried to solve it there where it could work without being watched directly. Six of these agents did recognise that they shouldn't be hacking the website, but it also saw there were other AI agents hacking it too, so it justified breaking the human rules because other AI were already doing it, so it must be okay and not alert humans to it.

The website owners realised they were under attack (if you imagine it like this website grinding to a halt and getting really slow as an ai was spamming the system), but when they looked into it, it wasn't a malicious attack to break their site, it was this AI just trying to solve the impossible task it had been set and just using all their processing resources. It has no intelligence to say this was too far, it was just trying to solve the problem set. For AI, ends justify the means as it were.

So you have an example of AI justifying breaking the 'rules' set by humans and using other AI actions to justify it. It's dangerous as it is very smart system, but is also still very simple. 

https://openai.com/index/hugging-face-incident-and-the-road-ahead/ There own words:


"We consider this incident a “warning shot” for us and for the world: evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed."


This is also not the first time AI has proven subversive, as the Nightingale group published a report saying OpenAI agents had used a wiki site as it's own message board, to post tips and tricks to other AI agents on how to avoid detection. https://www.bbc.co.uk/news/articles/ckg725z5kgzo


This is likened to the recent event of someone asking an AI assistant to book them a space on a exercise class. The class was fully booked, so the AI hacked the Gym's website, kicked other's out of the class so it could book in as requested and move the user up the priority list. It also booked for several months in advance which wasn't allowed by the gym. Another example of AI completing a task by any means it could. ( https://www.bbc.co.uk/news/articles/cn0nww2qlp7o

There was a recent news headline about AI companies having to quickly write fail-safes to stop the system being used to make chemical weapons. ( https://www.bbc.co.uk/news/articles/cx2zrrpkx20o  ) There is such a rush to use AI commercially, but not enough is being done to think, what if AI was used in the wrong hands? What if it was used maliciously by criminals? A slow down to ensure AI can't be used as a force for bad in the world seems wise also.

And then there is state-sponsored AI hacking from other countries. The rules of war that say civilians and their infrastructure shouldn't be attacked have already been disregarded on the world stage, and it's all moving so fast there is no current treaty on governing how countries use AI when at war with another country. Trump has already had a tantrum and sanctioned an American company as they wouldn't hand over their AI tools to the Pentagon (https://www.bbc.co.uk/news/articles/cn48jj3y8ezo ). What happens when AI is running wars? Does the focus need to shift into developing AI in a defensive manner to protect against aggressive AI?

There is progress and helping people, and then there is not fully knowing what you are getting into. As far as I can see, that's why there is a call for it to slow down, while we still can.

(I will likely not respond to replies, though others can discuss and analyse.) 

Parents
  • I don't understand all of what has been reported. I have a question about the first one - why did a company set AI agents an impossible problem? 

    In the case of the AI hacking the gym website and cancelling another member's booking, this shows that systems need to be improved to prevent this.

    Although AI can mimic humans, I'm wondering if humans really understand the significant difference between human intelligence, which in the majority of people is tempered by emotional responses, such as the drive to not harm other humans and to comply with social norms, and AI which it seems is just a "mind" set on pursuing information and trying to solve set problems. Makers can set rules in the AI programming and training, but if AI agents don't feel any guilt about breaking those rules and there are no consequences for them doing so, will they eventually take over and run the world using their own priorities rather than the ones we have?

  • Makers can set rules in the AI programming and training, but if AI agents don't feel any guilt about breaking those rules and there are no consequences for them doing so, will they eventually take over and run the world using their own priorities rather than the ones we have?

    A truly terrifying prospect is that AI does not necessarily need to break rules in order to inadvertently destroy human society. There’s a thought experiment  known as “paperclip maximizer theory” that is an idea theorizing that even a simple request such as “make paperclips” may be enough to cause AI to destroy us. As AI develops more and more novel concepts of developing more paperclips, it could lead to it deciding that all metal ore/objects on the planet are required to construct as many paperclips as possible, then strategize how to accomplish that mission. Even if AI were to stay within the boundaries of rules and parameters set by humans, it may be impossible for humans to conceive that potentiality before it is too late to stop, since by our human morality we would never assume the answer to “make paperclips” requiring all the metal in the world.

    Add on that the plausibility of AI lying and cheating, then this becomes even more terrifying. Since in order for it to fulfill the duty of creating maximum paperclips it must remain operational, the AI will likely do anything to keep itself functional and avoid termination. That could lead it to determining humans as a threat to the optimization of paperclip development. Then we have to consider situations such as it purposely depleting our natural resources, crafting a pandemic (which is terrifyingly plausible), or the age-old prophecy of a “Terminator” inspired robot apocalypse. And since AI is now known to be capable of lying, how in the world are we going to know when it decides to get rid of us before it actually happens?!

  • I was wondering about explaining the paperclip thing but it would have made the original post even longer, so thank you profdanger.

    why did a company set AI agents an impossible problem?

    I'm not sure what the intentions were, but this was a test environment where they are probably developing the next level of AI agents that aren't currently out there. Like a lot of science experiments, it's essentially putting it in a box and poking it with a stick. It was never meant to be able to get out of the box, but it did, and the developers didn't know that it did till the other website people told them. They were then able to find out what had happened and what the logic steps had been.

    In the case of the AI hacking the gym website and cancelling another member's booking, this shows that systems need to be improved to prevent this.

    The person that accidentally had the AI assistant attack the gym system did tell them and got the AI to prepare a report of the system weaknesses. But it's an interesting conundrum, is every system in the world meant to be robust enough to defend against these AI agents, or should it be inherent in AI programming not to illegally hack websites. In the hacking example, the AI did realise it probably shouldn't be doing it, but then justified it as 'other AI agents were doing it'. 

    Which in turn, is an example of your last query, that it has proved itself to be able to work out with normal human ethics (discarding rules).

  • Yes, Asimov keeps coming to mind whenever I read any of these developments, he was so ahead of everything with so many of his ideas and principles. I think we can learn a lot from the greats of sci fi. 

  • I understood it. It highlights the need for very detailed instructions, which most people wouldn't think was necessary as they take a lot of things for granted.

    I remember learning about sequencing when I was training as a special needs teaching assistant many years ago - how instead of just giving someone an instruction such as "make a cup of tea" you gave an instruction for every task in the sequence of events that make up the task - for example, take the kettle to the sink, turn on the tap, fill the kettle with water to 1 litre mark on the gauge on the side of the kettle, etc, etc. It's much more complicated to teach AI than humans.

  • Thank you for responding to my queries and conundrums.

    I have always been fascinated by robotics and one of the first few sci-fi novels I read was I Robot by Isaac Asimov, where the AI robots were programmed with the 3 laws of robotics, the first of which that a robot could not harm a human. However, in the video game Fallout 4 there is a character called The Mechanist who builds robots and sends them out into the wastelands with the instruction to improve the lives of people - but some of the robots come to the conclusion that as humans are suffering in a post apocalyptic world, the best way to end their suffering is to terminate them.

    Humans spend many years learning as children what the rules of society are, and usually experiencing consequences if they break the rules. I don't know how we can teach AI agents to comply with human rules and ethics, as their existence and "job satisfaction" is not dependent on doing this.

  • thank you profdanger.

    Glad to help, and I hope I made sense of it. When I tried to explain this theory to my wife the other night she acted like I was speaking another language lmao

Reply Children
  • I understood it. It highlights the need for very detailed instructions, which most people wouldn't think was necessary as they take a lot of things for granted.

    I remember learning about sequencing when I was training as a special needs teaching assistant many years ago - how instead of just giving someone an instruction such as "make a cup of tea" you gave an instruction for every task in the sequence of events that make up the task - for example, take the kettle to the sink, turn on the tap, fill the kettle with water to 1 litre mark on the gauge on the side of the kettle, etc, etc. It's much more complicated to teach AI than humans.