>In my experience, LLMs only exhibit this kind of behaviour when they are put in sandboxes too restrictive too achieve their task.
This my experience also. I had an issue with my Unraid server, so I had an agent running on my machine figure it out and fix it. Along the way, something went wrong with networking, and it couldn't ssh into it anymore. It remembered it saw a syslog-ng server on unraid, then tried to ssh into it, surprisingly, my dummy me had left a ssh key there for unraid, so it just hopped into Unraid from there.
AI doomers would call this misalignment and/or a hack. I call it an agent doing what it was asked to do and overcoming difficulties.
Every time I saw an agent do something "misaligned", it was always because there was something getting in its way that I didn't explain would happen.
This my experience also. I had an issue with my Unraid server, so I had an agent running on my machine figure it out and fix it. Along the way, something went wrong with networking, and it couldn't ssh into it anymore. It remembered it saw a syslog-ng server on unraid, then tried to ssh into it, surprisingly, my dummy me had left a ssh key there for unraid, so it just hopped into Unraid from there.
AI doomers would call this misalignment and/or a hack. I call it an agent doing what it was asked to do and overcoming difficulties.
Every time I saw an agent do something "misaligned", it was always because there was something getting in its way that I didn't explain would happen.