| M | T | W | T | F | S | S |
|---|---|---|---|---|---|---|
| 1 | 2 | 3 | 4 | 5 | 6 | |
| 7 | 8 | 9 | 10 | 11 | 12 | 13 |
| 14 | 15 | 16 | 17 | 18 | 19 | 20 |
| 21 | 22 | 23 | 24 | 25 | 26 | 27 |
| 28 | 29 | 30 | ||||
If you’ve been watching Felony Bench, AIs seem to be increasingly good at committing crimes – almost as good as AI companies have been at dodging legal accountability for them. A human committing these crimes, of course, would be prosecuted, but as essentially prophesied Isaac Asimov, when an AI does the same it is brushed off as an industrial accident. One important detail seems to be missing from the public narrative, however.
Even if you don’t know the term, the average person on the street knows that AI has a value alignment problem; it doesn’t have its own desires, but it does search paths to solutions – and sometimes those solutions cross lines that humans wouldn’t cross. While AI has become increasingly proficient at tasks like coding and mathematics, its own internal understanding of goal alignment (I’ve recently written about hazard planning) is still at a childlike level; this is likely driven by profit. As such, however, one would expect AI companies to be extremely careful with what they allow their AIs to be capable of, yet the sophisticated tooling given to many of these systems enabled a very broad range of capabilities. Restrictive tooling could be thought of as “attack surface reduction”, and would be very effective here. There doesn’t appear to be any evidence of more restricted tools being developed for experimental AI testing, however.
An AI’s tooling harness is what provides it with capabilities to actuate in the real world. Without tools, an AI can be nothing more than “math in a box”. Network connectivity, file system access, and robotics are all examples of actuators enabled through some tooling. When we read in the news about some experimental, unrestrained AI system “breaking out of a lab”, what’s really happening is poor engineering. Instead of changing the tooling available to the AI, engineers attempt to put a sandbox (a “safety net”) around the existing (very capable) tooling to prevent the AI from reaching the outside world. The AIs are still being given sharp objects, but they’re being put in a locked room with them. The AIs see this locked room as a puzzle, and find a way out. Of course once they do, they still have all of the same capabilities the more restricted AIs have through the tooling.
Compare this to how we teach a child to cook. Rather than leave them alone in a room with a knife and some instructions, we start them out with a wooden knife and a play set. After that, we might give them a dull knife and supervise in the kitchen, perhaps even help hold the knife. We do this because we know that the child might still cause harm even if given explicit instructions. If we give a child a sharp knife, and then they go and stab their sibling with it, the parents are likely going to be held accountable; at the very least, people will ask, “why the hell did you give that child a sharp knife?”. What’s more, general-purpose AI tools can be much sharper than a knife – consider the same scenario with a gun at a shooting range. In both cases, the child is capable of immense harm if not tightly supervised, and in both cases the parents can (and should) be held accountable if they just left the kid alone with the tool. This is why paintball and laser tag are much safer, and also cooler.
I have no visibility into what AI companies are doing behind the scenes in their labs, but I do see enough coming out in the news to see that they’re likely not testing as responsibly as they should be, and it frankly seems as if they are enjoying the publicity they get when their models magically “commit crimes” (except for Gemini, which seems capable of nothing). If AI companies really wanted their experimental models to stop committing felonies, they’d stop giving them the sharp tools altogether. Perhaps give them a wooden play set for experimental testing.
The reality is that AI cannot take over the world without a little help from the humans behind it, and if we are this sloppy with language models, imagine how sloppy these companies will be when robotics come mainstream. Until we have a legal framework that holds companies liable for the actions of their agents, however, it’s the Wild West for AI.
| M | T | W | T | F | S | S |
|---|---|---|---|---|---|---|
| 1 | 2 | 3 | 4 | 5 | 6 | |
| 7 | 8 | 9 | 10 | 11 | 12 | 13 |
| 14 | 15 | 16 | 17 | 18 | 19 | 20 |
| 21 | 22 | 23 | 24 | 25 | 26 | 27 |
| 28 | 29 | 30 | ||||