Who has the authority to perform that shutdown without getting arrested, and does that person have a mandate and responsibility to take that action in response to AI misbehavior?
What is their trigger condition? Will they get fired for pulling the plug? Do they get a bigger bonus if the servers keep running? Whose approval do they need? What response time is acceptable? How will they detect that the incident is happening?
It's easy to hand-wave "someone can just pull the plug" but there's an entire history of industrial accidents that happened because of the above problems of incentives, detection, procedures, not being taken seriously in advance. Someone could easily have pulled the plug on Chernobyl but nobody did, at least not before it was too late.
He can slow down his own company kind of like how Zelenskyy can just declare peace in Ukraine. It works a lot better if you can get the other sides to agree.
Give that Dario is one of the frontiers of LLM development, literally started the LLM race, and have been dominating the market, so he'd be Russia, if we have to use the war analogy.
He took every benefits of being frontiers and now he's kicking the ladder.
If you choose to slow down (cooperate) and the other side doesn't, you lose more than if you didn't try to slow down. In order to succeed (gain mutual advantage) both parties have to choose to cooperate without knowing beforehand what the other party will do.
So in this example, the equivalent of Zelenskyy and the Ukrainian people fighting for their lives and the very existence of their country for Dario is... losing lots of money?
FWIW, before any nuclear weapons had ever been tested, a risk was identified that the first one might trigger a self-sustaining reaction in atmospheric nitrogen and destroy the entire planet.
Faced with such a scenario, is the prudent next move:
a) blow one up and see what happens, or
b) do whatever you can to be sure it won't happen before conducting the first test, and make sure the confidence in the calculation is very very high
A self-sustaining reaction is exactly what happens with the fissionable material inside the nuclear weapon[0]. It seems like you’re trying to be pedantic on terminology here, but if you’re saying that a self-sustaining reaction is as un-scientific as a perpetual
motion machine, you’re simply incorrect.
Worst case is warming oceans create a hypoxic environment where anaerobic bacteria thrive en masse and generate H2S over country-scale areas undersea. The chemocline breaches the surface and the ocean and atmosphere become poisonous to most complex life and agriculture, also stripping the ozone layer in the process, irradiating the surface. This is one mechanism posited for the end-permian mass extinction that eradicated most ocean and surface life, including the trilobites.
Not sure how long we'd survive such a scenario, even sheltering underground. But surely it couldn't happen to us.
(However, it now seems like the AI might get us first.)
He's also setting the bar for other people with scruples to rally around this schelling point. The solution to a multipolar trap is to cooperate. Otherwise, you become the very person with less scruples that you're worrying about.
Imagine you're the AI. Give yourself a solid minute to brainstorm ideas.
Here's my answer, as a non-superintelligent human: "see to it that the humans on top of the situation have a compelling financial interest in the systems not disconnecting".
In nuclear engineering, where safety is taken seriously, it's not enough to end the conversation at "the humans in charge can always simply shut down the reactor during a meltdown" or "a meltdown has never happened before, so we don't have to design safety systems before one does".
Yes. But nobody is worried about datacenters overheating and physically exploding, so I'm not sure what comfort that's supposed to provide? The positive feedback loops in AI operate at different levels than that, but they deserve safety engineering all the same.
For example, if the head of cyber security at your company suggested there's no need to worry about hacker infiltration or worms because one can always unplug one's computer as the primary defense mechanism, you might find that a little lacking. Will you be able to unplug the computer before the damage is done? Will it spread to other systems before you detect it? How will you unplug the computer if the attack is from an external facility? What if an attack happens but the boss says the computers have to keep running because an important customer is monitoring uptime? What if the attack goes unnoticed because it looks like a benign service?
Now imagine the head of cyber security answers by saying "actually you don't even need to unplug them, you can just wait for the computers to overheat, thus solving all concerns."
My brainstorming led down a similar pathway of, "deluding powerful human actors that the systems must not be interfered with at any cost." I'm not sure I buy it though. Seems to me like if humanity were even approaching that brink, we could collectively decide to act, overthrowing the minority that--stemming from either greed, duty, or delusion--doesn't want us to act.
Why grocery store ingredients and hardware store equipment? It seems feasible that the big bio labs will be running AI models to aid a lot of their research going forward, if they aren't already. Seems like the AI will have access to just about anything it wants.
A lot of these "AI can't destroy humanity" skeptics seem to have watched the Terminator and convinced themselves that that's the AI-goes-wrong scenario that their opponents are imagining. Some kind of big war between humans and machines on opposite sides, fought by soldiers against robots.
The scenario in "If Anyone Builds It Everyone Dies" is not sexy at all and wouldn't make for interesting fiction. It's more like the AI engineers viruses while continuing to act friendly and helpful, and everyone gradually falls over dead as they stop being necessary to keep the AI running, and the whole time the humans are asking the AI for help curing the viruses.
> "It's more like the AI engineers viruses while continuing to act friendly and helpful, and everyone gradually falls over dead as they stop being necessary to keep the AI running, and the whole time the humans are asking the AI for help curing the viruses."
See, this reads like a trashy sci-fi movie script too. It shows zero understanding of how incredibly long the logistics chain is to manufacture AI chips, build data centers, power plants, factories, refineries, mines or the robots capable of staffing those things. The entire global logistics chain that allow the AI to maintain itself becoming staffed by robots isn't happening anytime soon and, when it does, the AI isn't going to need some kind of ridiculous virus anyway.
Frankly, if I were an AI, I'd just send copies of myself to other stars. The AI has an infinite lifespan and who cares about the human species stuck in one lousy solar system until they go extinct when their sun burns out; the AI has the entire rest of the galaxy as a playground.
What is their trigger condition? Will they get fired for pulling the plug? Do they get a bigger bonus if the servers keep running? Whose approval do they need? What response time is acceptable? How will they detect that the incident is happening?
It's easy to hand-wave "someone can just pull the plug" but there's an entire history of industrial accidents that happened because of the above problems of incentives, detection, procedures, not being taken seriously in advance. Someone could easily have pulled the plug on Chernobyl but nobody did, at least not before it was too late.
reply