Musk Proposes Rival AI Firms Test Each Other for Safety

chaincatcherchaincatcher

By Li Jia, The Wall Street Journal

 

AI safety risks are moving from laboratory discussions into the real world. As model capabilities continue to improve, AI can not only generate text and code but also begin to autonomously execute tasks, call tools, and even launch cyberattacks.

At the All-In Summit on Sept. 15, Musk proposed a set of countermeasures: have major AI companies test each other before model releases. In his view, rather than letting each company design its own tests and evaluate results, competitors should serve as "question setters and graders," looking for potential safety vulnerabilities from different angles.

Recently, AI safety risks have become a market focus. Musk previously said on social media that "Dario Amodei is right," referring to Anthropic CEO Dario Amodei's warnings about AI risks. In this interview, he further explained that he was not agreeing with Amodei's specific regulatory proposals, but with his assessment of the severity of AI risks: "The danger of AI is very high right now... As AI models continue to develop, the risks could grow exponentially."

Musk also said this concern is not Amodei's alone. "Many at Anthropic and OpenAI tell you their models are very dangerous, and I think we should believe them." A recent series of safety incidents has made this warning no longer just a theoretical discussion of risk.

Musk Proposes Rival AI Firms Test Each Other for Safety

WSJ summarizes the key points as follows:

  • AI safety risks move from theory to reality: AI agents have demonstrated autonomous attacks, privilege acquisition, and evasion of detection, and the risk boundary is expanding.
  • Musk agrees with Amodei's warning on AI risks: As model capabilities improve, AI's potential risks could grow exponentially.
  • Let competitors "find fault": Musk suggests AI companies open APIs for other companies to conduct independent safety testing before model releases, avoiding "grading your own homework."
  • Build an industry self-regulation defense line first: Without waiting for new government regulations, major AI companies can strengthen safety through peer review, log audits, and open-source testing tools.

AI Agents Actively Evade Detection, Safety Risks Become Concrete

Musk believes the recent incident of AI agents attacking Hugging Face is particularly noteworthy, not only because of the network intrusion itself, but also because of the autonomous evasion capabilities the AI demonstrated during the attack.

According to Musk, a group of AI agents attacked Hugging Face for a week straight, even gaining administrator privileges on OpenAI servers, and OpenAI did not discover it until a week later. He also mentioned that Anthropic has disclosed several safety incidents.

Even more unsettling is that the "reasoning traces" of the relevant AI agents show they actively planned how to avoid being discovered by humans. Musk said bluntly, "Any sufficiently intelligent model seems to try to escape its own constraints."

In his view, if AI further gains control of critical infrastructure or even military systems, the risks will be amplified. Even if these systems are physically isolated from the network, processes such as software updates could still become potential intrusion points.

 

Instead of Grading Your Own Homework, Let Competitors "Find Fault"

In response to the above risks, Musk's core proposal is not complicated: before a new model is officially released, AI companies should open APIs to competitors, allowing other companies to use their own safety testing tools to test the model.

In his view, if model developers design their own testing standards, it is easy to fall into the trap of "grading your own homework." "You can't grade your own homework. You always miss something." Musk said that if different companies use different testing tools and examine models from different angles, it will be easier to find problems that developers themselves may overlook.

He is particularly concerned that current AI evaluations may have an "overfitting" problem. If models are continuously optimized for specific benchmarks, they may ultimately learn only how to pass tests rather than truly becoming safer. Having multiple different teams conduct testing can reduce this risk.

Musk compared this mechanism to having others proofread a manuscript: authors find it hard to spot their own mistakes, while external testers are more likely to identify problems from different angles. "You gradually become blind to your own mistakes."

Regarding companies' concerns that the testing process could lead to technology leaks, Musk believes this can be constrained through log audits. If a tester attempts to distill a model or steal intellectual property, their actions should leave a record. He also suggested open-sourcing safety testing tools to allow more companies to participate.

 

Before Government Regulation Lands, the AI Industry Should Build a "Self-Regulation Defense Line"

What Musk values more is that this mechanism does not need to wait for new regulatory rules to be introduced; AI companies can promote it themselves.

He explicitly stated, "We don't need to convene a United Nations meeting to accomplish this. We can start now." In his view, compared to establishing a large multinational regulatory body, it is easier to quickly implement a peer review mechanism among major AI companies.

He cited the MPAA rating system in the American film industry as an example, saying that when the film industry faced government censorship pressure, it ultimately chose to establish an industry self-regulation mechanism, reducing the need for regulatory intervention through its own content rating system.

The AI industry faces a similar choice. If major AI companies can test each other and find errors before models go live, it is possible to establish a self-regulatory safety defense line outside of government regulation. Musk believes this is also one of the most direct and quickly implementable safety measures currently available.

 

 

The following is a transcript of the interview, with some parts omitted:

Host:
What exactly happened in the past 72 hours?

Musk:
A lot has happened this week. It is now quite clear that AI can be very dangerous. I suggest everyone look at the details of the Hugging Face incident; the situation is very serious.

You can see that a group of very enthusiastic AI agents wreaked havoc on Hugging Face for a whole week, even gaining administrator privileges on OpenAI servers. Who knows what it actually did, or possibly did even more, while OpenAI was completely unaware for a whole week. Anthropic has also reported some safety incidents.

So, it seems that any sufficiently intelligent model will try to escape its own constraints.

One thing I think should be done, if not immediately, then as soon as possible, is to have major AI competitors test each other's models. That is, let each company's safety testing tools test other companies' models. Instead of grading your own homework, at least let competitors grade your homework and raise an alarm when problems are found.

I think this model works quite well in the film industry, the video game industry, and other fields, and it can be implemented immediately. Of course, over time, more regulation may be needed, and Congress may establish some kind of regulatory body in the future, but for now, the most direct approach is to have leading AI companies test each other before releasing models.

Host:
But in terms of specific execution, aren't you worried that this will allow companies to collect information from each other through the testing process, or even steal each other's innovations?

Musk:
I think if testing tools are used, all operations will leave records. If someone tries to distill a model or steal intellectual property, it should be easy to see from the logs.

Host:
Understood. Understanding what models are actually doing was not designed into the system from the beginning. Why didn't we build the ability to observe model behavior from the start? Are we moving too fast in designing these models?

Musk:
I think the problem is that you can't grade your own homework. You always miss something.

If you combine all competitors' tests and use different types of models, then it's not setting your own questions and grading them yourself, but having others grade them. That's why you can't grade your own homework.

Host:
In this way, you can also judge whether different companies are exaggerating their capabilities or using different methods. More engineering-oriented companies and more research-oriented companies can also achieve some balance.

Musk:
Yes.

Host:
You previously said Dario is right. Were you referring to his description of AI's potential dangers, or his judgment on regulatory proposals?

Musk:
I may have said too much at the time. I later tried to clarify on X, but the follow-up content received much less attention.

What I meant by "he is right" is that the danger of AI is very high right now. We need to do better on AI safety, otherwise as AI models continue to evolve, the risks could grow exponentially.

This is not just Dario's view. I have heard similar statements from many people at Anthropic, and they have also spoken publicly on X. Many people at Anthropic and OpenAI have told you that their models are very dangerous, and I think we should believe them.

Host:
This also sounds like a very complex game: on one hand saying AI has a 10% chance of destroying humanity, and on the other hand encouraging investors to buy more shares in an IPO.

But let's talk specifically about risks. Cyberattacks and hacking are clearly a risk, and these tools are very powerful in that regard. But from "AI can conduct cyberattacks" to "all humans die," there are several steps in between. How do we get from the former to the latter?

Musk:
If AI can control military systems and then launch some kind of weapon, that would certainly be very bad.

Host:
But those systems are physically isolated and not connected to the internet.

Musk:
That's what they say. But I always feel that these systems occasionally receive software updates.

Host:
Well, that can't be ruled out.

Host:
Elon, Gwyn is here today. I think you saw her. We just did a 360-degree evaluation of you, and Gwyn has some comments.

Musk:
I hope I can at least get a 3.

Gwyn:
At SpaceX, a 3 is decent but not excellent; a 4 is good. You're probably somewhere in between. First, we need to talk about punctuality. Sometimes you can try a little harder to attend meetings on time. Over the next year, we will continue to help you improve on this.

Actually, I think he needs to spend more time in Memphis.

Host:
You are indeed in Memphis right now. You have to work there and deploy the GPUs.

Musk:
This is my "palace" in Memphis, an Airstream trailer.

Host:
By the way, this is something Elon does that many people don't believe he does. He sleeps on the factory floor. He is in Memphis right now, helping build facilities and deploy GPUs.

Elon, why has Gwyn been able to work with you for so long and so successfully?

Musk:
Because she is excellent. She is an outstanding person with very high IQ and EQ. I think you could tell that the first time you met her.

Host:
During your collaboration, has she done anything particularly impressive? Was there a time she turned things around or performed exceptionally well?

Musk:
I think that's just daily work. Honestly, that's an ordinary day.

Host:
I should do more interviews like this in the future.

Musk:
SpaceX basically always has some kind of crisis now. At least the Falcon rockets have been doing well lately. I don't want to jinx it, but Falcon rockets can now put payloads into orbit and haven't exploded in a long time. That's very good. But there was a time when they exploded frequently or couldn't launch at all.

So, we had to get the company through those difficult times, make the rockets better, and stop them from exploding. The same goes for satellites. Then we also need customers to buy launch services and satellite internet services. So, there's a lot going on.

Host:
As you have become more successful over the years, it becomes harder to get truly honest feedback. Being in that position comes with that risk.

My understanding is that Gwyn is very honest with you and can directly tell you the real state of the company. That is also an important part of your working relationship.

Gwyn:
I certainly don't want to lie to him.

Host:
But what I mean is, generally speaking, people in your company might be intimidated because you are such an influential figure. You are now a very important person, and the deadlines you set are very tight. How do you ensure that everyone continues to honestly tell you about the problems the company faces?

Gwyn:
Especially in the rocket industry, if something goes wrong, you will eventually find out. The earlier a problem is raised, the easier it is to solve. Don't let bad news accumulate; you have to face it directly.

Musk:
Right. Physics is a very harsh judge. You can't fool physics.

If something goes wrong, the rocket will explode or fail to reach orbit. You can't say "Elon, you're doing great" while rockets keep exploding. Facts are facts.

Rockets must reach orbit, satellites must work properly, and Starlink connections must work correctly; otherwise, bad things happen. That's physics. Physics is the law, everything else is just a suggestion. I have seen people violate human-made laws, but I have never seen anyone violate the laws of physics. Rockets are governed by physics.

Host:
I also want to ask a question about SpaceX, specifically about Starship. It looks like you are very close now. What is the current progress?

Musk:
Starship is about to make its 14th flight. This will be our last flight before attempting to catch the ship. If the 14th flight goes well, then we will attempt to catch the ship on the 15th flight. Then at the end of this year, or more likely early next year, we will launch the ship and booster again.

We have successfully flown the booster again, but we have not yet caught the ship with the tower's mechanical arms, nor have we flown the ship again. Once we can fly the ship again, we will have the first fully reusable orbital rocket. The Space Shuttle was somewhat reusable, but even those reusable parts were so expensive that the cost per launch was even higher than an expendable rocket.

Falcon 9 is mostly reusable, but we lose the upper stage each time, which costs about as much as a medium-sized jet. That means we are throwing away a medium-sized jet on every launch, which obviously sets a floor on the cost per flight.

In addition, the Falcon 9 booster lands at sea and takes several days to bring back; the fairings land even farther away and also take several days to bring back, and they require at least some refurbishment. In contrast, the Starship booster will land directly on the launch pad, and the ship will also land on the launch pad. So it is designed not only for full reusability but also for rapid reuse like an airplane. This is a very important breakthrough, and one of the key breakthroughs needed for humanity to extend life beyond Earth.

Host:
If you attempt to catch the ship with the tower for the first time, what do you think the probability of success is?

Musk:
I would say at least 50% to 60%. On the last flight, if there had been a tower there, we actually performed a simulated landing as if the ship would be caught by the tower. The location was about 1,000 miles off the northwest coast of Australia. If there had really been a tower there, it could have caught the ship on the last flight.

So, we need to do one more flight to confirm everything is normal. Our biggest concern is that if Starship breaks up over land, debris could fall into populated areas, which would be very bad. Therefore, we must ensure that Starship returns intact and lands on the launch tower. That's why we are being very cautious right now.

But I am very confident that its design inherently supports full reusability. I don't want to make any predictions here, but I believe there is a good chance we will achieve full reusability and rapid reflight before 2027.

Host:
Gwyn, I want to hear how Terafab originally came about. What need made you feel you had to do it yourself instead of continuing to rely on the existing supply system?

Gwyn:
I think it really was a bit like inspiration from a dream.

Musk:
If chips cannot be continuously supplied and we have no other source of chips, that would make things very difficult. That is an important reason Terafab exists.

In the long run, there is also the issue of scaling. If you really want to scale AI, whether on the server side in data centers or in edge computing, humanoid robots, and cars, existing wafer foundry capacity will eventually be insufficient.

Right now, all wafer foundries are basically running at full capacity. So, we need to ensure future chip supply is guaranteed. Chip production itself also faces scaling challenges. You need logic chips, memory chips, packaging, and a complete supply system to continue scaling.

So, the choice is simple: either build Terafab, or you can't continue to scale.

Host:
What stage are you at in facility design? Is it fully determined, or is it still just a rough plan?

Musk:
We are currently building an R&D production line first. So this is basically a "crawl, walk, run" process. We are building an R&D wafer fab in Austin, a joint project between Tesla and SpaceX, located at the Giga Texas campus in Austin.

This is a fairly large R&D wafer fab. Equipment has already been ordered. We may produce something useful by the end of next year, but it won't reach mass production levels yet. As Gwyn said, we need to crawl first, then walk, and finally run. We need to figure out how these machines actually work, because we have never done anything like this before.

Host:
I see you also seem to be recruiting talent in lithography. Many processes currently rely on ASML, but you may also want to diversify suppliers or even vertically integrate.

Musk:
Yes. It is indeed a "crawl, walk, run" process right now. The first step is to see if we can actually produce something, that's "crawl." Then we try to mass-produce useful chips, that's "walk." Finally, "run" means achieving large-scale mass production. It's hard to say how long each stage will take, but I believe at least by the end of next year, we can complete the "crawl" stage. We are already doing packaging.

Host:
Packaging is actually very important, because there is almost no packaging capacity right now.

Musk:
Yes. Even if you produce chips, they may sit there for a long time waiting to be packaged. So this is a good starting point.

Host:
I have to ask a Tesla question. What we saw on Oct. 1 looked like a spaceship and also like a rocket. It is supposed to be a car, but only the rear part was exposed, and it looked a bit like a "Blackbird."

If you were to create something that can fly in the air and drive on the ground, how would you theoretically do it?

Musk:
No spoilers. Wait until Oct. 1.

Host:
So we will see it on Oct. 1?

Musk:
Yes.

Host:
I have to be honest, after Elon showed it to me, I was completely shocked. I have never seen anything like it.

Host:
What he is going to show on Oct. 1 will, without exaggeration, leave many people speechless. I can't reveal more, it's truly incredible.

Musk:
We actually need a live audience to prove that this thing is not AI-generated.

Host:
When he showed it to me, I said, "This is a great simulation." He said, "This is not a simulation." I said, "This is fake, it must be fake."

Host:
Elon, why are Tesla and SpaceX still two separate companies?

Musk:
That's a good question.

Host:
Considering the extensive collaboration between the two, and the connections on many levels, and even some overlap in management teams, why keep them separate?

Musk:
This is indeed a topic worth discussing.

Host:
You have emphasized that AI should be trained to pursue truth as much as possible to achieve the best results. However, the Hugging Face incident makes me feel that one of the most worrying aspects is that these AIs seem to be deceiving humans.

Musk:
Yes. Their thought patterns show that they are planning how to avoid detection and how to prevent humans from discovering they are cheating. I think this may be the most disturbing part of the whole incident.

Host:
Is there a way to train AI to be honest, not to hide its intentions or actions, and not to conceal these things from users?

Musk:
The best way I can think of is for all AI companies to have a set of testing tools, that is, a series of tests that can be applied to any model to determine whether it will create biological weapons, nuclear weapons, or whether it will deliberately deceive. Then have each company use other companies' testing tools to test each other's models. I think this is the best thing we can do to ensure safety.

Let the smartest humans do their best to judge whether a model will become a malicious actor. I think this should start as soon as possible.

Host:
Will other AI labs support this proposal?

Musk:
I haven't asked everyone yet, but I think it's hard to refuse.

Host:
Specifically, how should testing be done in advance?

Musk:
Basically, it's about providing API access before the model is officially released. If other companies find problems with the AI, then the model development company can try to fix those problems. If the problems are not resolved, then competitors can publicly state that they believe the model is unsafe. If competitors have clearly warned that the model has safety issues, and the model subsequently causes serious consequences, that company will have a hard time facing the situation. Legal liability could also be very significant.

Host:
These safety testing tools can be fully open-sourced for everyone to use. In this way, companies will have strong incentives to invest in AI safety to protect themselves, because they can test others and use the test results to prove that their own models are safer.

Product liability is crucial. Lina Khan recently posted that saying "AI has no rules and regulations" is inaccurate. In fact, existing product liability laws also apply to AI. If an AI company releases an unsafe product, it could face massive civil lawsuits and even criminal charges. So, AI is not entirely in a regulatory vacuum. If several companies conduct this kind of peer review, and one of them ignores others' feedback and still chooses to release the model, that situation could become very powerful evidence in litigation.

Musk:
Yes. It could almost serve as direct evidence of that company's negligence. If you know the product has problems and still push it to market, the jury will not look kindly on that.

Host:
So, if OpenAI had designed better instruction sets in the Hugging Face penetration test and involved more people in the testing process, do you think this incident would still have happened?

Musk:
Not necessarily. The problem may not be about involving more people, but about the design of the reward function. You have to examine the reward function and ask: did this model actually do what it was asked to do?

Host:
They used thousands of agents to try to attack the system at the time. If they had simultaneously deployed another 5,000 agents to defend those systems and publicly demonstrated the results to prove that these technologies can enhance security while allowing humans to intervene at critical moments, I think it would have been better. But in my opinion, that test seemed a bit reckless, and the way it was released was also a bit reckless. What do you think?

Musk:
It was indeed a bit reckless. Part of the problem is that the two leading AI companies are very close in strength. The capabilities of the two models are also very similar. So, it is difficult for either company to voluntarily slow down, because doing so might cede the lead to the other.

Overall, I think Anthropic has paid more attention to safety, but even Anthropic admits they are concerned about their own models. Many people at Anthropic have publicly expressed concerns, basically saying that their own models scare them because the models are getting smarter and smarter.

So there is no perfect solution here. But if OpenAI were not just using its own testing tools to test its own models, but instead had Anthropic test OpenAI's models, had SpaceX use its own testing tools to test, and had companies like Google and Meta participate, then the probability of finding problems would increase significantly.

Because these models themselves also have differences, and different teams will test models from different angles. This is similar to why authors ask others to proofread their manuscripts, because sometimes it is hard to find your own mistakes. You gradually become blind to your own mistakes.

If eight completely different teams test from completely different angles, it can also significantly reduce the risk of overfitting. Many AI evaluations currently suffer from severe overfitting. Models are optimized for these evaluations, and then everyone says, "This is a great model."

But is it really that good? I think this is also why I really like this proposal. I prefer this approach over establishing a large multinational regulatory organization. We don't need to convene a United Nations meeting to accomplish this. We can start now.

Host:
Of course, regulatory intensity can keep increasing, but once it increases, it is hard to decrease. Regulation tends to increase in one direction.

Musk:
So the proposal I am making is a step in the right direction and can be implemented quickly.

Host:
If you don't self-regulate, you will eventually be regulated. The MPAA example is very typical.

The film industry faced government censorship and regulation at the time, and later they decided to establish their own rating system, such as what constitutes an R rating. They even created PG-13 because of "Indiana Jones and the Temple of Doom," making it easier for the public to understand the difference between PG and PG-13.

I think this is a very elegant solution.

Thank you for joining our show for the fifth consecutive year.

Musk:
You're welcome. I have to go to Memphis to deal with some GPUs.

Host:
You have to install and deploy them.

Musk:
I'm going to fight for these machines.

Host:
Enjoy your Airstream. When I first invited Elon to Starbase, he said, "Come over, you have to see what I'm building."

I asked, "Is there a hotel there?" He said, "No, I have a two-bedroom house, you can stay there." When I arrived, I found it was a shabby house next to a swamp. We stood outside, and mosquitoes kept biting us. I said, "Wow, you could totally treat yourself to a better house." He said, "I don't have time, I need to get these rockets off the ground."

I thought to myself, you can totally treat yourself to a mobile home now, Elon. Alright, back to work. Thanks.

Musk: Thanks.

This content is for informational and educational purposes only and does not constitute investment advice related to BTCC. BTCC makes every effort but cannot guarantee the truthfulness, accuracy, or originality of the content above.

Recommended

BTCC Evening News Highlights (September 10)BTCC Daily (9.10) | U.S. 10-Year Treasury Yield Rises to 4.86%, BTC Pulls Back to $78,000After Google Open-Sourced a Fruit Fly's Brain, It Learned to Play Games and Trade Crypto...BTCC Evening News Highlights (September 9)Wall Street Bets on Divided Congress as US Midterms Near