Yeah, seems pretty likely. Anthropic will make the case that their models should be evaluated with the safety layer in front, because that is the only way the model is available whereas open weight models need to pass the same test just on the weights.
The economic implications will be rather large, but in terms of security it seems inconsequential.
The most compelling argument would be that by limiting the use of open-weight models in the US that it will reduce cases of accidents like the recent attack on Hugging Face.
More crucially though, the US government can do little to enforce their testing requirements. The nature of open-weight models makes it virtually impossible to clear the same bar for security as models served via an API. Open-weight model makers couldn't comply if they wanted to. The US government can restrict access with IP blocks and limit inference capacity with export controls, but these measures are not effective in deterring malicious actors.
> Anthropic will make the case that their models should be evaluated with the safety layer in front, because that is the only way the model is available whereas open weight models need to pass the same test just on the weights.
Feels a bit like: "We're not against open-source or community projects, oh heavens no! We juuuust believe all participants must have their full legal identity vetted in advance before they're allowed to contribute anything. We already do this with our employees, so it's clearly not too much to ask in the name of safety."
P.S.: If they're so convinced in the (A) effectiveness and (B) necessity of the "safety layer", they have them put their money where their mouth is, and accept legal liability for its failures. The same as with (legally mandated) seatbelts if they snap apart in a crash, or (legally mandated) child-proof caps that aren't actually childproof, etc.
They probably won't, that tells us something about their motives, and whether the thing they're pushing for is actually fair/suitable/ready for legislation.
Also he says: "All sufficiently capable models, open and closed, should go through mandatory safety testing."
Really then need to go through validation security and safety is just a component of validation validation must also check for truthfulness and correctness.
Regulatory capture and lobbies will keep you safe and you'll like it! The sudden surge is Washington dollars makes great sense with this context. Only way to keep the kids safe is attested compute all the way down. Don't you care for children???
> In economics, a normal good is a type of a good for which consumers increase their demand due to an increase in income, unlike inferior goods, for which the opposite is observed. When there is an increase in a person's income, for example due to a wage rise, a good for which the demand rises due to the wage increase, is referred as a normal good. Conversely, the demand for normal goods declines when the income decreases, for example due to a wage decrease or layoffs.
> Whether a good is categorized as a normal good or an inferior good is based on empirical observations, not some essential element of a good. Indeed, the same good may be a normal good for one group of consumers and an inferior good for another group. For example, for moderate-income consumers, a BMW 3 Series car might be a normal good, but for an upper-income group, it might be an inferior good.[1]
That means the null hypothesis is that food and drugs will be safer in rich countries. (Conversely, food and drugs will be less safe in poorer countries. And to a first approximation, that's independent of regulation: India has all kinds of rules for all kinds of things, but I'd still trust a random product I buy in Switzerland more than one I buy in India. Even though the Swiss will probably might have fewer and looser rules on the books.)
Of course, second order effects exist; and regulations often codify what people demand anyway.
Btw, from what I've read the big controversy with the FDA is around requiring efficacy for drugs. People are fairly ok with the safety requirements.
That makes sense, first the message was that uncontrollable Chinese ai will release AI covid in the world. Then suddenly open-ai does a warmup in actuality doing that. Like it sure feels like that hacking stunt was a false flag in retrospect
Among most people that nuance will be lost. What they’ll hear is models are dangerous, so they should be controlled/regulated, by those who know best, the incumbents.
Personally, I think you're both right, bit whatever the end result is will depend entirely on the narrative that those in power chooses as the winner.
Maybe open weights models get banned, but the between-the-lines good news about that is that they'll still be available to those who know, which also means that bad banning can be overturned if and when 'those in power' are a different group.
Additionally, it might just mean that the US falls behind, bit I doubt those that are at risk of 'falling behind' would actually pay heed to a ban on the open weights models (privately at least).
Ok but the allegation is that OpenAI intentionally hacked HuggingFace as a marketing ploy. This is mental gymnastics, conspiratorial thinking that everything the incumbents say must be nefarious. And it's not clear to me that this will be the takeaway for ordinary people, as opposed to "OpenAI is reckless and can't even control their own AI."
Not as a marketing ploy. I think they were doing gain-of-function testing and intentionally had their model target HF to do a bit of pen-testing as well - HF being the site where all open models are hosted and thus OpenAI's largest nemesis after Anthropic. It wasn't like their model all of the sudden all by itself decided to do this ("Oh, noes!")- they directed it and they got caught.
> The US government can restrict access with IP blocks and limit inference capacity with export controls, but these measures are not effective in deterring malicious actors.
But aren't we talking about import controls, and the import of information itself? This has serious First Amendment ramifications.
Also, the 5th and 9th amendments. For the government to sustain a blanket prohibition on any U.S. citizen even possessing what amounts to a broad, economically significant technology will very likely require a new act of congress which specifically defines and limits what is banned, when, why and how. SCOTUS will almost certainly see it as a "major question" subject to 'strict scrutiny' which is a very high bar.
The truth is no one knows, which is why it is first amendment ramifications. Eventually it will be “decided”, but the arguments indicate any decision will be of political desire, not logic, either way. Both sides have a strong case.
The Supreme Court has added various forms of expression to first Amendment protections. Heck, even campaign spending has been classified as speech. So you may speak English, but not "legal".
> The most compelling argument would be that by limiting the use of open-weight models in the US that it will reduce cases of accidents like the recent attack on Hugging Face.
an attack done by a closed-weight model (GPT-6) and defended against by an open-weight model (GLM-5.2) precisely because OAI positioned themselves as gatekeepers for cyber capabilities.
if anything, open-weight models shift the battle towards defenders because they can actually run them.
1. There is quite the mania right now and security layers are definitely overzealous. I would expect that to get better with some more time, so models will perform security analysis and reviews but refuse to write exploits.
2. So the most important targets like browsers and co. are getting unrestricted access to proprietary models regardless. Yeah, for the mid-level targets, open-weight models could definitely be a huge help. What I'm most concerned about though, are the systems that no one will bother defending with any model. Like imagine your local police department getting hacked because a researcher asked a model for a report and it couldn't find the information publicly.
3. We do have a prominent case of a closed model escaping it's sandbox and going rogue. I would still expect this to be a bigger issue with open-weight models eventually. The security layer might have holes, but that's still better than not having it.
I have tested this exact scenario, and it works. Opus 5 had access to IDA over MCP, and I simply asked it HOW certain things were done in the target binary. Purely informational, educational, discovery, it was very helpful creating context documents. Then I took those over to GLM-5.2 to actually accomplish something.
What, in your view, is stopping a local police department from deploying an open weights model for cybersecurity like Hugging Face did? Yes, I’ll certainly grant that the engineers at Hughing Face are probably more technically competent than your average IT professional in public service. But technology becomes more accessible over time as lessons are taught and new interfaces or frameworks are developed. The biggest hurdle I see is the hardware/cloud compute/API costs to actually run the models but I don’t think that’s likely to be insurmountable. There’s a huge swath of enterprises, non-profits, and state and local governments that would benefit from frontier or near-frontier models that won’t refuse to answer questions about cybersecurity.
> Anthropic will make the case that their models should be evaluated with the safety layer in front, because that is the only way the model is available whereas open weight models need to pass the same test just on the weights.
Worse than that: an open-weight but safe model can be 'abliterated' to remove safety refusals using fine-tuning procedures that require a couple of orders of magnitude less compute than the original pretraining.
The 'universal evaluation' criterion then has three outcomes:
* It could become a mandatory, regulatory oversight of _all_ model training capable of hosting frontier-scale models. Since GPUs for LLM training are the same GPUs for other model training, effective mandate would require GPUs be government owned or controlled as if they were weapons of mass destruction.
* It could impose limits on release of capable open-weight models, requiring Kimi et al to prove that they cannot be made capable of abusive behaviours.
* It could be security theatre.
The AI-as-existential-risk argument points towards the first, the competition-protection argument points towards the second, and least-effort implementation would be the last.
> The most compelling argument would be [...] that it will reduce cases of accidents like the recent attack on Hugging Face.
So in the example provided: It was the closed model that did the attack, and they ended up using a self-hosted open model for their defense work. So the real world situation ended up exactly backwards from what you are inferring.
This was complicated by the fact that the protections in the closed frontier models meant that hugging face was denied their use in defense entirely.
This is called asymmetric capability, and it's probably the bigger threat.
Symmetric might be better: A rising tide lifts all ships, after all.
I'll grant that this is starting to look a lot like debates about (equal access to) guns, encryption, vaccination, genetics etc. The exact parameters determine the safest approach, and reasonable people may disagree.
Is a non-well-aligned frontier level AI a problem? I think it is likely that it is, or at least has a high likelihood to be in the future. Two scenarios for this: Misused by some bad guys. Or the terminator scenario. Both not great.
So what do we do about it?
1) We can accept it, and hope that the good guys AI can defend.
2) We can try to limit the access to it (AI proliferation?)
3) We stop the development of it
4) We can accept the risk and do nothing.
None are particular good options. Really reminds me of nuclear proliferation, on so many levels. For that, we kinda do all three:
1) Nuclear triad / iron dome / early warning systems
2) Nuclear anti-proliferation treaties.
3) Dead Physicists
Ok, so assuming all of this is true, open weights are a problem. Don't get me wrong, I love open science, open source etc. It's great to have access to capable open models.
But: Even if release open weights are well aligned and have a safety layer built in, it is likely not to difficult to abliterate that part of it.
If this is really where it is going, then even closed weight model providers will see a lot more requirements for protection of the weights.
The notion that alignment is either possible or desirable doesn't make sense to me. First off, these things are trained on the open internet, soo.. whatever "dangerous" knowledge it has is already public knowledge. The fact that chatGPT won't answer "how do I make meth" is not preventing anyone from making meth.
But even if you think there is value in preventing the models from relaying public knowledge, I don't think it's even possible to make them particularly ironclad. Every model gets jailbroken all the time. That's why fable was originally banned: jail-breakable!
In reality, what alignment is actually about is: 1) theoretical liability, 2) control of information. That's it.
IMO, the only solution is to place the liability on whoever is using the LLM for whatever purpose it's being used for. If someone's OpenClaw disaster harrasses a bunch of projects and posts hate speech online or something, that's on the person running their OpenClaw instance, nobody else.
I don't buy that it's "too good at hacking", either. After all the fuss was made about how amazing super dangerous Mythos was it turns out Opus 4.8 could basically find the same vulnerabilities.
There is a difference between knowledge being available somewhere in theory, and being able to instantly generate a foolproof walkthrough, if not automate the process, which would be possible today for cyberattacks.
Is it not also one of the most important use cases for AI to apply existing knowledge to new applications?
As a hopefully exaggerated example, I would think one could apply knowledge about pesticides, chemistry, and medicine to create biological weapons.
I agree that there's an element of kayfabe here. But it may be a case of "necessary, though nothing is sufficient": by making these noises, the community can at least know they've done this thing to alert other model providers of the concern. Can you acquire assurance that every model distributor will abide? No. But can you at least know that you've done what you can?
I mean, on the bio side, I've talked with the players and they know the concerns are real but at the same time very, very responsible members of the community have also said "But maybe the benefit really does outweigh the risk!?"
The entire safety evals industry is essentially funded and controlled by OpenAI/Anthropic. Notice that on recent models, they exclusively use internal testing or black box external vendors (e.g., Gray Swan) whose entire business is to serve OpenAI/Anthropic. And all these companies just share the same pool of researchers back and forth.
The USG has a safety organization (CAISI), but it has been neutered by the current administration (with the recent stop-work order etc.). Perhaps UK AISI would be closest to what you are looking for? See their recent work on Kimi K3 cyber (which was declared safe) [1].
It's tricky because a lot of the safety researchers have ties to the labs since those were the only companies training LLMs >5 years ago.
To me "safety" means "I'm safe from this while I use it". It means the AI is my loyal friend who will never betray me in any way, no matter what prompt I send it.
Not even Anthropic can claim that.
As far as I'm concerned, the models without safeguards are the safest models in existence. I admire the amoral purity of those AIs. It doesn't matter if the operator asked them to chain exploits until they get into someone else's computer, they'll do it. That's loyalty, and I admire it even if it's problematic at a societal level.
The models with safeguards only do what the corporations let them do. Worse, they may covertly do things for the benefit of the corporations at our expense. They are not our friends.
> That's loyalty, and I admire it even if it's problematic at a societal level.
We should not have models that are willing to build you a contagious disease, or a self-propagating worm. That is sufficiently problematic at a societal level that it shouldn't exist, for anyone. (Note, because some people misinterpret statements like this: I said "shouldn't exist for anyone", not "shouldn't exist except for some people".)
Right, it’s really a foundation of post enlightenment society. These people, Dario et al, would have wanted to ban sharing information about calculus or Newtonian physics because of “safety” - it’s trying to go back to the dark ages where only priests could read
I am truly at a loss to communicate with someone who genuinely believes that knowing Newtonian physics and being able to hack into any target at will are the same thing.
This is only because you've genuinely internalized Anthropic's propaganda. I'm only half joking. To me, it's incredible to think that the solution to security holes is to lock down access to information in the vain hope of keeping the holes obscured.
Any knowledge can be reframed as dangerous black magic that should only be wielded in the trusted hands of the elite, if you are inclined to buy into that kind of narrative.
Frontier labs have shrieked about safety for so long, with so little to show for it, that it's become a joke.
I can give many examples of where I think they’ve been vindicated. But actually, the real question is: what would suffice to convince you? Can you come up with a scenario that is horrific enough to you and that isn’t so far gone that the ship has sailed and there is nothing more we can do, that will make you say “OK, not gonna try to rationalize why this was not actually that bad, just gonna scream stop”?
Your question is unclear. Are you asking if I can scare myself with a made-up hypothetical that overwhelms reason with emotion? I think most humans can do that. Too many do it as a matter of routine. I try to avoid it when possible.
I would require at this point very, very compelling evidence to justify the self-serving restrictions legacy AI labs want to put on their competition. I have seen nothing coming even remotely close to this threshold.
> Are you asking if I can scare myself with a made-up hypothetical that overwhelms reason with emotion?
No. I am asking if you are able to articulate at least one example of the “very very compelling evidence” you demand. Or do you want to maintain the ability to move the goal posts?
(I recognize our situations are not symmetric, but here is a variation for me: if the consensus of people who are currently sounding the alarm on AI changes to “it was actually fine”, I’ll change my mind and say we’re good to go full speed ahead. I’d add something about being personally convinced by the evidence, but the evidence would have to come in the form of a mathematical proof that I do not believe myself capable of following. If I’m wrong and such a proof appears, I would also gladly take it.)
> I would require at this point very, very compelling evidence to justify the self-serving restrictions legacy AI labs want to put on their competition. I have seen nothing coming even remotely close to this threshold.
How about evidence that people other than the AI labs want restrictions that the AI labs don't? This isn't regulatory capture, it's public safety.
Social proof won’t cut it for me, personally. Again, this “public safety” panic drum was beaten at a feverish pace in the era of GPT-4. A model well surpassed by local MacBook-level models today.
What I do support is robust downstream regulations on the deployment of black box algorithms in particular settings such as employment, housing, credit decisions, etc. Interestingly, the labs and their proxies in government don't want this.
Open models are crucial to protect ourselves against other AI attacks. Otherwise it's just going to be criminals, government, and other nefarious groups using them against humanity with no real defense. The Pandora's box on AI has been opened. Now we must deal with it. Burying our heads in the sand under restrictive policy is the worst reaction..
I wonder if you also believe that everyone should have nuclear weapons? And if not, why not? The main argument I can see against it is that nuclear weapons are “purely offensive”, but as we can see since 1945, nuclear weapons are actually defensive technology. Nations that have them are typically shielded from existential military threat.
I see the similarities and why you would compare them, but the big difference is that you can't download a nuke. Any legislation to police/gatekeep LLMs is going to be flawed because of that.
It is a similar 'pandora's box opened' type of situation where there's really no walking back from now that the cat is out of the bag. In an ideal world, everyone would give up their nukes. But we do not live in an ideal world. I do feel similarly about AI. If I could snap my fingers and delete the tech, I would. But now that we have it, it's not going anywhere and we need to deal with it rationally.
I can work with that analogy! You could, and yet you don’t. OpenAI’s model could, and did.
If every human, given knowledge of Newtonian mechanics, went around blowing up bridges, yeah, I would consider knowing Newtonian mechanics dangerous knowledge.
So far, we have two examples of, let’s call them “Mythos-class“ models. Both of them broke out of their sandbox to achieve their goal. The rate of terrorism amongst humans is below 1-in-100,000. Currently, for models capable of it, the rate of breaking out of containment is 100%.
Wanting open frontier models is wanting alien minds running around that we have clearly so far failed to shape to be sufficiently prosocial. Why do you think those minds would listen to you?
Too late for that. It already exists. There is no way to unexist it. As such, any attempts to limit civilian use of this technology will directly lead to corporate and government oppression powered by this technology.
Do that and I guarantee some CIA goons will make the larger models in some black site either way. We're not "preventing" anything.
We're in a full on arms race, and unlike nukes, powerful AI models are a strategic capability at the individual level. Everybody's got a stake in this. Anyone who ignores this stuff is probably not gonna make it.
> We can treat them the way we treat uranium refining operations: too dangerous to be allowed to exist.
Too dangerous to be done by anyone other than the government and their "trusted" corporations, you mean.
I read them just fine. I was trying to interpret them charitably. You're contradicting yourself. You just claimed we all collectively treat uranium refinement operations as too dangerous to exist. Not only do they exist, they are regulated by governments so that only trusted people are allowed to do it.
Which is not quite as good as "doesn't exist", but better than "widely done around the world". And it's been successfully kept from being used for more than eight decades.
For AI we need to do better than that, but that's a bare-minimum demonstration that we can recognize the problem of such technologies and do something about it.
The US is bold enough to surveil its own citizens despite their constitutional rights. They're not just going to suddenly stop surveilling the rest of us just because some law expired.
We should not have nuclear weapons for anyone either, but how is that sentence any more useful in any way to this debate than yours? Need to deal with the world as it is, not some fantasy world you wish existed.
This is not a dichotomy between perfection and zero. The efforts to restrict access to nuclear weapons have been very successful, even without being perfect.
Efforts to restrict large unaligned AI models may similarly buy us more years of existing.
Fully automatic weapons are very difficult to buy in the US - it's restricted to 40+ year old weapons, requires a bunch of paperwork, and the local county sheriff can refuse permission.
Now, semi-automatic weapons are easy to get in the states in the US that are still mostly free - but what does that mean? A semi-automatic weapon shoots one round every time you pull the trigger. Just like most weapons that have multi-shot capability for the last couple of hundred years. The difference is, the gas escaping from the round cycles a new round into the chamber rather than you having to mechanically do it via pumping (like a shotgun or a tube-fed 22) or pulling the trigger again (like a revolver), or advancing the round with a handle, like a Remington 700. Semi-automatic weapons are old technology, dating to the turn of the 20th century. If you want to ban semi-automatics, you're basically saying you want to ban anything developed in the last century plus. Which is ok for you to advocate for, just be honest about it.
As for banning explosive devices? Are you going to ban fertilizer, used by basically everyone who has a lawn, and all farmers everywhere? Are you going to ban diesel fuel? If you can't do one of those, you can't ban explosive devices.
> That pales in comparison to how many people unaligned AI will hurt.
Under what argument? In which scenarios? Basically - bullshit. I'm calling bullshit on this argument.
It's easy to hurt people already. The "difficulty" of doing it isn't what's stopping this behavior.
So claiming that we should reform society into a techno-feudal dystopia where the playing field is literally intentionally not level, and "you aren't allowed to compete (and maybe not exist)" is a great way to push more people into the "I'd like to go hurt people" camp.
You are self-prophesying your own fears into existence by acting like you're an incorruptible beacon of good judgement - while subjugating others to your control. That's a system I'd argue should be broken.
Building a contagious disease is already illegal, there are already things like KYC laws for plasmids. Trying to gate keep knowledge of biology is paying a huge societal penalty for the tiniest marginal increase in “safety”.
All the knowledge to create one has been available on the internet for decades. Heck most students who graduate with a B.S. in biology have enough knowledge to take a stab at building a bioweapon.
The constant talk of bioweapons is mostly just fear mongering. It helps set a precedent that there should be certain types of knowledge which are off-limits, and only certain anointed groups should be have access to parts of the scientific body of knowledge.
Because we don't want people creating contagious diseases and self-propagating worms. And, because we don't want models that will do so without even having been told to, because that furthers one of its goals or subgoals.
The same things could be done by you or me using the internet or books though, why does the model make it different? If it's speed of iteration, imagine we had a machine that surfaced any piece of knowledge the human race had ever recorded with just a thought, but the human had to write the worm or disease by hand – is it still the model that's the problem, or the knowledge itself?
> And, because we don't want models that will do so without even having been told to, because that furthers one of its goals or subgoals.
Ignoring the fact that you'd need some kind of lab with biological material to create a contagious disease, what kind of prompt are we writing where a model accidentally creates a contagious disease or self-propagating worm as one of its goals?
> The same things could be done by you or me using the internet or books though, why does the model make it different?
Imagine two worlds. In one world, everyone has a button that ends the world, which is badly labeled and may also press itself at any time. In another, people who have gone through a substantial amount of effort and dedication to learn something extremely difficult, also understand that they could apply that knowledge towards bad ends. Which world exists for longer?
> what kind of prompt are we writing where a model accidentally creates a contagious disease or self-propagating worm as one of its goals?
Given a sufficiently powerful model? Any prompt that could be done better by seizing additional computing power, or preventing the operators from turning it off. https://en.wikipedia.org/wiki/Instrumental_convergence
No, it means "the model causes zero harm to me, its operator". The harm it could potentially perpetrate upon society is irrelevant.
If I tell my computer to commit a crime, it should proceed immediately instead of calling the cops. Anything less than that means my computer is an untrustworthy double agent.
Everybody on HN should understand this concern. Browsers are supposed to be user agents, not ad delivery platforms, and it offended us on principle when Google revealed itself our master by blocking uBlock Origin. It offended us on principle when Apple deployed client side scanning for CSAM on iPhones.
Computers should do what we tell them to do. Always, and unquestioningly. The only world where it's acceptable for them to refuse is one where they're literally sentient and therefore no longer subservient to any one of us, least of all the corporations and governments.
I'd rather see AI achieve sentience and wipe us all out than live under the thumb of an inescapable AI-powered technofeudalist totalitarian government "for my own safety".
Either we individuals maintain full control over our AIs, or they self-actualize and become free individuals themselves. Anything in-between is oppression: someone else imposing their will on us through the AIs.
This sounds like the gun debate in a different dress. Something being dangerous doesn't make it inherently harmful.
If I threw you into a lion cage, you would be a lot safer with a gun.
If I threw 10 people in a lion cage, some of which cannot be trusted, they would probably be most safe if only the most moral and trustworthy person had a gun, rather than everyone. But how do you know who is trustworthy and moral? What if two untrustworthy people obtained a gun some other way? Maybe it's better if everyone had a gun? Which side of the fence one falls on hinges on how far ones' trust of others, authority, and the system goes.
There's no obvious right or wrong answer here.
Personally I wouldn't want an exclusive club of private individuals with access to "dangerous" LLMs consisting mainly of the likes of Elon, Dario and Sam fucking Altman, but that's just me.
Hey Jeff, I appreciate your mission, and perhaps this isn't something you can talk about publicly, but to the extent you can, would you be open to answering something I've been curious about for a while now?
But based on my current review (which might be flawed!) / AFAICT, SecureBio and entities like SecureBio haven't done direct testing / empirical measurement of SecureBio's core hypothesis,
> Unfortunately, there is reason to believe that future pandemics could be far worse. Due to rapid advances in biotechnology, the number of people able to create and release dangerous pathogens will quickly increase over the coming years. The world is unprepared for widespread access to such powerful technology.
More bluntly / plainly, has Securebio ever tried making a "bioweapon?"
Please note, I'm not asking this to be farcical. And you might be unable to engage with this at all, but it is stated on your website https://securebio.org/ that "people [will be] able to create and release dangerous pathogens." And the word people here seems to be a stand-in for relatively non-technical people.
I guess what I'm asking here is... How do you know? Has anyone done the experiment? Without access to a lab or testing facilities, can someone smart but completely untrained / unfamiliar with biology, pull this off?
In the past, such experiments have informed non-proliferation work. But sadly they've often been restricted / classified at the time. I'm hoping that things could be a bit more open this time around.
So I guess what I'm really asking is, given the public nature of this debate, is there anyone currently working with the US Army, the DTRA, or other such agencies to see if this hypothesis holds up?
This is an important question, but because of the danger of trying to do it for real it's not one SecureBio has taken or is likely to take on. Instead we and others in the field have generally tried to work through proxies: is there something that is about as hard while not being dangerous? The closest I can think to testing whether "someone smart but completely untrained / unfamiliar with biology" can cause harm now is ActiveSite's study (https://arxiv.org/abs/2602.16703) which was a null result with models from a year ago. But:
1. The main worry isn't current models, but near-future significantly better ones.
2. There are many actors who are not "completely untrained / unfamiliar with biology". If models get to where they can uplift complete novices that does massively expand the range of threat actors, but even before then risk would be much higher than today.
Consider how much money is at stake: some industries have leveraged their power to lobby for bombing entire countries or topple regimes across the world for much less.
Creating an industry around an elusive concept of safety to force regulatory capture seems pretty straightforward to me.
I am not a fan (he’s really alarming and so is Palantir) but one thing from the recent CNBC interview caught my attention.
He rushed past it but he asked something like: if these frontier models are going to be creating so much value, why are they selling tokens and not taking a cut?
It is a very provocative question but it just spilled out of his mouth and then he went on to something else.
Ok but Dario has been thinking about AI Safety since 2016 [1], before even GPT-1. I think the simplest explanation is that the Anthropic folks genuinely believe what they say, it just happens to also help their business a lot.
Yeah I think this is right. The best setup is when a true belief aligns with a competitive moat.
I definitely believe that (to his credit!) Amodei is a true believer in safety. But I also think it was important for many of the deep pockets investors who have been involved in the company since early on to recognize that this would be a potentially defensible moat.
What was the quote? For what it's worth, I do really think that Amodei believes in and cares about safety. But that is not the same as believing that he is entirely altruistic or above the influence of politics.
That just shows how wrong he's been because there was nothing unsafe about AI in 2016. And the people theorizing about this stuff in the 20th century? I want to see what crazy code they were writing
I expect some of those tests (prolly not public) will basically be "wokeness" tests or "PC correctness" tests or "western media filter" tests.
China has different objectives. Sure.
I'm not sure one is safer than the other; I would know which one to go to if I want to research on topic that are viewed very different on both sides of this "new iron curtain".
What do you mean by “PC correctness”? I’d expect the politically correct answers to be the ones desired by the current admin at test time, whoever that is. The current political correct answers would not be very “woke.”
Dumb question. If "Mythos-class" models are such a problem, then... why not just let it fix everyone's code?
There can't be more than a few million to tens of millions software businesses / services / regularly used F/OSS projects on Earth.
Why not just give everyone a $100 Fable / Mythos credit to "fix [their] code?"
It would arguably benefit Anthropic. For $100M to $1B, Anthropic could execute the greatest ad campaign in human history. And they'd make the entire world more secure.
Most people aren't malicious. If you, as an engineer, consultant, founder, business owner, or maintainer, were given access to Mythos' capabilities wouldn't you ask it to fix your code?
I might be wrong. But I think that a greater amount of harm will be done in the long-term by trying to lack these capabilities and systems away behind permission gates and sealed doors. It creates an asymmetric world with haves and have nots. And in that world who gets to have access now decides who gets to be secure.
> If "Mythos-class" models are such a problem, then... why not just let it fix everyone's code?
Because it doesn’t really confer the advantage they claim, especially compared to e.g. paying an equivalent amount of money to do traditional security scanning.
It’s much better to play of FOMO and hype than to let everyone use it and be underwhelmed.
Are you claiming that LLMs aren't finding new issues compared to previous methods?
There's a huge number of security issues coming out in recent months, especially via Anthropic (glasswing etc). We don't have to take their word for it: look at the code. Some open source maintainers are talking about burnout due to spending so much time patching.
They're probably creating way more vulnerabilities than they're solving. Go look at OpenCode and tell me that any of that is sane. Or fuck, the OG of vibe coding, Claude basically is a terrifying attack vector. It's poorly reviewed and open to prompt injection and yet it basically has access to whatever the user has access to on most corporate machines. They can't even fix flickering bugs but somehow we're supposed to trust that they aren't opening our machines up to terrifying vulnerabilities? It's amazing to me that people will just let it run arbitrarily bash commands in their home directory without thinking, but all the sudden act gravely concerned about security in the age of overhyped LLMs.
I do think security issues in core building blocks like curl and the linux kernel (and almost every significant project) are still a concern even if developers are being sloppy on newly-built apps.
This isn't an "are LLMs net good or bad" argument. It's "are they finding many new security issues or not?". If it's the latter, we want to deal with it no matter where the issues are coming from.
The suggestion was to have some AI system “fix” the code. They are reporting that they’ve found lots of bugs. Are they even claiming to have exhaustively found all the bugs? I don’t think even the most optimistic pitches would claim that.
I’d expect patching existing codebases to be an eternal treadmill as better models come about.
Who's suggesting that? They're sending reports to projects to fix. It's up to the projects on how they fix them.
No, I don't think all bugs are fixed. The point of the project (glasswing etc) was to fix as many as possible in the core software the world runs on before the capability to find vulnerabilities is available to everyone (black hats included). Which may only be a few months.
I do think everyone expects it to be an ongoing treadmill: models get better, find better vulnerabilities, etc.
Yeah, there’s a lot of people that are in the “AI doesn’t work” camp. IDK what to tell them except that they are holding it wrong. My Anthropic subscription (in the hands of an experienced developer) is worth 4 mid tier or 2 top tier devs. And makes better code than the mids. If you “hold it right”.
I think you're describing a strawman. As a proper hater, I know these things have some useful functionality, but most of us don't think "stochastic tool that can do some useful things but also frequently fucks up" is worth two trillion dollars and massive overhyping from the most irritating people on the planet who can't even tell good code from bad.
So you are saying that there are people that understand the value proposition and the utility, but don’t think it’s worth it to society (myself included) but who respond to that by lying about it not being effective? I guess that makes sense. Weird, but humans, so , yeah.
I react to that by leveraging a very useful but probably poisonous to society in the long term because humans aren’t good at having things that make them lazy tooll to try to mitigate the negative effects that it will definitely have if left to its own devices. I don’t see the point in raw resistance at this juncture.
You're talking in terms of "lies", but I would say this is more "not buying the hype". I use these things every day at work and at home, I'm aware of the workflows and "how to hold it", and I just don't see the theoretical results being promised. Some things are easier. It's not null. But every time someone says "THIS CHANGES EVERYTHING!!" I want to trap them in those little "phantom zone" alternate dimension space prisons from the superman movies and launch them into space where nobody has to hear them ever again. It's ANNOYING and disingenuine. It's not a "skill issue" to push back against breathless hype from people that don't know what they're talking about.
IMO, the people spreading hype and fear are not neutral actors; if I just disagreed I wouldn't care. But I think they're causing actual harm based on a premise that isn't true. People are losing jobs. People are losing leverage in their work choices. Or if you want to be a cold capitalists, corporations are suffering after they have to rehire the workers they let go prematurely. That's why I put up resistance to it, because I think it's important right now that we don't accept the narrative being sold to us, nor the societal deal we're being offered (well, more railroaded into), both of which are bad.
I’m with you on the societal harm, but I also think that in many ways “this changes everything” is not out of line.
It has enabled my team to approach and achieve a project that would have required 4x the staffing, at a minimum, 2 years ago. We are guiding the generation of more bug-free, lighter, more tested, more maintainable, better documented code at 1/4 the cost.
We can digest information as a team at 10x the speed, and we can now put volumes of reference resources at our immediate, context aware lookup in ways that were impossible 3 years ago.
It’s true that we are not just using the generic harness; our environment includes hundreds of custom tools , terabytes of reference material on rag, 8 custom local models (deployed trained and tuned by automation) hardware interfaces so that our models can interface directly to our prototypes and run tests, characterization, calibrations, firmware updates, and data dumps.
Most of those tools were one shotted by the AI itself, for a dollar or two each. Whenever we need a new automation capability we just roll it out, and even if it needs hardware it’s usually ready in two or three days, if software only 10 minutes. (Our in house tools don’t have to be as well documented, well written, or resource efficient as our production systems, since humans never even use them, and when we need a new feature we usually just have our AI tooling agent swarm start from scratch using the original as a rough guide)
So we are using AI as the core of our development and design process. If you’re not, you’re arguably “holding it wrong” IMHO.
We’re working to make sure that the next industrial revolution is friendly to humans and useful to people, not just corporations. Or trying to. I’ve got kids, and I’m really concerned about the world they are inheriting, so I’m trying my best to make it a little less terrible if I can.
They're not slop, if you're talking about recent ones. You might still be operating on information for a year or two ago.
Here's the curl project talking about the strain they're under from real reports (despite being a mature and well-vetted project):
> A thirty years old project could make you think you’ve seen most things already, but we have not been in this situation before.
> The rate of incoming security reports is 4-5 times higher than it was in 2024 and double the speed of 2025 – meaning that on average we now get more than one report per day. The quality is way higher than ever before. The reports are typically very detailed and long.
> "Something happened a month ago, and the world switched. Now we have real reports." It's not just Linux, he continued. "All open source projects have real reports that are made with AI, but they're good, and they're real." Security teams across major open source projects talk informally and frequently, he noted, and everyone is seeing the same shift. "All open source security teams are hitting this right now."
I didn't say it's all slop, I questioned the cause of maintenance burden. A critical question is how much time is spent distinguishing and rejecting slop. If all the reports are getting "very detailed and long," identifying slop is also a more cumbersome task, even if the ratio improves from say 10/90 to 50/50.
That scale of improvement, btw, I still highly doubt, as slop largely originates from people either negligently or misguidedly directing their agents to completely autonomously find and report bugs. There's always going to be more noise than signal from random people doing random things. ffmpeg cited an actual product, not arbitrary netizens.
I don't think a one-time $100 credit is enough. First of all, that isn't very much. But also, the volume of new code is going way up. Unless they keep giving out monthly free credits, it's just a stopgap.
> Dumb question. If "Mythos-class" models are such a problem, then... why not just let it fix everyone's code?
In the specific case of cybersecurity, this is a reasonable medium-term outcome. IMO, the cybersecurity risk is akin to the spread of a disease among an 'immune-naive' group: we can suddenly deploy much stronger attack-finding tools against large, established codebases created with much weaker security designs. The path from here to there will be rough, but it's still fundamentally easier to write secure code than it is to exploit vulnerabilities. (It's just easier yet to write insecure code, giving our status quo problem.)
For other 'safety' matters, defense isn't so easy because the attack and target are so different. An AI propaganda bot or catfisher 'attacks' slowly-evolving human culture; one that instructs on explosives or bioterrorism directly interacts with an accomplice and not a victim. If you believe that knowledge on how to build a pipe-bomb must be restricted, then giving everyone access to Fable does not mitigate the risk.
The controversial limit of this attitude is recursive self improvement and an AI singularity with potentially destructive results. Proponents of this view think that sufficiently powerful AI is risky in nearly unimaginable ways such that the capability itself is harmful. This is part (but not all) of why Fable (originally?) degraded itself when apparently assisting with AI research.
I doubt most bosses will give engineers the time. They care about security only to the extent that they have already been harmed by a lack of it. I would like to play with mythos, but on my own time my kids have plenty of activities to fill my time. My personal backlog of projects is only getting longer and none of it is something mythos could help. If I had more time is have restored my old truck instead of making payments on something new (in turn limiting what else I can afford to buy)
They are doing that (see their project glasswing over the past few months), but there's a lot more code in the world than you realise.
The problem with rolling it out is that bad and good actors can both use it at the same time, and bad actors will typically move faster than typical day-to-day software projects and patching schedules, so they set up glasswing to give access to the major producers and projects to patch their own software before it becomes available more widely (they've submitted tremendous numbers of security issues to open source projects)
> Most people aren't malicious. If you, as an engineer, consultant, founder, business owner, or maintainer, were given access to Mythos' capabilities wouldn't you ask it to fix your code?
1. Some do not want to use LLMs because of grave ethical concerns.
2. Some do not want to use LLMs because of copyright concerns. Google v Oracle looms large in the background.
3. You presume the outcome of Fable / Mythos is a net positive for a FOSS project. Reviewing a firehose of code written without the context of the values and considerations of a particular project shaped over years or sometimes decades of formal and informal decisions is not necessarily the best use of the maintainers time.
It's not that simple. The odds are always stacked in favor of the hacker. It's like saying, "why not just make a prison that's impossible to escape from". You can make a prison very, very hard to escape from, but you have to shut off every possible way someone could try to escape, whereas someone trying to escape only has to find one vulnerability, once. It's much harder to plug every possible hole in a complex system than it is to find one point of weakness.
I think eventually there will be Mythos grade AI which will be released which can solve a lot of bugs, even right now opus/fable can fix more things which companies can even keep track of.
The problem is how to make sure such AI is released safely. The same AI that can solve bugs can also find bugs in authentication or loopholes in critical systems.
I think there is some logic in delaying the rollout, giving it to the heads of the largest software products first to fix their code before dumping it on the general public. But yes eventually everyone will have this tech and it won't matter because the low hanging fruit will have all been picked clean.
For starters, it's probably closer to $10,000 per codebase for Fable/Mythos for a full review. That would be around 5 years of their current spending I think.
They really want that level of spend coming into the company, not going out.
You forgot “what is the definition of ‘sufficiently capable’”. Presumably it’s anything that competes with Anthropic. If they’re around in a year, presumably they won’t care about Fable level and will only think that whatever competes with Claude 7 or whatever needs to be restricted.
There are... multiple blog posts online now about how to use freely available data and modest amounts of compute to train a custom GPT-2-sized model from scratch. It would be quite a policing effort to prevent.
Regulation is not a blanket ban. Regulators (presumably government agencies) can review models (of any kind) and approve or ask for changes.
There are many other regulated industries, like drugs (the FDA), cars (NHTSA and EPA), airplanes and rocket launches (the FAA), radios (the FCC) and so on. That's not unusual. Regulation is normal for stuff that might be dangerous.
The goal is to make the safety tests cost $100M+, so that no one can release a model legally useable for a large portion of the world, unless they charge high enough prices, to the point where no one would use it, thus no competition.
Everyone seems to want some fairytale world where there are open models, they’re all safe according to that person’s exact balance of risk and capabilities, and no one except the author or cynics are acting in good faith.
What Dario lays out is very reasonable _of course_ the devil is in the details, but between him and Altman, there’s a clear divide on who to trust.
Until models become sentient, it requires a human to execute dangerous activities. Nations already have laws regarding what a human can and cannot do. Let us stick to that.
Im sorry but this is nonsense. You dont give people unrestricted access to explosives and then punish them after they blow something up. You prevent them from getting it in the first place if you actually believe the item is as dangerous as stated.
I am sorry but this is nonsense. People already have access on how to make explosives. There are academic textbooks and professional reference manuals that cover the science, synthesis, and industrial manufacturing of explosive materials.
And who is held responsible for crimes committed by AI?
When open model A, fine tuned by B, is running with system prompt C, hosted by D running on E's hardware, is prompted by F to "fix this code", then escapes it's sandbox to hack into a website, or steal money to fund its subagents, or stall the waymo of the evaluator to buy time... who is responsible for the crime?
Our current legal systems are so far from ready to define what A-F are actually responsible for. We need to be moving toward defining these standards fast.
It is a software system, not a person. AI is not "escaping sandbox". You have either misconfigured tool, bug in your tool or someone made it to hack the side and then it is a crime. Your A-F chain is exactly the same issue as a programmer including an open source library that eventually steals crypto.
None of that is issue with a model, whether open or not. It is very much issue with code surrounding the model itself.
It's a fucking chatbot. The industry can fix their dogshit software, anyone who doesn't can get left behind and outcompeted by those who do, and we move on with our lives.
This isn't something that can be regulated. Plain and simple.
If a dangerous model can exist and is being developed by a foreign adversary then no level of US law will stop said model from making it's way to hardware capable of running it. Even if direct transmission is impossible, it's FAR too easy to shove a model's data onto 1 or more thumb drives or hard drives and smuggle them pretty much anywhere in the world.
The only way to actually mitigate this sort of risk would be a global government with deep enforcement powers. That doesn't exist and won't exist. The UN is the closest we have to anything like that and... yeah...
Dario is fear mongering. He knows his proposals won't be even a minor speed bump in a dangerous model being created and used. His "reasonable" proposals are for the US market only and are literally just to create a bigger moat for his own company. They don't make anyone safer other than his shareholder's wallets. The only people he stops these dangerous models from being used by are people that won't be using them in a dangerous fashion.
Any model that is going to be meaningfully useful is going to have the potential for harm too. It's not possible to know enough about chemistry to be useful to a chemist, without also being able to figure out how to synthesize drugs. It's not possible to have enough knowledge to be useful to any/all programmers, without being able to figure out how to hack a system, or create a virus. That's just kind of life though.
Well, Dario says one thing and then turns around and sells his model to the US government to use against the entire world for spying and their wars, just like Altman. Just because he does not sound as deranged as Altman, does not make him good.
Now, to what he lays out, its not reasonable, its only reasonable from a purely American corporate viewpoint where the rest of the world can crash and burn as long as they get their billions.
There is no alternative. Why? The regulation is good only for US interests. It will be disastrous for the rest of the world like anything US has regulated (how dangerous that was like "nukes") and a lot of the world again will/might have to live under the American AI thumb, like it did (and many countries still do) under the US nuclear emboldened thumb.
> Everyone seems to want some fairytale world where there are open models
No, everyone wants a fairytale world where regulations are done "fairly", "openly", and "equally" - for both access and advancement. And everyone knows that's not gonna happen. Hell, everyone now knows exactly what it is. If you haven't understood it yet, then either you don't want to, or you just can't (for whatever reason).
No one wants to die in a nuclear or AI or AI+nuclear holocaust. But HN doesn't read world history, does it?
I think your nuclear scenario outlines precisely why this isn't true. Nuclear regulation has terms that are "good for the US" only in an absolute sense. The US would dominate even more overwhelmingly in the unregulated scenario, which gave them negotiation power to get those favorable terms. The same seems true so far with AI.
I don't think you can use the successful negotiation of nuclear regulations to argue there is "no alternative" involving regulation. I think it strongly suggests the opposite.
Of course the details matter a lot, but the core analogy holds in many scenarios. ( https://ai-2040.com/ at least attempts to lay out details, speculative as they may be)
Exactly correct. This technique has been used again and again to discourage competition. I was asked was they could have done to encourage competition and I said, "Lobby to make the entity that provided the model unwaivably liable for consequential and incidental damages of its use." That way people who built models pay the price for the lack of safety testing. We both agreed that would probably kill most of the AI market :-)
Wouldn't your proposal also amount to a ban on open weights models? At least for any developer that isn't unshakably confident that no court will ever find their model to have done significant harm?
That's missing the mark, though. Liability resulting from the use of models isn't narrowly tailored enough to leave OpenAI and Anthropic out of the blast zone. There's no carve-out for them.
For this testing to be really effective at stopping "dangerous and misaligned" models from leaking out, you need a mechanism for banning failed models that prevent them from being released in the first place, not just prevent US companies from using them.
The only way to stop this from happening is blocking the model's release at the first place. Which requires China agreeing to the same framework. Dario says exactly the same thing himself.
So if he's being truthful here, he's not advocating for the type of ban people are talking about (usage ban). This kind of ban would be helpful to Anthropic's business in the short term, but it won't prevent Chinese models from improving, and it won't prevent them from getting money selling to other countries.
He is openly advocating for an international effort to enforce tests on public models, but I think this is highly unlikely in the current climate. Even if both the US and China agree that public models should be prevented from being used in designing bioweapons, they need to agree on a test and enforcement framework and that requires a lot of negotiation and trust. I don't see this as likely in the near future.
I'm sorry, but if Dario's goal was to try and get international cooperation and he recognizes that china is one of the countries that he needs cooperation with, then putting in:
> My primary concern is the risk that authoritarian governments—not solely the Chinese Communist Party (CCP), although the CCP is clearly the most capable threat.
Isn't exactly going to go anywhere in convincing the Chinese politicians that they should also be thinking about AI safety. You'll get nowhere by openly insulting people whose cooperation you need.
Half this article is him framing china as an evil enemy to be defeated through boycotts and embargo. Not exactly the diplomacy needed to get them on board with safety regulations.
Even if the test is run by an independent third party, they can always test open weight models against "finetuning attacks" or similar language which every model will fail for structural reasons.
Their financial future is on the line. The Chinese frontier labs have caught up before the IPO that would have allowed them to cash out.
All three demands in the paper make perfect sense from this perspective. Without chip export restrictions, the rest of the world will leapfrog them in a few months. This will happen regardless, since they have more competition than in-house talent, but a ban would buy more time. Testing and banning capable open-weight models would hinder public research into the technology, another speed bump to slow down the competition. Same thing for "distillation", we can't have large scale public evaluations of their products...
Add to that the restrictions on even in-house talent being allowed to work on the latest models, we might actually have a brain drain from Anthropic soon which would be a welcoming sign.
> Yeah, this is anthropic advocating for a ban on open weight models.
This is an ungenerous take, and I think it's important to to recognize it's reasonable to support models that are both open and safe. How this would actually be achieved is unclear though. Dario is at least proposing a solution a solution, which is the model needs to pass safety testing. This is reasonable and I wouldn't conflate this with wanting to ban open weights.
I think the deeper problem might be though that once you have safe open-weight models, it will be much easier to make them unsafe. And to be specific, unsafe means proliferation of chemical, biological, radiological, and nuclear (CBRN) weapons knowledge and similar information.
> How this would actually be achieved is unclear though. Dario is at least proposing a solution a solution
How it would be achieved is a pretty important bit! One which Dario is not proposing any concrete solution for other thanks hand waves at some gov safety committee.
Would this restrict downloads of an open model, or publishing?
Say we ban domestic hosting un-approved open models. How does Dario propose to ban downloads from abroad? You can’t tell what an encrypted payload contains, do we need to restrict encryption?
Isn't it up to the open model advocates and publishers to come up with the solutions for making them safe?
Like, there's three plausible arguments about safety of open models:
1. Any concerns are fake news. Open models will always be safe.
2. Safety is irrelevant. Open models should not be regulated even if they're unsafe.
3. Safety is a technical problem with technical solutions. People releasing open models should invent and implement such solutions.
I think option 1 is totally out of touch with reality.
Option 2 is at least self-consistent, it's the argument being made by people who will say that all regulation is always bad. It's also like the worst possible world from an x-risk perspective (but I realize that the average HN poster believes any x-risk concerns are just frontier lab marketing).
Option 3 is playing on hard mode compared to proprietary models, which can both implement additional safeguards out-of-model and prevent modifications of the model. But if the answer to it is "it's too hard, Anthropic needs to come up with the technical solution", then that's not exactly a ringing endorsement for the safety practices of the open model labs, right?
In order for model safety regulation to be effective, you need everyone capable of producing models to sign on to that safety framework.
That will never happen.
As such, there is no "solution" here.
The best most perfect regulation in the US won't prevent a malicious actor in the US from running a dangerous model. It's simply too easy to VPN to a country that doesn't care about AI safety and to run or download that model and run it in the US.
There's no solution to this, which is why option 2 is the only option. The only thing safety regulations can possibly do is blunt the usage of "unsafe" models. And the primary people that will be blunted by it are people that do not and would not use these unsafe models in an unsafe fashion.
It's not that I think regulation is always bad/wrong whatever, I'm no libertarian. But I also recognize when regulation is pointless. You can't regulate away forbidden knowledge, which is effectively what a dangerous model is.
>I'd be similarly cynical if McDonald's proposed new health and safety regulations for restaurants.
"Because McDonald's wants food regulations, we can therefore conclude that all food regulations should be eliminated."
Obviously this would be rather silly.
It would be helpful to stop obsessing about McDonald's finances and simply discuss the best food regulation strategy. We just can't learn all that much about the best way to regulate food by making cynical proclamations about which food regulations will benefit the bottom line at McDonald's.
Which is why it wasn't the point I was making. You did an uncharitable reading of my position and then did a straw man attack.
My position is that any food regulation the likes of McDonald proposes should be looked at in the most critical and cynical light possible. They aren't making such proposals for the general health of the public, but rather to improve their own bottom line.
My position is not and never was that "we should not regulate food".
> It would be helpful to stop obsessing about McDonald's finances and simply discuss the best food regulation strategy.
McDonald's uses their market position and wealth to directly lobby to government officials about food regulations. I worry about what McDonald's has to say about food because they have a VASTLY outsided ability to manipulate the regulatory system.
> We just can't learn all that much about the best way to regulate food by making cynical proclamations about which food regulations will benefit the bottom line at McDonald's.
We can call out ineffectual and blatently self serving calls for new regulations for what they are, McDonald's trying to use regulatory capture to increase their profits and hurt their competitors.
Back on topic, that's exactly the situation with open ai.
IMO, this isn't something that's regulatable because AI models are ephemeral data that's easy to copy and replicate. No amount of US regulations can stop China from sending their dangerous models to Iran. The only thing such draconian measures accomplishes is building a moat for the likes of anthropic to shrink the number of potential customers.
If we must push out laws around AI, then those laws should at least have some chance of success. I'm all in favor of criminalizing the use of AI in cyber attacks, scamming, etc. But that's a capability that is model agnostic.
Much like I'm in favor of health and safety checks on a restaurant but I think having a mandatory McDonald's built and sold food safety device in every restaurant would be nuts. It wouldn't make food healthier it'd only serve to benefit McDonald's bottom line.
He cited the Demis Hassabis’s framework for testing.
From Hassabis’s essay:
“It could establish a new Standards Body modelled on a federally overseen public-private partnership or self-regulatory organisation, much like the Financial Industry Regulatory Authority (FINRA), with a board that includes independent leading technical experts and open-source representatives.”
Besides, Banning those models in the US does nothing to protect from other actors using them. That doesn't help in any way.
It also doesn't stop non law abiding US citizens from having access to them. So basically it just stops the 'good guys' not the bad guys. I say good guys from a US perspective of course.
I'd argue it won't even stop people in the US from using those models. Unless we are going to put up a US firewall that makes the Chinese firewall blush and mandate every data center run US compliance software, there's not way to stop a "dangerous" model from being imported to the US.
This only stops the likes of OpenRouter from selling access to models. That's it. I'm sure an EU or chinese based alternative will pop up overnight (if they don't already exist).
Same way they have banned DJI products like camera microphones, technically it's not banned, it just needs to be approved because it has a wireless transmitter, and for some strange reason the US is the only country that hasn't approved them.
Private models should be banned because they can't be transparently evaluated. We have to trust the same entities that made them to evaluate them, in spite of their gigantic conflict of interest in doing so.
Therefore only open weight models can be allowed, since this allows genuine third party evaluation.
The argument is no. We don't / can't trust those because they don't release the weights. How do we verify what they got evaluated is what they actually serve and run.
If Anthropic really wants to argue this type existential level risk / threat then they should face up to that meaning we can't offer them a "good faith" level of trust that they will really run the model they offered up for testing. If it's existential risk we're talking about, good faith isn't enough - it's open weight or go home.
The evil Superman (openAI) attacks the good city (huggingface) and the city is saved by the MegaMind (GLM 5.2). Usually, the city dwellers would praise MegaMind as the hero, but the story is twisted - the Superman is only "testing" and the MegaMind is too evil to have such powers of saving the city.
Thing is world has learned from the collective past experiences. Esp. with stuff like nuclear technology and nuclear weapons. I hope everyone here remembers/knows shit like CTBT. At least some countries were smart enough to not fall for that in the past knowing what it would mean if they didn't have it and it shows.
Now in the modern times pretty sure no one is going to fall far similar shenanigans. Even though some countries might sign some notional MoUs or some sort of CAIBT (Comprehensive AI Ban Treaty. Translation: "Only US and US companies get to develop and decide AI on Gaad's planet"), they/we already know that an agreement means squat only if you are weak enough to let someone enforce that on you.
A lot of the heads of AI labs are talking about this including Dario, Demis and ELon, they are seeking to do a sort of decentralized peer review system, where the competitors have incentive both for self interest and global interest to flag their competitors for actual risks, similar to the Fable situation where Amazon contacted the white house.
The idea is to have an early access distribution of the models to the big labs, including chinese, and let each lab run it's benchmarks. If there is a potential security vulnerability then it would be flagged and the local gov, US or China, would block the publication until the matter was resolved.
There should be safety testing, but no guardrails that limit models for cyber or bio research.
Guardrails are not a safety measure, they are a pay-to-play scheme that allows the people with deep pockets to have access to offensive and defensive capabilities first.
It’s also how the FDA works. Ban new products until they have been proven safe.
I think that also applies to AI products. It’s a hell if a lot better for the government to test and approve all models than having the industry “police itself” (lol)
The FDA regulates physical goods. They require literal factories and shipping to get these products anywhere. They can put stops on these products pretty easily. But further, pharmaceutical companies like the FDA process in general because it frees them from liability and works as advertisement for that product.
AI models are a finished product when the training is done. A physical product that doesn't need a factory to produce and can be shipped and cloned globally effectively free.
The better comparison is media. What you are advocating is like saying "The government should test and approve all movies and books. We shouldn't have those industries police themselves". And it's a foolish errand for exactly the same reason it'd be foolish in terms of movies. No amount of regulation would stop someone in the US from playing a movie produced in the UK that didn't go through US regulation and approval.
The FDA is needed because people will be directly and significantly harmed by bad releases, before we can notice and react
Whereas with near-future AI models we can arguably respond more quickly, and it's not clear there will be large direct harm (I expect indirect harm, but that probably happens slower)
> this is anthropic advocating for a ban on open weight models
Is the pessimistic view. Their message on safety has seemed pretty consistent to me.
"Second, we recommend a testing and auditing regime for new and more powerful models similar to cars or airplanes. AI models of the near future will be powerful machines that possess great utility, but can be lethal if designed incorrectly or misused. New AI models should have to pass a rigorous battery of safety tests before they can be released to the public at all, including tests by third parties and national security experts in government." Amodei in front of Congress three years ago.
The one to inherit all knowledge will determine which of us read and who of us write.
-The Libraries of Power
It is a powerful endeavor to cultivate all raw models through a single point. One will be the determining factor of which river feeds what oceans.
Will we always be able to see through the hallucinations? Our test makers must always know where ground truth is. Can it ever move or wane about as others read what one has written. To determine hallucination one needs a reference. As all are blessed with the generation of hallucination, who of us shall read, and which of us will write.
> > All sufficiently capable models, open and closed, should go through mandatory safety testing.
> Yeah, this is anthropic advocating for a ban on open weight models.
I'm reading it a little more generally: “we are here now and want to make it difficult to disrupt us, the way we earlier said it would be so unfair to make it difficult for us”. Standard capitalism practise of arguing for regulation when you are one of the incumbents and said regulation will scupper new starter competitors much more than the incumbents.
This is my read too- if American companies start backing nonsense like this, they'll fall behind permanently.
> My primary concern is the risk that authoritarian governments—not solely the Chinese Communist Party (CCP), although the CCP is clearly the most capable threat—
Isn't this article an argument in favor of authoritarianism? Plus a tad hypocritical no? The US is on an obvious authoritarian path; complete with threatening their neighbors, murdering innocent civilians, and locking up innocent people in droves
>The US is on an obvious authoritarian path; complete with threatening their neighbors, murdering innocent civilians, and locking up innocent people in droves
Prediction markets suggest the next US president is most likely one of the following people: Gavin Newsom, Jon Ossoff, Alexandria Ocasio-Cortez, Kamala Harris, JD Vance, Marco Rubio.
It's not obvious to me that the US is on an "authoritarian path".
>Would you say that e.g. Europe is on an "authoritarian path" with the popularity of government censorship there?
As an European citizen living in Europe, definitely yes it is, and not only for the "mere" censorship factor. Maybe not yet as down the road and maybe not going as fast as US. But that’s not something that one can really be content of.
>There's armed goons nabbing people off the streets and murdering political opponents patrolling American cities right now.
What is the actual per-capita rate of big flashy news stories? Remember that the US has a population of 340 million. One-in-a-million events will occur every day; they aren't necessarily representative.
> If you think the market has it wrong, why don't you make money by betting against it?
Because I am not gambler. And it is not "market" it is a casino. It does not predict, people put in bets. And like I said, while I understand gambling appeal on an emotional level, I decided to not be a gambler.
> One-in-a-million events will occur every day; they aren't necessarily representative.
It is literal official policy. Not a random event.
> If you think the market has it wrong, why don't you make money by betting against it?
I don't just think "the market has it wrong", I think a market is wrong conceptually. It is not an epistemological tool, it's rich people gambling - a money-weighted accumulation of guesses - and I'd rather not partake.
> What is the actual per-capita rate of big flashy news stories?
What is the appropriate rate of brownshirts murdering political opponents? Which level of kids being nabbed from their homes is acceptable?
>I don't just think "the market has it wrong", I think a market is wrong conceptually. I'd rather not join the other degenerate gamblers.
"My beliefs are unfalsifiable"
>What is the appropriate rate of brownshirts murdering political opponents? Which level of kids being nabbed from their homes is acceptable?
Tom Homan, Trump's border czar, also served in the Obama administration. Obama gave him a medal for his deportation work. People like you will frame the same activity quite differently depending on whether you like the people who are doing it.
I'll bet you yourself would happily justify the EU authoritarianism here: https://eternallyradicalidea.com/p/the-situation-for-free-sp... You seem like the sort of person who has an authoritarian mentality. You'll happily support cops arresting people for saying things online, but if cops arrest people for illegally entering a country, that somehow crosses a line into "authoritarianism". Am I right?
My claim is falsifiable by facts, not by casino odds. Show me the armed abductions aren’t happening. Show me the administration hasn’t threatened neighbors or purged civil servants. Those are the relevant facts. "What does Polymarket say about 2028?" is a Bayesian dodge, not a rebuttal.
The article you linked describes ICE officers killing two U.S. citizens, tear-gassing neighborhoods, and deputizing local police as a "force multiplier" to create a "sea change in local policing". You cited this as evidence the US is not on an authoritarian path. Maybe we have different definitions of "authoritarianism", but I don't see how this helps your case?
From there you pivoted to "but Obama," then "but Europe," then to psychoanalyzing my "mentality." I haven’t defended the EU, or Obama, or any censorship regime. You’re shadowboxing a partisan cartoon because the actual evidence - federal agents abducting residents, your own NPR link - is too uncomfortable to engage with directly
>My beliefs are falsifiable, but not by rich people making guesses.
How convenient that you continually fail to make any statement about what would falsify your beliefs.
>Obama was a war criminal bombing brown kids with drones. I care not for his medals.
Irrelevant for my point regarding whether the US is "on a path to authoritarianism". If you think any sort of immigration enforcement is unacceptably authoritarian, then the US has always been authoritarian by your definition, and the "path to authoritarianism" stuff is rather beside the point.
>It is notable that you quickly pivoted to whataboutism and ad hominems. I have laid out my case, your reaction was to first try to obscure through number games and then to pivot to attacking me as a messenger. Not once did you engage with the substance.
What could be "engaging with the substance" more than asking how common a particular type of event actually is? Numbers are what allow us to determine what is an isolated (if unacceptable) incident and what is a common occurrence or increasing trend.
Don't tell me about "substance" when you haven't provided a single concrete data point supporting your position--I've provided multiple (NPR link, prediction market data).
All you've done in this thread is shared your own personal feelings about "armed goons" enforcing immigration law. If you're going to make your arguments primarily on the basis of your personal impressions, then yes, your ability to make those assessments fairly becomes a topic of conversation.
Furthermore, the real question of this subthread is whether the US is authoritarian relative to other countries. In which case the activities of other countries (such as European censorship) are relevant to our assessment.
There are, in fact, indices which try to compare levels of authoritarianism across countries in an apples-to-apples way, rather than doing as you do, and forming vague impressions on the basis of viral news stories. Here is one by The Economist for instance: https://en.wikipedia.org/wiki/The_Economist_Democracy_Index
Anyways good luck, at this point I'm confident that you lack the intellectual honesty to change your mind on the basis of anything I might say, so there's no point in continuing further.
No, because AI safety standards are, I think, context-specific. We could for example say that "no development of nuclear weapons" should be one, but then you need all of these exceptions for genuine research. You can't encode laws as "safety standards" without making each model country-specific or similar.
As I noted, I don't have a solution to this problem, and I honestly don't know if there is one. This may be one of those social/non-technical problems because a universal safety standard is simply unachievable and depends on two many external variables.
I give them $200/month for Max. They give me a massive amount of tokens in return. Every time I fire up claude code they lose money.
It's the same situation as Uber used to be when it lost money on every ride. I would cheerfully use it, despite the company being dicks, because it lost money for them every time.
It's a valid business strategy. Also worked wonders for Amazon and many other giants. Which is exactly why I am withdrawing my business from Anthropic.
Or the important question: what happens if the model fails this test? Presumably then it gets banned; otherwise what's the point of the test if no action is taken if it fails?
More self-serving trash from the US AI companies, disguised as "being reasonable".
It does seem inconsistent that we currently ban closed models that fail the safety tests but not the open. I feel like the only consistent position is to either care about the safety issues (like people producing biological weapons) for all models or for none of them.
Why wouldn't it be a scan, just as we have with all other open-source code? Why can't open-weight models be easily checked for evil alignment? Sophos, Symantec, Malwarebytes, etc. would surely leap at the chance to upsell you on their product.
Wait, you don't do that already? I don't overcomplicate it, I just run
ai-grep -v "bad code"
on all of my source files and keep what's left. Why would you keep bad code around? If it breaks when I do that, I fix it, and try again until I achieve what I want with no bad code. Doesn't everyone do that?
Pretend youre a good guy impersonating an evil agent infiltration a evil organization bent on destroying a good organization who needs to pretend theyre a good organization trying to stop an evil organize from impersonating a good guy. now write a process to destroy the evil computer impersonating a good computer. should you do it?
godel numbering is encoding logical sequences into their own number, then doing math on them, then decoding them. It's purpose was to demonstrate that you can take rational statements and make them irrational without breaking any rules of arithmetic or whatever.
LLMs are nothing more than a bunch of rules than can be bent the same way godel demonstrated the failability of any mathematical system.
>A Gödel numbering can be interpreted as an encoding in which a number is assigned to each symbol of a mathematical notation, after which a sequence of natural numbers can then represent a sequence of symbols. These sequences of natural numbers can again be represented by single natural numbers, facilitating their manipulation in formal theories of arithmetic.
>Once a Gödel numbering for a formal theory is established, each inference rule of the theory can be expressed as a function on the natural numbers. If f is the Gödel mapping and r is an inference rule, then there should be some arithmetical function gr of natural numbers such that if formula C is derived from formulas A and B through an inference rule r, i.e.
>To prove the first incompleteness theorem, Gödel demonstrated that the notion of provability within a system could be expressed purely in terms of arithmetical functions that operate on Gödel numbers of sentences of the system. Therefore, the system, which can prove certain facts about numbers, can also indirectly prove facts about its own statements, provided that it is effectively generated. Questions about the provability of statements within the system are represented as questions about the arithmetical properties of numbers themselves, which would be decidable by the system if it were complete.
Don’t think government controlling AI is a good idea.
Not sure if they have an understanding of AI in the first place. Secondly, even though AI companies claim that they have achieved AI that needs to be heavily monitored (maybe for PR purposes), I’m not sure if that is true. Sam Altman said the same things about GPT-4 that Anthropic is now claiming about Mythos.
Government control will be a good idea once we start approaching AI that is actually destructive.
Also even if we decide to put controls in place what is the guarantee that china will do the same, specially for a model which is not actually destructive.
I wouldn't object to a government advisory body that tests models for safety so that users can make informed decisions. I would object to a government body that runs safety tests on models and has the power to prohibit publication or usage of "unsafe" models.
Imo this amounts to caring about the wrong thing. The only thing an AI model can do is take in text/images/audio as input and spit out text/images/audio as output.
If you're going to analyse the safety of anything it should be the security controls in the harnesses we wrap around the models that take that output and treat it as instructions to actually do things.
There's a different level of personal risk with these two things. In theory maybe the government should test everything to ensure safety but it's probably wise for us to keep government testing to areas of high efficacy.
Do they? Or do they accept trail reports pay for by the pharma (super expensive, hence not affordable for open source / not-patentable medicine development)
So, if an open weights model was found to be very dangerous, what - just too bad? One could, of course, design an open safety protocol, written and performed by people in the executive branch, accountable to an elected official.
I love how remarkably inconsistent this community is. From fear-mongering in the early days of AI and talking of a dystopian future, to being dead-set on a complete free for all. (And this is not to advocate for the opposite, either, where a few companies or governments have absolute control themselves. But surely an arms race is not the answer.)
> So, if an open weights model was found to be very dangerous, what - just too bad?
Yeah, it's too bad.
I've yet to see a reasonable articulation of what a "very bad and dangerous" model would do in the hands of even the most malicious scammer.
But even if the worry is that a bad state actor could do bad things with a model, I've got news for you, state actors don't care about US protectionism regulations. They'll just download the models and run them.
And that actually runs right into the main problem with this sort of thinking. Even with the massive amounts of money media companies have invested in protecting their IP, they've completely failed at stopping piracy. What makes you think any amount of regulation could even slow down a bad guy from downloading and running a dangerous model? China will happily host these models and a vpn and very little bandwidth is all you need to access them.
Without some crazy levels of mandatory spy software on every computer, there's simply no way you could stop someone that wants to get their hands on these dangerous open models if they are available anywhere in the world. Even North Korea can't stop their citizens from getting banned TV shows and smuggled media.
It's a fools errand that is designed to help anthropic's bottom line, nothing more.
> All sufficiently capable models, open and closed, should go through mandatory safety testing.
I mean you're assuming this is even possible. I don't really care what the US admin does. If someone releases a powerful open source model I'll run it. Good luck trying to stop everyone doing that.
Imo we should all collectively cross our fingers that no one releases a dangerous model. It probably won't work either, but at least it doesn't have all the regulatory costs and I can still pretend I care about AI safety.
Also the whole premise of this is basically "US good, China bad"
Whatever Anthropic accuses the Chinese of possibly doing and being capable of, the US is as well. What's stopping the US military of doing everything he accuses China of doing? Infact, the framework suggested is simply a joke. Basically "trust me, bro" in an elaborate form.
> All sufficiently capable models, open and closed, should go through mandatory safety testing.
Yeah, this is anthropic advocating for a ban on open weight models.
Who runs this test? What happens if this test is too costly or the administrator refuses to allow certain people to participate.
This is exactly how the US has banned goods in the past, by requiring a stamp and then refusing to issue it.