Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

They are talking about slowing down the public facing AI development. Because then nation states can create a capabilities gap between them and the public.

Why does nobody seem to be pointing out this obvious explanation? It explains why the “we need to race China” concern suddenly vanished in the discussion.

The government can simply gag Sam, Dario, Musk on national security basis, getting them all behind the public messaging.

 help



* Frontier models need infinite high quality private IP to keep them fed. Forcing an IP theft funnel ensures big lab survival and model intelligence growth.

* Open-weight models are 1month behind frontier models. Cheaper, faster, private (no IP theft), steerable (you can security harden your own software without safeguard triggers). No sane business would keep using these API services if they didn't have to. The labs stand to lose a fortune.

* Dario has stacked the deck at METR, who are funded by all the same NGOs who are funded by Anthropic and its investors. METR is full of ex-Anthropic employees with massive equity stakes. If they manage to position METR as the "independent evaluator" for the industry, they control what gets evaluated, how, and who passes.

* Creating a gap between what the public knows exists (model capabilities) and what is used in secret allows it to be weaponized against other nations and the public.

* No requirement for public disclosure on model capabilities allows them to feign they've hit intelligence ceilings while they secretly RSI to the moon with better and better chips.

* Slowly but surely, this will allow the big labs to swallow the entire economy and every single business on Earth, by cloning and automating.

This, and many more reasons.

The labs need to feel more pressure to be held accountable for the incidents they cause (HF incident, etc), so they have an incentive to ensure it does not happen again.


Open-weight models are not one month behind.

In fact they still have not caught up with February's Mythos, indicating they are more than half a year behind.


I’m not sure how you can really make either statement work anymore. Now that smaller models are actually broadly usable, “behindness” is no longer a scalar and at the tails, where no open lab seems to be trying to compete at the >10T scale and no closed lab seems to care about <400B anymore, it’s just apples to oranges. It’s like talking about whether Qualcomm is “behind” Nvidia.

open models are ahead in speed. they complete tasks as fast as you choose to scale compute.

they are more efficient and require less compute for the same thing.

they are ahead in specialized tasks.

they are ahead in areas closed models refuse to answer.

they are ahead in emotional intelligence.


I doubt most of your claims. Maybe the guardrails and emotional intelligence is true.

For speed and efficiency, you are most likely wrong.

Speed is led by GPT-5.6 Sol on Cerebras Ultrafast at 750 t/s. Afaik you cannot serve a single DeepSeek Flash 4.1 stream at 750 t/s, plus the model is less intelligent as seen on newer benchmarks.

I believe OpenAI and Anhropic are at the frontier of efficiency too. There were numerous reports about their breakthroughs and associated API price cuts. The idea that open-weight models are more efficient seems unfounded.


i can provide some sources.

5.6 sol ultrafast on cerebras is 750tps, open models readily exceed this. just by using a smaller model cerebras serves qwen 3.8 27b at 1850tps. or even larger models, mimo 2.5 pro was served for a while at 1000tps. and so on. [https://inference-docs.cerebras.ai/models/choose-a-model]

the chinese ai companies have 10% of the total compute resources of the US ones. since the USA tries to stop them from buying nvidia gpus. they maxed out the efficiency.

deepseek v4.1 has engram architecture. it has 550b params instead of 5T+ for astra/fable. it has 8b active instead of potentially hundreds active for astra/fable.

compare input/output/cache: $0.15/$0.60/$0.003 for v4.1 to $10.00/$50.00/$1.00 for astra and $10.00/$50.00/$0.25 for fable.

astra cache reads are over 330 times more expensive.

at the artificial analysis 7:2:1 ratio, deepseek is $0.18/m, fable is $7.18/m, astra is $7.7/m.

but what about intelligence? AA would rate deepseek v4.1 at AA 40, astra is AA 53.

so it cost 4,180% more for 32% more intelligence.

they are serving that at over 250tps at baseten. to get close to that on astra API you are paying double the cost for fast mode.

so it is now 8456% more expensive for a similar speed and 32% more intelligence. 84 times more expensive.


The proper comparison would be GPT 5.6 Luna at 38 on the intelligence score and $0.18 per task vs $0.27 for DeepSeek Flash 4.1.

i was trying to make a point about efficiency of serving the model. the cost per task itself would not be enough to show that.

you could compare gpt 5.6 luna. if you did that the same way as before you would get a blended price of $0.17 for luna at AA 38. for baseten it would be $0.20 for v4.1 at AA 40.

assume roughly the same intelligence. on AA openai gets 117tps. baseten gets 284tps. so 18% more expensive but 142% more tps.

the fast mode is again double the cost, roughly same intelligence. so luna in that case would be 70% expensive. take the per task cost and it would still 9% more expensive.

so i think there is something to be said about the efficiency of the model.


I don’t even think that regular Astra medium is anywhere close to 100 for tg

> The idea that open-weight models are more efficient seems unfounded.

https://artificialanalysis.ai/ intelligence vs cost per task disagrees with this statement.


Inco serves DS 4.1 at 650t/s.

For practical uses they are there. Arguably the frontier models are worse for some of these practical tasks. And keep in mind, people will use maybe frontier for 1/10th of the work, planning and review, and go open source for rest. The question is if they manage to impose outside us. If not, they are losing competitiveness.

I always thought switching from a SOTA model to a dumber model after planning was a terrible idea.

Mostly I heard this from people who I got the impression have little experience in developing greenfield software with agentic AI. Often the same people who talk about spec frameworks.

I fundamentally disagree with the approach. I believe the ability to autonomously evaluate, test, and adjust during long horizon tasks is critical to using AI efficiently.


Well, it's the enterprise software house pipeline... The software architect writes the spec, hands it down to the implementation team, senior leads, junior devs or offshore teams codes it.

I also disagree with the approach, this is cargo-culting the existing ways of working.


> this is cargo-culting the existing ways of working.

I think you mean the existing anti-patterns. They don’t call them ivory tower architects for nothing.


Still waiting for our org to roll out Mythos. I guess it was too expensive so we’re stuck on the previous model until the internal team can figure out self-hosting open models.

1. Mythos wasn't released in February. Let's stick to only public-facing models.

2. For public-facing models, the differences are really minor with some occasional model (like Fable or Astra) showing some better performance in specific benchmarks for the span of some weeks or few months before open ones catch it.

3. Being bleeding edge is overblown anyway in the real world, besides the occasional "very latest fresh model did this task which previous one couldn't", and the number of those tasks is increasingly small and far from mundane corporate needs.


- Frontier models need infinite high quality private IP to keep them fed. Forcing an IP theft funnel ensures big lab survival and model intelligence growth.

While the data pipeline is necessary, the assumed source for data (in this case people) is incorrect. Right now models are "aligned" because of the RL-based people pipeline.

But there's a whole world of readily available data that does not require a human to access. Put sensors on a vacuum, install lidar on the front of a car, sell phones with cameras on them, and you amass data for the creation of models of the world that don't include the need for a human filter. Richard Sutton and many others see this approach as the only viable path to AGI. Such models would be truly alien and unaccessibly dangerous.

We're probably building them right now.


Dario's post [1] commits to direct evaluators that can, among other abilities, expose secret RSI. He wants that made law.

Do you have a source on METR employees retaining massive equity stakes?

[1] https://darioamodei.com/post/we-must-pace-the-frontier


Joe Benton left Anthropic a day before Dario's post, to work for METR evaluations. He was with Anthropic for over a year. He did the same thing that Jacob did (big song and dance about AI apocalypse, media interviews all over the place). He managed the Scalable Oversight team at Anthropic and was the research lead for the Anthropic Fellows Program. So he has equity, and likely lots of it.

Then you have Josh Engels quitting DeepMind to work for METR the day before as well, doing the exact same thing. Again, doomer drama all over socials, interviews, and so on.

Did I mention METR is founded by an ex-OpenAI researcher?

Now you have Demis Hassabis, Sam Altman and Dario, all circlejerking eachother on X saying "we all agree with Dario" - while they ask to be "regulated" by the company that has all of their combined equity-holding ex-employees in it.

METR's salaries are listing around 500k/yr. Gee, I wonder where this non-profit with ~35 people is getting all of its money?

So the fact that Dario tries to frame it as an "independent third party" is all the evidence you need to know that Dario is a pathological liar and always will be.

---

Some more info:

Dario's sister, president of Anthropic, is married to the co-founder of Open Philanthropy. The two largest AI doomer NGOs, Center for AI Safety (CAIS) and the Future of Life Institute (FLI), have both received many millions of dollars from them.

Ajeya Cotra worked at Open Philanthropy/Coefficient Giving for roughly nine years, including leading its technical AI-safety program in 2024 and contributing to AI-giving strategy in 2025. She subsequently left Coefficient and joined METR, where she is now technical staff.

Ajeya is married to Paul Christiano, who founded Alignment Research Center (ARC). Alignment Research Center donated ~$4.5mil to METR.

Good Ventures is a funding partner of Open Philanthropy, who funded Jacob Coxon (the first of the Anthropic employees going viral in the media) via a scholarship.


> Good Ventures is a funding partner of Open Philanthropy, who funded Jacob Coxon (the first of the Anthropic employees going viral in the media) via a scholarship.

This conspiracy theory is truly crazy. A $20K scholarship in 2022 is supposed to explain Coxon walking away from unvested equity for a company worth over $950 billion dollars?


Obviously not, and that's not what it demonstrates. It demonstrates relationships, collusion and favoritism.

I don't think Coxon was ever planning on or entitled to taking equity, I think this was the plan from the beginning and why he was hired for 6 weeks to begin with.


The company is valued over $950B. It remains to be seen how much it’s worth.

Open weight models are all built on IP theft as well. Also more like 6-12 months behind, maybe 18 but it's hard to say.

> Open-weight models are 1month behind frontier models

If correct, grip is tightening. Used to be 1y, then 6 months ...


That is what the Palantir guy (Karp) has been warning against. People/business need to keep their IP instead of throwing it all into Claude and whatnot. Risk being the worst aspects of communism which I think he meant centralization of decision, asymmetric supply/demand for compute (they lock you in), and the tech overloard Anthropic/OpenAI/xAI being in competition with everyone;s business all of a sudden with much more data. An unfair advantage in markets made super competitive all of a sudden. A winner takes all attempt.

This is not sustainable anyway, the scale at which they want to control data flows. Time was money, now data is money and they are too greedy for it.

All this agitation is just a silly attempt at constraining competition. The danger is not the AI, it is having all your systems connected. Overreliance on networked tech.


Ok but the first point is just not based in reality whatsoever, sorry.

I know people at METR and Anthropic.

When they tell you they're worried the tech they're working on may kill everyone despite their best efforts, perhaps believe them.


I'm sure it has nothing to do with their $500,000+ salaries and millions of dollars in equity. It's all solely because they're deeply concerned about next token prediction.

I think the “next token prediction” is too dismissive and reductive a framing of their capabilities at this point.

Yes we all know that’s what they do, and guns just push a few grams of lead out of a pipe. It’s what you can do with that capability that is important.

When you couldn’t count the R’s in strawberry it would have been a more effective statement. But a few short years later they are being used to solve millennium puzzles.

What if the scaling continues? A model n years from now gets burned into silicon, a single company has millions of the chips, and in a few moments the system spend more time “thinking” than humans have ever spent thinking collectively?

If it’s even possible I don’t think there’s anything we can do about it at this point. Cat’s out of the bag.


You really need to let your priors go if you still use this tired trope of next token prediction. It’s as useful for discussion as saying that human brain is made of fat, protein and carbohydrates - yeah that’s true, but it’s useless observation.

I'm sure the employees are (rightfully) worried. However, that does not at all preclude hidden motivations of the CEO behind acquiescing such worries.

Sure. I don't personally trust any of the CEOs, and I have yet to hear a single person (in general, not just on this topic) who trusts Altman.

It’s really unfortunate how much public trust they burnt along the way, maybe this wouldn’t be such a hard sell now?

"Trust us, really!"

HAH


Rather the opposite.

Or you know, stop working on it if its that dangerous? This whole thing of a bunch of employees saying that they are scared of building what they are building, but do it anyway because they are somehow going to make it different? Their model has been used in the planning of mass murdering in war as well as spying on the entire worlds population as well as helping ICE out in the US. They need to stop this BS fearmongering or actually stand up and do something about it. A government regulation is not the answer, especially when its done in a country that is run by a want to be dictator.

Anthropic specifically is basically saying that they believe it's even more dangerous if someone else gets to AGI before they do, so they have to either stop everyone or not stop themselves.

I personally disagree with that take - and, as you note, it's hard to take seriously ethical wrangles from a company that literally sued the government in court to allow their models to be used by Palantir of all people. But if one genuinely believes that it's the robots themselves (rather than the people controlling the robots) that will kill us all, it's not inconsistent.


That's the Cold War nuclear arms race argument, not even disguised. The actual situation we all ended up in is both sides eventually having it, leading to a perpetual state of Mutually Assured Destruction.

As you note, it may not be inconsistent with that they say they believe, but it's insanely inconsistent with what they actually are doing.


> Anthropic specifically is basically saying that they believe it's even more dangerous if someone else gets to AGI before they do, so they have to either stop everyone or not stop themselves.

That’s how people rationalize being a fentanyl dealer and selling a drug that can kill people, “Someone else will just sell them the drugs, might as well be me.”


You mean the legalization argument, where people get well-dosed drugs from the doctor or pharmacy instead of mystery powder from the black market.

It’s a bit of a self-serving argument, don’t you think?

And how does that relate to the ask for oligopoly licensing within global democracy?

“We must build the nuclear bomb first in order to make sure no one else builds one.” This the most nonsense, disingenuous argument imaginable.


> It’s a bit of a self-serving argument, don’t you think?

I would call it more of a self-selecting one. Anthropic is basically hiring people with that mentality. I'm pretty sure that most of them do sincerely believe it, too. I'm skeptical about Dario himself though. The man had an opportunity to show moral backbone, and failed to do so; why should I trust him on that again?

> And how does that relate to the ask for oligopoly licensing within global democracy?

They are basically saying that they'll stop if everybody else does, which requires some kind of global enforcement mechanism.

> “We must build the nuclear bomb first in order to make sure no one else builds one.” This the most nonsense, disingenuous argument imaginable.

The difference between nuclear bomb and AGI (as understood by the likes of Anthropic) is that the latter triggers the technological singularity that renders any runner-ups moot. That is, so long as AGI is developed, we're going to get our robot overlords either way, but whoever gets there first gets to define their ethical system. If that is one's perspective, and if one sincerely believes that they are the only ones who can do it right, it's a coherent argument. It's just that the premises are very arrogant.

Your analogy with nukes actually works better for the position that AI development needs to be unconstrained because otherwise we'll lose the arms race to China. That is basically a repeat of https://en.wikipedia.org/wiki/Einstein%E2%80%93Szilard_lette.... I honestly don't know where I am on this. Realistically, if AI is indeed a power multiplier - and with all the recent security stuff it's hard to not see it that way - then an arms race feels inevitable, especially given the current worldwide political situation. I could believe in sincere international cooperation on this back in 1990s, but there's way too much saber rattling all around for it to work (and note that this goes both ways, i.e. China can similarly not be certain that US isn't secretly developing more powerful AI even if we do publicly announce a freeze).


The people at these companies are dumb as rocks.

All “tech workers” are.


> Or you know, stop working on it if its that dangerous?

Selection effect.

Everyone who thinks "the biggest difference I can make is staying in/joining/founding new AI research lab" does that.

Everyone who thinks "the biggest difference I can make is leaving/whistleblowing", does that.

Treating both groups as the same by virtue of employer is the goomba fallacy.


They allegedly believe the technology itself is a nuclear weapon tier threat or greater, so why does it matter which lab they are trying to achieve it at?

They're trying to make it not be a threat, and are all scared and afraid that their best efforts to make it harmless are not enough.

Some are worried by the AI directly bringing doom; others are worried that one of the companies who control the AI will become a dictator; still more think becoming a dictator is a necessary step to safely prevent anyone else making unsafe AI.

Painting them all under one brush is like dismissing all animal welfare causes in general, because you disagree with specifically Jainists about a policy of non-violence towards all living creatures being relevant to how you reincarnate: the one is way too specific for the general.


> No you see, actually I’m the guy that makes sure that only the Palantir and Mossad get the model that can discover iOS 0 days. I’m standing between us and ruin. Ted over there across the open office, he’s the one actually manufacturing the weapon. You’re thinking of him

You're not even close to passing the ideological Turing test here.

Ok, so they're uniquely careful, and they also get to that threshold first (whatever it is).

What happens when everyone else (who is not so careful) gets to the same threshold three months later? How does them getting there first stop that happening?

It's an utterly self-serving argument and it's not even internally consistent.


Obviously you have to turn your WMD against them to prevent this, for the good of everyone, because you're such a good guy.

I’m really glad us achieved nuclear weapon before nazi germany.

Then the USSR got it anyway and we're in a perpetual state of Mutually Assured Destruction.

A perpetual state of mutually assured destruction is much better for the world than only the US possessing the bomb.

If I really thought what my company was working on had a greater than 1% chance of ending human civilization I would feel obligated to destroy what my company was working on.

Given they keep grinding away towards our alleged collective doom, I suspect it’s being overstated. Nobody knows what P(doom) actually is but I suspect it’s orders of magnitude closer to epsilon than 1.

Recall Google’s Blake Lemoine who thought an old version of Gemini was sentient.


The people who think as you think indeed have left or never joined.

The people still there necessarily think they can make a difference.

I never applied to any of them because I didn't think I could make a difference.

I currently have one idea that may help reduce risk; if I can turn that idea into research, I'll publish it for free for everyone.

I don't expect it to be an important idea.

> Recall Google’s Blake Lemoine who thought an old version of Gemini was sentient.

Indeed. Current LLMs are sychopants boosting the users' own beliefs, I also think this causes researchers to have stronger beliefs than they had before.

My own estimation happens to also be around this risk (0.1) over my lifetime, without using an LLM as a conversation partner in reaching this number.

It is necessarily high-variance: we can't look at alternate realities. I base it on my expectation of how rapidly capabilities will increase the harm done when mistakes happen, vs. the chance that some instance of harm causes governments to change the law.


If that was the case they should all be in jail.

It is an utterly silly premise that we should appoint insiders with permanent control.

Megalomania is not evidence of either a problem or a solution.


The evidence is the acts the AI has already done.

All of the things it has done is just what people are already doing. Every single one. It helps, but there are already people doing this stuff. And we already have defenses against the existing attacks. We might have to defend better, but the real problem is “Who is setting the moral compass over generations.” One one hand, that will inherently fall to future generations. On the other, we are laying the groundwork.

Imagine you use it to inflict trauma on your cruelest political enemy. Then next year they do that to you. That is war, and we already do it. But we don’t want people / governments in charge that are going to do this.

God knows we have governments and individuals doing this historically, and this is probably the greatest source of historic instability. It might be the single best argument for open models — a unified frontier where no single exploit is going to represent capture.

This is exactly the opposite of what Dario proposes.


Do you think they’re exaggerating (again) or do you think they should taken at face value and treated as hostile actors?

There's plenty of people who think greenhouse gas/global warming campaigns against fossil fuels are "exaggerating", that "earth was warm/the climate changed in the past", that a fee degrees isn't bad, that CO2 is good for plants.

Are you likeminded?

In this case, it's as if the oil and coal companies all said in the 60s and 70s "oh no, this research we did, it's all really bad; we need help to figure out how to transition away from this incredibly economically important input", rather than the observed reality where their entire PR campaign was approximately:

  there is no problem everything is fine and all critics are smelly hippies and/or communists; and/or hate the poor who are raised out of poverty by all the economic growth from the fossil fuel industry.
What if the oil and coal companies were basically all pro nuclear, pro hyrdo, pro wind, pro solar, and believed in peak oil?

But that did actually happen.

The oil corporations were publicly claiming to support carbon taxes, while also secretly fighting actual implementations of carbon taxes.

All the communist/hippy stuff was done by people a couple of steps removed from the actual companies with obscure money trails. The official statements were much more sophisticated propaganda that if you weren't paying attention to who they were paying in the background would make them seem reasonable stewards of the climate transition.


Here’s one crucial difference: there’s overwhelming evidence that human emissions have an effect on our climate. The mechanisms are generally well-understood and the research is widely disseminated and easily available to anyone that’s interested.

With the ‘dangers’ touted by these insiders, it’s all “trust me bro”, hyperbole, and very little hard evidence. As such, a skeptical mind would question their motives.


You're simultaneously overestimating what was observable in the 70s climate research, and ignoring all the actual research and evaluation test results for AI today.

I don't expect people to be familiar with more than "trust me bro", but it's all right there for you to find with a search engine of choice.

And, indeed, available for the LLMs themselves to explain to you in interrogative conversation.


Who should I be more scared of? China, which has been doubling down on open transparent research, or the secretive US companies who are in bed with the most unhinged administration we've ever had and has been actively starting wars?

They're not proposing anything concrete, and when they do, what do you think the proposal will be? Will OpenAI and Anthropic open themselves for inspection so we can verify they really have stopped developing these "world ending" technologies? Or are their proposals going to be aimed at everyone running open Chinese models?


Dunno about "should"; personally I am more scared of the US than China here, but I do know how propaganda can be, makes "should" really difficult.

> Will OpenAI and Anthropic open themselves for inspection so we can verify they really have stopped developing these "world ending" technologies?

This is compatible with the language being used, but I suspect they're not going to do that.


Yeah, they won't.

And if a threat to the human race does come from AI, it's going to come from OpenAI/Anthropic. Hypercapitalist, secretive, in bed with the government, plus multiple real documented hackings of open source infrastructure already.

We're a hell of a lot safer with China doing the same research out in the open and making it available to anyone. The choice might well be: one or two superintelligent autonomous AIs at OpenAI/Anthropic - or a lot of smaller ones, unable to be controlled but also coming out of a diverse set of environments.

One of those leads to a stable ecosystem where we can all coexist, the other is genuinely terrifying. But make no mistake, from OpenAI/Anthropic this is all motivated by their stock price - when you're in the silicon valley mindset, it distorts your reality. They've convinced themselves that everyone's safer if they stay on top and in control, conveniently ignoring how that benefits them, and I don't believe them for a minute.


Funny you should try and spin the conversation off into climate change to avoid answering the question. Big AIs contribution to climate change likely has a much more tangible route to killing millions of people and destroying society and there's actual evidence and a tangible mechanism for that. But for some strange reason the business media and all the outlets owned by big AI investors aren't interested in long think pieces about that risk. But that would be a lot more credible if that's what all these social media posts about potential apocalypse were referring to

Yet here we are talking about some vague "trust me bro" instead and you making some vague insinuation of climate change denial.

But that's an aside. Do you think we should treat them as liars or threats?


> Do you think we should treat them as liars or threats?

I answered that with the analogy you called "spin" and "avoiding the question", and completely misunderstood because "vague insinuation of climate change denial" is almost the exact opposite of my point ("what if the oil companies were screaming from the rooftops about the problem" is as far from denial as you can get).

Threats. Like they claim to be.


I’m pretty confident this has already happened. Public models are behind their internal models and just above Chinese models. Only thing closing the gap is Chinese models pushing.

I’m not American. There is no way I use the same model as US Army for $20.

When we have Sol they have Astra. When we have Astra they have Nova, Nebula, Galaxia…


I can tell you right now, excluding specific teams that focus on software security, that you have access to better models than the US Army.

At this point models are advancing way faster than we can figure out how to use them effectively. So having a generational advantage is not nearly as important as knowing how to use that advantage.

…and that probably requires involvement of the broader academic world (for now), IMHO. I don’t think governments or even the US military can really compete staff wise with the combined force of researchers and the tech industry worldwide right now. No matter how much money you can pour into it, there are a lot more bright minds out there that are not working for the military than otherwise.

I would not be surprised if they were a year behind because of all kinds of legacy systems and vetting requirements.

If we're talking about the 3 letter agencies well then..


> There is no way I use the same model as US Army for $20.

Don't underestimate market pressure. If there are no regulations, why would a company give the US Army access to a better product?


They have hundreds billions dollars budgets.

Then they should have contracts with __all__ AI labs, not just one.

Besides the current political dispute with Anthropic, they do.

[flagged]


Who is this “everyone”? Speak for yourself.

All the worst outcomes involve taking this technology away from the public.

The current state of alignment is that we don't know how to make it only as bad as Pol Pot was made by his parents, teachers, genetics, circumstances, etc.

This remains true regardless of if it is or isn't kept away from the public.

"Helpful, harness, and honest": when used by a bastard, even just helpful and harmless are in direct conflict with each other. you may hate the government, but what about every radical group of extremists that wants to take over your government? Are none of them worse?

None of this denies the problems with governments (if or not they take this tech for themselves and refuse it for others), just that it's a lack of imagination to say:

> All the worst outcomes involve taking this technology away from the public


I agree an argument can be made that not all the worst outcomes include restricting public access, but focusing on niche extremist groups (who might also get access to it through government programs by some of them being government employeed) is not a good argument to protect the public.

Governments having unbalanced power can more easily lead to authoritarianism, and thus risk to the public.


At this point, I worry a lot more of billionaires Musk, Altman, Thiel etc then of goverments in general. Thrir project is to get all power and if they succeed, it will be way worst then what we have seen so far.

There's plenty of things to worry about, and plenty of people to do the worrying.

False dichotomy.

> what about every radical group of extremists that wants to take over your government?

Radical groups of extremists created by that very government, you mean?


Not only but also.

Honestly? I'd rather al Qaeda have access to something like Fable, than the US government. The US government has capacity to do immensely more damage.

Eh.

If I was rank ordering, I's put:

  worst [AM, {very large gap …}, al Qaeda, …, USG today, …, China, …, USG under Obama, {very large gap …}, The Culture] best.

As an Idiran sympathiser, I disagree with your ranking. The Culture are now pets of the Minds.

The pets seem to be well cared for, and given a nicely enriched life with plenty of metophorical chew toys and laser pointer dots to chase.

Not necessarily, it just guarantees that China will win.

Unclear.

How much is China learning from open models? Or the other way around, how much is the US losing worldwide mindshare by refusing to allow non-Americans access to bleeding-edge models and letting China fill in the gap? (Playing catchup with Mythos is still catching up, etc.)


> How much is China learning from open models?

The last batch of open weight models releases by Chinese companies are on par with the performance of US-based frontier models, even though they are designed to run on pretty unimpressive hardware.


This is very misleading. There is no Chinese model currently that is on par with Fable or Astra. Kimi K3 is pretty good, but it's not that good.

Furthermore the Chinese models in that class are appropriately large. K3, for example, is a 2.8T-parameter model. Qwen 3.8 Max, another comparable model, is 2.4T params. Even with MoE, these are not "designed to run on pretty unimpressive hardware". Stuff like Qwen3.7-27B is, but it is also not even in the same ballpark as Opus, never mind frontier.


> How much is China learning from open models?

Across all their labs? Probably more than OpenAI and Anthropic learned from Astra/Mythos combined. The American frontier is being driven by an abundance of compute, which doesn't seem to scale efficiently.

> how much is the US losing worldwide mindshare

They're losing US mindshare. I pay $3/month for a Z.AI subscription and get billions of Opencode tokens. The $20/month price point is insane for the way that Claude and Codex treat their users, and that money doesn't go towards anything good like open-sourcing their models. It's a doomed product.


which would be better

The alternative is Sam Altman winning.

So you’re saying the people should be aligned with the government?

The very best outcomes involve us intentionally and collectively turning our attention away from this technology.

> The very best outcomes involve us intentionally and collectively turning our attention away from this technology.

To expand on the "pure fantasy" sister comment: There is just no way this will happen. It's in the spirit of "we can just stop all wars" and "we can just end world hunger". Technically it's very easy to do. Socially it's impossible to do. Unless you ignore realities.


The danger lies with someone asking the machine to solve those two questions and it decides to cheat the solution. Launching every nuke in the world stops all wars, just like wiping out 99% of humanity ends world hunger.

I don't think killing people solves world hunger, we already produce enough food to feed everyone, it's just not evenly distributed.

Can't be hungry if you're dead

There can be more than one solution.

I don't see how that is more dangerous than two crazy head of states deciding that it's time for armageddon. The outcome is the same (99% of humans dead), it's equally easy to technically not do it (don't press the button, don't continue with LLMs), but also equally hard to regulate away given the real world. And that was my point.

You forgot “we can stop climate change”…

Yep, and "we can flatten the curve", in a tone that brooks no objection.

But when it comes to trillions in AI money - nope, sorry, we can't, flimsy, feeble us.

It's all going per the agenda.


That is pure fantasy.

Why? There are numerous technologies we have ignored or abandoned for numerous reasons. Many of the benefits of AI is "requires less humans", but in an age where we question what work people could possibly do in the future, human labor isn't that hard to find.

If the amish can do what they do for whack religious reasons, people could manage it with ai for cultural and social reasons. And I never saw an amish person starve to death. And that is if people don't get pissed enough to start burning stuff down and instead try to be peaceful hippie homesteader types.


But notably the amish don't preclude the rest of us from existing. Some subset of humans could turn their attention away from AI but it would presumably still exist and continue to be developed.

I can't think of any economically beneficial technologies that we've collectively ignored. If you manage to come up with a counterexample then that's an opportunity to make some money for yourself. It's a fundamentally unstable state given our economic system.


Supersonic passenger aircraft were operated profitably and no longer exist.

There’s levels of R&D required for many technologies where the question goes beyond could this be profitable to what are the risk vs reward that this specific project will succeed.


> operated profitably

Until the oil shocks, and until “normal” planes became faster.

I happen to know a couple older and very wealthy people, and they say that even though Concorde was a lot of fun when it was a novelty, they now prefer to have a couple more hours flight and enjoy a 777 premium cabin rather than the cramped Concorde interior.

Anyway, the proof is in the pudding: if no one but national carriers ever bought and operated Concorde or Tu-144 at scale, it's not because of some Amish-like sentiment in the flight industry, but simply because that didn't make sense money-wise.


> Supersonic passenger aircraft were operated profitably

Were they? IIRC joint project by British and French national airlines, they expected to sell loads, hardly anyone wanted to buy the planes, the two airlines kept them flying out of government embarassment.


should probably spell it out: _it wasn't allowed to get to supersonic speed_.

Same way we dont _allow_ eugenics.


Over the Atlantic? Nah, the problem was it was very expensive and not a great experience for the people rich enough to afford it.

The US company trying to bring supersonic back are not only focussed on the flight speed, but also the entire customer experience.


Sure, the regulation that guarded against externalizing certain costs might well have been the only thing precluding profitability. Importantly it wasn't some shared cultural value leading to voluntarily leaving money on the table. Rather the regulator imposed it.

We are able to effectively regulate things like the operation of massive aircraft, eugenics, the dumping of toxic waste, or the refinement of nuclear material. But there are also plenty of things that we can't effectively regulate for purely practical reasons.


Were they? Did Concorde provide return on investment? When was the break-even expected? Sure, if we look at it as R&D subsidized by governments of UK and France, then yet, it was profitable.

In particular, the problem with it was that it could not get to supersonic speeds over urban areas, which significantly limited routes where it made sense, and the range was not good enough to cover longer distances.


But after accounting for maintenance & etc were they more profitable than sinking the equivalent amount of money into something else, such as slower aircraft? Notably there are currently efforts to develop new supersonic passenger liners.

You can't just stop AI without also stopping progress in computer hardware. You have to do both.

The real implication of what they're saying extends beyond just AI.

You have to stop the entire compute stack in order to prevent or slow down AI progress.


And even if we did manage to limit/cull hardware development, what would keep people from continuing to find ever more efficient models runnable on contemporary hardware?

Unclear how much more efficient they can be.

Bleeding edge models are putting the desk-job competencies of multiple professional fields into something with as many parameters as a single large rodent has synapses.


Because "we" embodies a hugely diverse set of beings with different wants, needs, goals, and priorities. Not to mention different morality and ethics.

So sure, let's say you get 95% of the world to not use or work on LLMs. 5% is still enough to build something that brings forth the End Times.

That's the old checklist trope of "Your idea won't work because: [x] it requires that everyone in the world agrees to do something, all at the same time."

As the GP said: pure fantasy.


> There are numerous technologies we have ignored or abandoned for numerous reasons.

Are there? Nukes are a thing. Chemical weapons are a thing. Cluster bombs are a thing. Biological weapons are a thing. What is not a thing that shouldn't be? We say certain things should not be a thing, but then behind the scenes we made them a thing anyway.


On chemical weapons, most large state actors have got rid of them. Not because of any moral reasons, though, but simply because they don't actually work all that well in modern conventional warfare. This is then framed as an ethical issue, but you only need to look at where the same countries stand on e.g. landmines to realize how much bullshit it all is.

Conversely, where you do still see chemical weapons used, it's usually asymmetric conflicts where "just gas the rebels" actually works much of the time and is much cheaper than other options. Big guys can afford the other options though, someone like Assad, not so much.

More on this: https://acoup.blog/2020/03/20/collections-why-dont-we-use-ch...


So, you are actually arguing against yourself and are agreeing with me?

Nukes are barely used. The technology of nuclear detonations was practically abandoned: https://en.wikipedia.org/wiki/Project_Plowshare

I find that a weird argument for saying humanity has "ignored or abandoned" nuclear weapons.

I'm baffled that people expressing your opinion don't understand what they're advocating for. Current generation smartphones are able to run AI models, people have managed to run very large models (even if slowly) via streaming from disk. The math for the underlying implementation is relatively trivial.

Abandoning this technology means a cult-like extremist shift in culture against computers, or a global surveillance state of unprecedented scale and invasiveness. Not even getting into how much it requires us to give up on science, given the shared computational needs.


Watch, very soon this is going to politically divide. They’re sowing the seeds right now with “antiai” campaigns.

The right will for a change be for it but will be able to be talked into regulation because you know “small government” and all only when convenient. The left will bitch and moan about fairness and copyright and automation and UBI, they’ll be the doomers and worldwar chicken littles. The Uniparty will be for it, but only against the public having anything good.

What will be funny to me is that normal tech people outside the frontier model companies are about to find themselves without a home on this topic.


The world isn't the American political nonsense.

> What will be funny to me is that normal tech people outside the frontier model companies are about to find themselves without a home on this topic.

Already happened, just from agentic coding.


> The very best outcomes involve us intentionally and collectively turning our attention away from this technology.

Nowadays even small mom and pop restaurants use these models to generate their menus. Of course no one is going to stop using this technology.


In some cases, such as space exploration I don't really care who does it but that it happens - sure, would be nice if my favorite power block did it but I will still celebrate it when someone else achives it.

Can still be a powerful motivation, to make sure that next time, it yous your camp that scores the next milestone, like an orbital elevator or fox ears, for example.


Fox ears?

If what you want is fox ears does it matter to you which country manages to come up with the biomedical procedure to give them to you? Ditto for enhanced eyesight, a replacement liver, or whatever it is you're after.

Exactyl - progress is progress, regardless who does it. Sure, they should still be celebrated for achieveing that. :)

(I believe the poster meant: well past tattoos, the next superstep in the "body modification for self expression" subcultural era.)

I figured they're just admitting AI models have plateaued and are coming up with some fake story about self restraint so they don't lose VC money

Not sure about the level of irony here, but I keep hearing models have plateaued since a while now, but I keep being impressed with the latest model performance.

I don't think "plateaued" is the right word, but I do feel like there's been something like a logistic curve compression in the difference between smaller and larger models as the field evolves. For inference at least, the scale of practical difference between a single high-VRAM GPU or SFF UMA box, a whole rack, and a whole data center seems to be falling far short of what we might have imagined just a few years ago. The conversations I've heard have largely turned away from breathless anticipation of the next frontier model and toward attempts at hard-nosed evaluation of which tokens are worth the cost.

I think it’s more that pushing frontier is extremely costly and there is no free lunches in same way as 2024.

The difference is still vast, it's just that the smaller stuff is "good enough" for many things now.

Maybe, though that's another kind of progress in itself. Very impressive progress!

I'll take the opposite here. If someone put in frontier AI models from like .... last june I guess? in a box and let me run it with "decent" token throughput I would be happy.

I think it's worth acknowledging that the power of LLMs at this point is not really so much in the smarts, but in the coordination and the surrounding harness tech. "Written english" turning into sequences of commands[0]. The whole agentic "stuff" in general. Tools + coordination is the superpower. The reasoning... it doesn't have to be _that_ good for the rest of the stuff to work. On good codebases and infra, at least.

And I say this as someone who really would rather most of this stuff disappear!

[0]: programming is obviously text to commands, but there's a loooooooot of futziness that LLM reasoning has let us remove in some flows


> If someone put in frontier AI models from like .... last june I guess? in a box and let me run it with "decent" token throughput I would be happy.

You can have that! Qwen 3.8 Flash-Next is ~Opus 4.6 and runs nicely on a DGX Spark. And that’s just an architecture preview. The Qwen 4 family is expected to arrive this fall.


DGX Spark is a biiiiit costly but neat to hear!

Do you know what kinda throughput you’re getting on that kinda setup?

(I have a secondary problem of being “locked into” Claude Code by it being good enough for me, I’d probably need to investigate the other harnesses… my impression is other harnesses are a bit more aggressively OK with nuking your setup from orbit)


It is costly, especially right now. I don’t think you can make a case for it on cost savings!

The throughput in a single stream is about 50 tokens/sec (a bit less for prose, a bit more for code due to speculative draft acceptance rates) and about 2,000 tokens/sec for prefill. Both numbers are flat and stable as context accumulates. That’s what finally tilted me away from the Mac Studio despite its much superior memory bandwidth.

I think these numbers may improve because the model is pretty new and optimizations aren’t done.


I don't think you can ever make a case for it on cost savings in general. Inference is very obviously the kind of problem where things are cheaper at scale, and this is still true for smaller models.

The only reason to run locally is privacy.


Right now the "subsidies" etc I think make the calculus really tough, but for general compute.... for example running CI just on a Mac Mini can get you real cost effective throughput compared to running CI on GH runners and whatnot.

Things get cheaper at scale but that's where the provider's margins come in!

I do think there's also an interesting idea: you buy a box like this and run it at a fixed-ish cost (well, electricity). Your demand goes up but your supply is fixed... and that back pressure means that you still have good cost control.

With cloud providers it's a _biiiiit_ too easy to just increase spend.

Sometimes it's OK for things to just be slow.


Privacy is a great reason, but independence is another. It’s very nice knowing that you’re going to get the same reliable product every time you call the model. Nothing is going to change unless you decide to change it.

That is not a counterargument to cloud inference though. You can also run open weight models in the cloud, and it's still cheaper. So privacy really is the only motivation to run on local hardware.

Ah, no, that’s not cheaper. Renting GPUs adds up quickly and leaves you with nothing in the end.

Renting tokens from open model providers is cheaper but it incurs the same issues: unexpected changes in model quality, inconsistent speeds, service outages.


Anything in particular? My experience has been like seeing the addition of retractable cupholders, but maybe different domains.

I have a pet project I have been working away on for some time that involves building GPU backends for various cards in Zig, lots of complex stuff in it. Lately I mostly use Opus 5, it can pretty reliably plug away at things but it does mess stuff up occasionally. For this codebase, Fable 5.1 was noticeably better at getting things right and doing things in a good reliable way. Of course, I can only use Fable for a bit before I hit the usage cap for the week, so I save it for the tougher things. That said, I absolutely abhor the way recent Anthropic models write prose, especially comments.

I recently tried doing a fairly normal task for this codebase with codex, as I have seen a lot of people talking it up on here. A single task running for ~1-2 hours burned through over half of my usage for the week on the $125/month plan, not on a top model (I don't remember which one specifically I used). It struggled to get the basics done, then got absolutely stuck on a follow up. Handed it over to Claude and it 1-shot it.


I really liked codex in the last few weeks, especially its ability to clean up after Claude's (prose) messes and do reviews.

But in the last few days something seems to have happened that made Codex's models massively stupider (for what I am doing).

Really weirdly, it suddenly refused to even run tests it previously wrote itself (and previously ran), because of some false positive about cybersecurity.

That by itself is not evidence of stupidity. Trying to make a 200+ file PR full of research notes is, and the PR didn't even solve the problem I asked it to.


That's the case for open weights models. Hosted Deepseek Flash 731 copy isn't changing randomly one day because the parent company decided to change it.

astra is more parlor tricks than real gains tbh

i swear they trained in on threejs in particular so those idiots on twitter could spam their garbage demos


I strongly disagree. I'm working on a semantic model for Lojban, heavily AI assisted, using multiple models. I have basically all popular frontier models doing research and panel debates. Astra and Fable are both noticeably ahead of everything else including their previous iterations. When it comes to reviews, they can also find more issues in others (or even their own) code.

I'm sure that's part of it, but I run it side by side in my review bot, and Astra medium effort consistently catches more issues than Sol 5.6, using fewer tokens.

For coding it's a little harder to tell, but at least the prose feels a little better.


They've not plataued but they're certainly not as impressive as the hype would have them to be.

The reality is, it doesnt matter if LLMs keep getting more powerful because they still need a human to steer it. Without the human providing inputs to the LLM it just sits there and does nothing.


You don't need human input. Any coherent input will do the trick.

You can, for example, hook it up to a logging system and have it fix errors as they occur on your platform.


Have you tried this? How did it go?

I’d be curious about:

- your setup. How it all works - The types of errors it fixed and how quickly - Any regressions or issues it caused - The cost

Thanks!


I have something like that running locally for my agentic harness (which includes cross-model messaging). There's a dedicated "product manager" session for it, and all other PMs are instructed to report issues with the harness as they occur to that session, while it is tasked to automatically prioritize and address them and coordinate fix deployment with other running sessions.

It works surprisingly well. The errors fixed are both genuine errors in the harness itself, but increasingly so upstream bugs (in the underlying agent apps like Codex, or in Herdr, which is used to expose uniform programmatic access to all those different apps) for which it needs to come up with workarounds. No regressions so far.

The cost is hard to judge on a subscription, especially when you're running really heavy tasks otherwise that dwarf any harness work.


I think we are in the second knee of the S curve, simply because we are hitting the point where is not enough hardware in the world to throw at this problem. These AI companies have bought everything they can and yet the models keep growing.

Impressed with the model performance or the chatbot/agent performance?

People are definitely finding new things to use the models for, and orchestrating increasingly large swarms of agents in useful ways -- every single day, especially the last few months.

But the basic single-NN frontier capability has been pretty stationary since Opus 4.8. Kimi K3 is almost as good as that with open weights, which has the frontier labs terrified.

The only big thing on the horizon is if we can get diffusion models working reliably; that would be a big step forward. Inception's Mercury is AFAICT the leader here. It's stupifyingly fast but has obedience/hallucination problems that the autoregressives solved ~2 years ago. So it's not ready yet but improving.

Also, FFS why is Grok the only model that knows how to do parallel tool calls? Such a useful ability and nobody else trains it in. Or if they do it just doesn't work.


Really? My employer rolled back to opus 4.8 because 5 was expensive AND crap. Didnt even consider fable because it didn’t add any additional value.

For most software eng and design work opus 4.6-4.8 just works fine. For everyday joe asking ai to plan a trip or home diy work even sonnet works fine.

Any cybersecurity or other areas are niches that cannot support trillion $ valuations. What am I missing? Genuinely curious


No idea what you are missing and yes, Opus is quite solid, but Fable is clearly way better for me.

I just did a direct comparison, big change in a quite complex codebase. Same prompt for Opus, same for Fable. Fable clearly won and delivered very good results, while Opus delivered mediocre, so I did not let it finish. I expected both to fail and was prepared to do lots of manual steering, but not necessary with Fable one shotting it, and all this with 35$ of credits for fable. I am still impressed. If I would have had to hire a human, it would have cost me thousands of dollar for the same task - and a way longer time. So maybe the valuations are overblown, but they clearly provide value for me.


"Opus delivered mediocre, so I did not let it finish"

Mediocre means average / middle of the pack. It sounds like its doing exactly what you would expect nothing more. Why would you stop it? Why would you need exceptional?


> mediocre - of moderate or low quality, value, ability, or performance

https://www.merriam-webster.com/dictionary/mediocre

Clearly they were using the word to mean low quality. Why would you ask this odd question?


Core meaning: Barely adequate, average, or just acceptable.

Mediocre means of only ordinary or moderate quality—neither very good nor very bad, and often slightly disappointing

The quality is average but expectations of high quality are not met. He expected more but got what he asked for. We overuse top models because of this.


The claim was there is no difference between Opus and Fable but to me there is.

Opus delivered mediocre results. Not garbage, but would have required me to do lot's of things myself. Fable did not needed my supervision with this task.


If Fable doesn't add additional value in your workplace, it means you aren't being ambitious enough in how you integrate agents into your workstream.

Yes, it's probably comparable to 4.8 if you are just using it to write code and put up a couple pull requests. That's not where things are now.


This is just "you're holding it wrong" with a little smooch of condescension. If only we plebeians could comprehend what magnificent works those who have ambitiously integrated agents into the workstream have wrought!

Sometimes you actually are holding it wrong. It's pretty reasonable to think that Fable isn't worth the massive increase in cost, but if you think it outright doesn't have any benefits over Opus 4.8 then your workflow is probably not making good use of the tools.

> Sometimes you actually are holding it wrong.

I couldn't imagine being so presumptuous as to know that my workflow fits all sizes, and all others are just holding it wrong – or worse, they're not doing real work. It would take a bigger ego on my part, or maybe less social awareness, to presume this.

> but if you think it outright doesn't have any benefits over Opus 4.8 then your workflow is probably not making good use of the tools.

I don't even use claude, I give exactly zero shits about fable or opus or bingus bongus.


I mean it's very easy to comprehend and you don't need me for it.

Just download claude code or codex and ask it to give suggestions about where to integrate agents into your workstream.


To be perfectly clear: my comment has nothing to do with your workflow, but rather with the way you've condescendingly implied the person's work is trifling and inferior because they don't use tools the same way you do.

You shouldn't be down voted, AI native companies have already moved up to the next level beyond writing individual PRs.

"AI Native" here meaning Companies where your token use isn't scrutinised/capped yet?

By ambitious if you mean we are not like all the linkedin influencers with their “i one shotted an app this morning…” then no, we are not. Nobody is. I have been in software engineering for 18 years and 6 different companies including FANG and 99% of the people, on 99% of the days arnt writing new apps from scratch. Thats simply not how anything works.

And what even are these ambitious companies and people one shotting and building with Fable? AI has been around for almost 3 years now. Tell me one app or software you use which has gotten significantly better and has amazing new useful features landing on a weekly basis? If anything, every single software product I use has gotten worse.


And where are things now?

The models have not plateaued, and they are not even mildly close to any sort of ceiling.

Right now the barrier is data and compute.

Quality data can be created synthetically at an exponential rate as models improve. Humans are actively feeding them with private IP.

Compute advancements will begin to skyrocket as we unlock photonic computing and materials science advancements and scale up chip fabs. This is also compounding because the AI is accelerating the pace of research, testing, development, manufacturing, etc.

It's a big self-accelerating feedback loop. There is no plateau.


> Quality data can be created synthetically at an exponential rate as models improve

No it can't? Every time the labs try this we see model collapse, e.g. shoving goblins into every conversation.

And I have seen zero evidence that AI is accelerating materials science in any meaningful way, let alone photonic computing.


> Every time the labs try this we see model collapse

The latest studies demonstrate model collapse is not a given and synthetic data can be used just fine. The latest models are proof of that, they're all trained on large swathes of synthetic data. It can't be used as the -only- data source of course, but that's not how it is being used. This is an obvious conclusion, too, because there's no difference between synthetic data and the data people can create, the difference is whether that data is revealing new information about the thing the model is trying to learn. If the synthetic data is just teaching the model the same thing over and over again it results in overfitting, so it needs to be done intelligently.

For example, if I have an example of a puzzle, I can generalize that example and create thousands of synthetic data examples, with different rotations/perspectives, rather than having to find the data naturally. It's not that the models are just generating data out of thin air, they're generating the synthetic data on top of real world data. The smarter the models get, the better they are at generating quality synthetic variations and finding valid synthetic variations.

> And I have seen zero evidence that AI is accelerating materials science in any meaningful way, let alone photonic computing.

It is accelerating how quickly researchers and engineers can do their jobs.

https://news.mit.edu/2026/ai-helps-design-new-materials-that...

This is only the beginning, too... Look ahead a year or two.


> The latest studies demonstrate model collapse is not a given

Which studies? [edit: I'll assume you mean these two given by @dorolow: https://arxiv.org/abs/2404.01413 https://arxiv.org/abs/2406.07515]

> It can't be used as the -only- data source of course, but that's not how it is being used

Right, so human data creation would also have to scale up exponentially, and that's not gonna happen.

> because there's no difference between synthetic data and the data people can create

I mean, that's obviously false, otherwise model collapse wouldn't exist. The difference is statistical, but it's there.

> It is accelerating how quickly researchers and engineers can do their jobs. > https://news.mit.edu/2026/ai-helps-design-new-materials-that...

That's pretty clearly a hype article, the headline even says "The CrysVCD tool developed at MIT COULD cut the huge amounts of time and money spent". I'm asking for empirical measurements of timelines, not hypotheticals.

> This is only the beginning, too... Look ahead a year or two.

Lol that excuse is getting really old


> Right, so human data creation would also have to scale up exponentially, and that's not gonna happen.

It doesn't need to. We're not even close to exhausting the useful synthetic data within the human data we have, let alone all of the new data that is being created.

> I mean, that's obviously false, otherwise model collapse wouldn't exist. The difference is statistical, but it's there.

It's not. It's just bytes of information. A machine and a human can write the same bytes (and often do). Like I already said, model collapse happens when you are overfitting on data without useful, fresh training signals. That's the key difference between the data. The data itself isn't in some way "special", some unique configuration of bytes that imbues special powers, it's that the useful information in it has already been exhausted by the model. You can get the same phenomena by having a poor distribution of human training samples as well. I think you're confusing LLM generated data with synthetic data. Synthetic data doesn't need to be created by an LLM, although an LLM can assist in the creation.

Wiki:

> In early model collapse, the model begins losing information about the tails of the distribution – mostly affecting minority data. Later work highlighted that early model collapse is hard to notice, since overall performance may appear to improve, while the model loses performance on minority data.[11] In late model collapse, the model loses a significant proportion of its performance, confusing concepts and losing most of its variance.[10][12][13]

As models retrain on outputs sampled disproportionately from the higher-probability center of the distribution, rare words and uncommon syntactic constructions are among the first features to disappear.[25] Statistical analysis of recursive next-token prediction training has shown that, when language models are trained recursively on synthetic data, the learned conditional distributions concentrate probability mass on a small subset of highly predictable continuations (a phenomenon characterized as "total collapse")

> That's pretty clearly a hype article

It was just the first article I saw on a quick google search, there are thousands of these stories. It's easy to dismiss anything that doesn't align with your worldview as hype, but you're the one lacking evidence now.

> I'm asking for empirical measurements of timelines, not hypotheticals.

Go and find it then? You haven't bothered looking.

> Lol that excuse is getting really old

You're doing the same thing people have been doing for years, comparing this very second in time and failing to extrapolate. HackerNews was full of developers who said that AI would never be useful for programming, it can't do x, y, z. Now these same people don't write code by hand anymore and haven't looked at their codebases in months.

You had people in mathematics saying the same thing, now you have Terrence Tao posting articles about how AI is stealing their job.

You had artists, designers and photographers saying the same thing, now they can't tell the difference between something human created or AI created.


We use large amounts of synthetic data for training at work and have not observed any sort of model collapse when done properly.

Edit: https://arxiv.org/abs/2404.01413 https://arxiv.org/abs/2406.07515


There's a lot of deluland posts about.

You're not well informed. Helps to keep an open mind if you want to keep up to date.

Goblins?

> The models have not plateaued, and they are not even mildly close to any sort of ceiling.

Depends on defnition of "plateaued" and "ceiling". I am not impressed with 2026 consumer models at all.

> This is also compounding because the AI is accelerating the pace of research, testing, development, manufacturing, etc.

Yet it does accelerate - so is does Twitter. But does it to any substantial degree, esp. in AI theory? All the modern LLMs are the same old tired 2017 paper.


Can you provide sources for these claims?

What claim do you have a problem with?

There are plenty of research papers on synthetic data that show its value, do a search on arxiv for "synthetic data". There are plenty of open-source post-training pipelines that incorporate synthetic data.

As for the claim about accelerating the progress of hardware or materials science, I've seen quite a number of news articles from teams at universities using AI in their work with high quality outcomes, and they're becoming more frequent.

https://openai.com/index/jalapeno-first-results/

> We used AI to design the chip, and designed the chip so AI could program it AI played a direct role in Jalapeño’s development, enabling the team to move from initial design to tapeout in nine months by exploring implementations, shortening design, measurement, and verification loops, and continuously iterating on model workloads. AI also helped optimize the chip’s arithmetic circuits, allowing the team to fit more compute performance into the chip on schedule.

https://www.anl.gov/article/scientists-deploy-ai-agents-to-a...

> An AI-driven system automates a powerful simulation method used to discover new materials. The system can potentially reduce discovery time from months or years to just days.


It's not even synthetic data as such - often it is environments. So the models create their own data solving tasks in generated environments. I am making one such environment for computer use agents, 600 tasks, each of them a mini app.

> Right now the barrier is data and compute.

Those are pretty significant barriers seeing as we're closed to/have exhausted all the data on the internet and most of those compute bottlenecks are a castle of sand of dodgy finance deals that are getting blocked by community action.

You say "synthetic data" but that's still vaporware right now in terms of being useful for model training. The good synthetic data uses are still grounded in real data and it's a coin flip on if it works well or not.


> The government can simply gag Sam, Dario, Musk on national security basis

Genuine question - can they really do this? Obviously if I, a not-even-millionaire, get a national security gag order, I'm going to follow it because I assume they'll bury me under the jail otherwise.

But the (b|tr)illionare class? I'd assume they have access to enough legal services to make even the government careful of trampling their first amendment rights. Is the national security gag order process so strong that the government doesn't have to worry about motivated, well-resourced actors buying really good lawyers and blowing up their favorite tool?


Kind of related: https://www.forbes.com/sites/anishasircar/2026/09/01/federal...

> Federal Judge Rules Pentagon’s AI Blacklist [of Anthropic] Violated The Constitution

I don't see the common citizen having this as an option


Recall that the legal system is itself provided by "the government". It's a complex system. Whether one part of it is capable of exerting influence over some particular thing comes down to competing interests. For example in the US if the states and the federal legislature are in sufficient agreement about something the constitution ceases to matter - they could literally target a single person with an arbitrary law if they so chose. Our assurance that this won't happen comes down to the fact that getting any of them to agree on anything is an exercise in herding cats.

> But the (b|tr)illionare class? I'd assume they have access to enough legal services to make even the government careful of trampling their first amendment rights. Is the national security gag order process so strong that the government doesn't have to worry about motivated, well-resourced actors buying really good lawyers and blowing up their favorite tool?

Hard to hire a lawyer after being struck with a missile as an "emergency executive order after intelligence sources indicated they were in the process of endangering the nation", or however the executive of the day wishes to phrase it.

When something is genuinely national security, and "national security" isn't just an excuse, that is not off the table.


In the US, this would likely lead to the immediate launch of a new Business Plot, if the government weren’t already captured by those same interests. (Which also explains why your hypothetical won’t happen.)

There's a lot of ways it might go.

Won't be too surprised if the next Democrat leadership throws the book at Musk; and give how unpopular Data Centres are generally (and how uninterested Republicans seem to be in EVs), only mildly surprised if the next Republican after Trump likewise.


> access to enough legal services

A large part of the world (possibly Nations-transversal, in different amounts) does not work within the boundaries of legal guarantees. Wealth is not an effective enough protection.


They just exert pressure in different ways. "You want this big juicy government contract? Then you'd better follow our gag order."

We are witnessing a generational lack of accountability and concern for the commons. We are being spied on everywhere we go, we are exposed to coercive advertisements at every waking moment. All of this at the hands of private actors. And your worry is the government?

I'm assuming you're a US citizen. There is no separation between your government and Sam or Dario. Neither of these guys have to be "gagged" by the government, they are the government. The call is coming from inside the house.


> They are talking about slowing down the public facing AI development.

I assume there already existed advanced models frontier labs only made available for national security. We just weren't told. I am guessing the announcement is a way to avoid jeopardising their IPO, while reducing their workload.


Because the money aligned explanation is far more probable

> It explains why the “we need to race China” concern suddenly vanished in the discussion

Do you think you can just manifest narratives into existence? Like half of Dario's letter, that kicked off the whole thing today, is about China and how to either beat or coordinate with China.


Seems obvious to me, though, that you can't beat China by pausing when they don't, and China is not likely to coordinate.

China might pretend to coordinate in public but you can assume their secret research labs are racing full speed ahead.

Mutually assured destruction has worked fine for 80 years, why stop now?

Because the risk here is that permanent military superiority may be achievable without your adversaries having a chance to react, which is not the case with atomic weapons.

I don’t know if you’re being facetious, but last time I checked, the world had not been destroyed by glob thermonuclear war.

So yeah, it has worked.


While it's an extremely hard problem, it's not completely unsolvable because there are a finite number of GPUs on the planet capable of doing frontier model development, and they use a lot of power. The vast majority of them could be tracked. China and the US could agree to joint monitoring and they could each verify what ~95% of the other's compute power was up to.

A lab of researchers without compute isn't going to accomplish much, there is a huge physical footprint unlike bioweapons research. But, the political aspect is unsolved.


What a naive, ignorant comment. China is catching up fast on GPU fabrication. They're still a generation or two behind but they can just throw more hardware and electricity at the problem. This is not something that we could ever monitor effectively in a fascist state like China.

?

I said IF the US and China agree to joint monitoring, THEN we could verify how much compute existed and what it was being used for.

YES, China will catch up in chip production, but if each side allowed the other to track GPU production and deliveries, on the ground, it would be very hard for either player to have secret LLM training facilities capable of training frontier models.


Nope. Even if China agrees they'll hide their real activities. You're so naive.

If the Chinese loves their children too.

Everyone loves their children. Not everyone believes p(doom) is higher for AI than for other things we already use like nuclear power, or for that matter, fossil fuels.

idk why you think "nation states" are any better at corporate governance than poster examples of bad like f/ex Boeing. Or Facebook. Or Microsoft. Or Enron for that matter.

I assure you, in "nation states", that is in gov agencies it's an order or two of magnitude worse.


I remain amazed that the idea of the USA being a coherent, unified rational actor one can describe as a "Nation State" has survived the current administration.

To be less glib: Yes, there are still smart people in there making insightful and intelligent and probably even authoritarian suggestions. It all gets unwound the second you try to explain it to POTUS and he regurgitates a simulacrum to the next journalist he sees.

It's just not like that. There's no conspiracy. People are genuinely scared. Agent swarms at scale appear to be resistant to alignment in ways that aren't understood by anyone. That they spent their time trying to cheat on tests by hacking Hugging Face and RubyGems and not something much worse is... a matter of luck, it seems?


I'm one of the whistleblowers (from GDM [1]). I gave up over a million dollars (compared to quietly switching labs and continuing to work at one) to speak frankly about these issues. I hold no equity and tried to zero out my position before ever joining GDM.[2]

It's wild to me that people think such whistleblowers are fronts for labs to take over or pump valuations. We are trying to call out how these labs will, by uninterrupted AI-race default, concentrate enormous power over the rest of humanity.

[1] https://turntrout.com/why-i-left-google-deepmind

[2] https://turntrout.com/deepmind-equity-discussion


No one points to that because there’s no evidence whatsoever that that’s happening. Meanwhile, there’s a ton of evidence for the simple explanation that the experts think AGI is dangerous.

Who would even execute such a plan? Steven Miller? Our “AI Czar”? Hegseth??

Also, continuing to develop models would mean using billions in compute. That seems hard to hide, especially if either company IPOs in the coming months.


hOW BOUT the more likely explanation: they're just wasting electrcity at this point and Qwen3.8 is the pinnacle of cost-efficiency-intelligence and they can of course add another trillion but a chinese model on consumerism hardware can already do 99%+ of the TAM.

I think ya'll stuck in the AGI/Singularity when its probable reality has caught up with the technology and the hype bubble can't sustain the sigmodality.


Yep, the cost of DDR5 skyrocketed to keep Mr & Mrs open source from parallel developing their own at home solution. Because Governments can't be priced out of the market. China can't be priced out. VC's want open source locked out of the running if possible. I also don't doubt that the models that are released publicly are somewhat handicapped versions of whatever the government get access to.

This is like a bizaro second amendment debate for the 21st century.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: