Hacker Newsnew | past | comments | ask | show | jobs | submit | credit_guy's commentslogin

I know it's unpopular, or unfashionable, but I agree with this letter.

LLMs are becoming so powerful that they are dangerous. We've seen last week with the OpenAI hacking (by mistake) Hugging Face debacle.

It is absolutely ok to have open weight models at the level of GPT-OSS-100B. That one was released one year ago, and I think it's still a strong one. GLM 5.2 is a whole new level, but it appears to still be safe. Maybe Kimi K3 will be ok too. But beyond that, things will start being dicey.

It's easy to dismiss this and claim that Dario Amodei is just looking to fatten his pockets. And, sure, if Anthropic manages to put the brakes on open weight models, that reduces the competitive pressure it feels. But that does not make what Amodei's argument incorrect.


> We've seen last week with the OpenAI hacking (by mistake) Hugging Face debacle.

If the biggest danger of LLMs is that they can hack traditional systems, there is no significant threat to humanity posed by releasing them in open-weight form. Security doesn't become less of a problem by making hacking even more criminal. That's what's an unsafe mindset looks like.


This is the "baby's first AI risk" tier of AI danger. There is no known upper limit to how powerful those systems get, and there might not be one.

Don't think "a smart guy". Think "project Manhattan and CIA put together, all in one server rack".

We're lucky to have "they can hack traditional systems" as an early warning shot. Clearly, it's wasted on many.


>If the biggest danger of LLMs is that they can hack traditional systems

This is doing a lot of lifting. If the biggest danger of LLMs is they could uplift bioweapon development, the situation is different. If the biggest danger of LLMs is they reach capabilities allowing for recursive self improvement, the situation is very different still.


> they could uplift bioweapon development

We've had LLMs for 5+ years, as well as Alphafold for 7+ years now. No novel bioweapon has been made with the technology that we're aware of. It's a farsical claim, there's no 21st century Aum Shinrikyo abusing the technology, after years of proliferation.

> they reach capabilities allowing for recursive self improvement

Again, you are predicating your entire argument on a hypothetical emergent behavior that we do not have any evidence for. I'm not worried about this whatsoever.


Ok. What's your stance on gun ownership then? And by gun, I mean naval guns.

My position is that anyone can own a cannon and shells, but you only get to fire it once before the feds step in.

Giving a naval cannon to the average person does not threaten humanity any more than giving them a gun or an LLM does. None of them are a panacea for anything.


It seems completely unhinged for civilians to have naval guns. You could level an entire neighborhood with one of those.

It's also completely unhinged to own a Shahed 136 or a ballistic missile, but both are fairly attainable and even legal within certain usages.

I don’t know what you mean man people obviously don’t own personal tomahawk missiles and that sort of thing. If you have a Shahed loaded with explosives on your property you will probably have to deal with LE

A mass shooter wakes up in the morning and thinks: "I'll fire my naval gun once, then politely hand it to the feds..."

I imagine that he'd be pulled over by the cops towing his artillery into battle, first.

On what grounds? You just said everyone should be allowed to own artillery.

but what is the point? A ban is supposed to make a certain thing less likely to occur. Does a ban of open source models do that? Presumably, the behavior you are trying to limit is the miss-use of these models but I don't know how many state sponsored hacking groups are going to give a ban a second thought.

Yes, that's how I read Amodei's post.

Imagine Kimi K4 will be as powerful as Mythos. Anthropic can work for months and months to set up guardrails on Mythos, so when the model is finally released, it will generally decline to help hackers develop and prosecute cyberattacks, and if they do, at least there would be a trace so the law enforcement can track the perpetrators. Let's now say that Kimi K4 is released after a similar effort to develop guardrails. But being open weights, someone can just take the model, and finetune it until it does not refuse to assist in developing cyberattacks, and moreover, those people can run the model on their own private GPU cluster, so nobody can track the attack back to them. The situation is actually worse than that, most likely. Guardrails might be just markdown documents which are added to the context like regular skills. Then removing the guardrails for an open weights model does not even involve any finetuning, just removing some docs from a harness.


I think the way to parse the current title "DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]" is that there was a leak that DeepSeek will pause fundraising because they perceive there is a compute gap with the US.

I am also guessing that the majority of the people who read this title will think that DeepSeek is pausing this fundraising because some comments they made about the compute gap were leaked. That is not the case.


Maybe: "Leaked Deepseek transcripts reveal plan to pause fundraising due to compute gap"

I don't know what "compute gap" means in this context though and it's not clear that that's why they plan to pause fundraising or if the title is conflating.


yep. the word they use is probably 克制 or self-restraint. no need to raise so much cash if you can't use it.

in his article he talks about the negative aspects of getting everything you want. (all the money, brightest minds, biggest share in AI) etc. he says that these are the things that will cause a company to fail.


I shouldn't win too hard, because then I'll lose?

Yes. Once you're in such a comfortable position that you'll keep making lots of cash regardless of whether you do well or not, there is little incentive to make good decisions, let alone take risks. This is known as the "curse of Oil" and is also what Intel's decline is attributed to.

Catching up without overtaking

Sudden availability of capital and the perceived need to be seen doing something with it can be a curse. WeWork comes to mind, eg. Stay lean and mean until you actually need the capital. A company like DeepSeek will have zero issues raising anytime.

Yep. Classic “giant series A, company gets amazing office space” vibes.

> I shouldn't win too hard, because then I'll lose?

Yes. It's a known thing.


Problem of the local maxim, extreme dependency is the same as evolutionary pressure to specialize for it, eventually the process owns you.

Interesting, could you elaborate please? What does dependency mean in this context? That you become over specialized for a niche?

“Raising cash” isn’t necessarily “winning” since you have to give up equity/control, usually.

Raising more money ≠ winning harder.

Unless they have a shortage of mathematicians, physicists, ... it seems a fundraiser could help bypass a compute gap by focusing even more on inference and training efficiency.

is this similar to Meta starting to rent out own compute as they cant seem to do much with it and monetising it is much better ROI?...

Sorry to barge in here. I couldn't find a good place to place

https://github.com/demo-zexuan/liang-wenfeng-investor-meetin...

The current link is 404, can mods update to above, detach, make sticky?

(No response from mods, understandable)


I'm not sure I understand the question but I've replaced the top link (https://github.com/demo-zexuan/liang-wenfeng-investor-meetin...), which was 404ing, with the link in your comment here. Does that help?

Yes, thank you!

It's now dropped off the FrontPage but I suppose I will just repost at some point with a less controversial title


email hn@ycombinator.com

The maturity of leadership in China seems to be on a whole different level from the U.S.

Lots of companies in the U.S. have fallen victim to that syndrome, but if you used the word "restraint" in that context in Silicon Valley most people would look at you like you're insane.


Yeah, wouldn't it makes sense to increase fundraising, so as to acquire more compute to close the gap?

I guess they likely can't right now due to no hardware available, either cause of bans or already all booked.

I gather this is the whole point of the original article. It seems to have been pulled, so I can't confirm.

At what point do the AI companies start building their own hardware, I wonder.


Other companies are investing and hardware design and manufacturing, so they won't duplicate the effort given the national drive it would be a waste.

Perhaps it's akin to a mineshaft gap?

> if the title is conflating

The title is certainly a great conflation. Any seeker of capital would want to regroup after an unfiltered leak of this magnitude, if for no other reason than to secure the forum from future leaks. The comments about the unlikelihood of enormous future profits were at least as consequential with regard to capital investment as anything else that was said.


I skimmed the doc and my impression is that your second listed interpretation -- DeepSeek is pausing investment because of a leak -- is the more correct one.

There's quite a bit of confidential information in the doc about the company and how it's positioning itself going forward to compete with US labs. I'd imagine they're not happy at all with this being leaked and are withholding investment as a punitive measure.

Not to mention the other interpretation seems illogical -- why would you pause fundraising if your perception was that you lacked resources compared to your competitors?


Because money isn’t free and if you know there is a chokehold in supply, why raise money ar current valuation when things are getting better by the day?

Maybe because if you need more money to catch up with your competitors than anticipated then your ROIC is lower and your valuation changes?

All the Chinese reporting I see point to the second (majority) interpretation. Liang being furious about his private investor talk leaked online is the news here.

e.g. https://x.com/_FORAB/status/2081034500101017616?s=20


Those could be subsequent developments, but that's not what the linked transcript was about. The transcript was a discussion of the DeepSeek founder (Liang Wenfeng) with investors, and he does not mention any leaks, or any frustration. He simply says that he is constrained by the supply of cards, and he has no problem of getting funding, but has no reason to raise further funding because he can't transform the cash into cards.

  > There is certainly no shortage of funds or resources --- in fact, all these are readily available [...]

  > Within our financial capacity, it's undoubtedly true that the more cards are always better. Our current strategy is to purchase as many cards as possible at a reasonable price --- exactly how many we can afford after using this funding round. The spending pace isn't predetermined; we'll buy whatever is available as long as prices remain competitive. In fact, I'd consider that a positive outcome if we spend the entire amount within six months. [...]

  > In reality, spending such a large sum is no easy task: you can't obtain enough cards, they're hard to come by [...]

  > Therefore, our only concern is whether we can obtain enough cards. If converting all funds into cards were feasible, we would undoubtedly do so without hesitation and are even willing to pay a premium for this benefit --- it's simply to cost effective. Even after paying the premium, however, achieving this goal remains challenging.

he's pausing because of the leak. the transcript content has nothing to do with it.

Thank you! That is indeed how I read it.

I wanted to post this which explains the wording but I thought the transcript was more interesting. Sorry. Maybe mods can help me to put what follows as auxiliary link. I don't know how.

https://www.bloomberg.com/news/articles/2026-07-25/deepseek-...

Update:

Less-paywalled word-for-word copy it seems at

https://fortune.com/2026/07/25/deepseek-liang-wenfeng-backer...

https://archive.ph/zpIrG


Most of that is paywalled, but this one paragraph in the Bloomberg article suggests it might be more to do with investors leaking information:

"The suspension stemmed in part from Liang’s frustration over online reports about his comments to investors during his first financing deal"

The part of the transcript I'd seen floating around online was this part from around 1 hour 26 min:

"With the largest models available today, we simply cannot afford to train them. Even if we spent all five hundred billion yuan, we still wouldn't be able to do so. Even if we could accumulate the resources, we wouldn't have the means to utilize them. The current largest model requires approximately 800 billion activations; domestically, we are still at a scale of several dozen billion activations, and even the largest domestic model may only require several dozen billion activations—a difference of an order of magnitude. To train a model of the same size as an AI system, we would need around 50,000 GB300 GPUs or Huawei 950 GPUs, totaling two hundred thousand cards. This is merely training; research has not yet been considered. Therefore, the biggest gap between us and the United States lies in resources."


amusingly ive been working on ultra sparse llm inference/ training/ model design because nature loaths a dense graph/matrix and cause i think it shoukd be possible. i actually stood up a 20-25 percent faster than sota causal fast attention kernel yesterday, will be standing up cuda/metal/armv8 kernels too and thats gonna be fun.

i genuinely think these models should be like 0.1 percent sparse for same capabilities we associate with them today, but theres no sane way to do that with extent tools. i built the right core tech for that in 2014 when there wasnt a market, but now there is and the experimentation velocity is wild.

amusingly llms really have a hard time using my simple apis because its not in distribution array programs. but i literally stood up cpu custom memory format and micro kernel for dense causal attention in less than 24-36 hours and outperforms the equivalent fused ggml/llama cpp fast oath by like 20-25 percent


I am not much of a math person, but if we look at the how the brain is wired, we see that the dendrites (the inputs) of a neuron are hundreds of micrometers in length, and the axons (the outputs) are millimeters, and very rarely can stretch to tens of centimeters. So they can sample only a tiny amount of internal state, and affect a much larger, but usually still small output.

In math terms, this means a layer of a network can be represented with a block matrix in the whole 'layer' matrix, which I think means its sparse as you said.

As I said, my math knowledge is rusty, but I remember that a lot of matrix optimization techniques center around decomposing large matrices into these smaller blocks, which are then evaluated, and the output is combined in a final pass. Which leads to a huge reduction on parameter numbers and the time it takes to evaluate the result


exactly. you certainly know more about the brain than i :)

Look forwarding your future releases

i definitely will be doing some drop of some faster attention kernels in the next few weeks.

like i can do all sorts of memory layout of tensors/matrices etc tricks that if you dont have the abstractions for it would just never happen. so i can optimize the kernel flops


Curiosity:

For most of the past five years, I've known ways to do better than Anthropic, OpenAI, and friends in many ways, at least on paper. I know I was right about many of them since many would show up 6-24 months later tools from the major providers, or otherwise become standard practice.

A central problem is the Mythical Man-Month. True, I could do those, beating then-state-of-the-art, but only given 2-5 years. I suspect many other people knew about them too and could do so as well. As I noted above, throwing people and dollars caused many of those to be built in less time than I could have regardless.

Other methods, I'm less confident about (>50%, <80%), but would lead to similar improvements orders-of-magnitude as you're predicting, but mine would need $$$$$ in compute and engineering infrastructure to build out. E.g. they need to not just theoretically work, but to try, I would need to convince someone to invest in them working.

So the TL;DR is that my knowledge was not at all helpful towards e.g. competing with OpenAI, Anthropic, or even building a small business.

However, where it was useful was in predicting where the industry was going. This is true in investing (but not easily, at least with my skill set), but in developing startups and systems, there were capabilities which I (correctly) assumed would be there, whereas there were many arguments that "AI will never be able to ____."

If I know how to do something, it will almost certainly happen, regardless of whether I'm the one who does it.

To be clear, my expertise is almost certainly nowhere as deep as yours. I'm not providing a direct analogy, or claiming others know what you do or can do the same. My point was really that if you believe you can have these models be 0.1 percent sparse for same capabilities we associate with them today:

a) You're probably right. They were built quickly for capabilities. A slower process can almost certainly lead to much smaller models too. That's a radical statement: Historically people claiming a 1000x improvement somewhere were crackpots, but that's very possible in an industry as fast-changing as this one.

b) Someone at Anthropic or OpenAI might be working on building out extent tools right now. Even if so, there are indirect ways to capitalize on that knowledge.

c) Critically, that predicts a future where Fable is $1/month instead of $100/month, and that's something which CAN be acted upon in planning.

It also suggests -- much less strongly -- the existence of much more sophisticated models at $100/month. There are open discussion in planning about whether models plateau, continue improving, singularity, or otherwise. That changes the biases there.


>Fable is $1/month instead of $100/month

will Anthropic (or OpenAI) lowers their price, or increase margin (to justify valuation)


thx for the kind response!

at the very least i have tools that let me easily hit better perf for fancy dense memory layouts, and the same tooling lets me experiment with frankly wildly wacky sparse and structured memory formats. the performance claims at least on the dense side are solid so far!

the sparsity angle is because i want magic in the world. like anyone with a really chunky computer like any of those mac mini pros or serious workstation / server tier compute should be able to train from scratch their one 31b equivalent model in a week or so tops is the goal post i have in mind


If you see the long 3rd paragraph twice, know that you are not alone.

My guess is they used the other Anthropic models extensively for synthetic data generation. The top most similar models for K3 are (lower means more similar)

Fable 5 -> 0.42

Opus 4.8 -> 0.45

Sonnet 5 -> 0.45

Opus 4.7 -> 0.46

Grok 4.3 -> 0.52

There's an obvious jump at Grok 4.3, and it would not surprise me that the similarity there is because Grok used Anthropic models for training too (you can get the similarity list for Grok and it does look like the top most similar models for Grok are either Anthropic models or some Chinese models).

The damning evidence that K3 used Anthropic models for training is that K3 is more similar to those models than it is to K2.6. If you look at the Anthropic, OpenAI or Google models, they are most similar with their own other models. Not so with K3, where K2.6 is less similar than 15 other models.

Now, why is Fable 5 the most similar to K3 and not Opus 4.8. I think it's quite likely that K3 did some fine tuning at the end, when Fable 5 became available. They probably had all the infrastructure in place, and Mythos had been announced for months, so they were probably waiting for the second the newest Anthropic model was released to start using it for synthetic data generation.


Note that until last generation all Chinese models were relatively small. This was mainly due to lack of training hardware, but as soon as the new Huawei NPUs started shipping some Chinese labs switched to larger models:

Deepseek: 670B to 1.6T

Moonshot: 1.1T to 2.8T

Xiaomi: 310B to 1T

The new models also use a different architecture so I would assume in tests they will look different from previous generations.

Z.ai (GLM) and MiniMax on the other hand have continued training existing models (with some surgical changes to improve long context memory) so they should score similarly in these tests.


I don't understand this point of view. They buy the book, they use it for training. When you read a book, you use it to train your brain. Is that copyright infringement?

Downloading from Pirate Bay is indeed theft, someone put the scan of the book online, you download it without paying, the author and publisher of the book do not get any royalties. Very different situation.


You could always just buy a used book or get it from the library, the publisher and writer don’t get paid but once and it switched hand multiple times.

When is this not plagiarism? As AI will just briefly paraphrase so it’s not exact word for word. They rarely give credit and only late add some sources but only the top sites despite gobbling down the entire web.

Wikipedia editors do the same, good luck on getting your site credited and a link on there for more than a day/week unless you have connections.


People who claim that the Chinese open weight models have some type of manifest advantage don't realize that the close weight models have a huge advantage as well: the researchers from OpenAI, Anthropic, Google, xAI, Meta are not dumb, they can read the white papers written by DeepSeek, Moonshot, etc, and they can inspect all those architectures and they can pick and choose the best tricks there are out there, and of course, they have access to their own in-house secret sauces.

Sure, any model that is not at the frontier can use the frontier model to generate synthetic high quality training data, so this can reduce significantly the training costs.

But at the scale of OpenAI, Anthropic and Google, it is quite likely that the (raw) training cost is very high anymore. Here's a few heuristics:

1. All the hyperscalers see a huge demand for inference. They can't deploy datacenters quickly enough to satiate all the demand they see. But, it's is impossible for the inference demand to be constant throughout a day or a week. If you use the times when the demand is lower than the peak demand (which is almost all the time) to dedicate the spare compute capacity to training, then your the cost of training compute is zero.

2. It is likely that increasingly a higher cost of the "training" is actually setting the guardrails, which is essentially post-training. As we've seen, without proper guardrails, the US Government won't allow you to serve inference. Anthropic was hit directly, but OpenAI delayed their 5.6 release as well to make sure the US Government is ok. This part of the training cost can't be reduced easily by using synthetic data generated by other models.

3. The frontier labs are also investing more and more in building an ecosystem around their models.

I am not a frontier lab insider, but take a look at the jobs posted on the Anthropic career page [1]. There are 74 jobs in "AI Research and Engineering" and by my count at most 15-20 are related to pure model training (of pre-training or RL type), and the rest are post-training, safety and security, alignment, interpretability, productivity and lots and lots of other things.

[1] https://www.anthropic.com/careers/jobs


People who claim that Postgres has some type of manifest advantage don't realize that Oracle has a huge advantage as well…etc

If the Google and meta engineers are not dumb how come they consistently trail behind the frontier labs and even the Chinese labs with a fraction of the funding.

Probably bad leadership


I always suspect they have the most to lose if legal decisions on copyright issues don't go their way.

Imagine a scenario (theoretically possible but increasingly unlikely) where a US court decides that using "pirated" copyright data to train models is illegal. Now the AI developer has invested hundreds of billions of capital into a thing that is declared illegal and has to be scrapped.

This risk affects existing megacorps more than "startups" like OpenAI and Anthropic (and Chinese companies), because the megacorps have much more to lose. They actually have the cash to pay damages if the flood of copyright claims arrive at the door. This will not only bomb their AI development, but also the rest of their established businesses as well.

And thus I strongly suspect legal issues are holding them back a bit. Megacorps want to win the AI race, but not to the extent they stake the rest of their established business, while the newer companies' only product is AI, so they have to go all in.

Notice for example how Meta's Llama performed much more poorly after they got smacked by a bunch of lawsuits claiming that they torrented a bunch of copyright data.

(Disclaimer: I'm an outsider and everything I base my speculations on is public knowledge.)


That plus they don’t distill so they have worse RL examples.

But they both spent tons of money on data collecting/labeling/generation, how is it bad compared to distillation? I thought their data are much better if they spent that much, and it seems they are stupid because with that much of resources putting in there with merely no output compared to the frontier models.

Creating a RL example by hand is hundreds of times more expensive than generating one using an LLM.

Of course the Chinese companies have incredibly talented researchers, and smaller, better organized org structures which account for the rest of the difference.


they, rightfully so, have no faith in LLMs.

>they have access to their own in-house secret sauces.

I remember some feature lauded by Gemini was reverse engineered by the open weights guys in < 30 days.

If they dont publish some technical information its hard to protect in the US, but conversely, once it is published smart people from outside the copyrightosphere can start working to reverse engineer it.

>3. The frontier labs are also investing more and more in building an ecosystem around their models.

Theres nothing there that isnt immediately replaceable.


> There's nothing there that isn't immediately replaceable.

Indeed. But that was not my point. My point is that we still have this old impression that training cost is dominated by compute and it is hugely expensive, and the Chinese labs can short circuit that by distilling the American frontier models. I don't think the training compute cost is a big factor anymore for the American frontier models, because of the reasons I gave. If the Chinese models can get the training compute cost down by a factor of 100, that's not going to make them 100 times cheaper, and not even cheaper by a factor of 2. Maybe 10% cheaper or so.


It’s giving desperate!

I think it's the opposite.

Kimi K3 has 2.8 trillion parameters. We don't know the number of parameters of ChatGPT 5.6 or Opus 4.8, but it's probably in the same region. Fable/Mythos are rumored to be around 10 trillion.

So, K3 is directly comparable with ChatGPT 5.6 and Opus 4.8, and the price is not so much lower:

K3: $3/$15 per 1 Mtok input/output ChatGPT 5.6 Sol: $5/$30 Opus 4.8: $5/$25

This is not a watershed moment. It's a competitor converging to the same capability and trying to undercut your prices, but not by a lot.

As for the open weights? For now, Kimi K3's weights are closed, and I don't expect the situation would change.


> As for the open weights? For now, Kimi K3's weights are closed, and I don't expect the situation would change.

It'll change on July 27 (based on https://www.kimi.com/blog/kimi-k3):

> The full model weights will be released by July 27, 2026


July 27th. But I agree with you that this is just normal competition. The only threat this poses is to Anthropic. OpenAI is more than capable enough to out-compete, their pricing is already reasonable. Greedy Anthropic will do their very best to try and stop this though, because they want to maintain the status quo of ripping everyone off.

And how much token and time you use to solve a problem? The price alone doesn't mean anything.

Example DeepSeek-V4-Pro (high) needs 10 times more token then GPT 5.5 (medium) and can compete only with the price.

The real price saver are the cache prices, the ting, that nearly nobody has on their radar.


Given how OpenAI got rid of their 5-hour limits and reset weekly limits so often, is Kimi really undercutting them on effective price?

The 5 hour limits are coming back soon right? I thought that was temporary

Maybe, but Tibo said something on Twitter last week making it sound like it might not be.

I’d also note that running a 2.8 trillion parameter model at scale efficiently is not simple. I would expect when open weights land getting it running fast, efficient, and at full capability will require sufficient resources it’ll be expensive outside of Chinese hosting. Which I think almost no western corporation would use for any internal work. You have to anticipate your use won’t just go towards training but will be actively mined for IP, trade secrets, MNPI, etc, or anything of use to the Chinese government or Chinese companies. I don’t say this to crap on the Chinese - but this is the playbook for the last 30 years.

That said I fully intend to use deepseek hosting for operational agents that are making decisions about non sensitive material. The economics are astounding.

Kimi? The economics aren’t that amazing to merit switching from 5.6. I expect fable will rapidly reappear in subscriptions. Competition is good.


Afaik, for a MoE model, total size doesn't matter that much for how heavy it is to run, the size and number of active experts at any given time does. Of course, you still have to store the whole model in fast memory, so there's that, but the reason these things are getting this large is because it doesn't really affect runtime that much.

The APIs for the frontier models via the US hosters do the exact same thing wrt saving the requests and responses for data mining. Let’s not pretend that pervasive surveillance is an eastern thing.

Come now. The purposes for which the data is used is relevant. I am much more concerned with my internal corporate IP being actively used against me than passively used to train the model. I would also note you and sign agreements that prohibit the collection of data for use as well, which is also one of the key selling points of bedrock. In the west you can actually enforce such an agreement in court and win.

I wouldn’t lump this into a west vs east thing as well. This is particularly PRC. I feel comfortable doing business in Japan, Korea, Singapore, Thailand, Malaysia, etc. But it requires some particularly strong willfulness to pretend the PRC isn’t actively and structurally built around economic espionage, and funneling IP through PRC for short term economic gain has been one of the primary factors in their growth over the last 30 years. This just scales it faster.

I wouldn’t expect the USG won’t compel AI companies in the US to disclose and retain data as well - however it’s not a simple thing, the companies are hostile to it themselves, courts are often unsympathetic to the government, and the “machine” for converting it into actionable economic advantage is non existent - and there’s a very significant human component in that all links in the chain are culturally uncomfortable with such things. While it happens and it’s possible it’s very difficult, fraught, and does not scale. The PRC is the opposite - the courts, government, and business culture are all aligned in the goals and processes.


I think it's safe to assume the Chinese models will try to steal your corporate IP and business know-how.

It's also safe to assume US models will do that too.


> Shrinking that into a 500 kg seed — or even Freitas’ original 100-ton seed — is not an engineering detail. It may be the entire problem.

How many AI tells can you count there?

But honestly (see what I did there?) the AI slop is reasonably cleaned up in this piece.

However, the essence of the argument has two deep flaws. One is that the time to complete an interstellar voyage is extremely long and you need some exergy, yada, yada, yada. We could start with sending self-replicating probes to the asteroid belt. There is zero chance that we'll attempt to send self-replicating probes to a different star system before we send them inside our own solar system. And the second error is this:

> Bootstrapping this loop [...] is a chicken-and-egg problem that no study I am aware of has worked through at the level of actual process flowsheets.

The fact that the current technology is not adequate, and nobody even attempted to solve such a problem is a weak argument. Three hundred years ago nobody had "worked through the process flowsheets" of making an injection molding machine, or a 3D printer, or a power drill, yet they are all available now.


The subheadings are all full of AI tells.

> The closure problem, honestly accounted

> A thermodynamic framing

All of those read like AI, especially considering that the subsections aren't consistent. Some are numbers, some are not, some are framings some are problem juxtapositions.

EM dashes everywhere, AI tells in subheadings, "It's not X it's Y" all over the text of the body. This is clearly AI writen.

Notice also the article has two by lines. At the top it's "by Paul Gilster" at the top of the text it's "by Peter Marinko"

Also note that the "metallugrist" they're interviewing that they claim "his current work explores the thermodynamics of technological civilizations" at Uppsala University, but the university's page for him says he's only involved in Animal Research Ethics Committee


> Notice also the article has two by lines. At the top it's "by Paul Gilster" at the top of the text it's "by Peter Marinko".

It is Paul Gilster's blog. It looks like all articles there are by him, hence his byline at the top.

In this particular article he writes an introduction talking about self-replication, then says "Right now I want to introduce Peter Marinko, who today weighs in on self-replication and the problems therein", describes Marinko, and then the rest of the text is Marinko's, hence the second byline after Gilster's introduction.

Something similar happens in the next article on the site. Marinko write a length response to comments on the first article. Gilster decided that would be better as a separate article to further discussion: "When Peter wrote recently with his thoughts on reader reactions, I asked him for permission to run it as a regular post rather than a comment, because I think this is a lively question and would like to see us continue to explore it".

So that too is a post by Gilster, and so with his byline, but after an introduction it just run's Marinko's text, so has a second byline for that section.

He does the same thing on this article [1]. He wrote a review a paper, some commenter had interesting thoughts, and Gilster posted an article talking about that, introducing the commenter, and then the rest of the article was the commenter's text. So two bylines, one for the intro and one for the guest text.

It looks like the other articles on the site are just Gilster, and so only have one byline.

[1] https://www.centauri-dreams.org/2026/06/05/observations-on-t...


300 years ago people also believed alchemy was a serious field of study


And today we call it transmutation - https://en.wikipedia.org/wiki/Nuclear_transmutation


That is true. What should we conclude from this statement?


That you shouldn’t assume just because in the past ideas have existed that became reality, that all current ideas are feasible or possible?


I didn’t assume that. The author makes the opposite assumption: that because nobody has ever proposed a path to technological feasibility, that thing will forever be technologically unfeasible. I simply find this argument to be very weak.


It’s a bit silly to be so sensitive to “AI tells” in phrasing. If you go look for them in original human writing, you are guaranteed to find them—just like the AI training did! That’s how they became “tells” in the first place.


But LLMs go through post-training that gives them a style distinctive from most human writing. It's certainly detectable.

This reads as heavily LLM generated/edited to me.


I find this line of reason to be incredibly irritating. So anything written above the level of "see Spot run" now must be AI slop? The author's piece is written well and reads easily. As a long-time user of nonrestrictive elements in sentences, I bristle at the idea that only AI is capable of writing sentences containing brief asides -- the things between the em dashes -- now.


The Hohmann transfer time between Saturn and Mars/Earth is around 16 years. So, all the ships used for such a supply mission would reach their destination at least 32 years after they’ve been built, assuming we build them on Earth.


That's not really a problem if the supply is continuous. Think of the last time you drank a 12-year scotch, that distillery had to be set up and start producing at least 12 years ago for them to label the product that way but they've continued production constantly since then which ensures there is a steady supply to be delivered to stores.


Spirits are not capital intensive. Building a rocket is very capital intensive. You'd like to be able to reuse it. But if one single mission takes 30 years, then you can reuse this rocket once, at most twice. Let's say you reuse it twice, you amortize the capital cost over a period of 90 years. Now, let's say someone builds the exact same rockets, but they do missions between the asteroid belt and Mars. Each trip takes about 2 years. Everything you can source on Titan, you can find an asteroid to source it from. By the time you reuse a Titan-bound rocket once, you reuse an asteroid-bound rocket many times.


I'm not sure what you're talking about, my initial comment was a (sort of silly) idea about building a mass driver on Titan that uses native materials to launch payloads on an inward bound trajectory, not to ferry them back and forth with a rocket.


That's an interesting idea, but I don't think you can use a mass driver to shoot something from an higher orbit to a lower orbit.


True, it would need to accelerate somehow. Maybe there could be some kind of on board propulsion that would be enough once it's in space. Solid rocket motors could be fabricated as part of the process but the timing and control systems would pose a problem. Maybe something clever could be done mechanically that leveraged temperature or pressure changes.


Sure you can — you just fire it went pointing ‘backwards’. But either way, you still need a burn to circularise the orbit when you get there.


> A given model has a shelf-life, which these days is measured in months, not years.

Not all new models are trained from scratch. ChatGPT 5.3 to 5.4 (and likely 5.5) was basically the same model, but probably trained a bit more, not a new model from scratch.

> The "someday" when frontier model providers can enjoy their current high inference margins without the burden of significant training costs is never going to arrive.

That is debatable. I believe the moat for the frontier model providers is the compute. At the level of 10 trillion parameters (that Fable/Mythos are rumored to have), you need serious compute to serve inference, and you also need serious compute to train. Will DeepSeek, Qwen, Kimi, GLM come up with a 10T new model anytime soon? I doubt that. People keep saying that the Chinese labs are catching up to the US big 3, and measured in months the gap is now only 4-6 months. I doubt a Chinese version of Fable/Mythos will be released in the next 12 months.


>ChatGPT 5.3 to 5.4 (and likely 5.5) was basically the same model, but probably trained a bit more, not a new model from scratch.

Then those models have an eroding moat and will be quickly driven down to commodity pricing. The only thing propping up inference margins are the cap-ex costs of training. That's the moat. That's why there's no way to win this game. You cannot have low training/infrastructure cost and high margins (such as would justify today's valuations).

>I doubt a Chinese version of Fable/Mythos will be released in the next 12 months.

I would take that bet.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: