Hacker Newsnew | past | comments | ask | show | jobs | submit | cootsnuck's commentslogin

Yea I don't get "load bearing" that much but "seams"... So sick of it.

If we fast forward say 5 years, I don't see how we don't end up in world where people (and enterprises) are more savvy with how they use LLMs. Meaning, more models, smaller models, weirder models, more specialized models, etc. And all of it running on a variety of hardware (edge devices, personal computers, on-demand cloud compute).

I don't see how NVIDIA can keep their spot as belle of the ball. If LLMs and friends are truly to become as useful and ubiquitous as everyone thinks they will, then commoditization is the only option.


We need to figure out what the real pricing is for a going concern. Right now, everyone is subsidizing and discounting to grow (or maintain) market share. The big question is whether the steady state, market derived inference pricing is above or below what we’re seeing today. I honestly don’t know. Anthropic had said that inference is profitable, but they’re clearly not yet profitable overall with training and buildouts still happening.


Has nobody from any of the companies hosting open weights models released detailed information on how much it really costs?


I’m sure someone does, but I’ve never seen anything other than vague statements like Anthropic’s “inference is profitable” comment. I suspect everyone is playing everything close to the vest because they aren’t yet public and they want to control the information flow to the street.


> I don't see how NVIDIA can keep their spot as belle of the ball.

FWIW, people were saying "ASICs will kill CUDA demand!" since the crypto mining boom. Then a few months later, CUDA found another niche application in LLM applications.

With the mounting demand for robotics, surveillance and autonomous weapons, I don't see how Nvidia couldn't keep their spot. They have their pick of the litter with hundreds of market segments, and unlike the rest of FAANG they're not afraid to branch out.


This is going to increasingly happen over the years to come. Big organizations will become more sophisticated with operationalizing their data, training and running LLMs will continue to be demystified and accessible, and over time we'll get more and more specialized / industry-specific models.

It's going to become another way to monetize your informational assets if you're a big older enterprise with troves of data. All you need is time to figure out how to make it useful for yourself and then eventually sell access to it however you want.

Think of all the data that big orgs have that isn't accessible to all the AI labs to suck up.


40M$ to get a marginally better model is surprising, why not just use the free weight models


Because "bigger = better" does not work for highly specialized expertise, where the knowledge is not available on the public Web. It is naive to believe it's all open out there merely because the Web is large and we have Wikipedia.

If you are a highly specialized professional intellectual property paralegal, a forensic tax investigator or a post-market pharmacovigilance analyst, you will need for-profit knowledge sources, and your answers will often require synthesizing and interpreting multiple sources.


I’ll bet against this because of bitter lesson. Generalised models with context engineering will be more reliable and cheap than training your own.


They're trying to find a moat in the AI era, for their relatively gigantic news business (and they're among the few still standing giants in news). Most of these organizations are very scared of what AI might do to them.


I'm pretty sure, given the reference to fiduciary grade standards that they are doing more to secure their moat in Legal, Accounting and related fields.

Reuters News is only slightly over 11% of their business. Legal and Accounting is closer to half.

on edit: just went and looked it up, adding in compliance offerings it is over 80% of their business.


they should unite to block all news from getting to LLMs without LLM providers paying per access


Because what these models are trained on matters.

China is fucking smart. They have technical expertise to know that by feeding the model certain data, they can shape the outcome of questions posed.

Yeah, in the end, maybe both US and Chinese models can solve Fizz Buzz, or some Erdős problem.

And they can answer our inquisitive minds when we wonder, "Why is the sky blue?".

But deep questions about what is normal in the world they can change the outcome of.

And "Why is the sky blue?" is a deeper question than we think. Because it isn't blue everywhere in the world right now. And whether it is the fault of the US automotive industry, or aggressive investment by China in their own manufacturing base matters.

Of course, both have been responsible for pollution at different times. Growing up in Michigan near Dow Chemical I know this.

But depending on how they select training data answers can be nudged one way or another.

The US is smart too, and they have people working on the same things. It's ironically easier for me to talk about how China's security services are likely to shape their models' perception of the world. Which is a damn shame, because we deserve unbiased answers so we can all help our country improve.

Or countries improve if you want to take a global shared world view.


CSAM and having the potential to randomly becoming unhinged due to strong alignment with the whims of the one dude? Surely puts enterprises at ease.


> The industry is very much moving towards one-model-does-all end to end trained

I've worked with hundreds of enterprises on voice AI and voice agent solutions. In my experience, this isn't true. Or rather I should say, the people actually paying for voice agents (i.e enterprises) are not moving towards STS solutions in a meaningful way. The composability, observability, and reliability profile of STS systems is not amenable to enterprise criteria. Not to mention costs.


exactly, we see the same thing, around 95% cases are still cascaded, even tho STS has been improving a lot


Which open source STT models have you had success with for fine tuning?


That made me feel relief until I realized it's my own local PBS station.


Any interest in connecting me with someone there so I can attempt to obtain a copy for the Internet Archive?


I would say even without rampocalypse there would still be the strong incentive to innovate at the edge and under more extreme constraints. The incentives are just even stronger now.

I'm looking forward to seeing what types of new things people create over the coming years once there is less obsession with massive unwieldy LLMs. I think the incentives are just too strong to ignore.


I think finding significant efficiency gains with LLMs and the like may lead to qualitatively better products. Looking at people's experiences to DSV4F makes me believe that even more than before too.

I don't think people are realizing that speed can allow for categorically different user experiences that are more than just "worse than frontier capabilities but faster".


Betting on innovation continuing to figure out ways to squeeze more out of less has historically been the right move. Look at Apple.

And I'd argue "hardware advances" are more proof of optimization.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: