Hacker Newsnew | past | comments | ask | show | jobs | submit | mbil's commentslogin

It's kind of antithetical to the tool's deterministic positioning, but have you considered making TERMy leverage an LLM for unseen or low-confidence queries, and then generate the config and update itself to make future similar queries deterministic?

This is such a nice idea! I could add a fallback towards LLMs, it was present but I removed it. Would you be interested to help me implement the auto-update? I must admit, the LLMs are very useful for this kind of work. I think that TERMY's design is now feasible BECAUSE OF the availability of LLMs. They make the dataset development feasible.

If you don't eat your biomass compound goo, you can't have any PlastiYum bacterial sludge



Bringing it down a dimension:

> The earth may be round, in which case we’re on the “surface” of a three-dimensional sphere. One much larger than what we’ve seen (since the ground around us is approximately flat), but nevertheless finite.

Of course maybe the 4D sphere of the universe is itself hung in a vast field of other 4D universes, that field is the surface of a 5D sphere, etc.


What year do you think it is


2026


Yes it seems like the thread is discounting that frontier providers are likely already baking new, stronger models. I agree that GLM and its ilk are quite good, but having used them I’m not convinced they’re on par with eg Opus in terms of things like tool calling. And they’re fast but less capable so I spend about the same amount of time with them, just with more hand holding. Maybe this is a harness limitation. I know on paper they seem comparable but anecdotally and qualitatively they’re not as useful as the frontiers’, so maybe there’s some truth to benchmaxing claims. For some workloads the distilled models may be good enough, and I suspect at some point there will be diminishing returns to spending a premium on frontier models, but I don’t think we’re there yet. That said I’m continuing to try them.

The question is whether this steals enough marketshare from frontier providers that they don’t have the capital to train the next model iteration. The open models are going to push down the unit price of an intelligence-token, but there will still be a market for a smarter bot. And as intelligence gets cheaper, the demand for it will rise (see Hank Green’s Jevons Paradox video). Not to mention there’s all kinds of other directions to go at the frontier (world models, robotics, video gen, etc).

Another thing, and this is pure speculation, but if the Chinese model providers already discovered the decrypting COT trick and leveraged it to do RL training, and assuming frontiers plug that hole, then maybe future distillation will be harder.


It’s not that frontier providers won’t keep on making good/leading models.

It’s whether you absolutely need the latest capabilities (at the cost of very high prices, sending your data to them, and being totally at the whim of 2 companies, that can shut you off anytime for any reason).

With how good LLMs are already, there’s tons of tasks where not being at the absolute bleeding edge doesn’t matter, especially when you add cost/freedom/supply chain risk/not leaking your data.

Even more - there’s increasing number of companies that give you ability to post train open weight model yourself, for your own use case. Given how many of the gains today are from post training, if you post train it for your specific use case, you’re very likely get model that you own, that works for you as good as frontier, at the fraction of the cost.

That’s not something for an average Joe to do, but for any bigger business with big spent it’s only natural thing to look into. Just one example - cursor composer - that’s fine tuned kimi.

It’s not whether frontier labs will stop releasing models. It’s whether they can generate enough profit out of them. 2 years ago (even 1) they basically had monopoly and combined with demand explosion as capabilities exploded - valuations grew to insane levels. But math now looks different - they no longer have monopoly.


There's no shortage of extremely valuable problems to solve.


> There's no shortage of extremely valuable problems to solve.

And there's no shortage of cheap models that are perfectly capable of solving them.


I’ve been conducting AI Coding interviews at my job. Candidate will screen share and use AI-assisted coding environment and tools of their choice.

I ask them to implement xyz thing. What I’m looking for is how they interact with the AI agent. Do they ask the agent to plan first? Do they review the plan? Etc. It’s pretty typical stuff that you might expect an experienced engineer to do if they effectively use such tools daily.

There are a series of follow-up questions about how to productionize the toy system, which gives some additional signal about how well they understand what they’re making. I sprinkle these in when there’s dead air waiting for the bot to think.

I think we’ll need to evolve and refine this problem and process as the models continue to improve.


>It’s pretty typical stuff that you might expect an experienced engineer to do if they effectively use such tools daily.

It's also the kind of stuff that somebody can learn in a week, so not hiring the right person who just didn't spend this week of time yet for whatever reason is a loss.


I used to think this was true until my company adopted a similar process, and the amount of very senior candidates who used it very ineffectively was highly surprising (eg staff candidate only using the default Google search AI), whereas the more junior candidates figured out you could one shot the problem but were unable to explain their code


Yeah it tells you nothing of value. It's the Recency Bias taken to its ultimate conclusion isn't it. Anything I learned about AI I learned in the last few months but somehow that's as important as all the experience of my decades long software career? Come on.


It signals you don't have enough time or will to chase the latest stuff, otherwise you would have learned it with other cool kids half a year earlier. It carries some signal, but I would not choose based on this alone. I would even consider selecting against this to a certain point, but I also work in a place that COBOL on a mainframe. YMMV.


> I would even consider selecting against this to a certain point

I agree. A technology professional who views themselves as "one of the cool kids" (read: easily manipulated by social media) is a legitimate security threat, as are many of the popularly promoted approaches to "LLM-assisted software developement".


I haven't considered looking at it from the security thread angle. Can you expand on this?


Leasing seems like it could make sense given rapid upgrades in the technology. Why spend $15k on a home robot that will be obsoleted in the next few years?


Leasing is out, but locked-in subscriptions are in!


Maybe it's akin to the Ballmer peak: improved performance at a specific level of relaxation


Some mixed cancer associations in animal models


Contact?


yeah, similar. Might be it, I guess? meh, apologies for being vague and old.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: