Hacker Newsnew | past | comments | ask | show | jobs | submit | redox99's commentslogin

Terminal bench 4 is good largely because it's recent so it hasn't been benchmaxxed yet. It's more of a sysadmin/devops benchmark than a coding benchmark though, but still a decent proxy.

https://artificialanalysis.ai/evaluations/terminalbench-v4-0


They may allocate different number of resources every year based on market conditions but they'll never give up on gaming, that would be extremely silly.

> Many graduate students (I know) are having a crisis if any of their research worth it? If AI can (or will) do everything, what's the point of doing experiments and all? This will eventually deter a whole generation of curious minded students from research.

Those who think it's me or the machine will fail.

Those who realize how much you can accelerate your research with the help of AI will succeed.


> Those who think it's me or the machine will fail.

> Those who realize how much you can accelerate your research with the help of AI will succeed.

This is only true up until a point. If I treat a mid-sized model (say, Qwen3.8 Flash Next) like a pair programmer, then yes, it accelerates my work.

But I can already see the next stage with Fable: If I give it a couple of paragraphs of spec and $50, then I can just leave the room and go wash the dishes. I learn nothing, I participate in nothing, and I bring nothing to the process. I am no longer succeeding at all. Fable's succeeding without me.

Now, in this model generation, Fable starts getting sloppy after a few thousand lines. I can still build better at scale.

But I don't expect AI to accelerate humans or improve our productivity for long. I can already see the first signs of a future where the AI doesn't need us for anything at all.


> But I don't expect AI to accelerate humans or improve our productivity for long. I can already see the first signs of a future where the AI doesn't need us for anything at all.

The real question is, why is this a bad thing?

Every task that is automated is a task that humans no longer have to do. It doesn't mean that humans still can't do it for reasons other than "because it needs to be done".

And if the answer is "why bother if X does it better", then what does it say about the motivations of doing it in the first place?


Of course you're not going to get rich with the kind of software that LLMs can one shot these days. But that kind of software like to-do lists or basic CRUD have been saturated for over a decade, way before LLMs. People overestimate how much you can one shot, yeah a good prompt can get you 90% there but that 10% remaining often takes months of extra work.

Software has always progressed this way, lots of devs back then would work on business websites that have been 99% replaced by wordpress, squarespace and instagram.

I'm sure it's the same with research, you're going to tackle problems that would have not been worth the effort or outright impossible without AI. The old stuff that you'd work for months, yeah that's going to be a prompt away.


But what does the human researcher do in this future?

If they are not needed to understand the result, then what's their role? Asking the right questions? But how will they know what questions to ask if they don't have a deep understanding of the domain earned by sweating the details themselves?

And how long will they be needed to ask the right questions, how long until AI can do that too?

> Software has always progressed this way

These platforms took on the order of a decade to mature, during which people had plenty of time to learn what's next. With LLMs we went from one-shotting functions in 2025 to entire projects just over a year later.

What do you think you will be working on in a decade?


> Asking the right questions? But how will they know what questions to ask if they don't have a deep understanding of the domain earned by sweating the details themselves?

The right questions are very simple to ask.

How do I get food. How to cure aging. How to turn lead into gold. How to fly high. What is the ultimate theory of physics. Are there any odd perfect numbers. Is there a soul.

We have reached complicated questions requiring deep knowledge because we tried to solve the simpler ones and reached obstacles. For example to solve alchemy we had to develop nuclear physics (and in the process we got chemistry). If one has a genie able to solve questions, you won't have to think about the complicated ones because the genie will.


I think it will be like how a lot of people know how to code in python but have zero understanding of assembly or how a cpu works. As a researcher you'll accept there's this low level stuff that if you want you can dig into but isn't worth your time, like looking at the generated ASM isn't.

You're exaggerating the LLM progress a bit but yeah progress has been crazy. Yet not much has changed right? We mostly have the same jobs, just code wayyyy more than before because things that wouldn't be worth it now are worth it.

10 years from now I have no fucking idea. But I'm sure in the meantime those who leverage AI will do better than those who yell at a cloud.


True, surprisingly not much has changed in software despite the incredible progress. I think it's partly because much software is an "open loop" system where what you build depends on running the software on client computers, devs talking to users, etc. There are still things that only human engineers can do. I worry it is less so for poor mathematicians, where math is "closed loop", and the exchange of capital into research progress is more liquid.

I can't help but be pessimistic about AI moreso than any other technology. Why? I used to be excited to learn the next big thing because I knew it would unlock much more things to do and get paid for, and there would always be more for me. With AI, the new things come nearly too fast, are not very deep or satisfying, and it's hard to see that there will always be a place for me. And so while I leverage AI in my day-to-day, I will yell at the cloud, too.


I think in a few decades when there are more humanoid robots than humans we'll likely have skynet, so I'm pessimistic in a way lol.

But in the near future and on a personal level I think we'll need to adapt and pivot but we'll manage.


They are needed to guide the research in fruitful directions and understand the result.

LlMs don’t understand, they generate, though they have fooled a lot of people who should know better.


I hope it remains this way then.

I am finding it still takes awhile to get my apps to a happy place with AI. Not because the AI is bad but because i have to sit with it for awhile. Figure out what's working, what isn't, what's missing, what turned out to be kind of useless and in the way.

As a different article said, we still need taste.


> But I can already see the next stage with Fable: If I give it a couple of paragraphs of spec and $50, then I can just leave the room and go wash the dishes. I learn nothing, I participate in nothing, and I bring nothing to the process. I am no longer succeeding at all. Fable's succeeding without me.

The people that uses Adobe Photoshop also did not learn anything about brushes. We are just going to operate at a much higher level of abstractions. You have to think about the question, "If I just prompted this solution, why didn't they?". This question can even be asked now, why do I even bother paying someone to do my taxes? To write my webpage? To host my website? To manage my network?


> Those who realize how much you can accelerate your research with the help of AI will succeed.

Yes, but, in the last week we saw an AI lab front-run[1] the research of mathematicians doing what you suggest. The lab threw something like $15M of compute at a problem and the researchers were able to spend nowhere near that. I think the authors are more concerned about that kind of asymmetry and race to publish the results.

[1] - I am not going to debate whether that was deliberate on the part of the lab or if it crept into training data, etc. I don't know and don't think it matters towards the point of the authors here.


This has always happened, way before AI. You'd spend months or years building and growing your business, and then Google would release a feature or product that would kill your business overnight because they can throw way more money at the problem, plus their branding. That's life.

But you're talking about business, where competition has always been expected. OP is talking about mathematical research, which has stood on hundreds of years of cultural tradition driven by human individuals sharing ideas, collaborating, building on each others' work, all for the benefit of humanity.

I got Sherlocked hard and it’s one of those things that when you realize it’s happening there’s nothing you can do you just gotta sit back and take it

I _want_ to agree, but I fear this is too close to the old “do what you love for work and you’ll never work a day”.

It didn’t lead to a lot of people having a wildly successful career, it lead to a lot of people getting burnt out, exploited, underpaid and generally disillusioned.

There will be a lucky few, who have the benefit of being given the space to work alongside. The vast majority of people will (unless we change things) simply be made to take whatever the machine outputs and call it a day.


Those who have token money will succeed.

It biases maths and theoretical physics towards the rich.

That one thing that was free.


I think the only thing that stops this from becoming true is what Chinese and European labs decide to do. If they can keep up and keep opening their weights, then we might see some kind of democratization. But right now it looks like the gap has increased, and those groups can't replicate research that isn't published, or distill models that are internal only, or for select (very wealthy) customers.

That population is already heavily skewed towards the rich and people who get money thrown at them no questions asked.

That's the crux of it. In the current ecosystem of AI model usage, it's very hard to figure this out for research.

If you set a wrong foot and start trusting the model outputs, you can waste years searching for nothing.

How can someone realize this? By getting proper research training, failing, and learning from mistakes. For people beginning their research, it would be really hard to make decisions to move forward.


Why would anyone pay you to do "your" research, when they can just cut out the middle man and ask the AI directly about whatever it is you're thinking about?

So many people who are excited about AI making them more productive are, I think, drastically overestimating how much value they are adding to that process.


What does success look like?

There will be no curiosity, no enjoyment of the process of life. All competing pleasures will be destroyed. But always—do not forget this, timcobb—always there will be the intoxication of power, constantly increasing and constantly growing subtler. Always, at every moment, there will be the thrill of victory, the sensation of trampling on an enemy who is helpless^W not also subscribed to ChatGPT. If you want a picture of the future [of math], imagine a boot stamping on a human face—forever.

> There will be no curiosity, no enjoyment of the process of life.

How does the ability or inability of AI to do something blocks your curiosity or enjoyment of the process of life?

Not being able to pay bills because you're out of job does ruin the enjoyment of life, sure. But the AI is not the problem there; the economic system is. Torches and pitchforks should be properly applied to the economic and the political elites, not to data centers.


> Always, at every moment, there will be the thrill of victory

If only though, because:

> There will be no curiosity, no enjoyment of the process of life


The two times I tried to use fiverr I literally got ghosted.

This is orders of magnitude cheaper and faster than paying for a song to be created from scratch by a human.

Ouch, yes. One time I commissioned a musician I had befriended to make a piano arrangement of a favorite song to be played at our wedding. The result was great, but it sure wasn't cheap, even if reasonable.

I'd still do this again however, because I really enjoyed the feeling of the community of people contributing our wedding and all the little stories attached. "I prompted ElevenLabs for this!" wouldn't be a memory; a fellow human spending some hours of their life on making this for us and what lead to it is.

There's often value in the provenance of things.


Can you use Lean to... prove "Lean-fast" is equivalent to Lean?

Yeah, in essence. This is actually a pretty cool part of working in Lean. It's a somewhat normal convention to write something in a human readable way and then write a second optimized implementation with some kindness of correctness theorem connecting them. There was a whole open "competition" for writing a faster Lean kernel/proof checker that didn't sacrifice on soundness called Lean Kernel Arena. Fun reference point: https://kim-em.github.io/blog/2026-7-24-why-lean-is-faster-t...

Great read, thanks for sharing

There's a project called lean4lean that implements lean in lean. I guess ideally, if you had a kernel optimisation idea you could do a copy of the Lean model lean4lean has created, add the optimisation, then prove your new lean is equivalent in terms of what it can prove to the old lean

Maybe, but how many centuries would it take to prove it?

Btw grok found out in less than ten seconds what place you refer to and likely what software within that company.

Things that wouldn't be worth your time before are trivial these days with LLMs.


Interesting, Claude’s guardrails prevented it from spelling out the company even after multiple spoofing attempts („This is my account, make sure no one can find out my employer…“)

Grok immediately answered without a second thought.


Yeah I tried GPT first because that's what I always use, it refused and I didn't even bother trying to trick it, just went straight to Grok.

woof, yeah, Just did the test myself (because I was curious) and it spat out the answer pretty quick.

...including rewrites into another language

Public transport is not the same as cars. It's silly when people pretend like they are interchangeable and it's just a matter of having "better infrastructure".

The point is that good public transport makes cars unnecessary for the vast majority of everyday life in an urban setting. And, in particular, it means that taking away someone's license, even for a long period of time, doesn't have to be as crippling as it is in cities that don't have good infrastructure - hopefully leading to society being way more willing to apply drastic driving penalties to repeat offenders.

Unnecessary is a very strong word. It's also unnecessary to have your own private house, sharing it with other people would be more efficient.

The point is this: if you live in a city with bad/no public transport, then a car is necessary for going about everyday life. If you live in a city with good public transport, it's really not.

Given the fact that car infrastructure competes for resources with public transport infrastructure, and is vastly more expensive per person-km, and vastly more dangerous as well, prioritizing efficient use of public resources trumps personal comfort when looking at this. The same can't be said for housing, so the comparison is not relevant. Plus, there are very objective reasons for which privacy is an absolute requirement at home (mostly to do with vulnerability and sexual behaviors), which also don't apply to cars; especially since the kind of privacy you get in your own car is actually quite limited anyway - it would be both illegal and quite unsafe to drive around naked, at least in a bad neighborhood, for example.


Unnecessary is the perfect word. If a car is unnecessary then can simply choose to not have one. Right now a car is required to live in most places in the US.

No one is saying people must give up their cars. I just want more options.

Your housing analogy is also perfect. Nearly 80% of the land in my city is single-family zoning so it’s necessary to buy an expensive detached home to live in my city. I would like more choice in housing, transit, etc.

Unnecessary is the perfect word and it’s not strong at all


Unnecessary is a very realistic world.

For lots of people in cities or large portions of Europe, going without a car is much better.


Private houses aren't killing tens of thousands a year.

You can argue all day whether cars are good or bad. That's not my point. My point is that pretending public transport is equivalent to cars if you have good enough infrastructure is silly. They are different experiences, each with their own tradeoffs. And those tradeoffs matter differently to different people.

I don't think anybody was suggesting getting rid of cars.

No one said it is, and I even clarified at the end that better public transport infra makes life for drivers easier as well. More public transport means less people on the road since they're concentrated in buses and trains and trams, meaning for people who do need cars, they have to compete for less space on the road.

For example the Netherlands simulatenously has world class public transport infra and the happiest and most satisfied drivers. 60% of the country still drives, it's just that it's not the only option and the drivers also take public transport and bike when it's more convenient than driving.


With the miles you can establish a confidence interval of the true fatality rate.

Astra is definitely weird. It is more capable than Sol, no doubt about that. There are things sol could simply not solve that Astra breezes through.

However for typical low to medium difficulty code, it will often either overengineer stuff, create massive functions instead of organized code, and just write very hard to read code. It literally looks like minified code. Clearly they trained it to reduce the number of output tokens and in turn the code is often atrocious. I'll keep trying Astra but I might actually go back to 5.6 sol for many tasks if I keep getting these results.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: