Hacker Newsnew | past | comments | ask | show | jobs | submit | dannyw's commentslogin

The way I see it, it’s all about how you use it. Let people have tools; their output stands on its own legs.

As pro software, the UI has a learning curve. Many YouTube videos are out of date, or unnecessarily drawn out.

Additionally, as a hobbyist, project settings, color spaces, clipping, etc can be tricky to miss.

I’ve already been using Claude to (1) help me with UI, (2) QA my delivery, and (3) occasional fixes for audio repair; waveform sync (can’t afford time codes).

I make all creative decisions - including most menial ones - myself, I wouldn’t have it any other way; I enjoy the act of creating. I welcome this integration.

I’m sure some users are gonna agentically grade or whatnot. I don’t like it, but I’m also don’t believe in telling other people what they should do with their own work and creative process.


"Video game" is the medium, the challenge and test is figuring out how to win; when given no instructions, and specifically designed to be private.

ARC-AGI takes it very seriously: they've never tested Fable, because they won't run on the eval set without ZDR.

Another possible way to look at this is that any benchmark reporting Fable scores is potentially contaminated.


I really, really don't find the Artificial Analysis Intelligence Index credible anymore. It's some weighted score of benchmarks, and benchmarks increasingly don't reflect how good a model is.

That should be obvious if you compare Gemini 3.8 Flash (which is an _excellent_ model especially for its price and TPS!! but 10min of prompting in any harness) will tell you it's nowhere near close to Sol/Astra.

But AA scores Gemini 3.8 Flash at 59, and Astra at 61.


> but 10min of prompting in any harness) will tell you it's nowhere near close to Sol/Astra.

I code in both every day a lot and it is not obvious to me 3.8 is far behind


Anecdotal experiences from my external early testing of Astra: if you love Sol (like I do) and wished it was smarter at everything, but especially better at high-level tasks and discussions; I think you'll LOVE Astra.

Astra retains the best parts and overall 'grounded collaborator and executor' of Sol in my testing (harness: codex CLI); while being a significant leap in capabilities & higher-level thinking.

When you prompt it like a technical collaborator, I've found Astra to be extremely consistent in staying as a collaborator, and not being over-eager, over-achieving or doing work that you haven't asked it to.

When you ask it to one-shot something, or explicitly ask it to make decisions, it will of course make its own assumptions and decisions, and generally very well.

Astra is also excellent at instruction following and respecting the guidance and steers boundaries you have.

^OpenAI does not review, limit, or tell me what to say; opinions are my own experiences.


This is quite exciting. Sol for me has been the absolute best model yet. I find myself using it 95% of the time even though I have access to Fable. Is the speed the same as 5.6 sol?

Fellow excited. Sol has been revolutionary for me. It’s been the first time I’ve let a model “go ham” on a production code base and it actually worked, and also didn’t bankrupt my company in the process.

I’m far more excited for the trajectory that OpenAI has chosen. I’ve been listening to mates whose companies adopted Claude wholesale, only for their AI use to become a double digit percentage of their salary.

I get that AI is a force multiplier, but that level of expense isn’t a path to mass adoption.

Honestly, this feels like a real revolution now in a way that 90s kid never really experienced. We grew up with technological progression, we never experienced the obsolescence of skill.

The internet revolution made for more skills and innovation, it didn’t obsolete entire careers. My kids are almost certainly going to grow up knowing less but being capable of more.

Imagine being a 1950s “human calculator” on the dawn of a computer revolution. That’s what it feels like right now.


> "Imagine being a 1950s “human calculator” on the dawn of a computer revolution. That’s what it feels like right now."

My personal / family history is a real-world example of that evolution. My grandfather was a "computer", my father was a traditional "programmer" (lots of Perl), and I'm a SWE / frontend architect / budding "AI Engineer".


> I’ve been listening to mates whose companies adopted Claude wholesale, only for their AI use to become a double digit percentage of their salary.

sigh...yep. that's us.


Off topic, but ooc what do you do such that you get early access to the models?

Precisely. BIS does not consider cloud compute services to be exports.

There are some rules around when you knowingly know that a customer is military or WMD, and you obviously can’t supply services to someone on the entity list; but KYC-through-credit card kinda works at small scale.


Are you using direct or via OpenRouter? I think OpenRouter Luna always uses the `flex` tier, which is quite a bit slower.

That’s incompatible with the file not leaving your browser, which you can trivially verify with the Network tab in chrome (or wireshack, etc).

This is just a C2PA metadata checker.


Why? The browser can see the file and contents so it can calculate and send the hash over without the actual file ever leaving the browser.

Strictly speaking, properly checking C2PA metadata requires network requests in the general case, because you need to check if the signing certificate has been revoked or not via OCSP.

But in anthropic's use case they can probably get away with just pinning their own certs in the verification webpage.


The tests measure Fable 5.1 (with fallbacks). The increased cost can come from Fable 5.1 triggering fallbacks less; which means less (cheaper) Opus when AA ran it.

The new cloth is different. It's softer and thinner, but the same size.

If you want to spend 12 minutes on what changed, AppleInsider made a video: https://www.youtube.com/watch?v=iL-4kCkvEvQ

This is actually the 4th Apple Polishing Cloth. Here are all the variants I know:

  * Original one (reference size)
  * New one (reference size; softer/thinner)
  * Apple Vision Pro (slightly smaller, square, different colour)
  * Nanotexture Display (much smaller, rectangular)

I believe I got this cloth with my new MacBook Pro with the nano texture display was very different than the original.

The narrative — ie being associated and seen as an AI-winner — is worth billions of market capitalisation at Apple scale.

I don’t think there’s a direct guerilla marketing division, but I would absolutely bet money that Apple’s marketing/comms team is encouraging creators, influencers, etc to promote Mac Mini and Mac Studio for AI.

And when those teams brief and speak to people, even on confidential terms, they almost expect some things to get leaked; or will keep sharing it with more people until a third party leaks what they want leaked.

So they may seed things like “super confidential, we were internally really surprised by the demand for AI”. Intentionally.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: