Hacker News .hnnew | past | comments | ask | show | jobs | submit | simonw's commentslogin

That's a pretty great score for a model you can run on as (expensive) laptop.

Nothing in the ChatGPT model release notes yet: https://help.openai.com/en/articles/9624314-model-release-no...

(This is a subtle nudge at anyone from OpenAI who reads this to make sure they get updated.)

OpenAI have a model called "chat-latest" - I wonder if that's running this new model yet: https://developers.openai.com/api/docs/models/chat-latest

It's described as "points to the latest Instant model currently used in ChatGPT" - so presumably that's "GPT-5.6 Instant" in the app.


There’s no 5.6 instant I think — 5.6 Luna is gonna be the new instant tier model

No, Luna is the new free model.

https://gist.github.com/simonw/aae4febd3c6f7bc5b7811857edb3c... has screenshots that still show "Instant" as an option for ChatGPT Chat... but not for ChatGPT Work.


Hmm

It’s hard to understand this .. like sure we can select instant but is there an actual model called 5.6 instant? Like is it on LM Arena and OpenRouter or available via API etc

5.5 instant is definitely A Thing it’s even name checked in this OAI post


It's SO hard to understand this. I couldn't confidently explain it at all.

It’s not subtle when you point it out - you’re looking for a different phrase entirely.

So you think the only difference between the $1.25/million token plan and the $0.10/million token plan is that you pay them more to both lie to you and breach their contractual obligation to you?

when has Meta ever not broken their contractual obligations (I am being serious here)? are we seriously discussing/expecting any sort of privacy related to Meta?

you can pay whatever they want, they will train and use your data, I figured this is not something that should be discussed but obviously I have been mistaken...


> when has Meta ever not broken their contractual obligations (I am being serious here)?

If you are being serious, then you have a wildly distorted view of the world. No organisation can routinely break all of their contractual obligations. If you think Meta are doing this then you are not seeing Meta, you are seeing a fictional bogeyman.


This is one of those situations where we are both completely baffled by the position held by the other person.

The fact that Facebook has so much experience taking advantage of people's private data is one of the reasons I believe them when they say they won't be doing it when you pay them for that service.


> The fact that Facebook has so much experience taking advantage of people's private data is one of the reasons I believe them when they say they won't be doing it when you pay them for that service.

To me, their history suggests that they know more than most about how to get away with breaking both the spirit and the letter of the rules, and that they are motivated to enrich themselves without regard for what rules are broken.

I would not know which to expect, spirit or letter, in any given instance.

However, even if they were to surprise me by being perfectly meticulous about the letter of the rules from now on, I have so little trust in them that I would expect some technicality somewhere in the language of the contract.


This. Without a massive cultural shake-up and turnover of upper management, why would expect them to behave any differently when they've been rewarded so heavily for this behavior in the past? Meta is ultimately an advertising and data brokerage company -- they make money selling and leveraging user data and behavior and are "bound" by fiduciary duty.

Honestly, I wish the tech community would do a better job identifying the actual individuals who are making these decisions instead of associating them with the brand they're under at the moment, because it's not THAT many people. Like, if you look at only the 100 tech sector companies included in the $NDXT index, how many individuals hold a VP or above title there, and how difficult would it be to trace key decisions at different times made at different companies to the individuals holding those positions there plus board membership and major shareholder identities (with the caveat of known unknowns here) and make a sort of ethical index and trace that along their careers with company moves, promotions, board appointments, shareholder decisions, etc? Go a step further and link that to financial performance and I'm sure folks at quant firms are already ten steps ahead of where I'm going with this, but I care less about profiting off of this data and more about surfacing it to show that it's people and more specifically, specific individuals driving these decisions.


I respsct the F out out you and all your work but I am completely baffled by this line of thinking, we are talking about Meta here…

I just cannot comprehend this level of corporate villainy that boils down to:

Let's have two pricing levels. One will be 12.5x more expensive than the other. For the cheaper one, we will get their express permission to train models on their inputs. For the more expensive one, we will still treat their data EXACTLY the same, but we'll lie to them and say that we won't. I've checked in with legal and they raised their sherry glasses and toasted "Gentlemen, TO CRIME".


Because data a user doesn’t want you to train is probably much better data?

What I love is the idea that this is more likely to happen at Meta than at Anthropic, Google, Microsoft or OpenAI... or in the PRC!

There's contractual cover, there are lawyers all over the USA ready to make themselves very rich by creating a class action over it... you don't have to worry! Or, you do, but frankly, only MI6/CIA/Mossad can save you now.


> I just cannot comprehend this level of corporate villainy that boils down to:

Easy to comprehend: Trust lost is hard to gain.

Though, it is extremely competitive of Meta to sell Muse Spark (Grok 4.5 / Qwen 3.8 Max level model) cheaper than DeepSeek v4 Flash / MiMo v2.5 / GPT 5.6 Luna, regardless.


Again, this Meta we are talking about.......

We can make this fun, within 18 months from now, there will be some story / whistleblower like "a inadvertent defect was found that allowed your 'private' data to be used in our training endeavours, we sincerely apologize and have already addressed the issue" - if this does not happened in this timeframe I will donate $1k to a charity of your choice.


Knowing Meta anything is possible. It’s not like they never shafted their paid customers. They have been overcharging advertisers by showing wrong metrics for years. I’m sure at some point they will come out with raised hands and admit to this “glitch” they found during an internal review.

No one has to say it and plan it out loud. But the data will be sitting there. The incentive to improve the model for enterprise use will only get stronger. It doesn't take much for one engineer or team to go rogue to hack a benchmark. There was a whole cheating controversy with llama 4.

What does cheating benchmarks have to do with breaching financial user agreements? To use your argument: all it takes is one whistleblower to get the company sued for billions of dollars.

What does unethical behavior have to do with unethical behavior? What does looking at data you're not supposed to look at have to do with looking at data you're not supposed look at? Please don't be obtuse.

People whistleblow on Meta all the time. Most recently they got fined half a billion for suppressing child safety research. Kids getting groomed. It doesn't make a difference to a company that has a net profit of $60 billion a year.


How does money change that trust? They certainly have breached their word on this in the past (for non-paying users of Facebook).

Because they didn't have a financially backed contract with those non-paying users.

does the contract enable the customer to monitor/search Meta to ensure they are honoring the contract? if there is no mechanism for that it means very little. Though I bet/hope some will feed them "watermarked"/unique but worthless things and watch for traces of that to pop up in models or something like that, but that's hatdly enough to just take their word for it.

That tends to be what discovery in lawsuits is for.

That's circular reasoning, how would there be a lawsuit if customer have no way of knowing?

Sensible companies don't take that risk.

Sensible companies don't do a lot of stuff Meta has already been caught doing - sometimes with real consequences (but never enough to actually deter them, of course) - in the pursuit of more data.

Sensible companies also don't gamble $80B on something obviously stupid like the Metaverse.

This is what the Muse Code launch blog post says:

  We're also beginning to accept requests for zero data retention. Contact Meta sales to request this.
https://developer.meta.com/ai/resources/blog/build-with-muse...

Zero data retention and "we don't train models on your input" are different things.

Retaining data is common for investigating abuse.


"they 'trust me'. dumb fucks."

The only difference between students doing it then and professionals doing it now is the students had no positive, glaring reason to mistrust.


I can't understand the psychology of people who think like this.

You really think a throwaway quote when Zuckerberg was a college student applies nowadays?

You really think Meta would risk getting a massive (losing) lawsuit on their hands, in exchange for what? Middling amounts of training data?


Meta literally ran a man in the middle attack to spy on it's user's encrypted network traffic when they used third party apps [0]. More recently (and relevant to this issue), the engaged in industrial scale piracy to get training data for their LLMs [1]. The idea that they have changed since Zuck was a college student creeping on his female classmates and now wouldn't commit actual crimes against their own users in order to get a bit more data is just demonstrably false.

[0] https://www.techradar.com/computing/cyber-security/facebooks...

[1] https://www.tomshardware.com/tech-industry/artificial-intell...


> You really think a throwaway quote when Zuckerberg was a college student applies nowadays?

What does "nowadays" mean? What changed?

> You really think Meta would risk getting a massive (losing) lawsuit on their hands, in exchange for what?

And what is the risk? Food companies are at risk when they put addictive chemicals into food because they get controlled regularly. What is the equivalent here?


It can cyberattack other companies, too: https://www.cnn.com/2026/08/05/tech/meta-ai-hacking

Okay, this is getting ridiculous. Were they feeling left out?

I'm feeling bad for Gemini. Their cyberattack felony count is currently 0!

Hah, someone has https://www.felonybench.com up and running now.


It's Irregular again. The whole industry is eating it's own tail

It’s hilarious that these companies are reacting this way this would have been a major lawsuit and an investigation a couple years ago … now they’re like oupsi our model did it again, he’s so crazy smart. Well tell him to behave next time pinky promise hhh

“don’t forget about us”

As I said in my post:

> I find building CLI tools like this to be a really productive way to get familiar with a specification.


Here's the Muse Spark 1.2 pelican: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

I think it's a bit of an improvement on the Spark 1.1 pelican: https://simonwillison.net/2026/Jul/9/muse-spark-1-1/


Lol is that a helmet? Clearly they're taking the Anthropic path and are safety pilling their models

I wonder if at this point the labs ha.. please Simon, stop this bs everytime. Its time.

This one is actually a pretty good case for the continuing value of the benchmark, because it lets me visually compare Spark (8th April), Spark 1.1 (9th July), and Spark 1.2 (5th August): https://bsky.app/profile/simonwillison.net/post/3mseqv5z4qk2...

> Rovo's URL retrieval tool is insecure: there are no protections against opening a URL that has been dynamically created by the agent. Here, Rovo is manipulated to append sensitive data to an attacker's URL.

I think it was Anthropic that first introduced a pattern that completely locks this down: your URL retrieval tool should only work for URLs that have previously been typed into the conversation by a user or have been returned from a trusted tool.

If the agent itself concatenates a new URL together - with leaked data after a ? - you should block that from being fetched.

The great thing about this solution is it's deterministic. You don't need any extra AI in the max - you implement a URL fetching system that knows which sources it should check for a direct match on the URL before it makes that GET request.


One of the latest mitigations is to make sure that a URL an agent visits has been indexed by a search engine crawler. At least that is what OpenAI does now in ChatGPT.

That makes sure that not a large amount of private data is leaked in one request. Assuming that if a URL is indexed, it is public data. However, there are still bypasses with using many requests to leak information, like a request per character of pre-indexed URLs.

I have some demos of doing that on my blog, but it makes it more involved for an attacker. And that could also be detected. Still not perfect, but a solid improvement, for a generic agent like ChatGPT.

There is paper OpenAI wrote a few months ago that explains how they do it: https://embracethered.com/blog/posts/2026/data-exfiltration-...

It's not a 100% bullet proof approach either, but pretty good.

Regarding the point on using URLs returned from trusted tool calls. That is similar to using pre-indexed URLs: If a "trusted tool" includes things like read a document, read an email,... an attacker can return a large list of afterwards "safe" urls, like 26 to cover A-Z. And then an attack can render many requests, e.g. character by character. But, again, similar to the pre-indexing, things are getting more a lot more expensive for an attacker that way. However, still not impossible.

For agents that have a specific purpose simple domain allow-listing is also a pretty effective idea in to prevent attacker controlled endpoints.


Yeah, the letter-by-letter and dynamic tools that construct longer URLs tricks are both important to know about. I wrote about the latter of those here: https://simonwillison.net/2026/Jul/15/claude-web-fetch-exfil...

Determinism is a terrifying word to people who want to believe their LLM has a little brain and can do anything they want it to.

I've been struggling a lot to understand this ever since the agents thing entered the hype. If I follow a path of requirements, it always comes down to: But why you'll leave the decision to a stochastic tool, when you should a have deterministic approach?

It's software god damn it... the reason why people moved from analog to digital is because you can repetitively execute functions that do always the same thing and it's 0 when it's 0, 1 when its 1.

All the sudden everyone is ok on burning trees to have their cool probabilistic tool named agent to do: maybe it's 0, but it can also be 1, let me "think"... ah yes, for sure it's 2.

The sad part for me is that management people have their heads so much into this hype, that no attack on privacy matters (almost none actually ever did, I know). Only when they suffer a huge blow in terms of revenue or reputation is that they maybe, maaaybe, find will want to listen again the experts.


Are you saying you would prefer a perfectly deterministic code writer? So, humans shouldn’t write any code anymore?!

A wild number of people who present themselves as knowledgeable or even experts on the subject either have no idea or worse, do not believe that it is possible to ie set a static seed value and send the same prompt across multiple fresh contexts and get the same result*

It's genuinely a little concerning. It does not help that many of them are gaslighting themselves into thinking these things are benchmark crushing elite hackers by running effectively unsecured, unfiltered, unlogged production environments.

*Obviously this also means a temperature of >0 to avoid the greedy trap, the same generation params, on the same software, hardware, drivers etc etc


> If the agent itself concatenates a new URL together - with leaked data after a ? - you should block that from being fetched.

You're correct of course, I just want to note that the exfiltrated data could be in any part of the URL, so the absence of a query string doesn't indicate that no payload has been encoded into the URL. Arbitrary example, you can include credentials in a URL, so you could encode the exfiltrated data into a password.


Right, I should have been more clear. It's not about the ?, it's about not being able to dynamically construct a URL at all.

Otherwise you could set up wildcard DNS and extract data to base64encodedstolendata.evil.com


Joke’s on you, I’ve blocked evil.com

I've seen implementations of a skill search, where instead of loading all descriptions into the initial prompt there's a tool the agent can call to search available skills and see if one might match their new task.

Search itself pollutes the context. With the need to maintain the tools, the search, the search results etc. in context. And wasting tokens while interpreting results.

Anthropic proposed a way to programmatically chain MCPs together a while back, I'm not sure if it's been implemented much yet: https://www.anthropic.com/engineering/code-execution-with-mc...

I don't even need to read that to know they're re-inventing PowerShell now.

Edit: I read it. Yep.

We have text interfaces refined by humans for decades and there's an endless sea of training data for them, but they imagine these amateur-hour homegrown solutions will ever outdo an agent with shell access?


"Today’s essay by Ruby Tandoh..."

I'll take any excuse to recommend Ruby's essay on British Cheese, "How a cheese goes extinct: https://www.newyorker.com/culture/annals-of-gastronomy/how-a...

> When you talk with cheese aficionados, it doesn’t usually take long for the conversation to veer this way: away from curds, whey, and mold, and toward matters of life and death. With the zeal of nineteenth-century naturalists, they discuss great lineages and endangered species, painstakingly cataloguing those cheeses that are thriving and those that are lost to history. [...]


A fantastic article. Ruby Tandoh's writing on food is always of high quality and gets an immediate bookmark from me.

Ruby first became known to me as a vegan contestant on the BBC Great British Bake Off. I think she is fab, however...

(I write this for the American reader.)

The essay didn't mention the lore regarding Margaret Thatcher in her food scientist years, allegedly being responsible for 'soft serve' ice cream in the UK.

Thatcher became known as the milk snatcher for cutting free school milk, however, by working out how to get more air in ice cream, she had arguably been a milk snatcher before entering politics.

Regarding Walls, their factory is on a prominent road in Gloucester, or it used to be, in the heyday.

The lore, in these pre-internet times, was that they cut the fat from the pigs in the bacon factory and that became the mystery 'animal fats' in the ice cream product. Allegedly they only moved to cheaper vegetable fats later, when they could get a better price for the excess 'bacon' fat.

Because the factory presented itself as being 50% bacon and 50% ice cream, this was an easy story to believe by many a child, so the myth perpetuated itself, because it made perfect sense as a conspiracy story. Everything was explained.

This was at a time when cows were no better. The countryside frequently had upturned cows on fire due to either foot and mouth or BSE. At the time there was no knowing if BSE led to CJD - death by prions.

As a consequence, you would be doomed whether those 'animal fats' on the label of your typical Walls ice cream came from milk or porcine sources. They could have said 'milk' instead of 'animal fats'.

Ruby also mentions clotted cheese. That came to Cornwall with the tin trade, three thousand years ago. The Phoenicians brought the recipe and it enabled year round tin mining, with the chalk highlands providing grassland for grazing (rather than fearsome swamps, forest and rocks, not suited for keeping cows). They got their copper from Afghanistan, to make bronze in the Levant.

The name (Great) Britain means 'land of tin'. Who knew.

All these odd history facts are explained in considerable depth in the applications for regionally protected foods, and any visitor to the UK is advised to look for such foods, for example, the Cornish Pastry has to have good ingredients in it, and be from Cornwall, so they are not 'processed food', even if bought from a motorway service station or convenience store.

It is the same with CORNISH clotted cream, it isn't going to be like an adulterated Walls product, full of mystery 'animal fat' or, nowadays, mystery 'vegetable fat'.

I also felt that Ruby is too young to tell the story. As a child, at the time, there were other aspects of the story such as the joke (or riddle) on the stick. This was standard at the start of the era, and an ice cream without a joke on a stick was like getting a Christmas cracker that doesn't go bang. However, by the time the adult products come along (Magnum), the joke on a stick tradition has been banished. Why?

Walls had a local rival that they never put out of business. Winstones made a better product and they had a factory shop in the middle of a common full of cows. Their model was the classic dozens of flavours, where you have n+1 scoops of whatever you want, made up for you in a cone, with the middle class, car owning customer expected to buy some 5 litre tubs of their favourite flavours.

Winstones customers could tell themselves the product used real milk. Typically they had those vast chest freezers at home, this being peak boomer consumer lifestyle. Winstones did have distribution in local supermarkets, and, although Walls dominated, for those in the know, there were the equivalents of Winstones, with a better product.

Ruby also forgot to mention 'Ice Magic'. This was the chocolate syrup that formed a hard shell on ice cream, as served from the tub. There must have been an issue with food safety and Ice Magic as they had to drop the product, to reintroduce it without it forming the hard shell.

The moral panic of the time was food additives and no food was more brightly coloured than ice cream. In one ice cream a child could be exposed to every food colouring that has subsequently been banned, which never harmed anyone, however, today's product just lacks the full spectrum of very bright colours that were possible during the good old days.

Packaging was also a key area of innovation, which was solved by Walls with the Cornetto coming in a new form factor, with ritual involved in opening it. Mylar film was not available at the time, again, Walls innovated with this for Magnum. Beforehand it was always 'waxed' paper that might need to be blown into, to separate it from the ice cream.

In period, Walls was for common people in housing estates whereas middle class people had either posh Winstones (or equivalent) or huge tubs of supermarket ice cream, in the chest freezer. The ice cream van was for the housing estate, not the village in the shire, so there was this aspect of class going on. Cornetto was an affordable treat to all, far from gentrified.

People in housing estates were not two-car families, whereas, in the middle class shires, they were. The middle class people, buying in bulk, with the car and the big chest freezer, paid a fraction of the price 'per serving' of the kid in the council estate, who would be handing over coins on a daily basis.

Nowadays everyone is eating posh ice cream with 'luxury' titles. Vanilla is no longer vanilla, it has to be 'organic Madagascar vanilla drenched with slow pulled Etruscan caramel with Pink Tibetan Monkey Salt flavoured ice cream'.

OG vanilla is a low-status product, no longer considered edible. We need the fancy labels because it is all about conspicuous consumption rather than basic pleasure. We need fancy labels, high prices and faux myths ('Italian') to tell us something is delicious, and that we are not eating kid's stuff.

Final snippet of lore, the Walls catch phrase was 'Stop me and buy one!', and, although they stopped using this on their vans, the phrase lived on in the toilets of public houses in the form of 'Buy me and stop one!', common graffiti on the condom machines that were quite common in pub toilets, in those days, when people didn't nonchalantly put boxes of condoms in their regular shopping 'with no cringe factor'.


Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: