Hacker News .hnnew | past | comments | ask | show | jobs | submitlogin

I can't wait to never see a reasoning button ever again. Why do I have to reason about what reasoning level to use?


I have the opposite problem. I'm not well-calibrated on when I'd want lower reasoning than what's available to me (and how to compare that to lower-tier models). OpenAI now has Luna, Terra and Sol, each at Low, Medium, High and Xhigh, with Pro/Ultra depending on harness and plan. That's ~15 possible combinations of model and reasoning level, and there isn't a satisfactory explanation of which one you want for any particular task.

I feel that work is basically split into two tiers, hard (which requires a good model and lots of reasoning by definition) and relatively easy (which won't consume much of my limits despite a great model and reasoning, so I may just as well keep it on Sol High).


> (which won't consume much of my limits despite a great model and reasoning, so I may just as well keep it on Sol High)

This equation significantly changes if you're paying API prices vs on a subscription. (Such as if you're integrating it into a different product, where it becomes worth it to figure out what the cheapest you can leverage is.)


I don’t see why they just don’t allow a smaller model to answer the question while letting the bigger one vet it. The vetting can be asynchronous and can be delivered after a few seconds (if it’s an easy query). If it’s a hard query, the UI can show the answer is currently being vetted or something.


if the bigger model can vet it fast it can also answer it fast.


Yeah, it's hard, there's so many permutations. But most are bad so it narrows it down.

Luna: Can be useful if price sensitive, always use with AT LEAST high effort. But codex plans are very generous, so just ignore it.

Terra: Forget it exists

Sol: Just use this. Medium can work for specific edits. Otherwise just use high or xhigh.

OpenAI basically agrees with this, and the slider gives you those options.

TL;DR: Just use sol medium/high/xhigh


Why not pro/ultra?


Pro: It's web chat only, I haven't tried it so idk. I think it's useful for Math and that kind of stuff, but I only do programming.

Ultra: Might be useful, the one time I tried it, it was extremely wasteful in terms of tokens. Definitely not a daily driver.

There's also max effort, I think it can be useful but the gains compared to xhigh are quite small.

I think max and ultra are only worth it when the others fail.


It’s very confusing to me. I’ve occasionally used pro to do high-level research and design. Then I ask it to create a prompt for Codex ultra. Ultra can do a lot of genetic benchmarking and testing to elucidate and resolve quandaries.

I don’t know if this is a good workflow.


Some people want quick results. Some people want it to keep searching for a new math proof overnight without giving up, and they have money to burn.

It seems like giving it a time limit or a budget in dollars would be clearer, though?

Or, keep searching until I come back to the computer and ask about progress.


The _Intelligence_ part of AGI should be able to guide the user through that without all the knobs.


OpenAI tried auto-routing with the initial GPT-5 release and it was immediately clear why that was a bad idea.


intelligence is not omniscience though


Asking questions to clarify user intent is a very low bar for intelligence. A bar that all SOTA models fail consistently at though. (It's both funny and legit infuriating when Opus, after having made a dozen wild assumptions without checking with you, then comes back with a request for clarification on some mundane topic).


but you can keep asking questions ad infinitum

very nontrivial problem knowing when to stop and making assumptions


Yes, that requires some amount of intelligence, that's exactly my point.


I don't understand this at all. Whenever I ask Gemini 3.1 Pro Extended, or Claude 5 Max something in chat, the most I ever wait is maybe 30 seconds. Is that really so bad?


"Wait 30 seconds" as a concept has been totally incompatible with the web, smartphones, etc for about 20 years now.


If you just want to know "When it's the next full moon", yes. Very bad as google can answer in 1s.


That I just type into Google.


That's also using AI :).

ChatGPT already competes with Google...


I didn't say it didn't? Besides, for factual queries that don't need personalisation like that they can aggressively cache them.


Auto-effort and similarly auto model routing suffer from a halting problem sort of issue: you don’t reliably know if a request is complex unless you use a complex model to make the decision.


Because the model can't read your mind and know if you want a quick answer, or an hour long deep dive.


Then it should just ask the user what they want if its unclear from the context.


So instead of just getting an answer, I have to wait until it asks me, and I have to type back a response? Extremely annoying.


That's what the reasoning slider is for!


Exactly! I posted much the same comment in another thread and there were lots of huffy complaints that amounted to "you're prompting it wrong"


How is asking better than a reasoning slider?


Because your incentives are opposed to the provider’s incentives


I think it's nice to be able to make the model reason for dozens of minutes when you want to go deep on a topic, even if the router thinks it's an easy question.


I absolutely need instant mode as this is what I use 90% of the time.

Sadly it's not available anymore on the updated desktop app (formerly Codex) and the previous desktop app (formerly ChatGPT) is abandoned.


> I can't wait to never see a reasoning button ever again.

I hope to see a memory-less mode that is not incognito. Want fresh contexts sometimes, but also want to keep the chats saved in history. Memory can spoil some creative work, it dials the model in too tightly.


just wait for the loot boxes, mark my words




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: