Hacker News .hnnew | past | comments | ask | show | jobs | submitlogin

The eval bar I want to see here is simple: over a complex objective (e.g., deploy to prod using a git workflow), how many tasks can GPT-5 stay on track with before it falls off the train. Context is king and it's the most obvious and glaring problem with current models.


This sounds like the kind of thing:

1. I desperately want (especially from Google)

2. Is impossible, because it will be super gamed, to the detriment of actually building flexible flows.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: