This could and probably is me projecting, but the enthusiasm was tepid. The Opus 5.5 mog cast a pall on the room.
The stuff that came off as potentially exciting was the ultrafast stuff available only to Pro 500 users.
The reset button stuff was an embarrassment. Everything that was released today, in respect to the hype that the team projected on twitter, was a failure. The internal OpenAI slack is a hell of an echo chamber.
I also have a feeling that $200 was where most users will draw the line. $500 looks like such a slippery slope in the field where all the recent SOTA models, including Chinese ones, are considered good enough for most work.
$6k/year/seat is not particularly expensive for a dev tool if (very big if) you can measure the value and demonstrate it's worthwhile. That's the missing piece. People look at cheaper options because they don't know if the expensive one is working.
If OpenAI can show that a $6k tool nets $50k in productivity or quality or speed they have no problem selling it. Companies will take the easy-but-expensive option over a cheaper alternative. The problem is that's really hard to measure and seemingly they've put no effort into trying.
If that's true that restores my faith in humanity maybe they're cheering for stuff their friends were working on. Otherwise, the consumerism has reached a point where giving up on society and waiting for the next flood seems like a reasonable option.
Feels huge. Sounds like Luna on their Ultrafast infra.
I did some internal benchmarking and found that GPT 6 Luna in batch mode at 20 records per batch was about 1.6x faster at 1.2x the cost of Jev with no tradeoff in accuracy. Tuning it up/down would make it faster per-record while also reducing the cost (assuming accuracy holds). Though this was only useful for offline processing.
It's another Jev copy, like we've seen so many over the last few weeks. But with no benchmarks or price comparison, which likely means it doesn't compare that well.
> But with no benchmarks or price comparison, which likely means it doesn't compare that well.
Doubt that's the case; this was only the announcement and the API isn't even released yet. As I noted, my own testing using a batching strategy (nothing more than "read 20 lines from this file instead of reading 1 at a time") yielded ~the same accuracy on an 800 row dataset we have with ~1.6x the throughput (at only batch_size=20) and 1.2x the cost. Didn't tune to find accuracy dropoff by batch size, but it's easy to see that the decision capability was there (understood that it's not the same without the probabilities, but even without it, the classifications were nearly identical).
Sign-in with sub accounts seems big. This is how I thought this was always supposed to be - the providers acting more like an OS/ecosystem-enabler.
Extensions into codex also seem cool.
The provider also eating the application stack at the same time, doesn't make too much sense. I don't want my team going all in on one provider versus the other. (things like the docs/slides/sites).
It’s simply ridiculous. None of their 20 new features seem genuinely usable for day-to-day life; I’m struggling to find any use cases for them. I have no idea what they were thinking with "Dots". The "Super Fast" addition is nice and important, but it’s just one small bright spot in a hellscape of useless features.
This is a wild take to me, and I don't even use Codex! I've been on Cursor for a long time, but this is finally making me consider switching (especially since OpenAI models are leaving Cursor).
- Cloud Agents are awesome, super useful, I can't go back to working locally only
- Dots have a cheesy name, but they remind me of a more polished version of Cursor Projects, which I just started using this week. You talk to a single agent who serves as orchestrator and organizer, and spins up cloud agents as needed to do research and implementation tasks. So you can just dump stuff into the thread without worrying about interrupting ongoing work. Works great.
- Sign in with ChatGPT: how is this not major news? They're launching with partnerships with a bunch of harnesses so you can use your subscription outside of Codex.
- Slack integration: now that they have cloud agents, this is a no-brainer, cursor's slack integration is endlessly useful to me. I'm constantly tagging notifications from various services that get dumped into slack and having a cloud agent dig into more details or fix an issue
I could go on, 6.1 sol, ultrafast, etc. Wild that you think all this is useless. I guess we just work very differently.
> Sign in with ChatGPT: how is this not major news? They're launching with partnerships with a bunch of harnesses so you can use your subscription outside of Codex.
Maybe the sign on process is new but OpenAI has explicitly allowed subscriptions to be used in 3rd party apps at least as long as I've been a subscriber (a few months)
"Cloud Agents" is just a blatant copy of Anthropic's offering. "Dots" is a copy of the Grok bot. Regarding the release of 6.1 Sol - I’m not dismissing it, and the model will likely be useful for many things, but come on: there is no way your flagship model doesn't compete with Sonnet 5.5. I can't say for sure, but there must be a reason they didn't show any benchmarks. The live demos were absolutely ridiculous. Sure, there are some useful features, but overall, it feels like they're missing the mark. Instead of simplifying AI usage, they're complicating it to the point where it's no longer clear where you're supposed to do anything.
I was talking about Cursor's Cloud Agents, which predate Anthropic's offering. Obviously they're all going in this direction, it's a no-brainer. I don't think calling it "useless" makes any sense just because Cursor and Anthropic were ahead of them there (if they were, I know nothing about Anthropic's cloud agent offering).
Well, it’s a bit of an exaggeration to say it was bad. It’s simply a matter of falling in line with the rest of the industry. I was expecting something of a higher standard or something new, so I’m disappointed.
> I just don’t seem to have any problems that AI could solve in my personal life.
My microwave handle broke. I looked up the part needed to re-attach it; it's a a tiny piece of plastic. $30 each, and I need two of them, so $60.
So I grabbed photos of the part from online, which showed it against a 1inch square grid. I gave the photos to ChatGPT, told it the grid size, and asked it to make a 3D model.
It did it. I made a test print. It works perfectly.
¯\_(ツ)_/¯
I guess for me, it's been partly a matter of trying to unblind myself to problems I've grown accustomed to solving in a certain way.
The reaction to Jev has been huge, as has been OpenAI's speed here is announcing a "not quite ready yet" competitor... the idea of Jev is one of those "Doh! Why didn't I think of it!" moments ... in retrospect it's obvious that business automation is all about decision making and the market for this is massive. Up until now most people just saw this as a use case for LLMs, until TypeSafe/Jev came along and said "actually you don't need an LLM for this ...".
I'm sure TypeSafe realized the idea is a winner and expected competition, although maybe not quite so fast?
The question now is how seriously OpenAI take this, and are they willing to compete with Jev on price, thereby massively reducing Luna margins (which OpenAI's "Decisions API" is based on)? I guess if you can't make $0.10/M for Luna, then $0.04/M is better than just giving that market to a competitor, assuming you're profitable at $0.04/M.
I suspect Jev was always intended to be an acquisition play and instead OAI said "nah, we're good" literally 2 weeks to the day after its launch.
I imagine investors/founders are horrified to discover there is no industrial-scale moat in no time at all... they're rumored to be in the early stages of a billion dollar round. Best of luck to that.
They say it's built on Luna, which costs $0.10M/in, vs Jev which only costs $0.04M/in, which is interesting ...
Based on how fast they rushed out a competitor (and the anti-Jev astroturfing), it seems that Jev is much more of a threat to OpenAI than they want to admit. Other than coding it seems the other high volume use case for LLMs is business automation, and for a lot of that all you need/want is indeed just a decision/classifier model.
None of this matters. The real news is the upcoming changes to the subscriptions.
OpenAI just rugpulled $200 plan users by slashing their usage in half and releasing a new $500 plan with the old usage plus a little more. Amazing, isn't it?
Half the usage for the same price, or slightly more usage for 250% of the price! Are you not entertained?!
This could and probably is me projecting, but the enthusiasm was tepid. The Opus 5.5 mog cast a pall on the room.
The stuff that came off as potentially exciting was the ultrafast stuff available only to Pro 500 users.
The reset button stuff was an embarrassment. Everything that was released today, in respect to the hype that the team projected on twitter, was a failure. The internal OpenAI slack is a hell of an echo chamber.
I also have a feeling that $200 was where most users will draw the line. $500 looks like such a slippery slope in the field where all the recent SOTA models, including Chinese ones, are considered good enough for most work.
$6k/year/seat is not particularly expensive for a dev tool if (very big if) you can measure the value and demonstrate it's worthwhile. That's the missing piece. People look at cheaper options because they don't know if the expensive one is working.
If OpenAI can show that a $6k tool nets $50k in productivity or quality or speed they have no problem selling it. Companies will take the easy-but-expensive option over a cheaper alternative. The problem is that's really hard to measure and seemingly they've put no effort into trying.
Didn't you see the crowd Wooo-ing and clapping when Sam Altman said "Today, we're announcing Dots"?
The crowd lost their minds before Sam even told them what Dots were.
He even called them out on it when announcing 6.1 Sol "You don't even know what it is yet"
Aren’t these crowds mostly people that work at the company? So you’re really cheering for the thing you built?
(I remember that being the explanation of why people would cheer at very lame things at Apple events.)
Yeah, the front dozen rows were reserved for OpenAI employees.
If that's true that restores my faith in humanity maybe they're cheering for stuff their friends were working on. Otherwise, the consumerism has reached a point where giving up on society and waiting for the next flood seems like a reasonable option.
> The crowd lost their minds before Sam even told them what Dots were.
Maybe they were cheering on The Muppets inspired AI avatars, resembling Kermit, Scooter etc.
https://www.theverge.com/ai-artificial-intelligence/1002033/...
It's all Anthropic's fault! :)
We pay them money so that they would put a new model in the bag.
No one mentioning Decisions API?
Feels huge. Sounds like Luna on their Ultrafast infra.
I did some internal benchmarking and found that GPT 6 Luna in batch mode at 20 records per batch was about 1.6x faster at 1.2x the cost of Jev with no tradeoff in accuracy. Tuning it up/down would make it faster per-record while also reducing the cost (assuming accuracy holds). Though this was only useful for offline processing.
It's another Jev copy, like we've seen so many over the last few weeks. But with no benchmarks or price comparison, which likely means it doesn't compare that well.
Sign-in with sub accounts seems big. This is how I thought this was always supposed to be - the providers acting more like an OS/ecosystem-enabler.
Extensions into codex also seem cool.
The provider also eating the application stack at the same time, doesn't make too much sense. I don't want my team going all in on one provider versus the other. (things like the docs/slides/sites).
It’s simply ridiculous. None of their 20 new features seem genuinely usable for day-to-day life; I’m struggling to find any use cases for them. I have no idea what they were thinking with "Dots". The "Super Fast" addition is nice and important, but it’s just one small bright spot in a hellscape of useless features.
This is a wild take to me, and I don't even use Codex! I've been on Cursor for a long time, but this is finally making me consider switching (especially since OpenAI models are leaving Cursor).
- Cloud Agents are awesome, super useful, I can't go back to working locally only
- Dots have a cheesy name, but they remind me of a more polished version of Cursor Projects, which I just started using this week. You talk to a single agent who serves as orchestrator and organizer, and spins up cloud agents as needed to do research and implementation tasks. So you can just dump stuff into the thread without worrying about interrupting ongoing work. Works great.
- Sign in with ChatGPT: how is this not major news? They're launching with partnerships with a bunch of harnesses so you can use your subscription outside of Codex.
- Slack integration: now that they have cloud agents, this is a no-brainer, cursor's slack integration is endlessly useful to me. I'm constantly tagging notifications from various services that get dumped into slack and having a cloud agent dig into more details or fix an issue
I could go on, 6.1 sol, ultrafast, etc. Wild that you think all this is useless. I guess we just work very differently.
> Sign in with ChatGPT: how is this not major news? They're launching with partnerships with a bunch of harnesses so you can use your subscription outside of Codex.
Maybe the sign on process is new but OpenAI has explicitly allowed subscriptions to be used in 3rd party apps at least as long as I've been a subscriber (a few months)
Oh gotcha, I didn’t know that. I thought it was both OpenAI and Anthropic that prohibited that.
"Cloud Agents" is just a blatant copy of Anthropic's offering. "Dots" is a copy of the Grok bot. Regarding the release of 6.1 Sol - I’m not dismissing it, and the model will likely be useful for many things, but come on: there is no way your flagship model doesn't compete with Sonnet 5.5. I can't say for sure, but there must be a reason they didn't show any benchmarks. The live demos were absolutely ridiculous. Sure, there are some useful features, but overall, it feels like they're missing the mark. Instead of simplifying AI usage, they're complicating it to the point where it's no longer clear where you're supposed to do anything.
I was talking about Cursor's Cloud Agents, which predate Anthropic's offering. Obviously they're all going in this direction, it's a no-brainer. I don't think calling it "useless" makes any sense just because Cursor and Anthropic were ahead of them there (if they were, I know nothing about Anthropic's cloud agent offering).
Well, it’s a bit of an exaggeration to say it was bad. It’s simply a matter of falling in line with the rest of the industry. I was expecting something of a higher standard or something new, so I’m disappointed.
Honestly this is me with almost everything the Ai companies release. I just don’t seem to have any problems that AI could solve in my personal life.
> I just don’t seem to have any problems that AI could solve in my personal life.
My microwave handle broke. I looked up the part needed to re-attach it; it's a a tiny piece of plastic. $30 each, and I need two of them, so $60.
So I grabbed photos of the part from online, which showed it against a 1inch square grid. I gave the photos to ChatGPT, told it the grid size, and asked it to make a 3D model.
It did it. I made a test print. It works perfectly.
¯\_(ツ)_/¯
I guess for me, it's been partly a matter of trying to unblind myself to problems I've grown accustomed to solving in a certain way.
Lots of good stuff flying under the radar.
- Subscription sharing is going to be huge for indie apps that don’t want to deal with token billing.
- Security scans built-in to Codex seems useful.
- Decisions API is a validation for Jev and the entire space it created.
The big takeaway is they announced decision model API for Luna
wonder how the jev team feels here
Validated, envigorated ?
The reaction to Jev has been huge, as has been OpenAI's speed here is announcing a "not quite ready yet" competitor... the idea of Jev is one of those "Doh! Why didn't I think of it!" moments ... in retrospect it's obvious that business automation is all about decision making and the market for this is massive. Up until now most people just saw this as a use case for LLMs, until TypeSafe/Jev came along and said "actually you don't need an LLM for this ...".
I'm sure TypeSafe realized the idea is a winner and expected competition, although maybe not quite so fast?
The question now is how seriously OpenAI take this, and are they willing to compete with Jev on price, thereby massively reducing Luna margins (which OpenAI's "Decisions API" is based on)? I guess if you can't make $0.10/M for Luna, then $0.04/M is better than just giving that market to a competitor, assuming you're profitable at $0.04/M.
I suspect Jev was always intended to be an acquisition play and instead OAI said "nah, we're good" literally 2 weeks to the day after its launch.
I imagine investors/founders are horrified to discover there is no industrial-scale moat in no time at all... they're rumored to be in the early stages of a billion dollar round. Best of luck to that.
Flattered, I assume.
They say it's built on Luna, which costs $0.10M/in, vs Jev which only costs $0.04M/in, which is interesting ...
Based on how fast they rushed out a competitor (and the anti-Jev astroturfing), it seems that Jev is much more of a threat to OpenAI than they want to admit. Other than coding it seems the other high volume use case for LLMs is business automation, and for a lot of that all you need/want is indeed just a decision/classifier model.
None of this matters. The real news is the upcoming changes to the subscriptions.
OpenAI just rugpulled $200 plan users by slashing their usage in half and releasing a new $500 plan with the old usage plus a little more. Amazing, isn't it?
Half the usage for the same price, or slightly more usage for 250% of the price! Are you not entertained?!
Why do those videos look like something out of Black Mirror?
i'm just glad i can use my Dootle Dots™ without context switching.
> Our friends at Hugging Face
Hahahahhahahaha
With friends like OpenAI, who needs State-sponsored adversaries?