Switching from GPT-5.5 to GPT-5.6 Made Me Less Productive

(vincentschmalbach.com)

4 points | by vincent_s 10 hours ago ago

5 comments

  • esperent 9 hours ago ago

    > My codex sessions now easily drain a subscription within 24 hours.

    I have the same problem. But I think that codex subscription limits have dropped a lot recently. They were doing a promotion through April/May of 2x usage which is when I started using it. A $200 weekly limit lasted me 5 days easily.

    GPT 5.6 has increased usage a bit but I think it's not that much compared to 5.5. My guess is that there's usage limit shenanigans going on as well.

    I do agree it has a tendency to go beyond scope of you're not careful but I find if I'm leaving it run overnight, calling the plan a "PRD contract" and a looping message reminding it not to go out of scope reduces that to manageable levels.

    But yeah... A $200 plan that burns through it's weekly limit in a day, sometimes in 12 hours, is not sustainable. It seems Qwen / Kimi latest are at least equal to GPT 5.5 so as soon as someone else is offering those with a decent rate or subscription, I'm bouncing.

    • vincent_s 9 hours ago ago

      It definitely feels like subscription limits have dropped since GPT-5.6 came out. But I couldn't really find any evidence for it, going through my history. All I could come up with is that 5.6-sol uses about twice the tokens compared to 5.5 (both on xhigh) [0].

      One thing I just did was to stop four long-running Codex sessions that ran on 5.6-sol and switched to 5.5 and asked it "Please check if you're really going against the actual goal or have you drifted away from that?" and all four replied something like "Yes: I had started to drift"

      [0] https://www.vincentschmalbach.com/gpt-5-6-sol-xhigh-uses-twi...

      • esperent 9 hours ago ago

        I always set thinking to high for the main session or subagents doing audits, medium for everything else.

        I think xhigh or max has a place if I'm doing something genuinely complex. But in general use it's slow and I have the impression the highest often gives worse results - over-thought, over-engineered. High or medium might give better results for less complex tasks.

      • vincent_s 9 hours ago ago

        I have seen that kind of drift in smaller models (e.g. DeepSeek V4 Flash) when setting the thinking too high. So less thinking would lead to better results. But that's not something I'd expect from a SOTA model. Higher thinking effort should lead to same or better results.

  • Supercrzy 10 hours ago ago

    [flagged]