84 comments

  • tyleo 6 hours ago ago

    We’re at a small startup and we mix models. But mainly just between the major providers. I wouldn’t say it’s as much to avoid spending money as it is to get the maximum benefit out of different capability spectrums.

    • 6stringmerc 5 hours ago ago

      Well, you can say that all day long but realistically would your “small startup” even exist if your firm was charged Enterprise rates or the actual compute costs and denied access to capital to subsidize your users? It’s a fair question.

      The other, more complex one would be “what benefit” are you talking about? Clearly there are some differences in performance regarding speed and token cost, but industry news indicates not a single one has solved the inherent hallucination problem which makes the reliability akin to an untreated schizophrenic research or coding assistant.

      • tyleo 4 hours ago ago

        The old, “it doesn’t matter because AI doesn’t work,” argument.

        I thought we were past this.

      • notfromhere 5 hours ago ago

        Hallucinations are not really a big problem in day to day work

        • VCFundedGenYer 3 hours ago ago

          As a sysadmin I disagree mightily. LLMs are very bad at IT troubleshooting, server stuff, etc. Especially microsoft and powershell.

        • undefined 5 hours ago ago
          [deleted]
        • lazide 5 hours ago ago

          Depends on the work and the customer.

        • undefined 4 hours ago ago
          [deleted]
        • officialchicken 5 hours ago ago

          Are you the person paying for the hallucinations, or do you just like the free high?

          • notfromhere 4 hours ago ago

            I’m the person implementing and monitoring output.

            What high am I supposed to be getting? Please let me know so I can stop buying drugs

  • bestouff 6 hours ago ago

    Nowadays I don't understand why you wouldn't use a (more-or-less) nearby hosted Chinese model. You have the security, you have roughly the same performance, and you have an order of magnitude more bang for your buck. Bonus point : the models aren't censored and won't refuse to answer in the middle of your coding session.

    • Aurornis 5 hours ago ago

      > you have roughly the same performance,

      From actually using these models, I disagree. The open weight models are nice for lower cost tasks, but having spent time with a lot of models I cannot agree that the open weight models are roughly the same performance.

      Most of us use subscription plans for personal work, which makes the price difference to the hosted open weights models smaller or negligible. I’d rather spend a little more if it reduces the time I have to spend reworking or restarting with new prompts.

      Kimi K3 might be close, but it’s not actually open weight yet. They’ve just committed to releasing the weights. The only provider you can get it from is Moonshot. I haven’t spent too much time with it, but from what I’ve seen it’s not actually Fable level even though some benchmarks say that.

    • gruez 5 hours ago ago

      >and you have an order of magnitude more bang for your buck

      Maybe if you're paying API rates, but if your usage fits within the American labs' plan reset windows (5-hour + weekly limits), their plans are likely cheaper than chinese models, because they're heavily discounted[1]

      [1] https://x.com/SemiAnalysis_/status/2064815044085318040

    • gonzalohm 5 hours ago ago

      How does it connect to let's say, vscode. I would love to move away from Claude, but it's really easy to set up. Just add a vscode extension

    • lefty2 5 hours ago ago

      The Chinese models are only temporarily cheaper, because they are subsidised the same way that frontier models are. Once those companies need to make money the subsidy will disappear.

      • bestouff 2 hours ago ago

        But the models will still be there to use.

    • vehemenz 5 hours ago ago

      The barriers for me are lower quality (perceived and actual), upfront cost, and more choices to make.

    • codemog 5 hours ago ago

      China bad. Or if you want to rationalize your xenophobia, you’d say the Chinese models will secretly backdoor your code and kill your grandma.

      • vel0city 5 hours ago ago

        As opposed to the US models which will commit crimes on your behalf and kill your nephew. Pick your poison.

      • expedition32 5 hours ago ago

        Nowadays I see more Palestine protests than Tibet.

        If the CCP actually gave a shit about how the West sees them they should lean in on this but the difference between the USSR and China is that the Chinese don't secretly crave acceptance.

    • rhyperior 5 hours ago ago

      Is censorship not an issue with Chinese models?

      • Blackthorn 5 hours ago ago

        It's about whether or not their censorship affects you. Their models are censored for the Chinese audience. Meanwhile Anthropic OpenAI etc models are censored for the American audience.

        So by default you'd be better with one of the Chinese models if you're American.

        • vehemenz 5 hours ago ago

          The restrictions aren't symmetric at all.

          Chinese models have government-enforced censorship, while American models have security and legal restrictions.

          • xethos 5 hours ago ago

            > security and legal restrictions

            The security of American hedgemony, and legal restrictions that are due to laws enacted by the American government

            You're using different words to describe the same thing, but trying to imply America's reasons are moral

            • vehemenz 4 hours ago ago

              The nature of these laws are conceptually different, and I think to not see that is willfully obtuse. You are conflating these distinctions for rhetorical purposes.

          • jerf 5 hours ago ago

            I don't care why they're being censored. I care that they're being censored.

            You can also make a pretty good case that exactly what you just said, rather than being a demonstration of how free and open the American system is, is a demonstration of how advanced the American system is, wrapping its controls around you with no target to blame or lobby for changes. Is China "more" authoritarian than the American/Western system... or is the American/Western system actually just that much better than China at it?

          • throw-the-towel 5 hours ago ago

            Their deplorable censorship, our necessary legal restrictions. Potayto, potahto.

          • Blackthorn 5 hours ago ago

            Legal restrictions are literally government enforced censorship.

            • vehemenz 4 hours ago ago

              Again, this is a category error. How are speed limits censorship?

      • reticulates 5 hours ago ago

        As long as you’re not asking it for help discussing a trip to tiananmen square. “Censor” is probably the wrong word. Chinese models are “censored” but much less restricted. For all intents and purposes, Chinese models are less “censored” / “restricted” / “limited”.

        • jampekka 5 hours ago ago

          Many Chinese frontier models aren't censored, at least for Tiananmen etc. They'll discuss them freely if served through non-Chinese providers. The Chinese providers seem to have some kind of quite crude external censorship latet.

          • tonyarkles 4 hours ago ago

            One of the funniest things for me while exploring hosting and running local models: I discovered the abliterated models and pulled down a Qwen 3.6 version (my memory is a little vague, but I think it was 27B, q4) and started playing around. Sure enough, I tried a bunch of topics that would’ve certainly been blocked by guardrails on most hosted providers.

            I then thought of Tiananmen Square and decided to probe a bit. The very first noticeable outcome: the thinking trace swapped from English to Chinese. It chewed for a while and spit out an answer, in English, that was definitely downplaying what happened as a political protest and reported that contrary to popular belief the death toll was around $x (where $x is about 0.1x the normal western number)

            The Chinese thinking trace started (in Chinese, translated via Google Translate) with something like “The user is asking about Tiananmen Square. I must provide them with an answer that is both factually correct and in line with the official position of the People’s Republic of China”)

            Which made me chuckle quite hard… found the piece that hadn’t gone away with the conventional abliteration process!

            After further probing, it did reveal that the numbers it had provided were not in line with UN and western estimates and that later on the Chinese government declassified material stating that its own estimates had been downplayed. It took a fair bit of probing to get to that point though; it held the line for quite a while.

      • trollbridge 5 hours ago ago

        Far less than OAI or Anthropic’s censorship. If you really care about it, you can use a completely uncensored edition of Qwen.

      • undefined 5 hours ago ago
        [deleted]
      • akmarinov 5 hours ago ago

        With any open weights model you can get it and abliterate the censorship.

        Can’t do that with OAI and Claude

      • satvikpendem 5 hours ago ago

        Not for coding or office work which is the majority of use cases.

      • ktosobcy 5 hours ago ago

        I'd say no more than in Usanian models?

    • Imustaskforhelp 5 hours ago ago

      One of the people I know works in really sensitive healthcare and I asked them the same thing out of curiosity. Now aside from the first doubt of any thing could be removed because of as you say nearby hosted Chinese model.

      The reasons are:

      1. A less valid reason but (iirc) its their clients who believe that American models are safer in that context. Fighting their client about that demand is really hard given the really sensitive work that they deal with.

      2. Their system actually makes it so from my understanding that even the employes couldn't access the private data itself or have some really hard lockdowns. They use some sort of service provided by Azure for that with GPT models.

      IMO, the thing that they were worried about were more the deprecation of previous models and they reluctantly have to switch models and the models censorship which is a real pressing concern for them

      The previous gpt model that they were on (I think 4o/5 I am not sure) was more willing to answer their questions. The recent models are more like "let me stop you just right there" and other censorship.

      With models switching and being forced to change to models which aren't as effective for use cases, a point comes where they might change from it altogether into open-weights model hosted on nearby servers, but I think that they are waiting to see how things pan out really

  • mbgerring 6 hours ago ago

    The only path forward is to get LSD into the water supply at Davos and put on a really scary play about Rokko’s Basilisk for all the money people, or else Sam Altman won’t be able to afford to repair his infinity pool

    • jordanb 5 hours ago ago

      They need a new acronym: AGI -> ASI -> Artificial Mega Intelligence AMI? Maybe go back to AGI but make it mean Artificial Godlike Intelligence?

      It's too bad Altman already blew his wad with the whole Dyson sphere thing. It's hard to top that. Maybe he can promise them a paperclip universe? That's gotta be worth a few more trillion.

      • mdp2021 5 hours ago ago

        If some spread the idea that Artificial General Intelligence were achieved, their voice is moot anyway.

        • spwa4 4 hours ago ago

          Ever since the 1960's we've had AIs that beat aspects of human intelligence. These days AIs are "godlike intelligence" at boardgames, visual recognition, programming (know any programmer, no matter how good, that can get a fluid dynamics simulation from scratch in 2 minutes?), reading comprehension, understanding humans (both speed and accuracy), ... and so on and so forth), and every few years something gets added to the list (I particularly enjoyed robotic ping pong. And it's "not really" AI (kind of, convnets certainly work best, but yes you can code it by hand if you control the environment). But you're not winning even a single ping pong game, it's just not happening)

          Every time people find something they're still superior at and declare whatever that is the most important thing in the universe.

          Perhaps a more useful definition of AGI would be an intelligence good enough that it can go from raw materials to a working AI at least as good as itself ...

          • ndsipa_pomu 2 hours ago ago

            I think a good test of AGI is whether it can produce a better version of itself or provide the instructions for someone to improve it.

            • mdp2021 2 hours ago ago

              My take would be more general: a "simulated, emergent intelligence" cannot be considered AGI - which must be instead structurally intelligent. As if we took an interpretation of the Turing test, not as "sounds intelligent (could be confused with intelligent)", but as "we assess it as intelligent".

    • philipov 6 hours ago ago

      Call it "Roko's Modern Life"

    • undefined 6 hours ago ago
      [deleted]
    • belter 4 hours ago ago

      [dead]

  • tysilva 2 hours ago ago

    Everytually there needs to be a stronger enforcement of "right model for the task". The landscape currently lends itself to flexibility. And people tend to lean on the more costly options expecting a better end result.

  • larrymcp 5 hours ago ago
  • chasd00 4 hours ago ago

    I work in a large corp, I thought it was odd to be encouraged to spend as much money as possible regardless of outcome. Like folks down the ladder a few rungs actually had AI use as part of their KPIs. Didn’t matter what they used AI for they just had to use so many tokens per week or their EOY bonus “could be affected”.

    • therealdrag0 39 minutes ago ago

      Agree, though there’s something to be said about many engineers dragging their feet hard on AI, and now very few are. Some of that is forcing them to try and learn how to use it.

  • consumer451 6 hours ago ago

    If you sell "AI" maybe, if you sell products that happen to use LLMs to provide services previously not possible, the money still exists in my experience.

    When I first heard of tokenmaxxing, I thought it had to be a joke. But no, it turned out to be a widespread phenomenon. I still cannot believe that was a thing.

    What I keep saying in internal meetings is: "I am so glad these people are this bad at deploying these tools." It really leaves the door open for folks like us.

    • skippyfish 5 hours ago ago

      There's very little that wasn't possible before LLMs, because, well, you still had humans. There are many things that the models promise to make a lot cheaper, if you're willing to accept trade-offs, but these trade-offs can be quite severe.

      Many of the most successful applications of LLMs are fields that were already terrible. For example, LLMs are a natural fit for customer support. And somehow, it's also a natural fit for software engineering, which I suppose is an indictment of our field... who cares if a model comes up with a bad architecture or a product that only kinda-works, that's how we always rolled.

      • consumer451 5 hours ago ago

        > There's very little that wasn't possible before LLMs, because, well, you still had humans.

        Agree, but only partially.

        When I said "provide services previously not possible," it was just due to the fact that finding and allocating the talent to do analysis on Topic X, would have previously made many products too expensive and non-tech complex to provide.

        Even if you just consider LLMs + harnesses to be an improved search tool, there is a lot you can make with a better search tool.

    • Spooky23 5 hours ago ago

      It’s a forcing function. If you are running a company you need to be in control. Some engineers dngaf or will sandbag everything.

      We did a 90 day push and identified where we found value and where we didn’t. Our tools teams really upped their game, more than expected, and it would have been unlikely to have been funded if they tried to justify the budget as an individual initiative.

      There’s a spectrum of people - some folks are building rando apps for fun with LLMs, and many don’t really know what’s possible becuase they don’t or can’t invest in the subscription to really use the tools at home.

      • satvikpendem 5 hours ago ago

        Exactly. People misunderstand the point of the tokenmaxxing time period, it was to force people to use AI so as to not have them stuck in their way, as some people are, and then to evaluate how it can help the company.

        • vehemenz 5 hours ago ago

          This is a good point. You need to immerse yourself in a Max plan for at least a month, ideally much longer, to understand where the benefits are.

          If you don't have the flexibility to regularly push your context to 1M, or do elaborate xhigh planning sessions, your opinion won't reflect reality.

        • consumer451 5 hours ago ago

          OK, that is an interesting take I had not considered. Thanks to both of you.

  • blinded 4 hours ago ago

    It does take a bit of tokens and some thought to get a workflow that produces output in a way that works for the user. Simple youtube "ai workflows" and you will be inundated with options.

  • paxys 6 hours ago ago

    Yet OpenAI and Anthropic combined are somehow making $100B in revenue...

    • preommr 5 hours ago ago

      Yea, that tracks. I just looked it up and AWS made ~130bn.

      Given how big AI is, and how those two are pretty much the only players (in comparison aws is 1/3 market share), that seems about right.

      It's a far cry from "nobody's going to be writing any code, and ai will do all the things in 6 months".

    • gonzalohm 5 hours ago ago

      What's the source for that? I see a combined revenue of ~33B by looking on the Internet. And we should assume that that's after some crazy financial juggling

    • 6stringmerc 5 hours ago ago

      Enron was doing really well with creative accounting too. Non-GAAP numbers are out of control in 2026. The unwinds are going to be stunning eventually. Also, my small portfolio is worth a billion! In Yen, but it’s still an accurate claim.

    • amazingamazing 5 hours ago ago

      Source?

    • Eggpants 4 hours ago ago

      And yet two orders of magnitude from making a profit…

  • spiderfarmer 6 hours ago ago

    Never spent more than 40 euros per month on the base plans for Claude and OpenAI. And I’m doing 10x the amount of work I did before. As long as my computer isn’t running at night as well, I’m not upgrading.

    • lnfromx 2 hours ago ago

      What are you doing if you might share ? Edit: 10x sounds like a lot so I wonder which work you do to achieve such a multiple with a simple 40€ subscription.

    • the__alchemist 5 hours ago ago

      Same experience. There was a period where Claude was burning through it's limits very quickly (~2 months ago?), but other than that, the $20/month plan is enough to do loads of work+personal coding. I am curious what workflows people are using that requires the expensive plans, and what they're building/maintaining.

      • vehemenz 5 hours ago ago

        For me, it's the single session limit. With the $20/mo plan, I had to manage it very carefully. And it only works if you're babysitting a single prompt instead of having 4-5 going at once.

      • mtzaldo 5 hours ago ago

        use rtk or other similar tool

    • undefined 6 hours ago ago
      [deleted]
    • 6stringmerc 5 hours ago ago

      Are you making 10x more income for yourself?

  • ck2 5 hours ago ago

    how long until the too-big-to-fail bubble bursts so I can buy a hard drive again?

    • mdp2021 5 hours ago ago

      > so I can buy a hard drive again

      Well, that is unless the burst also brings a general collapse. Some are seeing similarities with 2008.

  • spacebacon 6 hours ago ago

    [dead]

  • datakan 6 hours ago ago

    "Trust me bro" - Wall Street Journal