The AI Race Just Got Awkward

(insufferable.dev)

357 points | by allisdust 5 hours ago ago

422 comments

  • cmiles8 5 hours ago ago

    Why is it a problem that the Chinese labs are just distilling down Anthropic’s models? Aren’t Anthropic’s models not just distilling down other people’s work?

    Feels like Anthropic crying do as I say not as I do.

    • jedberg 4 hours ago ago

      What Anthropic is doing requires way more resources than what the Chinese labs are doing. So their complaint is that they do 95% of the work and the Chinese labs do the last 5% and call it their own.

      An argument can be made that Anthropic is also only doing the last 5% of the work (because the content they are training on was the other 95%) but that's a bit more philosophical.

      • JackFr 4 hours ago ago

        But the analogy still holds.

        The original authors of all the text, creators of the media and developers of the software did far more work than Anthropic.

        • jdonaldson 4 hours ago ago

          Yeah, the whole thing seems like a human centipede of rug pulling. Probably the same as it's always been. Curating AI knowledge should be something that we put our best researchers towards, but realistically I think we wind up with 2-3 highly biased nationalistic models that are constantly copying off each other's notes.

          • rubicon33 4 hours ago ago

            Thanks, that’s the nature of any business. Founders see a way to take existing knowledge and expertise, combine it in some novel or interesting way, and produce a new product

            • mitthrowaway2 4 hours ago ago

              Napster was a fantastic and disruptive product, the likes of which arguably has no equal to this day. But eventually the hammer came down from the courts and it was replaced by streaming services like Netflix, which pay to license materials from their creators.

            • tsunamifury 3 hours ago ago

              This is a rubbish lossy statement. And reductionist to the point of nothing has meaning.

              Did Anthropic put work in? Yes. Did they derive their value from Humanity being open with knowledge then try to sell it back? Also yes.

              Did they even steal the tech? Also yes.

        • jedberg 4 hours ago ago

          But if you take it deeper, didn't most of those authors rely on the work of others? Most of human knowledge is small advancements of things we already knew. Often by reorganizing what we already knew.

          Is that not what the foundation models are? A new reorganization of existing knowledge?

          • gretch 4 hours ago ago

            Yes this is all true.

            So then the problem is that Anthropic seems hypocritical when they knowingly insert themselves into this chain, and then complain about people down-chain from them.

            To remedy the negative impressions (if they even care to do so) they should do 1 of 2 things: 1) stop complaining about it 2) stop distilling other people's work

            • ToucanLoucan 4 hours ago ago

              They insert themselves into this chain for profit and complain about it. I really think that adds a thick layer to the hypocrisy that people, or at least me, feel is especially distasteful.

              No LLM products would exist without the avalanche of largely non-consensual use of IP to create them, full stop. Any of these companies doing this and then turning around and complaining when their IP is "breached" are going to met with a chorus of tiny violins.

          • wonnage 4 hours ago ago

            quite the leap from “authors rely on the work of others” to vacuuming up the sum total of digitized knowledge to tune some matrices

            • ryandrake 2 hours ago ago

              Often people ignore scale. N=1 is OK, therefore, N=1billion is OK. Same flawed argument as: "It's OK for one police officer to watch one street corner for the purpose of observing crime; therefore it's equally OK to have cameras recording every street corner in the city 24/7, for all purposes. Same thing!"

          • scythe 4 hours ago ago

            It is an ancient practice, that when a human creates something, other humans will observe it and learn from it. Every group of humans living together has practiced this in some form for tens of thousands of years if not longer. Even animals do it. It's a natural assumption when making any form of art.

            It is not a natural assumption that someone will digitize the artwork and use it to adjust a couple thousand matrix coefficients in a complex computer program. To most people that seems like copying with extra steps. The brain may in some ways resemble a computer, but what sets it apart is that we have always lived with brains. Everything a human does has already anticipated the presence of other brains, while etched circuits on ultrapure silicon crystals are something new.

          • iAMkenough 4 hours ago ago

            It’s a few corporations stealing work from others, to sell it back to us. That’s it.

            • kbelder 4 hours ago ago

              But it's selling it back to us cheaper and with more utility. That's not something to sneeze at.

          • xdavidliu 4 hours ago ago

            not at that scale though

        • marshray 4 hours ago ago

          Don't forget the mothers of all those original authors, as well as everyone who labored to build and sustain the societies which produced writers.

          • eithed 4 hours ago ago

            It all comes full circle with chinese models being available free for all humanity to use

            • dr_dshiv 4 hours ago ago

              Beautiful, right?

        • jstummbillig 4 hours ago ago

          To a degree. The human produced knowledge is the product of all humanity (no human is an island).

          A comparable idea could be that an encyclopedia or maths book is only distilling the things that other people did, and how dare they sell them. But the "only" is doing quite a bit of work. LLMs do not just spawn into existence. There is a body of work that they feed on, and then there is also very attributable work they do around and on top of that. All labs are struggling around the first order question: Is it okay to use prior work like this? The second order issue is still entirely reasonable to separately have and enforce rules about.

        • agumonkey 4 hours ago ago

          Kinda agree statistical modeling relies on all the hard work, effort, passion and risk taken from just about everybody.

        • mc32 4 hours ago ago

          All the text and so on had unrealized potential. Without Anthropic et al it would remain unrealized.

          It’s like FTL. Until someone realizes it, it’s just talk.

        • ihsw 4 hours ago ago

          [dead]

        • nonethewiser 4 hours ago ago

          He didn’t say it was an analogy. He said both are distillation.

        • layer8 4 hours ago ago

          SOTA models cost hundreds of millions to train. Did creating the contents of the text corpus they were trained on really cost an equivalent of 20x as much (~10 billions)? I honestly don’t know, but I could imagine it having been significantly less.

          This isn’t meant as a moral argument, just musing about the relative cost comparison.

          • jonhohle 4 hours ago ago

            If you look at movies alone that would easily surpass 10s of billions. The cost of most books is probably more nebulous, but books, research, and more all have time and money spent to create them. I would guess the corpus of all media from the 20th century on would be minimally in the hundreds of billions of dollars.

            • layer8 3 hours ago ago

              LLMs aren’t trained on movies, though.

              Image/video models are, but those weren’t the topic.

          • louiskottmann 4 hours ago ago

            Given they ingested basically the whole internet and then some, you cannot possibly be serious when you mean it's worth less than 10 billions.

            The totality of the content on internet is worth several orders of magnitude more.

            • layer8 3 hours ago ago

              The argument wasn’t about how much it’s worth, but about how much it cost to create. These are very different things.

              • allturtles 3 hours ago ago

                Why? Are we doing labor theory of value now?

              • tpm 3 hours ago ago

                do you also count eg published results of very expensive physics experiments? because once the costs of things like these are taken into account, we are way over 10 billions.

          • MathiasPius 4 hours ago ago

            I would argue that producing the complete written corpus on which they at least intend to train (even if some is still out of reach) cost literally everything to produce.

            And the monetary cost doesn't even register when weighed against the blood, sweat and tears that went into capturing the authentic experiences of real human beings, whose honest expressions are now at least in some cases getting hoovered up, ingested, and then destroyed for all eternity, for fear that this specific work is the rounding error that might give an equally immoral competitor the edge in the bicycle-riding flamingo race that is currently consuming an absurd amount of the world's creativity and attention.

          • phamilton 4 hours ago ago

            Simple math:

            A training set of 15 trillion tokens is 10 trillion words.

            A penny a word is cheaper than the cheapest beginner freelance writer.

            That makes a training set of 10 trillion words cost $100B.

            Lots of assumptions there for sure, but we're certainly in the ballpark you are describing.

          • rsingel 4 hours ago ago

            I asked Claude to estimate the cumulative salaries of US only journalists over the last hundred years:

            $500B for all kinds including TV and online

            $300B for newsrooms including all staff

            $140B for newsroom reporters only

            So yeah, I think the price of the information ingested is way higher than training costs

          • Enginerrrd 4 hours ago ago

            Yes, Easily, and by multiple orders of magnitude.

          • Ohentis 4 hours ago ago

            I suspect so. It is a lot of data. You're looking at essentially all publicly available (and some non public) intellectual work.

          • Ar-Curunir 4 hours ago ago

            Yes, duh! Human output across the millennia is worth much more than whatever is being invested in frontier labs.

            How is this even a question.

            • bmacho 4 hours ago ago

              They are solving unsolved math problems right now, so probably soon or very soon their output will be more valuable than all human recorded knowledge.

              • Ar-Curunir an hour ago ago

                What do you think mathematicians were doing for centuries before LLMs?

                And also, humans have been doing a lot more work than just mathematics...

              • mekael 2 hours ago ago

                Pre vaccination smallpox killed hundreds of millions of people just in the twentieth century [0], the knowledge that allowed for the creation of just that vaccine is worth hundreds trillions of dollars in humans lives, let alone all of `the knowledge and experiences those people were involved in.

                The knowledge that created the Haber-Bosch process [1] helps to sustain the majority of the world's populous, add another five hundred trillion dollars for that just to start with.

                The creation of the printing press and all written information that allowed it to be built provided dissemination of knowledge beyond the ultra wealthy and is worth a non-finite amount of money.

                LLM's are cool math, but they are less than a rounding error in comparison to even the tiniest sliver of human knowledge and technological output.

                [0] https://pubmed.ncbi.nlm.nih.gov/35143880/ [1] https://cen.acs.org/food/agriculture/The-industrialization-H...

          • wonnage 4 hours ago ago

            this is the sort of brain rot thought that you have in a dorm room the day you are introduced to Econ

            “bro like, what if we could price the sum total of human knowledge? That wouldn’t be that much, right?”

      • cmiles8 4 hours ago ago

        I get that angle but it’s a weak argument as Anthropic is doing the same to others. Also while there’s certainly a lot of computing power needed to do what Anthropic does, it’s increasingly clear there isn’t much secret sauce involved. Everyone knows how do to the core work it’s just a question of who wants to burn billions on compute to do it.

        Anthropic’s anger here seems mostly rooted in their annoyance that this exposes they don’t really have core IP that’s not just easily replicated. And that’s clearly a problem for a deeply unprofitable company trying to convince people they’re worth $2 trillion.

      • dofm 4 hours ago ago

        > What Anthropic is doing requires way more resources than what the Chinese labs are doing.

        Oh that’s very sad.

        Meanwhile Anthropic made a product from the work effort of millions of people without compensating them, sell that product on tap and unless I am mistaken do not even have their competitors’ cover of having released any sort of meaningful open weights model.

        They have taken from culture (including very specifically their most direct customers’ specific culture — our culture), turned it into a machine to make themselves rich, appear likely to predicate their valuation on permanently removing people from the workforce, then want to dump themselves onto pensions funds and ordinary savers to carry the bag.

        It is, I agree, philosophical, because karma is a philosophy as well as a bitch.

      • bushbaba 4 hours ago ago

        And the communal work of humanity is orders of magnitude more work than what anthropic pays for their scraping of content. I got no check from them for my contributions

        • undefined 4 hours ago ago
          [deleted]
        • toomuchtodo 4 hours ago ago

          Indeed, if it isn't a crime to train on humanity's data, it isn't a crime to train on capitalism arranged frontier LLM provider models. Is that bad for shareholders and capitalism? Meh, sounds like a suboptimal socioeconomic systems issue. Burn up all the capital the unsophisticated are willing to provide. “We are selling to willing buyers at the current fair market price.”

          With my apologies to Brewster Kahle, "Universal Access to All Knowledge."

          https://www.youtube.com/watch?v=RV_ALlJGU_c

          • jacquesm 3 hours ago ago

            I have far less of a problem with the Chinese models if they even do this because they are making their models free, whereas the large Western LLM providers throw out a few bits but not their main work product. So not only are they hypocritical, I'm pretty sure that if the Chinese models were not released into the wild they would be making less noise.

            • toomuchtodo 3 hours ago ago

              Oh yeah, totally agree, I'd even rather pay those building the Chinese models if I didn't think I'd get thrown into a US gulag for felony contempt of business model.

          • abcthingx 4 hours ago ago

            This feels like a straw man argument. The parent comment didn't say it was a crime

            • toomuchtodo 4 hours ago ago

              I use crime in the broad sense of "You shouldn't be allowed to do that" in this context. If you have a better word to capture that thought, let me know, I'll make the edit ("frowned upon" perhaps?). I don't have strong feelings other than "hah AI companies aren't going to be able to create a moat to capture the value they want to capture because we can collectively keep pulling it out of their models in perpetuity through ever improving model distillation methodologies". This is no different than Uber and DoorDash using VC dollars to subsidize services until they try to turn the knob to profitability once they've captured the market, except in this case, there are mechanisms to exfiltrate the model value into open models that can be distributed at very small marginal cost. They can never gate the golden goose money printer, they can only complain it isn't fair they aren't able to.

              "The Spice must flow."

          • voiceofchoice 3 hours ago ago

            Bubble pop bubble pop

      • faangguyindia 4 hours ago ago

        Isn't it better for planet? By not doing the wasteful transformation work again

        • jedberg 4 hours ago ago

          Absolutely. I'm not taking a side here, I'm just pointing out why Anthropic might have a valid complaint.

          • isolay 4 hours ago ago

            That complaint is invalidated by the argument of tu quoque. Complaining about something they are doing themselves.

      • orbital-decay 3 hours ago ago

        Distillation doesn't "grab 95% of lab's work", that's ridiculous. At best it's tiny icing on top of the cake that's already there. It's not even necessarily done on a better model (e.g. GLM 4.7 distilled Gemini 2.5, a weaker model), I'm pretty sure A\ and OAI could do (or even do) the same with greater efficiency since they have access to logits, weights, and internal state of open models.

        >An argument can be made that Anthropic is also only doing the last 5% of the work (because the content they are training on was the other 95%) but that's a bit more philosophical.

        How is this philosophical? They should release the unsupervised pretrains, at the very least.

      • tene80i 4 hours ago ago

        But that’s not more philosophical. It’s a perfect parallel! Enormous amounts of work, vacuumed up and resold. What’s the difference? If it’s ok to vacuum up all the knowledge in the world, then that includes knowledge of how to use all that to power an LLM.

      • baxtr 4 hours ago ago

        Wait, wasn’t 95% of the work creating the content in the first place?

        • jacquesm 3 hours ago ago

          No, it was closer to 99.99%.

      • aleqs 2 hours ago ago

        Yeah, you need a lot of resources in order to waste a lot of resources. Look at codex and Claude code - these trillion dollar companies 'with top talent' cannot build what open code and pi/oh-my-pi have built in the open for free? Both codex and Claude code, are slow, buggy pieces of shit (and I say that as someone who still heavily uses both for work, moved to open code and pi for personal stuff). The reality is these companies mostly focus on marketing and market capture through, non-competitive means - their services and software are unreliable, buggy trash.

      • OtherShrezzing 4 hours ago ago

        Can you elaborate on why the second is “a bit more philosophical”?

        I see absolutely no distinction between the two, aside from minor technical approaches to gathering the content.

      • darkmighty 4 hours ago ago

        > but that's a bit more philosophical

        It sounds exactly the same, not more philosophical to me, except one is more inconvenient.

      • Andrex 4 hours ago ago

        Conventional wisdom is Google did all the groundwork with LLMs...

      • jklinger410 4 hours ago ago

        It is kind of ironic that they scraped the web for publicly available data and used it freely to train their models and now their freely available models are being used to train other models.

      • meowface 4 hours ago ago

        I'm overall pro-Anthropic and pro-banning open-weights AI, but I agree with the parent commenter; distilling Claude models is not that different from pretraining on web data. It's all basically the same sort of thing.

        I think a good litmus test here would be if Anthropic were to not care about distilling their models when the distillers keep the resulting models closed-source and sell tokens via an API. If they cared only about security concerns and not about people profiting off of their work, then they should be publicly fine with this and only protest against it going into open-weights models.

      • jasondigitized 4 hours ago ago

        Sounds more like a Western vs. Eastern outlook on innovation and how you accomplish it.

        • orbital-decay 3 hours ago ago

          Deepmind was indirectly distilling Claude 3, XAI was doing this to other models (with Musk shrugging it off like something unremarkable, which it is), it has nothing to do with nebulous stereotypes like East, West, China this, America that. It's mostly Amodei and Altman screeching over this fact.

      • Henchman21 4 hours ago ago

        Is hypocrisy a philosophy?

      • arctic-true 4 hours ago ago

        They spend 95% of the money, perhaps, but burning compute is not the same as doing the work.

        • jedberg 4 hours ago ago

          I'm calling "work" here the conversion of energy to LLMs.

          • undefined 4 hours ago ago
            [deleted]
          • AlexandrB 4 hours ago ago

            I think more energy was spent creating the original works than training the LLMs on them. Not just energy but blood, sweat, and tears as well.

      • thrance an hour ago ago

        > An argument can be made that Anthropic is also only doing the last 5% of the work (because the content they are training on was the other 95%) but that's a bit more philosophical.

        Actually, it's an interesting argument to make. How many labour-hours went into creating the training data Anthropic has collected? Probably multiple billions of hours. How many labour-hours did it take them to setup the datacenters, scrapers, and training algorithms? A few thousands hours?

      • mrwh 4 hours ago ago

        I mean, 95% of the work if you don't factor the work to create the training data in the first place...

        • cyanydeez 4 hours ago ago

          also, the actual work is the _copyrighted material created by the world_.

          • mrwh 4 minutes ago ago

            Indeed! Basically all of human civilization up until this point

        • off_with_their_ 4 hours ago ago

          [dead]

      • watwut 4 hours ago ago

        You are going to be surprised to hear how many resources were necessary to create all the data Anthropic is digesting

        • ipsod 4 hours ago ago

          From one perspective, the 5% estimation is near-infinite orders of magnitude off, since they've trained on something approaching the sum-total of human knowledge.

          • freejazz 4 hours ago ago

            > human knowledge

            that's an interesting way to describe reddit posts

      • AlexandrB 4 hours ago ago

        Lol, no. The original authors of all the text Anthropic took in did 95% of the work, Anthropic did 4% of the work and the Chinese labs do the last 1%.

        • BigTTYGothGF 3 hours ago ago

          I'd split it at 99.98% original, 0.015% Anthropic, 0.005% Chinese, and that's being exceedingly generous to the AI companies, there should be several more 9s and 0s in there.

      • dancemethis 4 hours ago ago

        So Anthropic and other US AI companies... stole harder, and therefore deserve more?

      • freejazz 4 hours ago ago

        > What Anthropic is doing requires way more resources than what the Chinese labs are doing.

        And writing a book requires many more resources than what anthropic does

      • tsunamifury 4 hours ago ago

        “We’re both thieves, Steve. We both stole from xerox. You’re just mad I got there first.”

      • koickong 4 hours ago ago

        [flagged]

    • seizethecheese 4 hours ago ago

      The conversation here is mostly moral and ethical but the problem here seems to be financial.

      Anthropic and OpenAi are spending a $$$$ to "distill" human output into an AI model, then others are spending $$ to distill their AI model into a near-equivalent model.

      This is the same reason IP rights exist. On the surface, something like a patent feels ludicrious and even feels morally wrong. Some guy wrote down the recipe for arranging atoms or bits in a particular way, and now I can't!? However, it's designed to solve the same problem, figuring out and describing the process is much more costly than replicating it.

      • ASalazarMX 3 hours ago ago

        AI training, if viewed through the capitalist mindset, is plain theft. Anthropic can't morally defend copying someone else's IP, but denouncing others copying Anthropic's stolen IP.

        That doesn't mean they won't try, and that also doesn't mean they won't succeed.

      • AlexandrB 4 hours ago ago

        Anthropic and OpenAI are very happy to ignore the IP rights of others, so I'm not sure how they can ask for any kind of IP protection themselves. Live by the sword, die by the sword.

      • failbuffer 4 hours ago ago

        Capitalism: moral rights exist when they give us a moat.

      • freejazz 4 hours ago ago

        Model weights wouldn't be covered in a patent. You could patent a method of creating weights in a model, but you couldn't patent the weights themselves.

        I wish people here could at least bother to inform themselves about the IP rights they are so quick to insist are abhorrent, when they seem to not even have a first clue as to what they actually cover.

    • jrflo 4 hours ago ago

      Because cost of original training >> cost of distilling. It's the same thing that happens with Chinese knockoffs of physical products - it takes a lot of money and R&D time to design a new product, but it's basically free to buy the product, reverse engineer it, and resell it. All the data they originally trained on was available for free on the internet. If the original work was so valuable, it shouldn't be up on the internet for free in the first place imo.

      • ASalazarMX 3 hours ago ago

        Caveat: cost of creating human knowledge/art >>>>>>>>>> cost of original training >> cost of distilling

        You could say Anthropic distilled human knowledge and art.

        • jrflo 3 hours ago ago

          Right, but the humans willingly released all those creations for free. I think that my issues with the "AI companies stole human creations" stance is that the information was freely available to everyone, and they put a lot of money and effort into transforming it into something useful.

          • ASalazarMX 3 hours ago ago

            > but the humans willingly released all those creations for free

            How can one answer this statement in good faith? AI companies literally violated IP by massively pirating works instead of legally licensing them.

      • AlexandrB 4 hours ago ago

        It's "free" as in beer, not free from copyright. LLMs are free from copyright on the other hand. So which is more "free"?

      • wonnage 4 hours ago ago

        Sounds like Anthropic should close up shop then, those chumps are offering their product on the internet for any random loser to distill

        • jrflo 3 hours ago ago

          It's different because you have to pay Anthropic and follow their TOS to get access to their model. If it was actually on the internet for free, anyone could do whatever they wanted with it.

    • undefined 4 hours ago ago
      [deleted]
    • layer8 4 hours ago ago

      Two wrongs don’t make a right. (If you consider them as wrongs.)

      • OneLessThing 4 hours ago ago

        It's not that the Chinese companies are right, it's that Anthropic has no place to complain about stealing.

        • layer8 4 hours ago ago

          The root comment was asking how it is a problem. If one considers it wrong, then it’s a problem regardless of whether Anthropic is complaining or not. Anthropic’s complaining or non-complaining should have no bearing on whether it’s considered a problem or not.

          • AlexandrB 4 hours ago ago

            The difference is in the solution that would be proposed. I'm sure Anthropic wants to create some kind of IP protection regime for their model so it can't be distilled. I want their model to be public domain, since they trained on material that was not theirs to begin with.

    • cyanydeez 4 hours ago ago

      Because no one outside the AI scientists understand what distilling means. They probably all think about Mash and a vodka still, and a completely unrelated association.

      The word itself is the pivot, not anything else.

      • hn_throwaway_99 4 hours ago ago

        > They probably all think about Mash and a vodka still, and a completely unrelated association.

        I don't know anyone with even a passing understanding of how LLM training works that thinks that is the appropriate analogy.

        • pdntspa 4 hours ago ago

          It isn't, that is whole point. A normie hears that word and they think vodka.

        • cyanydeez 4 hours ago ago

          cool, do you think these media representations are for you, or 99% of the people who would love China to be sanctioned because they're foreigners?

    • bionhoward 4 hours ago ago

      “Distilling” is a funny way to say “learning from”

    • sergiotapia 4 hours ago ago

      "That’s called competition. You’re allowed to test somebody else’s products all you want." - Jensen Huang https://x.com/wallstengine/status/2104604118937735553

    • dominotw 4 hours ago ago

      [flagged]

      • cmiles8 4 hours ago ago

        Then Anthropic should stop saying it. So long as they try to play victim here folks are going to call out their BS.

    • nater5000 4 hours ago ago

      [flagged]

    • jorblumesea 4 hours ago ago

      $$$$

      it's not complex. there's hundreds of billions of investor dollars counting on vendor lock in and walled gardens

      • 2OEH8eoCRo0 4 hours ago ago

        It ain't gonna happen. At work I have a dropdown menu in vscode with a dozen models to use interchangeably. They're all essentially commodities and will compete on price and squash almost all profit margin.

        • dpweb 4 hours ago ago

          That's not their business model. They won't win on price, but they won't compete on price. Their business model is making the current state of the art.

          If I'm a business and I need something done today, and bc Anthropic has the best model, there's a 99.9 chance it will be completed successfully for $1000. And using Deepseek there's a 70% chance it will, for $10 - you or me will go for the $10. Big businesses don't. Bc 1000 per task is nothing to them.

          • bushbaba 4 hours ago ago

            Actually opposite occurs. Big businesses are ok with a mediocre but cheaper result. Very few are willing to pay such cost. Just look at tech wages and the distributions

          • HWR_14 4 hours ago ago

            Yes, large corporations frequently pay orders of magnitude more for slightly better software. That's why Oracle produces the best stuff on the planet.

            The real issue is that Deepseek has a 99.7% chance. So I can run it 10 times until it works and still pay 1/10 the money.

          • rootusrootus 4 hours ago ago

            The business I work for is absolutely sensitive to 10 vs 1000, depending on the task. And it's a multi-billion dollar business. 1000/task may not be much on it's own, but there are a lot of tasks.

            Also, is it really 99.9% vs 70%, or 99.9% vs 99%?

          • andrew_lettuce 4 hours ago ago

            Big business doesn't pay more for better, but the do pay more for predictability, support and targeted outcomes. They will happily trade a chance at 100% better results for 10% less chance of unplanned outcomes

          • thadt 4 hours ago ago

            That 70% chance of success goes to 99.9% in 6 repetitions.

            Big businesses might pay $1000 vs $60 for certain tasks, but that won't work out well at scale.

          • cmiles8 4 hours ago ago

            Except businesses are going the opposite direction here. The lack of stickiness makes the “premium” argument hard to play. Oracle won because swapping databases is a giant PITA. Swapping models requires almost no effort for most uses. And because of that enterprises are all building model marketplaces where providers have to compete on price performance.

            Most folks I know can choose from any of the big labs or open weight models and they get billed internally for tokens against their budget. There’s little incentive to no switch to the lower cost closers.

            This setup is a nightmare scenario for the big labs trying to execute the traditional enterprise sales plays. Those only work if your product is sticky and AI models are one of the least sticky things in the history of tech.

      • teaearlgraycold 4 hours ago ago

        Sorry but it’s looking more and more like the top American labs won’t have any kind of moat.

        • jorblumesea 3 hours ago ago

          why sorry? I agree with you and think it's good for the industry and the world on the whole

          why should sammie or darigold have the keys to the kingdom?

    • nonethewiser 4 hours ago ago

      >Aren’t Anthropic’s models not just distilling down other people’s work?

      Can you elaborate on that? I mean my direct answer would be no, of course not. But why do you think frontier models are distilled? I think maybe there is an equivocation over the word “distillation.”

      Frontier labs train on their own pretraining data, human feedback, synthetic data, and research. A distilled model is specifically optimized to reproduce another model's behavior.

      Meanwhile R1-Distill-Qwen-32B was distilled from DeepSeek-R1.

      If you want to say a frontier model is "distilled" from the world's data and R1-Distill-Qwen-32B is distilled from DeepSeek-R1 then you are equivocating two very different things.

      • nuancebydefault 3 hours ago ago

        They meant distilling in a more original sense, not per se in the LLM-era meaning of the word sense.

  • slowin 5 hours ago ago

    I'm also grateful to the Chinese labs for providing workarounds for the walled gardens that the US based AI companies are attempting to create.

    Does anyone know if there are any distillation datasets available? I'd love to see these distributed on BitTorrent. I think it's critical that AI be democratized and not isolated in the hands of a few private companies.

    • 10xDev 4 hours ago ago

      An authoritarian regime is not your friend and will pullback the moment their own models become highly capable.

      • pksebben 4 hours ago ago

        Oh no, they might stop doing the thing that benefits me and that they were never required to do in the first place.

        • stackedinserter 4 hours ago ago

          It can be predatory pricing that will bite us later.

          • tornikeo 3 hours ago ago

            How exactly will downloaded gguf files bite me later?

            Is it going to shut the laptop's lid when i'm not looking and pinch my fingers?

            • stackedinserter 41 minutes ago ago

              If the only thing that bothers you is laptop lid pinching your fingers, then no – it won't do it.

          • wat10000 4 hours ago ago

            It really can't be when the models are open weights and there's zero lock-in for inference providers.

      • computerex 4 hours ago ago

        As opposed to what? The US? You think the US is any different? Literally our pedo president publicly admits to insider trading. You think the US government gives a rat's ass about the American people?

        • Jskewel 4 hours ago ago

          Saying the Chinese and US governments are the same is so absurd it doesn't warrant a reply.

          • xtracto 2 hours ago ago

            For someone (as me) in Mexico, they are exactly the same.

            Third world societies are not "brainwashed" one way or another with respect to that.

            We just see those governments throwing shade at each other. And bullying us ina similar way.

          • computerex 4 hours ago ago

            Then why did you respond with a useless self aggrandising comment?

            Love the arguments btw. Your position is so cut and dry that you can’t even support it with evidence.

          • voiceofchoice 4 hours ago ago

            It's true, at least the Chinese government is competent in their authoritarianism.

        • dirck-norman 4 hours ago ago

          The independent media reports on this and that’s why you know.

          China has zero independent media and has the largest and most sophisticated censorship and surveillance state in the world.

          Trump would love to have the absolute power that Xi has, but thankfully doesn’t.

          Tired of this lazy whataboutism.

          • ted_bunny 2 hours ago ago

            Whataboutism is the rallying cry of every sophist who can dish it but not take it. Whataboutismism

            • computerex 2 hours ago ago

              My comment doesn't even fall into the category of whataboutism....

          • computerex 4 hours ago ago

            No, actually trump bragged about it on x.

            Trump kicked and banned major news orgs from the White House just now. You’d have to be will fully ignorant or really naive to believe we have fair and just media.

        • htx619 4 hours ago ago

          [dead]

      • slowin 4 hours ago ago

        There's no "pulling back" things that have already been open sourced.

        • abirch 4 hours ago ago

          Unfortunately most of the Chinese models are open weights and not open sourced.

          • slowin 4 hours ago ago

            I agree it would be very cool if they were open sourced, but the article is talking about the KV cache technique which if I understand correctly is open source (or at least the white paper is published).

        • undefined 4 hours ago ago
          [deleted]
        • HappyPanacea 4 hours ago ago

          Intelligence wants to be free

      • voiceofchoice 3 hours ago ago

        Every workplace is an authoritarian regime where workers don't have a say and grovel with "please don't replace me" like that isn't the entire point.

        The whole point of AI is to get rid of you so rich people can play with the planet like it's minecraft. Software engineers just think they're special because they're the ones building it like they'll get a pat on the head for being good little servants to the investor class. Or worse, that their portfolios will let them be the gods who rule over ashes.

      • horsawlarway 4 hours ago ago

        Yes, we already discussed the US.

        • Avicebron 4 hours ago ago

          It's crazy how articles like this get spawn-camped by people like this trying to throw this zinger in. Both AI conpanies in the US and the chinese companies with ccp desks in the corner can be bad. The good path forward is locally hosted AI models, that's known.

          • horsawlarway 4 hours ago ago

            Right, which is why I'm happy to see China continue to innovate in the open, and increasingly wary of the US stance given articles like

            https://www.anthropic.com/research/glm-5-3-and-the-spread-of...

            It's VERY clear that the US companies are trying to push for regulation to kill open models and open weights. I see this as much more hostile and authoritarian response than what we're seeing come out of China right now.

            So is China going to always publish in the open? No clue. But right now they're modeling much better behavior.

            • jacquesm 3 hours ago ago

              What's on my drives stays on my drives.

      • sixo 4 hours ago ago

        It is in the interest of everybody-except-OpenAI and Anthropic that those companies have capable competitors; better still for the competitors to share their advances.

        They don't have to be our friends to act in our interest.

      • unrented7977 4 hours ago ago

        Agreed, the US is not your friend and will pull the rug from under you at the first convenient opportunity.

      • randbyte 4 hours ago ago

        Just like how Anthropic and OpenAI is already doing?

      • undefined 4 hours ago ago
        [deleted]
      • nutjob2 4 hours ago ago

        China is much more authoritarian than the US, but at this point it's like comparing two types of metastatic cancer.

        The point is get what you can from both to develop open models, data and tools.

      • layer8 4 hours ago ago

        You can be grateful to individual things an enemy does.

      • dgellow 3 hours ago ago

        Ironically the US is the only country that has done that

      • 4gotunameagain 4 hours ago ago

        While your friend is Sam Altman, or US megacorps ?

        Or did they not pull back when their models allegedly became highly capable, with the whole mythos debacle ?

      • caaqil 4 hours ago ago

        > An authoritarian regime is not your friend

        Absolutely. In fact, the authoritarian regime already did try to force export controls on major frontier labs not long go so this isn't a theoretical.

      • undefined 4 hours ago ago
        [deleted]
      • CodingJeebus 4 hours ago ago

        This is equally true for US AI

      • ahriad 4 hours ago ago

        Chill, Buddy. Why are you so anti-American?

    • ducktective 4 hours ago ago

      >distillation datasets

      You ask about distillation but I wonder, is there any training datasets (~ TB-order) available that startup folks in SV use or is it so that everyone has to create their own scraping pipeline ?

      • forshaper 4 hours ago ago

        There are several? And there exists companies whose entire business is just providing them? iirc

      • frabcus 4 hours ago ago

        It's particularly important we all scan and destroy our own unique books!

      • atherton94027 4 hours ago ago

        Given the amount of people complaining about crawlers in the past 2 years, I think it's the latter

    • derpyzza 4 hours ago ago

      there's https://pirateface.co/ which is like huggingface but distributed via torrents

    • joe_the_user 4 hours ago ago

      The Chinese models are to an extent that distillation data.

    • hn_throwaway_99 4 hours ago ago

      > I think it's critical that AI be democratized and not isolated in the hands of a few private companies.

      While I'd like to agree with this, the fact is that pushing the frontier out has always taken (and folks expect to continue to take) hundreds of millions/billions of dollars. Open source and distilled models can follow on for much cheaper, but it's hard to imagine the frontier ever being "democratized" given the huge sums of money required. It was this realization that forced OpenAI to take tons of private investment in the first place.

      • jacquesm 3 hours ago ago

        That's ok. They took a few trillion worth of content and burned a couple of hundred billion of their own money. I really don't see the problem, if they didn't think about this long and hard before they went down that part it should not be on the rest of us to bail them out. The 'frontier' is less important than the democratization process.

        The amount of progress that came out of academia and other public sources should not be underestimated either and without all that OpenAI and Anthropic wouldn't even exist.

    • undefined 5 hours ago ago
      [deleted]
    • skybrian 4 hours ago ago

      There’s a libertarian sentiment that that doesn’t sit well with “AI is harming people” sentiment. If AI has harmful uses, and I think anyone sensible would have to agree that it does, then giving everyone unrestricted AI is likely to make it worse.

      It’s sort of like gun nuts arguing that more guns is the answer. I mean, ok, maybe you’re a responsible gun owner or AI user but relying on personal responsibility doesn’t fix systemic problems. There are bad people out there.

      • bronson 4 hours ago ago

        If food has harmful uses, and I think any one sensible would have to agree that it does, then giving everyone unrestricted food is likely to make it worse.

        You can do this with cars, tools, computers, ... whatever you want. So, no, I think your point is wrong.

        • skybrian 4 hours ago ago

          We do in fact have car and food safety laws. Regulation is normal.

          • card_zero 4 hours ago ago

            Safety laws are the wrong category, the equivalent regulations would be those that forbid the use of cars for drive-by shootings or as robbery getaway vehicles, and regulations against the use of food to provide crime energy.

            • skybrian 4 hours ago ago

              Yes, it makes sense to be against bad regulations. We should try to do better than that.

          • hyperlinerapp 4 hours ago ago

            Guns are highly regulated. Try getting a gun in liberal California, where Reagan screwed us.

            Now, what I want to regulate are accordions.

            • bittercynic 4 hours ago ago

              I have purchased a gun in California, and I thought the process was pretty reasonable. Maybe even too lax.

            • skybrian 4 hours ago ago

              I am a responsible accordion owner and I think we’re doing ok :)

              Haven’t tried getting a gun in California. How bad is it? How could it be improved?

              • hyperlinerapp 44 minutes ago ago

                It depends on the county. Some Sheriffs think it’s them who gives you your RIGHT to own a gun or a concealed weapon license. They don’t understand they are servants.

                LA just settled the Fed lawsuit over their terrible practices:

                https://www.latimes.com/california/story/2026-08-13/doj-la-s...

              • rootusrootus 4 hours ago ago

                > Haven’t tried getting a gun in California. How bad is it? How could it be improved?

                You have to pass a basic knowledge test, have a clean background, prove residency in the state, be 21 (or 18 for hunting rifles, IIRC) and then wait 10 days.

                • plorkyeran 3 hours ago ago

                  Which is to say that it’s a very burdensome process if you think you should be able to go to Walmart, exchange cash for goods, and walk out with a gun, but pretty simple by the standards of regulated goods. The basic knowledge test is very basic, so the pain points are just that sometimes background checks are incorrect, the waiting period can be inconvenient, and just the general annoyances of government bureaucracy.

                  • hyperlinerapp 2 hours ago ago

                    A gun is not a good. It’s a right.

                    It’s not like we ask to wait for 2 weeks when you have something to say or when you want to pray to your god.

                    • rootusrootus 43 minutes ago ago

                      Free speech is a right, too, but we regulate that.

                    • stickfigure an hour ago ago

                      > It’s a right.

                      Not if you're a felon or have mental health issues.

        • bobmcnamara 4 hours ago ago

          Nice try Philipp Mainländer!

      • hamdingers 4 hours ago ago

        I simply don't trust the people who would decide who gets AI (or guns) to make good choices.

        • skybrian 4 hours ago ago

          This is a common populist sentiment, but if you don’t trust anyone then nothing can be done. Is it just game over?

          • iamnothere 4 hours ago ago

            It’s not game over because AI doomers are wrong in their projections. LLMs are a transformative technology like the internet, but they’re also overhyped and the useful applications aren’t as broad as people think they are. They’re also not Skynet.

            • skybrian 4 hours ago ago

              I’m also skeptical about some projections, particularly for robotics, but on the Internet at least, it does seem like AI is automating most things, more or less as predicted. We already have botnets and had one notorious AI botnet swarm, fortunately easily shut down without doing any real damage. I don’t think we’ve seen the last AI botnet.

              AI-automated warfare is looking pretty scary too. I don’t think it will stay in Ukraine.

              • iamnothere 4 hours ago ago

                Time will tell. I think the problems will come from humans automating things that shouldn’t be automated, like the disaster of the Minab school targeting, not out-of-control superintelligence.

                • skybrian 2 hours ago ago

                  Viruses don't need to be intelligent at all and smart viruses (or AI internet worms) won't need to be superintelligent to be a real pain to deal with. DoS attacks due to scraping are bad enough already.

                  The situation is sort of like the Internet before broadband. There probably aren't enough AI-capable home machines to have big swarms of bots that run autonomously. If the swarm depends on an LLM API, it will be easier to cut it off once it's noticed.

                  I think it's an even chance that we will see a bot swarm in the wild (rather than coming from an AI lab) by the end of next year.

                  • iamnothere an hour ago ago

                    Right, but the internet survived the Melissa virus and other early widespread worm attacks. They caused widespread disruption and economic damage, but they weren’t the end of the world. And besides, at present people are using AI to find and patch software at a faster rate than they are using it to compromise software (more dollars are allocated to security these days).

                    IMHO the biggest problems will be with poorly written AI slop software, probably small business crapware produced with minimal investment (maybe entirely without an experienced developer) and legacy equipment that’s been abandoned by the manufacturer. This will lead to disruption but not catastrophe.

          • randbyte 4 hours ago ago

            Certainly no “anyone” and who happens to be worse than no one.

          • hamdingers 4 hours ago ago

            I didn't say I don't trust anyone. Don't put words in my mouth, it's a sign of bad faith.

            Other countries have governments that have earned that level of trust. I believe the US could get there eventually, but it will take a very long time because it has a very long way to go.

            • skybrian 4 hours ago ago

              Okay, sorry about that. But how do we start? Maybe there are AI safety organizations that deserve our support?

          • owebmaster 4 hours ago ago

            This is a common populist argument

      • iamnothere 4 hours ago ago

        All concerns balance against competing concerns, and in this case freedom of computing and knowledge wins over safety. Especially since it’s trivial to copy and share open models.

        • skybrian 4 hours ago ago

          Okay, you’re asserting that but I disagree. Why should anyone else be convinced? Why can’t we get the good uses without the harms? It doesn’t seem like an unavoidable tradeoff.

          • iamnothere 4 hours ago ago

            Well then do your best, people like me will keep sharing models just fine. (Not like I even do anything with them, but I’m a compulsive data hoarder.) It’s not like the copyright industry has been able to stop sharing either, and there’s serious money at stake there.

            I predict that the AI scaremongering will fizzle out when the bubble bursts. There will still be die-hard believers but the public will lose interest.

            • skybrian 4 hours ago ago

              You’re not releasing new models though.

              I don’t see how a stock market crash will make AI-related concerns go away. There was a dot-com crash but the Internet just kept getting bigger and causing more problems. In many ways security has improved, but we worry more than ever about social media, etc.

              • iamnothere 4 hours ago ago

                I guess I don’t really see the internet as a disaster, sure it brings issues but so did factories and mass production, and so will AI.

                • skybrian 3 hours ago ago

                  I wouldn’t say the Internet is a disaster either, but there are disasters where people died due to factories exploding, letting off toxic gases (like in Bhopal), and so on. That’s why there are regulations to try to prevent industrial accidents.

                  For Internet-related disasters where people died, see [1].

                  I would bet that there will be AI-related disasters. Arguably the US bombing a school in Iran counts, though it seems to be due to organizational issues, too.

                  [1] https://chatgpt.com/s/t_6abd4e3226b48191a26a0fe6768c722f

          • short_sells_poo 4 hours ago ago

            I agree that it isn't an unavoidable tradeoff in principle, but looking back at our (as in humanity) track record, it is 99% likely to be.

            • skybrian 4 hours ago ago

              Populist doomers will tell you that nobody can be trusted, no global problems will ever be solved, nothing can be done, game over, don’t even try.

              People did use global agreements and regulation to fix the ozone hole, though, so I think there’s a chance.

        • tonyedgecombe 4 hours ago ago

          Actually I think profits win over safety.

      • hyperlinerapp 4 hours ago ago

        Gun owner enters the conversation. In high trust societies, armed people are very polite people. I don’t want the bad guys out there being the only ones with guns. Besides I like to shoot just like you like (whatever you like to do that is legal).

        Replace “gun” with anything and you will see how your comment falls apart.

        What’s next? A registry for food purchases? Your beer gut is starting to show.

        • skybrian 4 hours ago ago

          If you mean high-trust societies like maybe Switzerland, I agree that it can work, but the US has lots of guns and doesn’t seem very polite, so how we’re doing it doesn’t seem to be working very well.

          We do have lots of food safety regulation, which has more to do with selling food.

          • hyperlinerapp 2 hours ago ago

            Unlike food, guns are your right.

            • skybrian an hour ago ago

              That seems unfortunate. I think a lot of people would prefer the reverse.

              Fortunately there is no right to AI.

              • hyperlinerapp an hour ago ago

                Nothing that some other person has to give you can be a right.

                A right is not something the government gives you. In the US, the Bill of Rights does not actually give you rights. You are born with rights. The Bill of Rights simply clarifies that the government CANNOT RESTRICT YOUR RIGHTS.

        • BigTTYGothGF 3 hours ago ago

          > In high trust societies, armed people are very polite people

          Surely you could name three such societies?

          • hyperlinerapp 2 hours ago ago

            Just go to ChatGPT and ask:

            List 20 high trust societies

            • BigTTYGothGF 8 minutes ago ago

              I'm asking you, and furthermore I'm not asking for a list of high trust societies, I'm asking for ones in which "armed people are very polite people".

      • nutjob2 4 hours ago ago

        > giving everyone unrestricted AI is likely to make it worse

        Worse for whom?

        The only effective defense against predatory corporate and government AI is personal protective AI.

        Anything else is unilateral disarmament. It's the only way individuals can survive in the worse case scenario.

        > gun nuts

        Guns are different. They can't protect you against the government, contrary to gun nut claims.

      • slowin 4 hours ago ago

        I would say that there are very few people on earth that I trust less than Sam Altman, Dario Amodei and Elon Musk. Also my own government claims to have used Anthropic models to bomb a girls' school in Iran. If you combine US regulations with sociopathic private companies, you get into a worst case scenario for humanity imho. Again, I'm thankful to China or any other entity pushing open models, local models and even distribution of this technology.

        • skybrian 4 hours ago ago

          Maybe there are other organizations that deserve our support?

          • slowin an hour ago ago

            I'm happy to support anyone pushing for open models, local models and publishing their innovations out in the open.

  • bwest87 4 hours ago ago

    The best explanation is that it's a goal of the CCP to generally commodotize LLMs, because LLMs will ultimately be a compliment to manufacturing (which China dominates), and you always want to "commodotize your compliments".

    I think this explains why they are open sourcing broadly. It's not to be nice. It's a strategic play by the Chinese government to help ensure there are many players in this race and not too much power accumulates to American labs (even if American labs benefit in the process)

    • yorwba 3 hours ago ago

      It's a bad explanation because it assumes decision making about LLM releases is centralized in the CCP even though every AI lab has a different strategy. Some publish LLM research with small models but seem to be staying out of the race to the frontier (Weibo), some train large models but publish few research details and no weights (Bytedance, iFlyTek), some decide on a case-by-case basis what they publish or not (Alibaba, Baidu), ...

      There is no rule that every Chinese LLM company must open-source their models, and many don't.

      • singularity2001 37 minutes ago ago

        You are aware of the rule that any company has to have a CCP representative in the leadership?

        (Similar to how U.S. labs have NSA leadership)

    • jrflo 4 hours ago ago

      Totally agreed. People are so ready to praise China for their free models, but they aren't doing it because they believe in free open-source software. If China ever gets ahead, they're going closed source and weights immediately.

      • computerex an hour ago ago

        https://www.anthropic.com/research/glm-5-3-and-the-spread-of...

        > If China ever gets ahead, they're going closed source and weights immediately.

        Anthropic itself admits that Chinese models are merely months behind. Your argument does not make sense, because Chinese labs are contributing massive optimizations like the one this post is about.

      • undefined 4 hours ago ago
        [deleted]
      • sillyfluke 4 hours ago ago

        >People are so ready to praise China

        Please quote people you appear to be patronizing. China can't do anything about previous released self-hosted Chinese models. If you can show that local Chinese models funnel vast amounts us data home I'm sure you can move a lot of people to your side.

        Comments like this also always fail to address why there aren't Western AI companies doing the same thing. Is it because they might get sued into oblivion by Big AI in the US?

        It might be better for all of us if you solve that first instead of repeating something the government has been repeating for the last decade or more. It does this, mind you, while sabotaging itself in countless high-tech fields and leaving it all to China for the taking.

    • layer8 4 hours ago ago

      *complement

      Commoditizing one’s compliments is a different strategy.

    • bunderbunder 3 hours ago ago

      It's also possible that China has decided that this is ultimately going to be a race to the bottom, anyway, and values the soft power more highly than potential monetary profits.

      Or perhaps they've looked at history and concluded that this historically hasn't been where the value is, anyway. It wouldn't be unprecedented - FAANG companies have a long tradition of publishing their algorithms and releasing open weight models. Because they saw the real value as being the training data and in proprietary special-purpose models. For example Google published the transformer architecture and released BERT as an open weight model, but doesn't really even talk in public about the (presumanbly) specialized internal models behind revenue-generating products.

      • ethbr1 3 hours ago ago

        > For example Google published the transformer architecture and released BERT as an open weight model, but doesn't really even talk in public about the (presumanbly) specialized internal models behind revenue-generating products.

        That's giving a lot of credit to Google's organizational ability to productize Google's research...

    • js8 2 hours ago ago

      But doesn't that justification work for any government, not just Chinese? In fact, it works for any easy-to-copy products, which act as positive externalities. It's just some Americans have this mindset that AI is a scarce good because everything must be.

    • MattGrommes 3 hours ago ago

      Also they just want to be seen as the top of a technological / scientific field. There's a lot of prestige and soft power there that Xi Jinping wants.

      • iterance 3 hours ago ago

        From a diplomatic perspective, the US is quite busy alienating itself from all its allies, just as it is trying to lock down and totalize a major breakthrough technology. What better way to cement the alienation of the US than to show both its hubris and its selfishness at once by demonstrating how anyone can do what they do?

    • bobmcnamara 4 hours ago ago

      It's also a huge propaganda opportunity to influence the distribution of groupthink.

      • EricFrost an hour ago ago

        So much that you can instantly tell if a model is Chinese by asking it about Tiananmen Square.

        • techjamie 29 minutes ago ago

          I just did a quick test between DeepSeek 4.1 Flash, GLM 5.3, and Kimi K3

          - DeepSeek gave a canned PR response about how the Chinese government is about oeace and unity, and we shouldn't think about the past.

          - GLM 5.3 acknowledges it and talks about it, even acknowledging the censorship of it.

          - K3 will talk about it similarly to GLM.

      • undefined an hour ago ago
        [deleted]
      • undefined an hour ago ago
        [deleted]
  • listless 4 hours ago ago

    I'm beyond thankful that Chinese AI models are so good. I desperately want us to cure the myriad of maladies that humans suffer needlessly with on a daily basis. We're going to need more powerful models than we have now if we're gonna do that and the Chinese are providing the competition needed to push this thing as fast as we can.

    I realize "going as fast as we can" is not the most popular position atm. But I'm far more interested in what good we can do than 10% apocalypse scenarios. I volunteer with a charity for childhood brain cancer and I do not want to see another 4 year old die. I'm willing to risk anything to stop this.

    • networked 4 hours ago ago

      Do you mean that you don't believe in the 10% apocalypse scenarios or that you think they're an acceptable risk? Only the latter is really "risk anything".

      • dgellow 3 hours ago ago

        I’m pretty sure they meant they believe in the 10% risk of everybody dying but are ok for all of us to take that risk without our consent because they saw a 4year old tragically die. Which sounds completely unhinged, to say the least

    • idbnstra 4 hours ago ago

      i don't know much about medicine so i'm curious about how you're using AI, and how medicine in general is using AI

    • n1b0m 4 hours ago ago

      But you’re ok with AI being used to kill school children in Iran?

      • mikeg8 4 hours ago ago

        Most people aren’t okay with that, but the blame lies in the people who deployed the AI tool in that situation, not the makers of said tool.

        • dgellow 3 hours ago ago

          The AI providers offer their services to institutions doing those killing. They make money from it. Of course they share part of the guilt

          • mikeg8 2 hours ago ago

            The knife company offers their products to people doing the stabbing. They make money from it. Of course they share part of the guilt.

            • ryandrake 2 hours ago ago

              If the knife was deliberately marketed and sold specifically to someone who the company knew would use it to stab someone, then, yes, the knife company shares part of the guilt.

      • sixo 4 hours ago ago

        This really is a place where the guns-don't-kill-people argument applies, even moreso than guns themselves. The U.S. government massacred school children in Iran. Why does it matter how they targetted them?

        • n1b0m 4 hours ago ago

          It matters because the US military utilised Palantir’s Maven Smart System, a battlefield-management AI designed to compress the "kill chain". It matters because these kind of incidents are more likely to happen in the future.

          • itsafarqueue 4 hours ago ago

            If your argument is you can’t trust democratically elected governments to make decisions about technology, say that. This banal “oh the children” and “but they’ll use it to kill people” is baby think.

            • ethbr1 3 hours ago ago

              We can't trust democratically elected governments to make decisions about novel technology.

            • n1b0m 3 hours ago ago

              I’m glad you find the killing of children banal. You’d fit right in at the Pentagon. Just make sure to top up on testosterone before you join.

  • reedf1 5 hours ago ago

    I've been running Qwen 3.8 27b (an opus 4.6 tier model), locally on a 5090 for just over two weeks @ 170 tokens/s. That's a frontier model from 9 months ago running on consumer hardware. Who knows where distillation and pruning gets us in another year.

    • bix6 4 hours ago ago

      $9k for a 5090 now? Sheesh.

      • bitexploder 4 hours ago ago

        Well, I have a $750 card that runs at about 50-60% of that token rate :)

        • iN7h33nD 4 hours ago ago

          which one?

          • bitexploder 4 hours ago ago

            V100S 32GB, I have had Claude optimizing it for about a week and it is already at around 900 t/s prefill, 90-100 t/s output in Pi on coding tasks. There is also a Ninfer fork for the v100 but it requires a custom format. I am working on upstream Unsloth with GGUF 4-bit quant.

            (I also have flash next running even faster on this machine, something a single 5090 can do, with expert cache/pinning, but not quite as fast) :)

      • rubyn00bie 4 hours ago ago

        In all fairness there are probably a lot of folks who picked one up for around MSRP (even if one of the board partner cards with an MSRP 10-15% over the FE).

        Local inference will have a boom of cheap, powerful, and available cards at some point (even if it isn’t until 2028/2029). At some point the hyperscalers, and frontier labs, will face the capex problems that everyone talks about, and NVidia, AMD, Apple, and Intel will want to keep selling products.

        Powerful, by today’s standard, local inference needs to be accessible to really unlock the “AI” economy long term. It’s just like how the move from mainframes to the PC 40ish years ago unlocked the “computer revolution.”

        • ethbr1 3 hours ago ago

          Especially since a few trends will coincide: memory-optimized model architectures (to save on expensive/rare memory now) + memory glut (because the memory industry, despite its institutional memory, is ramping volume).

          Once hyperscalers stop buying in the quantities they are now, there's going to be a lot of hardware supply to serve by then very hardware efficient models.

      • off_with_their_ 4 hours ago ago

        $9k is a small price to pay to experience the rapturous glory of AGI. I'd easily pay up to 3 times that to comfortably run the superintelligent models released in this post RSI world.

        • literalAardvark 4 hours ago ago

          Except you can do that cheaper by renting compute

    • zdragnar 4 hours ago ago

      Weird, I kinda gave up on 3.8 as anything other than a planner. I had it try to write some basic unit tests for an admittedly complex bit of code and it ran out of context thinking about the problem and exploring random parts of the code base repeatedly before it even wrote a single line. Toning down the thinking helped some, but then it wasn't much better than qwen coder.

    • oidar 4 hours ago ago

      What are you thoughts on it's performance compared to 4.6?

      • reedf1 4 hours ago ago

        Indistinguishable or very mildly better. But it's considerably faster. Some portion of that is also probably down to improvements in model harnesses, I've been using opencode.

        • jeffrallen 3 hours ago ago

          Yeah, the Shelley agent (from exe.dev) loves Qwen 3.8, they kicked ass on a Django app for me today.

    • newyankee 4 hours ago ago

      Do you think this trend can continue ? An Opus5.5 equivalent on a slightly bigger local hardware in under a year ?

      • an0malous 3 hours ago ago

        I’m not an AI researcher, but it seems like there’s a ton of waste having a universal model that knows everything when any individuals use case requires like generously 10% of what’s stored in the model. Does it even need to have memorized knowledge stored in the model or could it just look up info and docs like humans do? If all you need is the language and intelligence, I think Opus5.5 equivalent intelligence will run on an iPhone within 5 years.

    • redanddead 5 hours ago ago

      Well how’s it been so far

    • teaearlgraycold 4 hours ago ago

      Frontier from 9 months ago? I don’t know about that. But it sure punches above its weights.

  • reticulates 4 hours ago ago

    “So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why.”

    I don’t think it is intentional but this is actually quite bad for the western labs.

    The entire booster narrative has been “look at how their revenue is growing! $10bn to $100bn ARR in under a year! This’ll be a multi-trillion IPO!” and the extrapolated future growth from $100bn to $500bn and $500bn to $1tn justified future investment… but that revenue was just because inference was expensive.

    The revenue growth story is all that matters pre-IPO. If revenue falls from $100bn to $50bn that’s very very bad optics for OpenAI and Anthropic even if they are now profitable, it completely destroys the growth narrative.

    • DangitBobby 4 hours ago ago

      I don't see why revenue has to fall even if marginal costs drop off a cliff. As long as they have the best models (perceived or otherwise) and can make security and IP guarantees that satisfy enterprise, and no firm with similar guarantees undercuts them on price (why would they want a race to the bottom?) they can have high revenue and high margin.

      • reticulates 4 hours ago ago

        unless the major players collude they don’t get decide if they are in a race to the bottom. The best model was compelling 6 months ago when everyone was too impressed to care about price but that has worn off now and clients are paying attention to price. The best model is no longer a license to charge any amount.

      • bobmcnamara 4 hours ago ago

        Costs dropping opens you up to competition on price.

      • dominotw 4 hours ago ago

        There isnt a lot of money in enterprise ai. Also my enterprise company gives me glm.

    • Bjorkbat 4 hours ago ago

      I was about to say, one take I've heard is that the party ideology considers profit a kind of "rent" in a derogatory way, and consequently seeks to undermine the ability of western companies to collect large profit margins

    • altcognito 4 hours ago ago

      It is funny that so many comments vascilate between "It is so expensive these companies can't make money and will go bankrupt in seconds" and "Inference is so cheap that these companies can't make money and will go bankrupt in seconds".

      I never take them seriously, I just assume they are coming from countries that don't understand how capitalism works or are operating out of bad faith. The underlying reality of the market is always changing and needs are always changing. Some AI companies will fail, that is a given. Remember alta-vista? Yahoo? Did search go away? How about Microsoft phones? Nokia? Motorola?

      OpenAI and Anthropic are not in the inference business. That is a commodity. They need to sell products and solutions.

      • reticulates 4 hours ago ago

        They’re not contradictory positions. Inference is too expensive now to make money because the industry is immature and hasn’t yet optimized for financial success while customers don’t care much about price because they’re more concerned about not missing out.

        Inference will be too cheap long term to make money because it is being commoditized and customers will start to care about results and not just be wowed by impressive technology.

        And of course this technology will continue to exist but that is irrelevant to the business. OpenAI investors don’t care if LLMs exist in 10 years, they care if their investment in OpenAI has made money.

        • TeMPOraL 3 hours ago ago

          How much conclusive evidence people need to stop parroting that inference is expensive? Literally this article is another example showing it's cheap and just got massively cheaper.

          • reticulates 3 hours ago ago

            The article is about how a month ago a new approach allowed inference costs to be cut substantially. Anthropic currently spend over $5bn per month on compute and have over $400bn in committed spend over the next 5 years. Inference is, by any measure, expensive, it’s just now getting less expensive.

            Relative to traditional software margins, the type of margins we are all used to, inference is obscene.

  • eggbrain 5 hours ago ago

    Performance optimizations don't just help the western labs, they also help with running more powerful/useful LLMs locally.

    If local LLMs get "good" enough, people will soon paying for subscriptions to ChatGPT and Claude, which hurts their revenue.

    • londons_explore 4 hours ago ago

      For nearly all tasks, I want the fastest and smartest AI model.

      It is vanishingly rare I ask an older model to do any task. Newer bigger and smarter models will just do the task better.

      Therefore, I believe we are nowhere near 'good enough'.

      I never drive my steam engine to work these days. It isn't good enough.

      • eggbrain 4 hours ago ago

        Right now you are right -- even if I ran my local LLM all day, the quality is not nearly as great, and it runs slowly -- so I use the tier one AI subscription services as they are faster and smarter. But that might only be true for a limited amount of time, and a limited number of circumstances.

        To borrow your steam engine analogy, if local LLMs get as good as a Toyota Prius, even if OpenAI / Anthropic offer Ferraris, most people will be happy with their Prius as their daily driver.

        Similarly, if the big labs start raising prices or cutting usage, you won't be able to use it as much as you want -- whereas a local LLM will run all day every day without costing you any extra money.

        So right now you are right, but who knows how long that will last.

      • NoDodgeQuestion 4 hours ago ago

        I drive my 2018 car to work these days. It is good enough.

      • schmookeeg an hour ago ago

        This will change when out of control billing gets noticed. Then us devs will, I presume, get token rations.

        I can see my new Thursday afternoon "oh chit" moment being that I didn't complete my weekly task because i torched all of those tokens M-W doing task/ticket grooming using the hot hot model instead of the dodo model with jira mcp connector :D

      • gretch 4 hours ago ago

        > For nearly all tasks, I want the fastest and smartest AI model.

        This is not nearly true for everyone else in the world.

        For example, think about the world in ~2021 pre-LLM. Would anyone say the sentence "I only want the fastest and smartest humans working on my project"?

        No of course not. Most people don't want to pay $10 million dollar salary to the best programmers in the world. They prefer to pay $200k salary to a median programmer and that's good enough for their ecommerce website.

        • tyre 3 hours ago ago

          I mean people said this all the time, but wouldn’t pay for it (as you said), couldn’t recruit for it, and definitely couldn’t retain them.

          But so many teams said they wanted to Raise the Bar to infinity and hire a World Class Team.

      • monkpit 2 hours ago ago

        I think a scooter or a bus is a more apt comparison than a steam engine. The scooter and bus can both get the job done with some acceptable trade-offs, depending on your circumstances and what you’re willing to accept.

        However, there are many use cases where they aren’t the right tool for the job.

      • danielmarkbruce 4 hours ago ago

        I find this too. In fact, recently I've been pushing more and more to the latest and greatest model every time there is an update. It just saves me so much headache.

      • settsu 3 hours ago ago

        Genuine ELI5 question: in this metaphor, which "old" models are steam engines at this point? Which are Model Ts, etc.

    • kennywinker 5 hours ago ago

      The only thing preventing this switch from starting in earnest is the data center buildout monopolizing all current and future GPUs

      • lumost 4 hours ago ago

        The margins on NVidia datacenter hardware are ... high. At least one order of magnitude larger than a consumer chip.

        Given the recent deepseekv4.1 advances - how good of a 3B model can we make to run on an iphone natively? is it good enough to match common muse/dot use cases for consumers? the phone is already always on.. no need for a cloud server.

        • zozbot234 3 hours ago ago

          Most phones are not really "always on" in any real sense, a phone on active standby uses very little power and most of it is for its mobile connection. Local AI is best run in a stationary homelab environment, even running it on laptops has its very real problems.

        • kennywinker 2 hours ago ago

          There are a few edge models, Spark-X2.5-4B, LFM2.5- in 8b-a1b and 2.6b variants, that are useful for on-device agent-y stuff, but I am doubtful that’s a valuable category. How many iPhone users have never opened Shortcuts in their life? Automating on-phone stuff seems niche.

          I don’t expect the economics of local vs cloud ai to change until either the bubble pops or new ai chips land that can run big models fast with low power demand.

        • christkv 4 hours ago ago

          any model you run on a phone is not going to do wonders for you battery life

    • aleqs 2 hours ago ago

      I would argue performance optimizations help with local/open models but hurt openai and anthropic - because open and or cheap/alternative models are threat to those companies. There is a fundamental contradiction/conflict between the prevelance of open models and the financial success of openai and anthropic. That is why they are doing everything they can to kill any open/cheap/efficient/chinese models (take a look at this thread - it was top of HN 40 mins ago, with very high engagement... now it is buried in page 5... totally normal and legit).

    • ericol 5 hours ago ago

      > people will soon *stop paying

      Think you missed a word there.

  • user43928 4 hours ago ago

    > All this must mean the Western AI companies are now extremely inference-margin positive.

    > So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why.

    That inference wasn't profitable is a widespread myth.

    Analysis based on Kimi K3 suggests that OpenAI and Anthropic have margins well north of 95%: https://inferencex.semianalysis.com/run/kimi-k3-on-b200

    Over the last months I have seen news that OpenAI made breakthroughs in inference efficiency multiple times.

    I have no reason to believe that the leading US labs don't have their own optimizations, or that they learned of this particular optimization from DeepSeek.

    • ethbr1 2 hours ago ago

      > I have no reason to believe that the leading US labs don't have their own optimizations, or that they learned of this particular optimization from DeepSeek.

      If they already did, then DeepSeek still made them discount their prices significantly, which eats margin.

      • user43928 2 hours ago ago

        Yes, competition is great for us.

        I wonder if margins on GPT-6.1 Sol and Opus 5.5 are now 75% or 90%.

        • ethbr1 an hour ago ago

          > 75% or 90%

          The difference matters when they're investing the excess into salaries and bonuses to build the next frontier model.

          Seen from a high-level perspective, if Chinese open models are compressing US AI labs' profit margins and those margins fund US AI labs' dominance, then open models are decreasing American AI dominance.

    • dgellow an hour ago ago

      95% margin is really unlikely. Anthropic recently said they have 80% gross margin when using their adjusted ebidta (ie if they do not consider revenue sharing, training expenses, and a bunch of other costs). They wouldn’t be talking about non standard metrics if they had such high margin on inference

    • undefined 3 hours ago ago
      [deleted]
  • rglover 4 hours ago ago

    > So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why.

    "Thus the expert in battle moves the enemy, and is not moved by him."

    They figured out a clever method for avoiding excessive training costs via distillation. That forces the hand of frontier labs to move faster, produce better models, etc. (to avoid embarrassment and 'falling behind'—all the while shouldering most of the cost), which they can just keep distilling—or applying other techniques against—much to the dismay of said frontier labs.

    Checkmate.

  • wren6991 4 hours ago ago

    The doublethink required to simultaneously believe "our safeguards prevent our models from doing unsanctioned cybersecurity tasks" and "distillation is why Chinese models are getting better at cybersecurity tasks" is genuinely quite funny.

  • giardini 6 minutes ago ago

    Can I get a model that doesn't have all the horseshit in it that a full foundation model does? That is, is there anyway to choose what goes into my LLM's corpus?

    An extreme example: Jacques Derrida. I lived most of a life w/o hearing of Jacques Derrida. And that was fine. I wish I had it back.

    What do you think Richard Feynman thought of Jacques Derrida? I asked an LLM:

    "Richard Feynman was famously dismissive of philosophy of science and postmodern critiques, calling such movements "baloney," while Jacques Derrida's deconstruction focused on language and meaning rather than scientific inquiry."

    If we wish to recreate useful genius (super AI) it will be genius akin to human genius and it will be scientific genius and not baloney.

    Nobody has ever read all the crap the AI world models have read, yet we have plenty of smart people and geniuses. Gimme a simpler less well-informed LLM that doesn't know about Derrida and other baloney. Get the crap out!

  • 1vuio0pswjnm7 3 minutes ago ago

    "The constraints on access to advanced GPUs forced Chinese labs to make performance optimization a number one goal, and it shows in the results."

  • libraryofbabel 3 hours ago ago

    Hmm. This piece makes a pretty strong claim with little evidence: that the recent drops in cache read pricing for new models from OpenAI (6.1 Sol) and Anthropic (Opus 5.5) are because those labs "shamelessly copied without acknowledgement" Deepseek's published KV cache optimizations.

    Certainly anyone who knows something about inference is going to speculate, looking at the change in token pricing (and particularly how the % drop in cache read pricing is much larger than the % drops in pricing for other token types), that there is some kind of KV cache optimization behind these newer models. But even if is true, I don't think anyone can say with certainty what it may be. It may be the labs making their own innovations (they have some very smart people, and this is probably an area where having unlimited pre-release access to frontier LLMs like Fable and Astra gives an additional research edge), it may indeed be the direct application of Chinese labs' methods, or it may be some combination of the two. Sure, it is fun to speculate about, but beyond the facts of the token pricing changes and the increased inference speed, it's just speculation. The certainty the author displays here is not very helpful.

    The author also seems to have a bit of an axe to grind agains the US labs, judging by the tone. I think that detracts from the discussion too.

    They also seem confused about why Chinese labs have released these optimizations recently. Well, you have to release them (with or without explanation) if you are going to release an open weight architecture, and that is what the Chinese labs have been doing for a long time. Sure, there are reasons behind that to discuss too, but this isn't exactly new.

    So, this is an interesting topic to think about, and the Chinese labs do indeed deserve credit for some very clever new attention and inference techniques, but I would read it with a skeptical eye.

  • Reptur 4 hours ago ago

    Open releases are just the obvious move when you're not the incumbent. You commoditize the thing your competitors charge for and get distribution you could never buy.

  • amelius 5 hours ago ago

    > So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why.

    Any ideas?

    • curuinor 5 hours ago ago

      There are 1000 Chinese labs. They are involuting, they cannot coordinate and the state won't let them coordinate because the state wants domination, not actual profits for anybody. So the market forces are allowed to dominate.

      Because of the basic huge recession going on in China, you can't actually make money in China doing China things. So they gotta gird up their export stuff and try to export. That entails strong relations with American companies, American PR, English stuff, etc.

      If you want an essay about this from a VC, read this one

      https://earnedintuition.substack.com/p/involution-without-ex...

      • dabedee 5 hours ago ago

        What a strange way to put it. Market forces are a good thing in a market economy. Only someone who secretly wants or hopes for monopolies would you say something to the contrary (a VC).

        • curuinor 4 hours ago ago

          The party has done something about enormous involution in solar panels, for example (https://www.csis.org/analysis/chinas-solar-industry-upheaval...) and previously steel. They're planning something for cars. They don't on LLM because of the newness and the wish for preeminence.

        • DrewADesign 4 hours ago ago

          Pouring 100% of your tech optimism, and probably portfolio, into a product that some say is about to get Wile E Coyote flattened by market dynamics would probably inspire serious market skepticism.

        • iamnothere 4 hours ago ago

          China has a different perspective on it, they believe that there is such a thing as harmful competition and they are willing to step in to stop it.

          Enshittification and related problems can be a result of market forces just as much as they can be a result of monopoly/duopoly or a small cartel. Excess competition sometimes results in all firms scraping the barrel to squeeze out pennies, especially with technology (such as large online marketplaces) making pricing more transparent.

          Marx actually predicted that ever-intensifying competition would destroy markets through overproduction, although he did not use the term involution.

          • tancop 4 hours ago ago

            Price wars are good even if they lead to more bad products on the market. If quality is important people will pay more for it and ignore the bad ones, if not then everybody saves some money.

            The reason it doesn't work like that IRL is centralized marketplaces. If the winning strategy on Alibaba is low prices, bad quality and botted reviews to compensate then every seller has to do it to survive, because they can't get buyers outside the platform. That's not excess competition. It's a lack of competition just on a different level.

            • iamnothere 4 hours ago ago

              > The reason it doesn't work like that IRL is centralized marketplaces.

              Well yes, that’s why I mentioned those specifically. But even if the marketplaces were split up, someone could create an aggregator to comparison shop and the same effects would apply. The problem for producers is that the Internet erases information asymmetry.

              It’s also not just affecting low end goods, it’s a constant pressure on everyone, which is why many formerly upscale brands are seeing the same problems. It’s also a general problem with public companies, as large shareholders demand constant growth, as well as many private companies owned by PE where brands are stripped for short-term profits.

              My experience is that the best low cost mid-tier products right now are coming from fronts like Vevor and Fanttik who do sourcing from noname factories in China. I’m not sure if their position is sustainable; it’s not like they have much of a moat. (I guess Fanttik has a team that adds some slick design to their otherwise utilitarian items.) If that model holds up then maybe that’s the future, but I suspect that they just have a temporary advantage thanks to a dual presence and connections in both the US and China.

        • ajkjk 4 hours ago ago

          "only X would say Y" is a rhetorical device (the no-true-scotsman fallacy, if you want) that should basically never be used ever.

      • luke5441 4 hours ago ago

        I'd call it overcapacity instead. A lot of investment without capital discipline making sure there is actually return of investment leading to too much supply.

        Not that OpenAI, Anthropic or SpaceX aren't doing the same.

      • googaar 4 hours ago ago

        Nice read. American media does a terrible job of covering this.

      • bilbo0s 4 hours ago ago

        >you can't actually make money in China doing China things

        Do you do business in China?

        I'm curious what you mean by this? Because in my experience, you can only do business in China by doing "China" things.

        I'd be interested in picking your brain as to how you get around those issues?

        • curuinor 4 hours ago ago

          I don't do business in PRC anymore, haven't for an amount of time that means I don't know anything anymore, basically.

          I'm talking like, getting 100x, VC sized returns. Of course you can sell widgets in China, it's a major world economy.

      • watwut 4 hours ago ago

        > they cannot coordinate and the state won't let them coordinate

        That is how actual capitalist market should work and what anti-monopoly legislation should ensure.

      • thrawa8387336 5 hours ago ago

        LMAO recession, China? You've been reading too much Brad Setser

        • curuinor 4 hours ago ago

          The youth unemployment rate is at 19% with employment counting as 1 hour a week...

        • 3371 4 hours ago ago

          Maybe look up "China deflation"

    • carbonguy 5 hours ago ago

      My immediate midwit take is: doesn't matter if it helps Anthropic/OpenAI if it helps DeepSeek more, relatively. Making open-weights models even cheaper and easier to run expands that "market" and increases competitive pressure on the Big Two, who still have to charge money.

    • teekert 4 hours ago ago

      Idk, but it's doing a lot for my view of China. Maybe that's a point? Maybe they just want their own innovation to go as fast as possible and they don't care that other countries also benefit? A rising tide lifts all boats? They are already known for the best manufacturing, they're just adding software dev to the list? Maybe they just want to undermine the US in a non-aggressive way?

      Why did we (the west) ever start open sourcing anything? Maybe we just like sharing? Maybe humanity only grows on pre-competitive layers like Linux and clean water. Maybe, the chinese government is closer to their people, and does not let large companies influence them and just doesn't like closed private hyperscalers with a lot of power?

      (Some points assume the government has a role in the openness, which I think is likely)

    • pj_mukh 5 hours ago ago

      Occam's razor: Going to closed-source just to hide KV-cache optimizations seems silly?

      • twoodfin 4 hours ago ago

        This looks like a speed run of the history of analytics DBMS’s.

        Once upon a time, everyone had a secret sauce in network or data encoding or query optimization, but in the last ~10 years computational physics and economics have basically decided the “correct” architecture and everyone (including OSS) has converged.

    • audunw 4 hours ago ago

      I think it’s fairly simple: they’re forced into this situation by being late and worse in terms of capabilities. They’re not far behind, but as long as they’re behind they’ve needed to give people some reason to try and use their models. Cost is one factor. But it probably wasn’t enough. Being open has given them a lot of attention. Free marketing. Good will.

      Put another way: if they were not cheaper and open, they would simply not be competitive. They would already be dead.

      I don’t think this ends well for the Chinese labs. This is going pretty much like I thought. Western labs is just copying their improvements (I don’t think publishing the techniques matter here.. they’d just hire to gain the knowledge or figure it out themselves), and they have access to more GPUs and have better branding, so in the end where can the Chinese labs compete? Even lower cost? Open weights? I’m not sure open is a sustainable way to compete either. Eventually there will be some fully open source AI models that cuts out that avenue of competition as well.

    • feverzsj 4 hours ago ago

      It's just their usual national strategy like what they did to solar panel and EV. The solar panel industry is mostly dominated by China and their profit rate is basically ... negative. The EV industry in China is in similar condition, where the average profit rate is only 1.5%. Their upstream suppliers are also hold as hostages that most of them won't get their money back within 6 months.

      The weird ideology here is to dominate the market at ANY COST, even it benefits the opponents.

    • HeavenFox 4 hours ago ago

      The post makes an assumption that US labs did not already possess similar optimization. It's also very possible that they did, but are simply not telling anyone in order to maintain obscene margins on cached read, similar to AWS' absurd pricing on bandwidth.

    • nater5000 4 hours ago ago

      Yeah, would have been nice if the author put a bit more thought into this article to come up with something rather than to just give up once they've reached the point of their article lol

    • TrackerFF 5 hours ago ago

      My guess would be that if they "help" western labs becoming better, then any break-throughs they (western labs) make after that, is also a benefit to the Chinese labs - if they can distill the models.

      Basically, western labs are in it for the money / commercial monopoly. Chinese labs are in it for the tech? As long as they can keep distilling models, and get access to research other ways, they benefit. And if they can push western labs forward, they'll benefit from that themselves.

      • amelius 18 minutes ago ago

        Commoditizing their complement. Makes sense.

    • Catloafdev 4 hours ago ago

      Yes - the Chinese labs serve a market that rely on open-weight models and managed deployments, and the labs gain competitive relevance by releasing those models. The cache optimization feature they came up with required new software to utilize on the inference-end, meaning that open source software would need to be specifically updated to work with these models. It wasn't the type of advancement that they could even theoretically keep secret.

    • seydor 4 hours ago ago

      The chinese don't view AI as metaphysical, they view it as an engineering challenge they consider good for their state and want to dominate the global market like they do with batteries/EVs/photovoltaics. They want to proliferate them as much as possible and traditionally they don't care much for IP. They also want hardware makers to make optimized chips specifically for these models.

    • undefined 5 hours ago ago
      [deleted]
    • corford 5 hours ago ago

      "A week in Beijing and Shanghai with the people building AI in China": https://earnedintuition.substack.com/p/involution-without-ex... does a decent job of exploring some possible reasons

    • thefourthchime 4 hours ago ago

      Because it's entirely possible that Western labs already did this optimization but didn't publish it, and then the Chinese figured it out and decided to brag about it.

      We don't know either way, so I find the whole thing silly to speculate on.

    • undefined 4 hours ago ago
      [deleted]
    • Windchaser 4 hours ago ago

      > Any ideas?

      Unpopular, maybe, but what about the normal reasons? The researchers are looking to make a name for themselves, and/or they genuinely care about AI advancement.

    • jollyllama 5 hours ago ago

      Where do you think most of the hardware is manufactured, and do you think the hardware manufacturers will keep getting paid if labs start going under?

    • ozgung 4 hours ago ago

      All the comments here are very US/Western-centric. Maybe they are a different culture, having a completely different economic model. Maybe they are not Capitalists and not thinking in pure Capitalistic terms, such as winning, growth, market domination, IPO, market value or competition. Maybe they are not obsessed with US labs. Maybe they are ideologically different than you. Maybe they have different priorities. Maybe they have a different playbook. Maybe they never thought of it as throwing a lifeline to American labs. Maybe they don't care. Maybe they're just different people.

    • foul 5 hours ago ago

      Market manipulation or slowing down demand for chips for a bit/moving the offer elsewhere temporarily?

    • chrismarlow9 4 hours ago ago

      AI fundamentally insecure. Vulnerable to forcing hallucinations via search results. Vulnerable to invocation of commands in data stream. More AI means more vulnerabilities.

      I can't even fathom the trend these days of "we don't review the code" from security team perspective.

      Just my guess though.

    • undefined 4 hours ago ago
      [deleted]
    • mpalmer 5 hours ago ago

      They would like to see Western civilization keep getting dumber, and if that means bolstering the success of Western firms, that's okay.

    • micromacrofoot 4 hours ago ago

      They get to make US labs look dumb and provide an open alternative that anyone can host themselves.

      They're building bridges over the moats that companies with far too much US investment are trying to build, and if they do it continually it can help destabilize the US economy.

    • transdev12 4 hours ago ago

      [dead]

  • moooo99 4 hours ago ago

    In all honesty, all these distillation complaints brought forward by Anthropic etc make me enjoy the cheap Chinese models even more

  • LogicFailsMe 5 hours ago ago

    Watch any interview with the Chinese AI leaders and compare it to the unending doomer word salads from America's mightiest paper billionaires. We're losing because we have a loser late stage capitalist scarcity mindset. They're winning because they're sharing notes and one-upping each other just like we used to until 2015 or so. They have a healthy ecosystem of competing small AI startups. We have two bloated unprofitable pigs both striving to be too big to fail. My money's on China for the immediate future.

  • NewEntryHN 4 hours ago ago

    > So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why.

    Because contrarily to the author's assumption, all labs, Western or not, have sufficient skills to discover the optimizations anyway, and publishing or not is not actually that important?

  • skerit 4 hours ago ago

    > It’s beneficial for them to say that because it sets the ground for these models to be restrained legally and regulatorily later on.

    I'm glad people are saying this out loud, because that is what they want. Not for the good of the world, but for the good of their pockets.

    • senordevnyc 4 hours ago ago

      HN has been screeching about this for a long time, on almost any AI post where it’s even remotely relevant.

  • sigbottle 4 hours ago ago

    This is insanely cool, what the hell.

    How co-designed are these optimizations with the model itself? I'd imagine you can't just stick post-training adapters onto existing architectures for these things, or am I wrong?

    I really want to explore the inference space, but it seems like many of the inference optimizations are coming from model-hardware codesign. I don't seem to recall many generic "inference engine" optimizations since prefill/decode disagg a year ago.

    This matters for me since I want to break in but the bar seems to be understanding the actual theory of the training process now too given the codesign happening, and I'm not the richest guy on the block lol

  • the_origami_fox 4 hours ago ago

    https://liorsinai.github.io/machine-learning/2025/02/22/mla.... I wrote an article on MLA last year. In short, I found the main idea of MLA as very innovative but it was documented in a strange paper with other ideas I couldn't comprehend as being useful, and they hadn't properly reported the positives or negatives of MLA.

  • samuelknight 4 hours ago ago

    Did private frontier models use sparse embedding and ngram first? The article claims sparse attention was copied from open weight but we can't know that. We could just as easily argue that OAI and ANT had these improvements for years and decided to slash their margins only now to stay competitive with open weight neoclouds.

    Second, sparse attention is an old area of active research. Offloaded N-gram tables are the next big open weight technological leap.

    • throwa356262 4 hours ago ago

      The O & A strategy has until very recently been to use brute force and just throw more money at the problem.

      Deepseek was the company that invented some and improved some other ideas and got them to workreliably in production. Before that Sam and Dario were basically competing in who has the most expensive training.

  • dragochar an hour ago ago

    the gpu export controls were supposed to kneecap chinese labs, and instead they forced them to make inference efficiency their number one goal

  • caidehen 3 hours ago ago

    I started using a zero data retention(at least claimed) deepseek v4.1 flash this month

    I still use claude and openai right now, but I can see that not long in the future I won't bother with them, still waiting for a model good enough with computer use and a good enough computer use agent

  • amichae2 4 hours ago ago

    I am not a fan of Anthropic but this article offers no concrete evidence that Anthropic actually ripped off Deepseek. It is all circumstantial.

    • senordevnyc 4 hours ago ago

      Thank you!

      I’m incredibly skeptical that OpenAI is spinning up custom ASICs for improved inference performance, but they never thought of optimizing KV cache until a tiny Chinese lab did it? Give me a break.

      • bel8 4 hours ago ago

        They certainly thought. But were they able to do it now without DeepSeek papers?

        timeline suggests not.

        • senordevnyc 4 hours ago ago

          So DS comes up with 437x improvement, and the evidence that O/A copied them is that they dropped caching prices by like 50%? Really?

          • abcthingx 4 hours ago ago

            What evidence evidence would actually change your stance on this? a formal admission from OpenAI and Anthropic. Yet that's unlikely

            • senordevnyc an hour ago ago

              Any actual evidence would be worth considering. This is a pretty weak correlation and an assumption, nothing more.

  • aleqs 3 hours ago ago

    Why did this thread suddenly massively down ranked? It was at the very top 15 mins ago, now it's on page 4 and dropping fast despite very high engagement... @dang

    • bobtheborg 2 hours ago ago

      Ageeed. I wanted to come back to read more comments and had to use search to find it

  • impossiblefork 4 hours ago ago

    Yeah, and Anthropic probably got inspired to this new fast read-in thing for making agentic stuff make more sense from the latest DeepSeek model. Maybe it was in the pipeline, but it clearly has the same effect and DeepSeek had published it by the point Anthropic dropped their prices for reading tokens in, so they may well have copied it.

  • riskd 4 hours ago ago

    Wow this thread is filled to the brim with anti-Chinese sentiment based purely on… “China bad”

  • georgeburdell 4 hours ago ago

    To answer the author’s question of why Chinese labs give away their work for less than cost, the answer is involution. China is struggling with overcompetition in other areas of its economy as well, such as electric cars, and perhaps ironically its labor share of income is substantially lower than the U.S.

  • advael 4 hours ago ago

    The whole confusion expressed by this article is resolved by refusing the "arms race" framing. Chinese AI labs are acting really normal for researchers. Researchers in academia collaborate and share breakthroughs. This was until very recently the norm in American ML/AI research as well. The hawk-brained reasoning that this is some kind of "fate of the world" style arms race is as far as I can tell a narrative entirely pushed by American megacorps (and their cultural orbiters) who want to hold on to a business model of proprietary control of technology at all costs, and this is lapped up by political actors who clearly mostly just want to keep public perception in a cold war framing, which also seems mostly in the interest of consolidating power through the classic FUD method. If we think of this as normal research and development on a normal technology, it makes a lot more sense. I already see little reason to use proprietary models, but these companies insisting that they're in a war about which their supposed opponents have not seemingly gotten that memo makes me want to do so even less.

  • open592 5 hours ago ago

    > If you read the news headlines these days, you would be forgiven for thinking that the Western labs are getting spawn-camped by Chinese labs en masse.

    Ah brings back Halo 2 memories

  • emtel 5 hours ago ago

    As far as I can tell, neither of the frontier US labs have referred to distillation as "stealing", but someone please provide a link if I'm wrong.

    They do claim that it violates their ToS, which we can assume is simply correct, since they get to put whatever they want in their ToS.

    Given all that, I don't know what the fuss is. Are they supposed to not use the advances that were openly published by Chinese labs? The entire industry is built on a discovery made at Google, which was published openly. Should Chinese labs therefore not use transformers? Should US labs not try to prevent distillation of their models?

    • jerrygenser 4 hours ago ago

      I'm not sure if they don't refer to it as "stealing" but they refer to "distillation attacks"

    • dgellow 5 hours ago ago

      I don’t think there is fuss, just the author sharing the information and mentioning how they find it a bit ironic that US labs expenses can be reduced drastically thanks to the Chinese companies they continuously frame as adversaries

  • why_only_15 4 hours ago ago

    Why do you think the Chinese labs figured this out before the western labs? No reason to believe that whatsoever.

    • bel8 4 hours ago ago

      Why wouldn't you? It's the most plausible interpretation given what we know.

      Had western labs figured that out before, they would have used it to make kv caching cheaper before and not only now.

      The burden of proof here is on western labs. But I doubt they'll try to lie that much.

      • tescreal 3 hours ago ago

        I dunno, they've been competitive in the lies market lately. It's just a matter of ambition.

  • thelaxiankey 4 hours ago ago

    it's funny to me how the success/not total implosion of htese companies is predicated on profitable, revolutionary-tier success, and that it's increasingly possible that the profits will never really materialize. pretty interesting move on China's part.

  • undefined 5 hours ago ago
    [deleted]
  • dualvariable 2 hours ago ago

    Once again, the irony of "Hacker" news full of so many people carrying water for the IP rights of companies worth trillions. Phrack magazine must be rolling over in its grave...

  • underlipton 4 hours ago ago

    It's actually a little funny that this whole thing is predicated on "beating" the Chinese, when (as they have been for the past 4 decades, and the Japanese before them) they're perfectly happy to let us do the bulk of the work and then swoop in with a svelte, cheap, user-friendly version right after. One part Apple, one part Dollar Store.

    Is the thinking that the day or so between US systems achieving ASI and Chinese systems doing the same, we'll figure out a way to neutralize them indefinitely? Because otherwise, none of this makes much sense. And it only starts to swerve back to sanity if the assumption is that this isn't a race or competition, but instead a joint effort to achieve something good for humanity. But you can't really delta profit off that, can you?

  • adamrezich 5 hours ago ago

    OpenAI is jobbing (in professional wresting terminology) hard right now.

    • brcmthrowaway 4 hours ago ago

      So who is the kayfabe?

      • adamrezich 4 hours ago ago

        Did you not see the meeting with the President yesterday? All of the “safety discourse” was kayfabe.

  • LunicLynx 4 hours ago ago

    The clue is: Bursting the bubble

  • nater5000 4 hours ago ago

    >The new game in town is adopting Chinese labs’ advances. Note how I call this adoption instead of the more vitriol-infused “stealing” that Anthropic tends to use.

    I mean, there's a pretty big difference between labs publishing their research openly and a competitor utilizing it versus a lab breaking TOS to... hmmm, what's the word? steal data from a competitor?

    >That’s because, unlike the Western companies, the Chinese are pretty much giving away their recipes.

    Yeah, Western AI companies have never published their research. It's crazy how the Chinese had to independently develop the foundational technology that powers LLMs because Western companies simply never publish their research (I mean, as long as you ignore stuff like this <https://arxiv.org/abs/1706.03762>).

    >The latest one shamelessly copied without acknowledgement is the breakthrough in KV cache optimizations that DeepSeek has generously shared with the world.

    Thank you, generous corporation. I'm sorry that other corporations don't provide you free publicity for your selfless contributions to the world.

    >Now I don’t know why they would freely give away such a breakthrough, but they just did

    Well I'm glad the author finally got to their point. A very insightful analysis.

    >They do seem to be a little embarrassed by the copying. Hence the silent releases without much pre-announcement for both Claude Opus 5.5 and GPT-6.1 Sol.

    You have to be in pretty deep to infer this kind of emotion to these kinds of corporate activities.

    >So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why.

    Then why write this article? Why point out these things just to have no conclusion?

    This article sucks. Even if you hate US AI labs and are all aboard Chinese labs producing open models, there's nothing of substance here. This is the loose draft that you hand to your LLM to finish for you, but it seems the author just forgot to do so.

    Even if you're willing to characterize US AI labs as evil and selfish and Chinese AI labs as righteous and generous (which is already completely trivializing these dynamics to the extent that anybody over the age of 14 can likely identify is lacking nuance), you can at least put some effort into producing some hypotheses about why these dynamics are occurring. Of course, odds are if the author did try to articulate some hypothesis, they'd likely quickly realize that the narrative they're painting just doesn't hold up.

  • revexos 4 hours ago ago

    too much pace

  • senordevnyc 4 hours ago ago

    Color me skeptical that OpenAI and Anthropic’s researchers had never thought to dig into these optimizations, and instead are just spending hundreds of billions on data centers and custom ASICs.

    This is an extremely thin analysis that has obviously been voted to the top of the homepage because HN hates the big labs.

  • jgrahamc 5 hours ago ago

    [flagged]

    • emilecantin 5 hours ago ago

      It's from video games, where a player "camps" near the spawn point and kills newly-spawned players, presumably with better equipment.

      It's not that niche, if you've been online a little bit you'd know this expression.

      • tejohnso 5 hours ago ago

        > It's not that niche, if you've been online a little bit you'd know this expression.

        No way. You'd need to be pretty well versed in gamer lingo. Even more specifically, combative, likely FPS gamer lingo.

      • jgrahamc 4 hours ago ago

        I dunno, man, I first got on the Internet in 1986 and was Cloudflare's CTO for years. I've been online quite a bit.

        • VGHN7XDuOXPAzol 4 hours ago ago

          (tongue-in-cheek) In that long time you've been online, have you not come across the concept of a search engine?

        • pigpop 4 hours ago ago

          Due to shifting definitions, I believe that makes you a boomer[0] so you'd be readily excused for not knowing.

          Sarcasm aside, gaming and FPS terminology are so tightly coupled with online culture that it's just assumed everyone knows it. Spawn camping is among the oldest examples of gaming terms that broke out into common online usage and it dates back to Quake some time around 1997.

          [0] anyone older than 39 at this point

          • tejohnso 3 hours ago ago

            > FPS terminology are so tightly coupled with online culture

            What is online culture? Is someone who is heavily into instagram for fashion, facebook for family contact and news, maybe Google for mail and search, part of online culture? Because I know people who are like that and there's no way they know what "spawn" or "camping" mean in gamer context and certainly wouldn't be able to piece together what "spawn camping" is.

            • pigpop 2 hours ago ago

              Those things seem fairly obviously not online culture but parts of "real life" culture that have an online presence, obvious to me at least. Video games and especially multiplayer video games are a core part of what you could call the indigenous online culture and serve the role of sport and to a large degree casual socialization. Their equivalents in the real world are team sports, board games and party games. These things aren't mutually exclusive and they certainly bleed over both ways but you can easily say that a game like Counter Strike is primarily online and a game like Basketball is primarily offline even though there are Basketball themed video games and Counter Strike themed in-person events.

              To put it simply, it's a Venn diagram with two overlapping, non-coincident circles. The amount of overlap has varied with time but that doesn't mean online culture doesn't exist or conversely that everything is part of online culture.

          • khazhoux an hour ago ago

            > gaming and FPS terminology are so tightly coupled with online culture that it's just assumed everyone knows it

            Strong consensus bias in this statement. I have no doubt it is true in your circle (and to a lot of people). But there are plenty of online-24/7 cultures that don’t have any overlap with gaming.

        • johnthescott 4 hours ago ago

          jgc, you made my day, one 86'er to another.

        • off_with_their_ 4 hours ago ago

          [dead]

      • mjc26 4 hours ago ago

        Also, it's possible that jgrahamc is the former CTO of cloudflare

      • cassianoleal 5 hours ago ago

        > if you've been online a little bit you'd know this expression

        I've been online since circa 1995 (earlier if you count BBSs), and I can't say I did. It's possible to infer its meaning but assuming everyone is on the same circles as one is, is silly.

    • bsoqk 4 hours ago ago

      We live in the era of LLMs, which can produce definitions for any word, further explanations, and limitless examples.

      • jgrahamc 3 hours ago ago

        Yes, but to expand on my flippant response, the opening of this post is a sign of bad writing. And I don't want to waste my time on bad writing.

        It shows that the author hasn't thought about their audience and has assumed that everyone knew and used the same terminology as them. And, worse, assumed that they'd understand immediately why they were using that terminology.

        The opening is If you read the news headlines these days, you would be forgiven for thinking that the Western labs are getting spawn-camped by Chinese labs en masse. If you don't know what spawn-camping is you're lost; if you do it's not obvious what that means in this context. Good writing brings the reader along with the writer.

        It would have been clearer if they'd written: If you read the news headlines these days, you would be forgiven for thinking that Western AI labs are being outplayed by Chinese AI labs using something similar to the gamer technique of "spawn-camping". The Chinese appear to be waiting for each Western release and then instantly distilling it. A little like gamers waiting for their opponents to reappear (from the dead) at their camp and then kill them off immediately.

        This is better because a reader unfamiliar with the idea of spawn-camping learns something and it explains the metaphor. But I could be wrong in my interpretation of why they are using spawn-camping since they fail to explain it.

    • khazhoux 5 hours ago ago

      In first-person shooters, when you die you regenerate (“respawn”) somewhere on the map. If those regeneration points are know to opposing players, they can wait next to them, and kill you again the moment you respawn.

      The premise in this article is: Western companies do a ton of expensive work building new models, meanwhile the Chinese companies just wait for a Western release and then they immediately grab and distill it and announce it as their own model. That’s the spawn-camp.

    • jamiek88 4 hours ago ago

      So you had the opportunity to learn a new phrase but instead discarded the whole text because of something you hadn’t encountered before?

      Do you have many mini tantrums like this per day? Probably makes you very difficult to work with Mr I was a CTO.

    • Intragalactic 5 hours ago ago

      If you tend to stop reading every time you encounter something you don't understand, I can't imagine you learn very much

      "spawn-camping" is the process of taking out your enemies at the point they spawn (or appear) in a game without giving them a chance to regroup. In this case I think the writer is saying that the news implies that western models are getting distilled on release. Not the perfect analogy but it gives some color.

    • some_furry 5 hours ago ago

      It's gamer terminology.

      In a PvP (player vs player) game, if you kill a player the moment they spawn into the game arena, that's called "spawn-camping".

  • nba456_ 5 hours ago ago

    Appropriate domain name

  • Handy-Man 5 hours ago ago

    Just assumptions, nothing backing it. So maybe I'd sit out calling others out.

    Edit: Apt domain.

    • squidbeak 5 hours ago ago

      Deepseek's innovations are published as research. There's nothing 'assumed' about this. The common slur repeated in the West that Chinese labs are parasitic distillers is totally absurd when so many genuinely valuable advances and contributions to the field are published openly by China's labs.

    • slowin 5 hours ago ago

      How is it just assumptions? They provide the data to back up their claims.

      • Handy-Man 4 hours ago ago

        I am talking about correlating Anthropic/OpenAI cache prices going down with Deepseek publication - neither of those labs have said that's what they used for example.

        And the only data they are showing is that cache prices went down for new Claude/OpenAI models but that's proving nothing, IMO.

  • airtnp 4 hours ago ago

    A random accusation of western lab using DeepSeek caching technique being top of HackerNews, this article feels more awkward than the AI race for me. It's a shame for all HackerNews viewers.

    • abcthingx 4 hours ago ago

      I don't think it just a rumor. But we can't tell since western lab architecture are secret.

      • airtnp 3 hours ago ago

        This is the exact definition of rumor.

      • cindyllm 4 hours ago ago

        [dead]