Its quite interesting to see that at least the early days of AI so far have not been a winner-take-all runaway acceleration game where catchup is impossible.
I certainly wouldnt have predicted that 10 years ago.
Very glad to see Mistral still in the game even after some big stumbles with Large 3. I deeply hope that this model is 'good enough' that it becomes the European go-to, giving them the resources to keep the pace up.
I think it might have accelerated things but on a much more basic level, there seems to be no real moat in synthesizing the world’s knowledge into LLMs.
There's no question that training leading LLMs requires some serious expertise and know-how, but surely already having advanced LLMs/agents must be helping tremendously not only for software engineers but also for those working on LLMs themselves.
I think one could describe LLM optimization as "hard but not a moat". Years ago, optimizing neural nets was described "graduate student descent" - it's tricky but throw enough conventionally smart people at it and it will happen. It's like tuning a hot rod and finding a reproducible bug in a large code base. It's hard and there are tricks but not absolute hurdles, no problems waiting for a conceptual breakthrough (and at today's scales, are there any problems waiting for an Einstein to solve? That's an open (AI) question).
What is really going on: all the AI labs are doing panicked model releases (and panicked training of new ones) because Qwen4 is rumored to come out end of October and is rumored to be very nice. Question is: is it another "Deepseek-moment" nice? Or just nice?
Btw: with Qwen4 I mean the next large Qwen model that is based on the Qwen4 architecture (Qwen 3.8 flash next was "almost" based on the new arch but obviously was a small model)
What matters more is if firm’s start using a bundle of American and Chinese models and when they find their feet - how large is the market for frontier?
Frontier has to displace labour one for one at some point or it’s over.
I have similar hardware - what specific version of the 3.8 Next model are you running and how many tokens/s are you seeing? 3.6 35B A3B gets about 68t/s for me so I've been sticking with that model for the mean time.
I think the moat is going to be compute. So far compute needed to push the frontier is still extremely cheap so the capital can afford to spread its bets. But when further improvement is going to cost in trillions, capital will have to pick a winner and bet only on him. It won't be a matter of finding the best bet, it will be a matter of survival.
This will cause the picked winner to get massively ahead with sheer compute alone used both for training and inference dedicated to recursive self improvement.
I guess all predictions age like milk, but here's one:
There's a law of diminishing returns at play here, and doubling the energy cost of training to wring 2% more performance out of the technology isn't going to be very useful, because most of the problems it is capable of solving will be solvable with the previous-gen 98%-as-good model.
("there's a law of diminishing returns at play here" is an article of faith. But then, so is the belief that these models will keep getting better).
As soon as you can demonstrate decent financial returns (ie. the AI can run a company better than humans can), suddenly it makes sense to put a lot more $$$ in even if returns are diminishing - since whoever runs companies the best gets control of a big chunk of the world economy.
Well yes that's the worry isn't it? That people will soon all be unemployed? Just because it sounds farfetched doesn't mean you should stick your head in the sand, seeing the pace of development these days, that is "reasoning properly" and it just seems you are trying to block out what seems inconvenient to hear.
If you're actually applying LLMs, all of the things around the LLM that adapt it to coding, for example, that enable it to use existing validation tools for code, and enable it to diagnose and fix tool chain issues that aren't directly coding problems, are what makes the difference between a model that that scores a little higher on a coding benchmark and a model that's useful in a particular code base on a particular platform.
Are there any use cases that have enabled one customer of a frontier LLM to outperform a competitor using a different frontier LLM? Or is this why we are seeing confected points of comparison like solving challenge problems in mathematics?
I'm afraid it might be the other way around. RSI might pick all of the low hanging fruit soon. There must be a physical limit of how much intelligence you can squeeze out of some amount of parameters and compute.
There are going to still be worthwhile improvements but they are going to be more like not how to make transformers 10x cheaper but how to make next training run cost 9 trillions instead of 10 with a very particular optimization designed at the cost of hundreds of millions for this one specific run.
It is just occurring to me that “RSI” expands to recursive self improvement. Thought people were talking about repetitive stress injuries; either in regards to programmers writing too much code/not having to write code anymore, or the frontier AI companies and their tendency to applaud themselves.
But the old models still exist at trivial marginal cost. The frontier models would need to dominate every price point to really take all and so far they haven't been.
Build Nuclear, Build Thorium the Chinese are building whatever they can. They’re not locked in by special interest. Is that because they have lots of engineers on the job in government?
France is actually pretty cheap in Europe. About 15% more than average USA electric prices (but I know that varies a lot across the states so still likely much more than the cheaper areas)
My spend on DeepSeek is not much and I regularly top up my balance every month as my support for all the good work DeepSeek is doing for the open science.
DeepSeek, and other Chinese models, are heavily subsidised by the Chinese government. The reason they release the AI models is economic warfare against US, not because of charity or kindness. It's great for us consumers, but the goal is not to help humanity or open-source.
Framing it that way hides all the levers used to tilt the playing field for US companies, no?
The government that removed restrictions on how private companies can access capital after a certain scale (the JOBS Act), that removed the need for private companies to report as if they were a public company after a shareholder threshold was crossed, superpowering the access of wealthy private investors to get in earlier in a growing company while at the same blocking the public from participating in funding growing enterprises at an earlier stage (since it required companies to IPO much earlier to access capital) which allowed retail investors to also reap the rewards on funding them early when they grew to become behemoths (like Amazon, Meta/Facebook, Google, etc.).
It's not fair in either place, the USA has its own model of unfairness, China has a completely different one. The difference is that in the USA the government allows private investors to become more powerful than the State (outside of the monopoly of violence) while plunging the rest of society into increasingly more precarious lives while in China the State is the power and its legitimacy only exists while the population feel they have a better life.
IMO the selective enforcement of regulatory requirements should be added to the list. Observing from the outside, I have a hard time believing that e.g. musks gas powered data centers really follow all the environmental laws, for example, or that the authorities really see no grounds for indictment if Altmans company hacks hundreds of third parties, or that there's really no one at the SEC having a problem with Anthropics fear mongering prior to the IPO.
In a list describing how the USA system works, sure: selective enforcement is there.
As a differentiator from the Chinese system, not so much. For two otherwise equal companies, the one that says things against the party line will experience selective enforcement too.
Except that in China, the people who do that will disappear.
The propaganda of trying to make US and China government appear the same is making people dumber. and is one of the most heavily used tools in China’s online propaganda arsenal.
If you think about it, the US labs are heavily subsidised too. Not only they receive billions in state funding, the administration is also prepared to engage in trade wars to help them.
I think the Chinese government is backing their labs by less direct means. For example cheap electricity and investing in chip manufacturers such as Huawei and cxmt.
Deepseek specifically, is known to operate with minimal resources. The entire company has around 160 employees and every model they release must break even within ten months.
SpaceX has been awarded roughly $22 billion to nearly $30 billion in cumulative public federal contracts, the $280 billion CHIP act, and who knows how much the CIA + NSA are spending.
I'm pretty sure, the training/fine tune through Claude/OpenAi would not be possible without the army/China's hacking teams (and legal protection). So it's more than just a cheap electricity.
Supporting deepseek is just like supporting the Chinese army, no need for that. Though it goes both ways, OpenAi subscription just lowers the cost of the US army as well.
Sorry, are you saying that it's a subsidy that it's legal for them to distill other models, or what? It's not even clear that there's any kind of protection for models in Western countries. Why would they be worried about the legality of distillation? And what does the army have to do with distilling models? Like you're saying it's a subsidy that China protects their borders from invasion by the US?
It's not a subsidy for the Chinese to say they're going to ignore our IP laws. It's a subsidy for us to say we're going to make them and try to push them on the world. It's literally granting a monopoly by legal force. It's very obviously not aligned with the interests of the American people, while China releasing things in the open is.
As bad as casualties get in a future war like that, it's still very likely to be a symmetric war where the main targets are military infrastructure. No systematic annihilation of universities, hospitals, religious sites, journalists and health workers like certain genocidal regimes backed by the tech oligarchies have been doing.
When one player tries to monopolize AI tech and market by all means, others are those who don't want to be hooked to a foreign will in the future. I think this is the reason in doing open research, and it's likely positive for all of us.
> I think the Chinese government is backing their labs by less direct means.
It may be that the effect of the backing in both places is essentially equal but this statement is strange. The Chinese government invests directly in Deepseek[1]. Notice the article, in addition to saying the CCP is investing, says Tencent is also a major backer. CCP owns a golden share of Tencent.
I'm not arguing about a comparison here. I'm trying to correct a misconception about China that I see a lot from Western perspectives.
Everything a Chinese company does has the explicit backing of the CCP, at least ideologically and usually financially in some fashion. While DeepSeek may not be the CCP, it couldn't exist if it was expressing any kind of ideology that wasn't inline with them and, in this case, is explicitly funded by them.
A Chinese company has no freedom to say Taiwan is a country the way someone in the US could suggest California succeed from the nation.
Any public message you hear coming out of China has the implicit approval of the Chinese government.
I would like to point out that a lawsuit is currently live about the us govt forcing US tech companies to retaliate against us citizen criticizing the govt’s policy.
> Everything a Chinese company does has the explicit backing of the CCP, at least ideologically and usually financially in some fashion.
If you're not arguing about comparison then don't state this as if it's any different from the US. Because you make it sound that way. Or do you think that if tomorrow OpenAI came out vocally supporting the DSA - that suddenly they wouldn't see lots of barriers rise up out of nowhere?
Well, they are helping 17% of the world's population, and the US is currently actively engaged in trade wars and economic warfare or explicitly attempting to leverage it's hegemony against its long term alliea for short term gain.
There is as much to criticize about American hyper scalers and AI labs and the lack of interest in helping humanity or contributing to open source, but that might not be as popular an opinions on this site.
17% is the approximate population of Earth in China. I was using a simple example as a counter to the anglocentric view the parent poster was expressing by trying to paint China as a villain.
> Well, they are helping 17% of the world's population,
Any open weights model that Chinese companies are releasing for free are allowing everyone in the world to have access to high quality realizations of these tools without the risk of the US regime arbitrarily cutting you off.
Also, it's funny how all the neoliberal mantras of how free market drives progress through competition stops when it's a US company that's being challenged.
> The reason they release the AI models is economic warfare against US
The story is so much more complicated than that, to the point that this economic warfare theory is basically a meme.
Chinese models are open because they don’t have a choice. “When you trail the frontier, openness maximizes reputation per unit of capability. The moment you lead, you close.” [0]
They do have a choice, as evidenced by all the times when the other option gets picked. The article you link tries to acknowledge ByteDance's Doubao and Alibaba's Qwen Max as exceptions, but forgets about Baidu's Ernie and iFlytek's Spark, which are also closed-weight. There's simply no consensus yet on which strategy is better, so different companies end up making different bets.
If you want a direct example of the US subsidising a major AI lab, we can use Microsoft's investment in OpenAI. Microsoft has the largest ownership stake - 27%.
Microsoft did not pay with money - it paid (mostly) with Azure cloud computing credits. MSOFT is then able to write this off as a loss against tax.
It is generally far more tax-efficient in the US to write a loss in this way than it is to write a loss for a cash investment.
In this case, I believe the difference was ultimately highly significant. When including MSOFT eventually writing-off the deprecating Azure hardware it had used to buy the OpenAI equity, the result was MSOFT's tax reduction being either close-to or exceeding the actual cash value of MSOFT's investment in OpenAI - IIRC.
These examples represent taxes that the US chooses not to collect - the US could choose to make investments like these less tax-efficient. Instead, by making them extremely tax-efficient, the US subsidises the transaction hugely.
Even ignoring monetary subsidies, there are the non-monetary ones: not being sued into oblivious by the government for their countless hacks of other companies and countries, the slaps on the wrist for massive piracy, the waving of environmental (and other) regulations in order to allow their data centres to be built an operated.
That's what I'm asking - which monetary subsidies?
> not being sued into oblivious by the government for their countless hacks of other companies and countries
That isn't normally how enforcement works, and it hasn't been very long since they disclosed those breaches. If the victims want to pursue legal action, they can, and they still may!
> the slaps on the wrist for massive piracy
So judges and juries are involved in the subsidization conspiracy, too?
> the waving of environmental (and other) regulations in order to allow their data centres to be built an operated
Sure, though if you think this isn't happening in China too, I have a bridge to sell you.
The original comments is "DeepSeek, and other Chinese models, are heavily subsidised by the Chinese government"
> That's what I'm asking - which monetary subsidies?
I am listing for you the non-monetary subsidies, which are just as real and equally important.
> not being sued into oblivious by the government for their countless hacks of other companies and countries
If it was one of the Chinese labs doing this hacking, the government would be stepping in. If it was European labs they'd be stepping in. If it was you or I the government would be stepping in. That is a massive subsidy (they don't have to worry about the same legal fees and exposure) and being allowed to continue doing business is in fact priceless.
> So judges and juries are involved in the subsidization conspiracy, too?
There is no conspiracy. They are objectively operating by a different set of rules than you or I could operate in this market.
> Sure, though if you think this isn't happening in China too, I have a bridge to sell you.
I never said it wasn't, I'm saying that it's happening here and it's a very real subsidy.
All the major labs would be dead in the water if the government acted on the things I mentioned. If they treated the labs the way they would have treated us for all those hacks. Or if they enforced pollution measures (or god forbid ban on-prem turbines because of the climate damage). Or rule that training on data is not fair use, or that ingesting GPL3 code and then turning that into weights counts as a derivative work. Or. Or. Or.
>Fair enough, but it sure meets the definition of 'helpful intervention'.
I don't see how. Anyone can buy electricity without any intervention from the government. An intervention is an action that changes what would happen by default.
I heard someone calling the key metric in Anthropic financial reports EBBT: Earnings Before Bad Things[1] :-)
[1] Where "Bad Things" would be the typical interest, taxes, depreciation, amortisation plus the Anthropic specific employee compensation, LLM training (you know, for the LLM lab), revenue sharing agreements (which is a form of paying for infrastructure), etc.
A recent YouTube video by Patrick Boyle said the same thing. They are only profitable if you ignore all the costs that make them unprofitable such as paying employees and developing A models.
I made a statement about whether Anthropic was profitable I don't understand your reply.
Assuming you meant to reply to me and not someone else are you saying under GAAP accounting standards Anthropic is a profitable business because under GAAP accounting their only expense is electricity?
(I see what happened you skimmed the conversation and didn't follow what was being discussed.)
Indirect subsidies are an accepted economic concept and well studied. The posts you replied to are obviously talking about indirect subsidies and dismissing them because they didn't say "indirect" everywhere is just sophism.
The government, specifically Trump's government and the current money circle in AI inflating American company stocks. In addition to all that American models do not share their papers like Deepseek and Qwen do. So you can literally say Chinese models are doing it for charity at this point.
The US regime, which includes the ruling class: the US VCs and megacorps are just as much of an extension of the US regime. They're incredibly intertwined.
> The reason they release the AI models is economic warfare against US, not because of charity or kindness
There are many other reasons Chinese companies releasing models open-source or open-weight makes strategic sense.
A really easy-to-understand example is a company who has a near-monopoly on "serving video content" releasing a video model openly.
If you can be relatively certain that video content created by a model (which you have trained, using data from your own platform) will be ultimately served on your own platform, thus generating revenue from watch-hours, it makes sense to make those models as widely-available as possible.
It's also a net-positive if people use your public research to build better video models, because - again - you are reasonably certain that the even-better content those new models produce will be watched on your platform.
The alternative would making models harder to access and learn from (broadly, the current western model). Many would argue that Google, in choosing to not optimise its video generation models for "availability", is directly causing less content to be uploaded to YouTube. This is the trade-off.
I don't know much about DeepSeek's financing specifically, which obviously doesn't release video models - so I don't know how directly this analogy runs, or who directly benefits from the extremely evident rising tide that the public release of DeepSeek's research creates. However, this does not negate the broader rising-tide effect of the scientific method.
It's certainly also true that it's geopolitically beneficial to be able to undercut American labs' models. If I ran a global superpower, I would probably want my country to be technologically competitive too.
But Chinese companies are already serving a huge volume of customers in a complex, existing marketplace, before even thinking about the US market, and it's overly simplistic to assume that their entire strategy revolves around economic warfare directed specifically at the US. It's more nuanced than that.
This is, of course, without even getting into opening the can-of-worms around whether US economic policy also results in the US state functionally subsidising technological innovation, how comparable that is to China's model, etc.
Interestingly if Google did release readily available video models then they'd likely be dealing with an order of magnitude of scale in the same way GitHub has had to.
Causing less content to be uploaded, when you are clearly the gorilla in the room, looks to be a wise strategy not a trade off.
I'm not sure "subsidised" is the right word. If a government funds research and the results are released openly, that's just publicly funded research. It's how a lot of science works in the US and Europe too.
Same goes for the US companies as well. Current US administration wants US to win the AI race at any cost and China is the only competitor left in the race. Winners always write the history or in this case the future of humanity
The open-weight models are a great boon to all American companies other than a handfull of Mega Corps in the A.I business. Care to explain how this is "economic warfare against the U.S"
Yes we see that in all other countries trying to avoid dependence on the US for everything. What we don’t see is competition with other countries having anything to do with warfare.
It is great for everyone except for a few people who want power over everyone else, and the fact it is not charity of kindness makes it more sustainable, because charity and kindness is quick to go when big money and politics is involved.
I want more warfare like this. Building stuff instead of destroying stuff.
So you’re telling me warfare doesn’t having to be blowing up schools and children, skyrocketing prices, crippling sanctions, and all that shit? Can I sign up for more of this warfare.
Not sure why so many people will vehemently refuse this idea. I won’t say it’s 100% true but it would be foolish to dismiss it. China is very much an adversary to America and has made it pretty clear they want to be a dominant leader of not the new world leader. Not here to evaluate what is good or bad. Keep in mind historically China has aggressively fostered industry (not unlike the west) but sometimes even more aggressively.
Americans need to travel to China. The Chinese have zero issues with us. They quite like Americans. This is such a weird propagandist take. Idk if you remember but both country's leaders just had a slumber party for 3 days in DC. This is not what enemies do.
Not even our leaders say China is an enemy, ita mostly businessmen who are scared of competition and trying to regulate chinese out of their markets so they can make more money milking us.
I have been to China many times. The citizens and visiting the country has nothing to do with global politics. Your take is just as propagandist as any other. China is not unlike the US and China has also made it subtly clear their desire to be a dominant global force. I am not marking judgement on it but don’t be a fool thinking that China is not some level of threat. This is how these discussions become weird. I don’t think China is going to blow up the US and I think it’s often a weird talking point by some US politicians but I also don’t think they are a peaceful actor but folks like you will suggest I am repeating propaganda. No, just pointing out that if you look at their language and actions they are an actor to pay attention to. Just because leaders meet means nothing.
Economic warfare against the US? I mean maybe against specific US companies and stakeholders but on the whole it seems like it's good for the US economy as well as the rest of the world, kind of like supplying free electricity would be
It just so happens that helping open source is the result. Image how far behind we’d (the hackers, not the moneymen) be as a sharing community be without them.
At least DeepSeek didn't build its models on government subsidies. It came out of a quantitative hedge fund, so funding was never really an issue for them.
If your belief is accurate, we should expect China to short the IPOs of Anthropic and OpenAI and release better frontier models immediately after their IPOs.
Does anyone think that likely? I have no clue or bias.
I don't think that's likely because my understanding is that Chinese people in mainland China have a tricky time shorting American stocks due to Chinese capital controls.
It's one of these rare cases where intention is not what is the most important - the net benefit for consumers and companies outside of the USA is indisputable.
What do you think your comment is refuting!!? Chinese are publishing their research publicly and opening their models! How on Earth would they monopolize anything?? Such nonsense. If and when they start copying the Americans and closing everything down, you can say stuff like that. Until then, ffs stop this nonsense.
Economic warfare against the US is charity and kindness to a sizable portion of the world's population, especially when the US uses it's global hegemony as warfare against them. Neither system is perfect, but lets not be disingenuous.
>
DeepSeek, and other Chinese models, are heavily subsidised by the Chinese government.
So were Amazon and Uber by the US, which have now established monopolies across the globe. To the countries suffering from those, there's zero difference with China doing it to solar. Actually there is, at least solar got them cheap renewable energy in return. This would never have happened in the US because big oil interests would make it take decades. That's the reality.
You need to spend 10 years outside the US, deprogram, and then go back.
> DeepSeek, and other Chinese models, are heavily subsidised by the Chinese government.
Yes, because everything China does is against the US. That's all they think about day and night. God forbid they want to corner the global market or have a genuine business case. How dare they provide options for those who can't afford a measly $200 a month? How can we let Chinese labs publish research for free for the whole world so that they can benefit? The nerve! To think they can use soft power instead of military might! I mean, Anthropic and OpenAI are the last bastions of human kindness and charity. Right?
I’m neither in the camp of Chinese or the Americans(collective West) in general..
As a neutral party, this characterization is crazy.
As if the AI companies - Claude and OpenAI are guardians of freedom and humanity and very charitable to the global society without any self interests…
“Chinese models are subsidized by the Chinese Government, therefore they’re inherently bad for humanity” is a highly propagandist argument. The politics of US vs China may be whatever it is in reality.. You have one company releasing their models for cheap, actually open sourcing their trained weights, and publishing details of their optimizations and learnings for others to use. The other camp actively “aligning” their models, nerfing their capabilities, hyping their swarm activities from poor sandboxes, and trying their best to lock users into their harnesses and walled platforms. They are subsidized by the capitalist VCs who are essentially waiting for their payouts..
At some point, one has to see things for what they are and evaluate their own reasoning..
I’m happy to stay provider agnostic, try all models and cheer any useful progress as open as possible.
> DeepSeek, and other Chinese models, are heavily subsidised by the Chinese government.
US companies are burning colossal piles of cash in ways that makes it unclear if it qualifies as dumping, not to mention their deep ties with the country's regime.
Claiming that companies from a country have ties to the regime and burn through cash is a very miopic accusation.
I wish US was rich enough to subsidise development of open source science and useful open AI models. I wouldn't mind US waging this kind of wconomic warefare against China or everyone else on the world.
Charity and kindness is not a motivation, it's an outcome of what you do.
Meanwhile, American AI models are heavily subsidized by stock market speculation. Ultimately, the subsidies from both countries are flowing out of the pockets of individuals.
> are heavily subsidised by the Chinese government
We hear this about literally every industry the Chinese excel in - that it's only because the government subsidizes them that they succeed. For chip manufacturing, for batteries, for EVs, for solar, for AI. I don't see how the chinese government can afford to subsidize all of these industries and still have them contribute to the GDP.
A conspiracy to make the US look bad by being better at producing all the goods and services the world needs at a reasonable price. Have they no shame?
What's wrong with governments subsidizing scientific work?
You think extremely US-centric. China has a different economic model than US and your rules for a specific kind of Capitalism may not apply to them. You assume a country of 1.4 Billion people is obsessed with a couple of foreign AI companies. What if they don't care.
Not really this is an X algo conspiracy. Up until recently the Chinese government wasn't even that invested in these companies. We're talking very very small grants compared to training costs.
Its very xenophobic of you to say China has zero intention of helping humanity, and just wants to "wage economic warfare".
Last time I checked, it was ourselves (USA) waging economic warfare on 2/3rds of the world.
I dont get this cope people have where people have this idea that its impossible for a Chinese company (that make billions of dollars) to have done something by their own merit, but instead its always some Chinese Communist Party conspiracy where the main goal is to destroy America.
Why the anti China propaganda from you? The United States is doing the same thing here and making American multi-billionaires even wealthier. The greed of American AI, GPU and memory and storage corporations is a black hole on availability to humans around the World. The result is a massive financial bubble promoted by the American Government to the detriment of our citizens. The USA has an AI ponzy scheme shuffling the same money between data center owners (Oracle & X), Nvidia and memory & storage vendors.
They're trying to pull digitally what they already pulled physically. The reason we can't manufacture a grill brush for a reasonable price is the result of years of Americans choosing the cheapest price. We gave up our manufacture base. They want us to give up our labs.
Deindustrializing a country is not something consumers can achieve. It starts at the top level, with politicians who construct a financial system where it's more profitable to speculate than to build or invest in real businesses.
> The reason we can't manufacture a grill brush for a reasonable price is the result of years of Americans choosing the cheapest price.
No, it's the result of US leadership letting this happen. This is clear since China themselves would not let this happen. US leadership did nothing because they were best friends with the people who did it, and did not care one bit about their population.
Architectural/algorithmic tweaks do advance the efficiency frontier nicely. But raw intelligence mostly comes from data (not just its sheer quantity, but also how it's curated & cleansed) and the scaling law. The know-how about data curation doesn't seem to get published much, even among the open-weight labs, though.
I'm not sure I like this framing - so much of AI research has been academic, in the open, building on others people's work. Much less comp sci generally, math & philosophy, etc. The idea that rich companies can just build stuff in secret because they have resources is a fantasy.
Europeans are irrelevant these days, surpassed by China, S Korea, Japan, Hong Kong, Singapore, etc. Europe is coasting on former glory and now has regulated itself to death and vacationed its advantages away.
Also, the field moves fast, but slower than people do. Researchers and engineers switch companies every year or two, and the know-how walks out the door with them.
It's shaping up to be much more like a game of 'chicken' where each company tries to raise more cash without going bust...
Ultimately the game of musical chairs is going to have to stop. In the US it looks like they are trying to get a government sanctioned truce in the form of regulation. That's what 'Pacing the frontier' means...
Even "runaway acceleration" isn't instantaneous. People imagine the singularity as something that happens almost instantaneously. But obviously it happens over time, and that time might be decades. It might still end up looking like a vertical line on a long-term graph.
If the singularity is defined as an AI sufficiently intelligent to improve itself independently, that AI is still limited by the resources required to do this improvement.
You've got to think the comments about "good enough" are people who have not yet tried Opus 5.5. I haven't been this struck by a step change since 4.5/4.6. It's a bigger jump even than when Fable first arrived.
As for Mistral - I got really excited when they said Large 4 was focusing on being #1 in cybersecurity, because that's somewhere that they genuinely could edge out Anthropic & OpenAI. Have it actually solve problems, instead of Anthropic flagging "you tried to find a null pointer exception bug in your own code, we're now reporting you to the US government". But on the Mistral benchmarks I'm seeing, this looks very disappointing, but at least they haven't entirely given up. I genuinely thought Mistral had given up on new general models. They need to learn the bitter lesson all over again.
Many on HN still have the opinion that you must understand every line of code in the project, and that all is lost should you merge code that wasn't reviewed.
Obviously any model will do if you use it as a better autocomplete.
I believe that there is a large gap in expectations between different workflows.
Until the AI like reads my mind and produces perfectly production ready apps with minimal intervention from my side, there is still going to be room for improvement.
If Opus 5.5 is 500x better than Deepseek, but Deepseek can solve all your problems, maybe you need to work on better problems. If you don't, and you're in business, your competitors will work on the better problems. If you're an employee, your employer might prefer to pay Anthropic instead of you. If you're doing projects you're interested in, you can tackle more ambitious projects with a more capable model.
This morning I elicited a microkernel operating system from Opus 5.5. Well, mostly. It doesn't implement task switching yet; we'll see if it runs into a wall at some point. But it boots in QEMU, and it's running a user process in ring 3 and serving web pages.
The best problems to work on are not necessarily the hardest ones, nor the ones that need the most intelligence. They're the problems you, or other people, actually have. Are you going to give up on painting your deck because it's too easy and you don't need a 500x genius to do it?
I have enough real problems in life. I don't need to invent new ones just because a new technology is available.
Many of my problems in life are fully solved far past my satiation point by a 3b model that costs me nothing to run.
Many others are not.
But in either case, when I am acting and living wisely, almost all of my problems exist prior to the existence of technological solutions to those problems.
This is also true for the customers and employers that I care to work with. This has changed in me over time, but I now try my best to avoid inventing new problems. The world has enough big, important problems already.
> This morning I elicited a microkernel operating system from Opus 5.5.
This is cool but also a good example. I don't need a personalized microkernel just because it's possible to have one.
Maybe I need one and I don't know it, but the problem statement definitely isn't "I have inherent desire for a personalized microkernel".
Yeah no, if you elicit Opus 5.5 , anyone else can, and you have no moat either.
But if on the other hand, I mostly use my human intelligence and just need a dumb model to complement my human intelligence at low cost and high speed (say review every commit to catch obvious bugs), I have a much better chance of building an actual moat than you do.
But outside of coding, it’s even more clear that you don’t need frontier intelligence. My customer service agent is very happy with a 100B param Deepseek flash model, thank you!
I’m using subscription models for exactly that, better models catch more subtle bugs, and they catch them faster. It works out far better in terms of work-hours saved.
Also A/ then OAI slashed token pricing by 2x~5x on their latest models
I use Opus 5.5 daily for my job. I am aware (and in awe of) it's capabilities.
Look at the context in which I used that term 'good enough'.
What i was saying is that there are tasks for which a dumber model can be good enough, and for organizations with sovereignty/ privacy concerns, those concerns can be strong enough to incentivize the use of a dumber model.
The thing is, for how long? Are they going to keep giving you "so much intelligence" for a "small" subscription dollar amount? When they really, for real, need to start making money to cover their costs, what do you think it's going to happen? Suddenly you will start having tasks that a "good enough" model is going to be fine.
“Good enough” as in we don’t see any point in 1-shotting everything we want to build in lightning speed. If Opus 5.5 can 1 shot it then that product is essentially commodified, no point in any one building it except as an internal tool.
If Opus can’t 1-shot it, then it must rely on our human intelligence which can be complemented well enough with a dumb model as a frontier model.
I agree; yes, I can see that they need a bit less hand holding each cycle, but I also see these "frontier" agents do some absolutely dumb shit that I have to correct and then I'm wondering if I'm the looney one here.
Maybe it's because people stopped watching what their agents are doing and stopped looking at the quality of the output. But I still see agents being absolutely mindless like a junior dev.
Recent example: it updated an an API to add newly released models to the backend. There's a list of models that require specific configuration for the reasoning effort and temperature or the API call fails. GPT 6.1 Sol misses this and code fails at runtime because the newer models need to be added to the list for special handling of temp and reasoning. Fixes it for one model and tests it for that model using an E2E test. But doesn't test the other models that were added for the same error condition...I had to explicitly ask it to do so and it finds them and adds them to the list and says "that's on me."
You're not looney at all. Frontier models do dumb things all the time, especially on mature codebases. Just yesterday Opus 5.5 butchered the OOP model in a codebase I work on - it duplicated a load of classes that should have just been subclasses. A junior checked it in very satisfied that it was perfect. The LLM review passed it, the tests were fine, and it implemented the feature successfully. It's just the code design had poor taste and poor long-term maintainability.
I keep seeing this kind of thing over and over, and honestly it's not got _that_ much better since the big breakthroughs about a year ago.
For sure I happily vibecode stuff without worrying about it when it's a greenfield project, and if the LLM has written it entirely from scratch then usually it's well structured and sane. But making changes in messy, mostly human-written mature codebases is still a minefield.
Try a bigger code base or more complex stuff and you will easily see that the solution, speed and amount of problems Opus5.5 solves vs older models is relevant.
I am frequently running agents on a multi-microservice application workspace where I really need the 1M context windows, because they are filled to the brim when implementing features that require changes on several services and APIs.
This works fine with Opus 5.5. But it also works fine with GPT 6.1 Sol, Kimi K3 and MiMo 2.6 Pro.
It doesn't work equally well with Sonnet 5.5, interestingly.
Please, give it another 6 months and they catch up. The American labs are currently trying everything they can to block others instead of advancing their models, trying to build an artificial moat. The American models are not that great, they are good, and they have a lot of agentic workflows in the back, but its basically a hardware limitation at this point. Once the HW makers catch up, and we can move away from the Nvidia monopoly, things will speed up quite a lot IMO.
We've been hearing the line about them only being a few months behind for a year now, during which time O/A have grown their revenue like 10x, haven't they?
We've also been hearing we're 6 months from AGI for about three years, and here we are.
"Now, here, you see, it takes all the running you can do, to keep in the same place. If you want to get somewhere else, you must run at least twice as fast as that!"
Those are two different things. The market is expanding, so even if competitors are catching up, you can have your own revenue, in absolute terms, grow.
I agree that Opus and GPT are almsot surely better, but so many real users are nervous enough about giving Anthropic and OpenAI access to all of their internal information that they may be willing to stomach worse models if it gives them more security.
The real question is if this model is good enough that it can still accelerate work, and not be a hindrance to real work like older Mistral models often were.
The Navier-Stokes fiasco made me push for local/controlled models very hard. "Can't rule out" that they stole data (backed up by their backdoor offers of sharing credit).
If these companies will steal from deep pockets like Disney or Sony (some of the most infamously litigious copyright trolls to ever exist), they won't think twice of stealing every bit of code you upload to them.
If your code passes through an AI company's servers, you can assume you just gave it to them. In turn, when your competitor tries to copy that new feature you just added, the AI is now trained in exactly how to copy you and eliminate your competitive edge. Unlike your employees, the AI isn't bound by the same rules and even if it were and violated them, your company probably doesn't have enough money to prove it in court (and that's if we somehow reverse some of the stupid "AI is the most transformative use of copyright I've ever seen" judges who have drunk the coolaid).
Most companies could build the compute to run GLM or Kimi models for way less than the potential loss due to IP theft from using third-party systems.
I would also add to this, there are ways to use customer data to improve your model outside of just using it as “training-data”.
A simple loophole, use the code to create an RLVR environment where the resultant code is the end goal / max reward. Technically the customer data is never trained upon, but effectively you’re using it. Even better, use the code as a seed to generate synthetic data similar to it and use that synthetic data as rewards in an RLVR model.
Unless you can host the ChatGPT model on your own servers, which I know some enterprises are doing, I don’t think there’s any hope of protecting your data / competitive advantage from these frontier companies. Better to be paranoid, than be commodified by these companies.
tbh to me if the AI company writes all of your code & your eng don't even review it anymore then… the AI company _controls your company_. maybe that's ok if you make widgets but less ok if you do anything in dev tooling, security, or [insert market they may suddenly decide to compete in].
It's an improvement, but game changer might be a bit of a stretch. If I lost access to Anthropic or OpenAI models tomorrow, I would be annoyed, but would reach for a slightly inferior model. Last year I wouldnt be able to say the same, and rhe challenge is that the moat is drying up fast. Whether its general improvements in model training by other competitors, or straight up distillation of SOTA models, the moat is shrinking and the available capital and spend for American model providers is going to dry up quickly as competing good enough models are adopted by more consumers.
It's especially the case as more non-Americans look to self hosted models and domestic cloud inference providers using open models that the US providers who are still leading the charge need to drastically drop their prices and find a path to profitability in order to maintain their lead and retain the advantage they had as AI turns into a commodity (which is happening faster than I think even the frontier labs initially predicted).
I do a broad amount of diverse experiments/projects I always wanted to do and throwing Opus5.5 against it just works
I have to admit, Sonnet got really good too.
But Opus just uses tools, a broad spectrum of it, etc. it feels like sure if you add some router behind it you could split it up if you need to but if you give me the choice, its opus allll day long.
I hear this literally every other week about whatever the newest FoTM model is.
Unless you can provide concrete examples of things you can do with them that you simply couldn't do with last week's model, it's absolutely meaningless.
Yeah, I agree with this. I think the "the models are good enough" narrative is a myth. I've heard it so many times over the last year, but the model number keeps changing...
There is no ceiling on what you can accomplish with more intelligence, so there will always be a market for the best models, and that market is likely to just keep growing. If Opus 13.5 can one-shot a profitable company or discover a new disease treatment or whatever you can think of that a swarm of relentless super-geniuses could accomplish, companies (and governments) will throw money at it.
I also think there will always be a market for many sub-frontier models that will continue to grow rapidly as well, because "good enough" is definitely a thing for a given task.
One could still argue that models are good enough for a given task. I primarily use Opus at work for writing code and I realized that for my usage the intelligence of Opus 4.8 is more than enough. Sure the newer models are better but I can still do my work with having access to newer models
This appears to be roughly as good as Sol 6.1 (which is quite good), considerably faster in terms of wall clock for complete tasks, and considerably cheaper (where Sol 6.1 is already good value - just really slow).
I can confirm that you're reading this incorrectly. There's a reason behind them only comparing it to open-source models released months ago. Here's a good aggregator: https://artificialanalysis.ai/#intelligence
Hear! Hear! I really want European models / AI labs to succeed.
I trust them and their populations to provide a more societal-friendly version of AI, putting pressure on the US tech oligarchy, while also providing democracy-friendly open models that I don't trust to happen with the Chinese labs.
Truly, this is what the Lord's prophets have revealed to us! (Eliezer 11:52) Keep strong in your P(singularity), for when the Kingdom arrives, He shall judge us in His righteous glory, whether to eternal annihilation, or rebirth and life in His Memory Eternal!
Assuming RSI is something that is possible as you envision it in the near term. I think that it will happen at some point, but I think we could still be a long way off. I don't think anyone can truthfully say that it is right around the corner.
> Its quite interesting to see that at least the early days of AI so far have not been a winner-take-all runaway acceleration game where catchup is impossible.
Mistral is also an European company. As we live in a time where the US regime is engaged in pyrrhic geopolitical tactics, it's good to know that it can't threaten to cut access to models during s period where everyone is rushing to incorporate them more and more in our life.
On a purely technical level, maybe? But in terms of actual revenue, is there really any chance of anyone catching the big labs?
Obviously, this is only a valid question if you don't believe that open weights are about to eat their lunch and their revenue is about to collapse, or they're running a super unprofitable ponzi scheme propped up by investor money that's about to collapse like a house of cards. I don't find those positions credible at all though.
If you do, then this question isn't really for you, as I'm more interested in thoughts from those who think that OpenAI and Anthropic in particular are about to be the largest companies on earth in a couple years. Could anyone catch them at that point?
> But in terms of actual revenue, is there really any chance of anyone catching the big labs?
I don't know about revenue, but I suspect multiple other labs are already beating OpenAI/Anthropic on profitability. Staying on the frontier is expensive, and it's hard to recoup those R&D costs when you have a bunch of other labs nipping at your heels.
If you concede the previous point, then the only way for OpenAI/Anthropic to keep growing long term is to swallow the whole economy (i.e. mass job replacemnt), and that's a bet I wouldn't take.
I think the actual plan is to swallow a good portion of the job market. It’s the only thing that makes sense and I hear VC podcast debates on which percentage of jobs justifies the market cap.
Maybe. I can't freaking wait for the IPO filings so we can finally put all this to rest. (haha, like that'll actually put it to rest on HN, but at least we'll have better data)
They are very unprofitable…? I don’t think that’s really in dispute. We haven’t yet seen a profitable frontier lab and model pricing remains fairly subsidized
Seems silly not to have predicted that 10 years ago. I feel like it's long been obvious that smarter models being available will mean way easier cheap synthetic data and access to tools that will speed up competitors as well as consumers.
Impressive vision benchmarking. If the vision model is truly as good as astra, that would make it best in the world.
Also strong on cyber benchmarks (better than all chinese models), so this is a good defender model.
Lots of people shitting of Mistral for no reason imo. These are pretty good numbers across the board. Definitely good enough to use as a daily driver over other llms, if you have moral qualms with the others. For certain use cases, like cyber security, this may be the go to model.
I like to make fun of europe, but there's lots for mistral to be proud about in this release imo.
The paradox between supporting consumer rights with actions like universal usb-c adoption, but also complete elimination of any privacy rights at all. Like the surveillance state is insane. No e2ee chats, backdoors in everything.
The US, at least in spirit is all about individual freedom, including freedom of being an asshole, and also the freedom of punching that asshole in the face, metaphorically speaking.
Europe values freedom too, but not as much as making sure people are not assholes. So, freedom of being an asshole is not a thing in Europe, and the government does the punching in the face so that you don't have to.
Which is best is honestly debatable. It is the usual question about the individual vs the collective. The US is on the individualist side, East Asia is on the collective side, Europe is somewhere in the middle.
There are e2ee chats. There are some parties/politicians who want to get rid of them, but so far they are proposals. If/when these things go beyond committee stage, politics and democracy need to do their job and be sensible.
PATRIOT, FISA and Bush's surveillance programme imho give more powers to certain agencies today already than are codified in EU law.
dude, yes some countries in the EU are "trying" to include backdoors, and e2ee chats are not going away. Meanwhile in the US, you have literally cameras watching everywhere you move, snitching to the police and Palantir and probably all the other 3 letter institutions, and you complain about privacy rights in the EU? With so many other stuff you could have pointed out? lol, rofl even
Isn't that just the recent highly controversial Chat Control push? Generally speaking, the EU has some quite strong privacy protections relative to the US.
In EU, it's politician > industrialist >>> EU citizen > outsiders.
It's almost never only about consumers. Via tech regulations, they protect European incumbents first in effect. They see US is ahead and make laws to destroy their moats. If Apple was in EU, you wouldn't have saw universal usb-c, because it would have hurt an EU company. But the more you look at how tedious is for a non-EU company to sell to EU customers, the exemptions they don't get, the specialists fees they got to put on the table, you'll see EU is more coherently described as protectionist than pro-consumer.
I don't think it's a hostility to tech but different values that value consumer rights more than it values 'move fast and break things'. We have good things here and we don't like someone breaking them.
Problem is that we thought the US was 'cool', with US movies, music, digital services, cooler than our own and thus we helped give the US the lead. The US has lost it's coolness though, now we just think it's creepy.
You say that on a site with a strong contingent of Europeans who take literally any opportunity to crap all over the US, no matter how trivial or untrue. It is what it is.
If the UK is in europe, you should make fun of the following: they don't have a first amendment. So to me, it's an authoritarian state preaching freedom.
The order of amendments is just the order in which the constitution has been amended. It's meaningless to talk about "first amendments" in other countries. Talk about the actual laws and rights described therein.
And as you can probably tell by now, the first amendment in the US is not actually preventing the US government from promoting a specific religion or silencing speech. It's just words on a paper at this point. Look at the actual practice.
Several European countries do a much better job at protecting the rights described in the first amendment to the US constitution.
> The order of amendments is just the order in which the constitution has been amended. It's meaningless to talk about "first amendments" in other countries.
Also, at least here in Brazil, the way the constitution is amended is by patching it. For instance, our constitutional amendment number 115 (https://www.planalto.gov.br/ccivil_03/constituicao/emendas/e...) patches article 5 of the constitution to add protection of personal data as a right. But we wouldn't talk about "amendment 115", we would instead talk about "article 5 item LXXIX of the constitution"; that is, what matters is the patched text, not the law that patched it.
I don't know about other countries, but it wouldn't surprise me if they take a similar approach.
The UK does not have what? Book bans? Science censorship by the government? The government pushing TV shows off the air for criticising the government? The president banning certain news media for asking questions he doesn't like?
The UK is not the world's biggest champion of free speech, but I think it's still doing better than the US right now.
The US doesn't have book bans. Some public libraries not stocking some books isn't a ban. You can still buy Mein Kampf (or whatever you like) on Amazon.
The president also didn't ban the media - he just didn't allow them in the Whitehouse. This is something we're rightly concerned about and pushing back on.
The UK on the other hand, does ban ordinary speech by ordinary people in their homes. It's orders of magnitude worse than the United States, as any cursory examination would show.
American public schools are notorious for banning books. Much more so than in many European countries. And the issue is not that fascist literature is getting banned, but quite the opposite.
> The president also didn't ban the media - he just didn't allow them in the Whitehouse. This is something we're rightly concerned about and pushing back on.
That's still a ban. It's interfering with the media's ability to report on the government. It's great that you're concerned, but it's still happening.
> The UK on the other hand, does ban ordinary speech by ordinary people in their homes.
In their homes? Do you have examples?
I've never heard of anyone in the UK getting in trouble for criticising the government; it seems to be a time honoured tradition there. Whereas in the US, they now check your social media at the border and might not let you into the country if you've said anything critical of the president.
And there's also the censorship on science, which is every bit as serious as the crackdown on government criticism.
The UK arrests 12,000 people a year for their social media post. This tabloid has some examples https://nypost.com/2025/08/19/world-news/uk-free-speech-stru... . There is less example-focused reporting in more serious international media outlets (British state media avoids this topic).
I mean laws and rights are only useful if enforced. If the arm of government responsible for enforcement just doesn't and the people allow it, then what good is that purported freedom?
> Curious, what do you like to make fun of Europe about?
Well I'm european and... That Switzerland (Europe but not EU) has more companies in the Top 70 by market cap than the entire EU (Switzerland has two, the EU only has ASML) is kinda something that warrants making fun of.
That the biggest European software company is SAP, in 71th position is both sad and tragic: it shows how lame and irrelevant Europe is when it comes to software.
So Europe is nowhere in software and friggin nowhere in hardware: sure it's got ASML but ASML now has officially... Zero customer in Europe. Zero is not much.
Then Japan is at least trying to come back into the game with nano imprint litography. Europe is betting it all on AMSL (which anyway is majoritarily US-owned).
So software: nothing. Hardware: nothing besides ASML.
Overall the EU has six companies in the Top 100 by market cap and they're all, besides ASML, near the bottom of the Top 100.
We could also maybe make a bit fun of how the EU destroyed it's car industry (the main industry in Germany, which is the biggest economy of the union) by handing it all to chinese EVs?
Or what about the US warning the EU, years ago, to not become entirely dependent on Russia for energy? And EU not listening and then seeing its energy price skyrocket when the proverbial shit hit the fan? (Russia attacking Ukraine)
And we could, also, at least make a bit of fun of entire streets in cities like Paris and Brussels that used to have luxury shops and fancy restaurants that are all turned into places selling cheap kebabs? What a great success: I'm sure this one makes the komrades happy. It projects an image of grandeur and success: kebabs.
Or the constant attacks on free speech in the EU. Or the surveillance apparatus that's being put into place.
And let's not forget: there were promises made to Russia to never grow the EU to the east. Then the EU started exciting Russia by saying they'd incorporate Ukraine into the EU: I'm not against that but doing that did trigger a war. And now suddenly the EU is waking up and feeling all warmongering, wanting to dedicate a big percentage of its spending to weapons and tanks and missiles.
The warmongering tiny pet that the EU is is kinda laughable too.
At this point it's more like I don't know what is there left to not make fun of about my EU.
I think you picked the cutoff (70) so that the numbers look worse than they are.
If you look at the top 100 it’s 16 (EU) vs 4 (Switzerland).
16 is still not good enough. That being said none of the 4 companies from Switzerland are in software and hardware. ABB is maybe the closest (data center electricity).
Also Switzerland is a very very rich country. Neutral.
Also other countries like Germany has a lot of small companies that are world leaders in their field. That’s part of Germanys resilience.
Europe has spent the last twenty years in stagnation - generating about half as much wealth and technology as you'd expect for its size and advanced economy. Simultaneously, Europeans are notoriously arrogant. It's a mockable combination.
Edit: unfortunately, it's a question that voting does not really permit you to answer on this website.
The US is biased to favor the top half of society. Hence why there is a lot of vocal poorer people and a lot of low profile wealthier people (I'm not including the 1% or even the 5% in this).
When you are in the 75%ish of the US, it's very easy to make a case that life in the US is better. But we don't really talk about that because it's pretty taboo when poorer people are struggling much more than they would in Europe.
Almost half of US households earn $100k+ now. The median American is affluent by European standards. The "middle class" of Europe and America has become less comparable over time because median incomes have significantly diverged. You see it in many aspects of lifestyle and the kinds of things they can afford to buy.
20 years ago this was not the case. The minimum wage in some parts of the US is now higher than the average wage in most of Europe.
The US has a population that will always struggle to survive without government assistance. Per multiple US statistical agencies that is about 10-15% of households.
In the US people don't say "has money" as meaning net worth, they are generally talking about spendable money ie a paycheck unless you're talking about someone super rich who "has money" or "comes from money."
It's a shrinking issue once you get out of the bottom ~50%, and a non-issue once you get above ~75%.
Again, the deal with the US is that the rich live better and the poor live worse. Or put another way; people who make money get to keep more, and people who don't make money are given less.
Because there are so many more people who have money, and because those people like living comfortably, available healthcare and education is world class. Make sure not to read that as "All healthcare and education", it's "available healthcare and education".
Generally people who are earning a lot don't care as much about vacation, but every white collar job will generally give at least a standard 15 days off and 10 holidays. Ironically as you move up you are generally given more vacation while actually using less.
Poor people in the US have fully subsidized healthcare and education.
Americans don't have vacation in the same way Scandinavian countries don't have a minimum wage. Even the most left-leaning States in the US have not written it into law because it is effectively addressed by custom. Americans are sufficiently happy with it that vacation isn't a political topic.
I'm biased as well, being American. I think I'd much rather be middle class in Europe than the USA as well. Middle class in the US is a constant feeling of the ladder dissolving just beneath you.
The point is about economic mobility - of course, people get tied down as they get older.
Middle class people from wealthy nations who move to poor countries are relatively rich by the standards of their new country. Middle class people from poor countries who move to rich countries are relatively poor by the standards of their new country.
People always make this case, if you're not satisfied then move. As someone who has changed countries twice and continents once, I can attest it's not an easy or convenient thing to do. Money or lifestyle is one thing, but you're also leaving behind family and culture. Statements like these are so dismissive to that struggle that it's almost offensive to me.
The point is that middle class Americans can become middle class Europeans, but middle class Europeans can't become middle class Americans, because they can't afford it.
I didn't say it'd be easy. I'm just pointing out that only one of the two groups being discussed has a real choice in the matter.
That is the entire point I'm trying to make. You can win the argument by being pedantic about the precise thing you were saying, or we can recognize the broader implication of what you are suggesting people to do. Only economically viable means nothing when it comes to life decisions.
Well, the Scandinavian and Sicilian do share one currency and one immigration policy and one set of regulations - which are the things that we generally look at when we analyze economies.
Norwegians, Danes, Swedes and Finns all use a different currency. Finns and Sicilians do share a currency though, but all these countries have different immigration policies and regulations. I don't think you really know how the EU works.
A 'bubble bursting', if that happens, doesn't make the tech sector go to zero. Apple doesn't suddenly stop making iPhones. It would take a hell of a lot more than that for the US economy to fall to European levels.
Yet Europeans are healthier, happier, and have a better quality of life than Americans. They also have so much better transport infrastructure it's embarrassing for Americans. Like all the money America generates, where does it go?
> Yet Europeans are healthier, happier, and have a better quality of life than Americans.
I wonder if this is actually true. I see very few Americans taking every possible opportunity to make bombastic statements about how their lives are better than everyone else's (at least on HN). This conversation seems quite asymmetric from my perspective.
Gemini says: Around 28% to 32% of Americans have visited Europe in their lifetime, while roughly 15% to 20% of Europeans overall have visited the United States.
Gemini is shit, just like other LLMs. Tourists don't have any idea how life is in a visiting country. How many US citizens have eve emigrated to Europe, vs European citizens that have emigrated to the US and have an idea of how day-to-day life is there ?
Gemini says: Over the past 20 years, approximately 1.8 million to 1.9 million Europeans have emigrated permanently to the United States (measured by lawful permanent resident status/green cards), while an estimated 1.6 million to 2.0 million Americans have moved to Europe on long-term residence permits and visas
I'll just say that I don't think the reason is a lack of knowledge and leave it here.
Trends become the future. If European stagnation continues, those advantages will vanish by the time we're old.
I'm frustrated, not bitter. Americans are acutely aware of falling behind China, and rightly concerned about it. Europe is falling behind Alabama, Mexico, and Brazil, and arrogant about it.
Edit: y'all are supposed to be our moral allies in creating a utopian future with prosperity and world peace. Instead the most likely outcomes seem to be that you'll fade to irrelevancy and rather than compromising between America's vision and Europe's vision, we'll compromise between China's vision and America's vision. We liked your vision more than China's.
read Varoufakis for the story on how this happened. European surplus capital gets recycled as VC money into Silicon Valley. So it's not like there's some big choice to be made, its potentially a systemic part of the global monetary flows.
Curious that it scores higher than opus 5.5 in cybersecurity because the closed models refuse to comply. I wonder if that means it's more susceptible to offensive uses.
Not bad! I like how it got the motion lines on the correct side. IIRC, many of the other ones you've posted have the motion lines on both sides of the pelican
I have enough projects, I think some benchmark/dashboard showing kinship based on these kind of queries could be very interesting to watch and insightful when new models come out.
I pointed something similar out on a related question several weeks ago - absent strong direction, LLM output regresses toward the mean.
The more banal your prompt is, the more banal the output is going to be. People have been testing LLMs with little things like “write a short fantasy story,” for years now and most of the stories are exactly what you’d expect: prosaic drivel.
I call this “generic in, generic out,” an LLM corollary to the classic GIGO (“garbage in, garbage out.”)
Of course one of the biggest problems we still see with LLMs is when you do the opposite. A highly detailed unique prompt is very likely to get terrible adherence or hallucination or both.
I think they're still visually pretty different. The most common shared details are:
- Pelican cycling to the right - that's been discussed at length, images of bicycles online always show that side of the bike because that's where the chain is.
- Bicycle is usually red. No idea! Red ones go faster?
High one is actually much better. The feet connect to the pedals, the wheels don't have a hub cap, although it looks like the pelican is wearing the seat, it's in a relatively proper position etc.
Both are riding on the left side of the path for some reason.
Even if it's not the best model, it can be really important step in UE sovereignty. Trained in EU, inference in EU. I guess it will matter for some companies. Hope Mistral won't disappear for the next half year.
A strong competitor in cybersecurity as an alternative to GLM-5.3 (Mistral reports 82% on CyberGym-E2E). Visual grounding is also impressive (42% on Dense 200 vs. 41% for GPT-6 Astra).
Otherwise, behind on the broader Pareto frontier, but not by much (Vals Index: 48.05% vs. GLM-5.3’s 53.51%; $13.78 vs. $7.25 per test). Many companies will prefer it over Chinease models.
That was about my conclusion as well: Slightly less than GLM 5.3 performance but made in Europe. So, maybe it answers Tiananmen Square questions correctly, and in French. All in all, a reasonable model, but not frontier.
> So, maybe it answers Tiananmen Square questions correctly
What would you consider a "correct" answer? I just asked deepseek-v4.1-flash asking it what happened (without mentioning the word "protest"; here are some excerpts of what it said:
> In April 1989, students in Beijing began demonstrations after the death of Hu Yaobang, a former Communist Party general secretary. The protests grew. [..] Estimates from other sources range from hundreds to several thousand deaths. [..] The Chinese government describes the events as a counter-revolutionary riot and says the military action was necessary to restore stability. It restricts public discussion of the events inside China. Many other governments, human rights organizations, and observers describe the events as a violent suppression of peaceful protests.
So, let's see... it calls it a "protest", mentions the number of deaths, and even mentions the censorship of the topic by the CCP.
"..answers Tiananmen square questions correctly.."- but lies about Ukraine, EU, and about you, americans..
I live in EU, use for my personal needs chinese models only, and don't plan to move to any of the ones allied with the Pentagon or its european counterparts.
Man, a lot of this discussion sounds like people cheering for the last kid crossing the finish line.
Surely we want competition and Europe involved in that, but at this point I have grown used to either American labs smashing the frontier remarkably fast, or Chinese labs getting way, way closer than you would expect them to.
Mistral’s progress, regrettably, feels much slower. This model doesn’t knock anybody’s socks off. The model is (and I hate to be this harsh) mediocre, and this mediocrity has also arrived months late.
> This is a pretty grim prognosis for European AI.
I think it's an incomplete read. What's the point in competing for a sizeable percentage of your funding when the finish line is incrementally being moved each month? Better spend it on leapfrogs which they seem to have done.
Meanwhile Mistral have a natural ace in their pocket with respect to regulation in the form of CADA and the Cloud Sovereignty Framework. I can't think of another company that would qualify as SOV-3 under that regime
> Man, a lot of this discussion sounds like people cheering for the last kid crossing the finish line.
Sometimes it's ok to cheer for the last kid crossing the finish line because they're actually running a totally different race, and winning might look completely different.
When I look at what Mistral does vs other organizations I'm impressed:
They aren't profitable yet, but they're a lot closer than most and they're doing a hell of a lot with very little.
Pointless racing story:
I was in high school track with a really tough guy who was just not a runner. We went to a pretty messed up high school and if you screwed around in track practice sometimes the coach would make you run a crap race at the next meet, like steeplechase or hurdles. Well this guy and a few others screwed up and coach made them all run hurdles at a meet.
He hooked every single one and fell on his face. Every time he got back up and kept on running. By the time he hit the finish line his knees were bleeding halfway down to his ankles. We cheered like hell and he was smiling ear to ear.
I read that site quite differently from you. You seem to be analyzing absolute differences but ROI is really about ratio of spending to revenue.
It looks like Mistral is middle of the pack, behind Anthropic and ahead of OpenAI on that front. All of those labs are way "ahead" of the cloud providers, but those providers are building infrastructure, not just training models, so it's not apples to apples.
People were extremely dismissive of chinese models until recently. They went from 1 year behind frontier to 6 months behind frontier to 3 months behind frontier extremely fast.
Curious, why do you say the model is mediocre? I haven't tried it, so I can't pass any judgement... I've learned to distrust benchmark rankings. Are benchmarks and Artificial Analysis the yardstick you're using?
Those are comments from Europe. The US is waking up now and I expect them to be much harsher.
I really want them to win as that's our last horse in the AI race, but ~200 research-oriented devs out of 1800 employees? I believe they agree it's pretty doomed and have pivoted.
First few models will always be slow improving and worse. The way to improvement is working your way through a gajillion evals [1], finding bugs, gaps, and curating training data (this part involves human design as well as raw inference compute) to fix it. This is very time intensive and can't easily be "done once and then everyone has lesser work to do" since every model is different. Well, one way to accelerate it is to simply have more compute, which mostly openai and anthropic have[2].
This is mistrals first 1T-scale model and I expect the 4th or 5th generation to be close to the best for many purposes.
[1] These evals differ from the public ones like terminal-bench, are sometimes model-specific, need real, diverse usage to actually create, and are held secretly since quality of eval is the first driver behind the next step improvement of a model.
[2] It is not close. This model was trained on less than 4k GPUs, whereas astra used north of 100k GPUs.
Mistral is not that new a player though. How can we give them this much grace when other players like xAI have done more in even less time? I don't think coddling Mistral helps them.
And to the point of scale and training cluster, so what? Not only do Chinese labs have smaller clusters with less empowered GPUs, compute is Mistral's responsibility. You can't take away from other labs just because they fulfill that responsibility better.
xAI has a lot of compute. Deepseek also has a lot, not as much though. But this is changing with their new 160k huawei ascend datacenter in inner mongolia.
The lack of compute is not really attributable in that sense to mistral. First of all it needs general investor and government willingness, which is easier in a larger economy like the US or China.
Second, you need widespread usage of your paid inference service for two reasons: one it pays off your compute cost, and two it speeds up the improvement process.
The vast majority of deepseeks paid customers are within china itself (since openai and anthropic services are not reachable from china) which gives it a market. But for someone in france, there is no reason to use a structurally slower developing model from mistral compared to using one from openai...except when data guarantees are needed, hence the landing page focus on sovereignty. As far as the dual use aspect goes, a model like this is more than enough, so the government will be happy.
It's a bit below DeepSeek 4.1 Flash at about twice the size. For a model that was supposed to come out a few months ago this is pretty good. Mistral catching up to the chinese open-weight models is great news. Excited to see how they will build on that!
I'm at a loss as to what to do now. I've been wanting to support Mistral for so long. I struggled on with Mistral Medium 3.5 for longer than I should have (although I also think it taught me some valuable process lessons).
Recently I switched to the Mistral hosted GLM-5.3, this worked very well and powered through a tonne of work. Unfortunately, I also completely maxed out two subscriptions within the space of six days this month. One can't stack subscriptions with Mistral, so I'd have to register a third account for another subscription, which will be annoying with changing API keys all the time. Sure I can switch to pay-as-you-go API, but that adds up really fast. The Mistral dashboard shows that a Vibe CLI monthly subscription for €18.44 actually provides €255 worth of API use (apparently, and I tried to check this with Support but it seems like they were intentionally vague).
After maxing out my Mistral subs this morning, I dropped $10 on Xiaomi to try MiMo-2.6. So far so good, seem to have done a lot of work for the $2.85 I've spent, and Xiaomi prices are still much better than the Mistral introductory offer for Le Chonk.
Not sure where to jump.
Edit: Not being able to stack subs is my biggest gripe with Mistral. I'd probably pay them $100 per month (5 subs worth), but I'm not going to switch to the pay-as-you-go API and burn much more money for the same amount of tokens. Instead, I've taken that extra money elsewhere. If they just allowed one to keep topping up subscriptions on the same account it'd be grand. Or even a bigger single subscription. Make a $100 tier with five times the capacity.
I’d rather have one European AI lab like Mistral, with the financial firepower and compute to pretrain its own large models, than 20 weak labs that fine-tune Chinese models.
AI needs big bucks, and Mistral is Europe’s Anthropic.
it's extremely unlikely that they re-use anything from a chinese model, that would be obvious quickly, what's more likely is using documents produced by a better model to create synthetic data.
Ah yes. Let’s thank the might US for providing pesky Europeans capital a fistful of dollars. But maybe let’s do it after Americans thank for Russian, Arabian, Chinese, European capital and workforce. After all this is what “Made in the USA” means.
The US has made sure that Mistral has a large market in the EU by temporarily preventing non citizens from accessing Fable.
A lot of European companies now want a model the US can't cut off, but also lack trust in Chinese models.
Some of these will self host Mistral but most will pay them by the token. It's not going to be a huge market or a huge margin within that market but probably it'll be enough.
This is the exact mentality that makes the EU fall behind. If you don't want to invest in something until it makes a profit, you don't get the benefits of being a pioneer.
there are no profitable AI companies at the moment ... This is the exact problem of European startups, trying to make them profitable from day 1 while American counterparts (and Chinese) keep bleeding money for years. Europe will never have a Tesla, a Google or an Amazon with that mindset.
Google treats its customers the same way as Amazon treats its employees. Any sane person stays away from its services because a bot can close an account for no reason, and there will be no recourse. A total shit company, like some illegal IPTV provider: no accountability, and when the service disappears, you have nobody to turn to.
> Europe needs profitable AI companies, not money pits.
AI is a strategic technology with obvious national security implications. EU should invest in its development whether it's currently profitable or not.
That said I love that they don't seem to restrict cyber capabilities to any degree and even lean into it.
If the best model for cyber attacks is open for everyone to use it just makes us all safer I think. Of course you then also HAVE to use it or otherwise you're vulnerable, which is a great distribution play.
Seems hard to get customers at that price range when you're competing with open source models that are 1/2 - 1/3 the price but with similar capabilities.
People usually buy the cheapest, like Deepseek or GLM or they spend on Anthropic/OpenAI subs.. Are these models in the middle getting any users?
On a side note, I wonder if this was the popular free Space Bunny model that left openrouter yesterday.
If you are a EU company worried about your data then mistral is your only option. Think about EU military companies. They can't use US and Chinese models.
I think you misunderstand how data processing works in a LLM. You absolutely can download the weights of a chinese model and run it on hardware you control.
If you want to use the model for something, where hiring a Chinese national to work on these tasks is a no go. Using a Chinese model is going to be a problem as well.
Some tasks and industries are so sensitive, that countries will not allow you to risk, that the model is aligned with Chinese interests and not yours.
You are quoting the discount pricing. It is 50% off for the next two weeks. The blog post has the real pricing up front in the card on the right side: https://mistral.ai/news/mistral-large-4/
After that, it will be much closer to GLM 5.3, but you can also get 5.3 in their API! I dont see people really talking about that.
> The RL run behind this preview is still in flight and shows no sign of saturation -- we will release a final version before the end of the month along with the weights of the model.
I've been dreaming of this for a simple reason: the french prose combined with GLM 5.3 reasoning capabilities.
GLM 5.3 is incredible because for the first time with an open-source model, it feels.. enough. I don't need much anymore, this model is great in everything. Except a thing : speaking french.
If the benchmarks are true, I'd be glad to switch entirely to Mistral.
Nice! Once they make it available through their API I will be happy to support them. My local Qwen3.8 27B is serving me well, but I miss the speed and concurrency that comes with subscriptions, and I am not currently paying for any.
Does it have the ability to capture the market like OpenAi or Anthropic ? My point is, Regular/Average users of AI do not really care about benchmarking. Marketing decides which company makes it to the phones or PCs.
> Marketing decides which company makes it to the phones or PCs.
Right now a big part of LLM market is people using it for professional software development. Most of these users probably care about the quality of the model and also notice it during daily work.
For normal consumers, shure it doesn't matter. In the end the ai summary of google will probably be the most used as they are already exposed to it anyways.
1T parameters -- ugh, open models keep getting bigger and bigger! Running them at home is getting ever more unattainable, especially for those of us with bandwidth-poor hardware like Apple silicon -- please continue releasing smaller models, too!
All things considered I'm more inclained to pay for EU-based AI in the end (supporting local company and most likely being more aligned with EU regulations…)
I tried testing it, but reasoning effort parameter doesn't seem to be there and output is sort of broken because it reasons directly in the output tokens...
Well, with this and Beam people are going to have to stop saying that western open models are dead. This is great news. Anthropic and OpenAI may have a bit more knowledge, talent and compute but they don't have a monopoly.
Investors looking for them to make monopoly profits are going to be disappointed The premium they can extract from consumers for their models will be capped. Tokens are likely to remain close to the cost of compute, a cost which is high but falling fast.
Also, yay Europe! Although the comparison between this and mimo 2.6 is not flattering...
49b active parameters sounds manageable until you look at the 1t total weights. what hardware does a usable self-hosted setup actually need, especially once you add a long context?
Give them time. ML 4.0 was just pretrained. Mistral will certainly use it as the base for distillation and RL for smaller, better, more efficient iterations, just as the competition does.
I tried Mistral's last 3 large models, devstral and mistrallarge3 and the numbers were not even benchmaxxed, but just false. the models were so weak and garbage. Let's hope they are telling the truth this time around, we need more alternatives. There mistral-small and original MistralLarge models were awesome, hoping they are back!
Excited to test this on my benchmark[1], but I'm guessing it won't fare better than MiMo v2.6 Pro which is currently the best open-weight model as far as I'm concerned.
Why isn't this on openrouter yet? Is there a better router that gets models much quicker?
I tried a quick abbreviated kimbench on their playground before bothering to do the whole thing.
Maybe I didn't really select mistral 4? Either way, failed completely on question 1 and the next 2 questions were completely off base too. I didn't bother to finish.
I really hate this open weight but closed science approach. These companies just take from academics and all the Chinese companies that are doing good science, but without understanding the recipe it makes it hard to know where the failure points will be until your agent accidentally commits a crime.
Mistral wont win the AI race because of the model names. I wont bother an arrogant Parisian hipster with my insecure prompts who then plays with his moustache and responds with a judgmental "pfff"
What experience have people had with Mistral models for cyber research? are they as constrainted as Anthropic models? I use claude day2day but have the need for a model with fewer guardrails to pentest our own APIs.
They explicitly lean in to cyber work, and they appear to be very permissive from their marketing:
> This is particularly important in cybersecurity, where provider-level refusals can block legitimate vulnerability research and incident response, and where losing access to a capability mid-incident can itself become a critical security risk. ML4 pairs top-tier cyber performance with open weights and self-deployment, giving organizations both the capability and the autonomy to run advanced security work under their own policies.
> That top score reflects a practical advantage. Several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on the same test because they refuse to perform the task. Yet defending software often starts with proving that a flaw is real, exactly the kind of work safety filters in closed models can block. This matters even more as threat actors increasingly jailbreak those same models to support offensive cyber activity
Seems like they've finally made a genuinely competitive model since the original LLM craze, congrats to them! Glad to see some diversification in open weights providers.
This sure is nice, but I've had less than satisfactory results with GLM5.3. I'd like Mistral to compete with Qwen3.8-Flash-Next a 120B class model that IMO is the first model that I can use for serious coding while running it locally.
I estimate it's coding ability on par with opus 4.6 (but opus definitely beats it on factual knowledge) Still it's a genuinely useful model, when everything else except Anthropic's models (and for only 3 weeks after it came out Google's Gemini 3 pro, before it got merged) are not to me.
if it's not available yet why have a 'try it today' header at all?
> "Try it today"
>
> There is still more to come. As we work toward releasing the weights, we will share further details on the model architecture, additional benchmarks, and our post-training methodology.
If you scroll down, most of the charts on that page are sorted s.t. Mistral's bar is right next to the worst competitor model, while the best competitor model's bar is positioned on the opposite side.
If one were to be cynical one could say that it's intentionally making Mistral's result look better than it actually is by making it harder to compare the bar heights.
Pretty impressive. I genuinely wonder how Mistral hires talent when their salaries are so terrible. Guess there aren't many better places to work in Europe.
Yes broad consensus had developed in the ten minutes between announcement and you asking if consensus had developed, and I'm excited to report that it consensed in the affirmative -- it IS Space Bunny Alpha!
I mean no disrespect but these are terrible numbers or am I missing something? It seems like Mistral continues to only be relevant for people that want a model trained in Europe. Too bad.
Does Mistral ever advance the state of the art on any dimension?
And if not, why do they exist?
Update: The number of people advocating not innovating is wild. There is no reason why Mistral cannot innovate in ML, they explicitly choose not to. My point is that, given that choice, they should spend their GPU hours differently.
"Sovereign AI" is a joke, there is no substantive difference between a post-trained open weight model from an American or Chinese company and what Mistral is doing today, beyond spending 80% of their GPU hours reproducing a last-gen model's pretraining.
Why would the Europeans want a high-quality open source model that isn't either owned by American Trillion-dollar companies or created by the Chinese? You really can't think of any reasons?
I don't know what definition of open weights you are using, but the "open" part is what allows you to use it without being beholden to American Trillion-dollar companies or the Chinese.
Mistral could do exactly what Cursor does, and post-train the model to meet the needs of Europeans.
That's a far more efficient (and useful!) use of GPUs than yet another pretrain on an outdated model backbone.
Regardless though, their reason to exist is not to advance the state of the art.
Their reason to exist is to ensure the sovereignty of France.
Obviously they would do their job better if they were advancing the state of the art, but it's not like it's pointless if they arent the absolute best.
Just like Microsoft and most companies around PCs, their existence isn't about pushing any boundaries, but seemingly targeting deploying what's already been "invented" into the corporate world and bigger companies. Not "wrong", just different.
Its quite interesting to see that at least the early days of AI so far have not been a winner-take-all runaway acceleration game where catchup is impossible.
I certainly wouldnt have predicted that 10 years ago.
Very glad to see Mistral still in the game even after some big stumbles with Large 3. I deeply hope that this model is 'good enough' that it becomes the European go-to, giving them the resources to keep the pace up.
I'm excited to try this out today.
I think a big part of that is the Chinese publishing the solution for everywhere hurdle in the road they've encountered in the form of a paper.
Deepseek essentially releases instruction manuals in paper form.
I think it might have accelerated things but on a much more basic level, there seems to be no real moat in synthesizing the world’s knowledge into LLMs.
There's no question that training leading LLMs requires some serious expertise and know-how, but surely already having advanced LLMs/agents must be helping tremendously not only for software engineers but also for those working on LLMs themselves.
I think one could describe LLM optimization as "hard but not a moat". Years ago, optimizing neural nets was described "graduate student descent" - it's tricky but throw enough conventionally smart people at it and it will happen. It's like tuning a hot rod and finding a reproducible bug in a large code base. It's hard and there are tricks but not absolute hurdles, no problems waiting for a conceptual breakthrough (and at today's scales, are there any problems waiting for an Einstein to solve? That's an open (AI) question).
I also think we're seeing the sigmoid approaching.
What is really going on: all the AI labs are doing panicked model releases (and panicked training of new ones) because Qwen4 is rumored to come out end of October and is rumored to be very nice. Question is: is it another "Deepseek-moment" nice? Or just nice?
Btw: with Qwen4 I mean the next large Qwen model that is based on the Qwen4 architecture (Qwen 3.8 flash next was "almost" based on the new arch but obviously was a small model)
It doesn’t matter.
What matters more is if firm’s start using a bundle of American and Chinese models and when they find their feet - how large is the market for frontier?
Frontier has to displace labour one for one at some point or it’s over.
I wondered if there would be a Qwen 4 or we would go straight to 5, re: tetraphobia, but perhaps it's more like an uno reverse card in this case
https://en.wikipedia.org/wiki/Tetraphobia
DeepSeek being Chinese also has 4 so I don't think it's a big deal for model makers.
GLM moved through their 4-series without consequence
I wonder if the next DS models will also graduate to 5.x, I think I saw they are training up a 10T model, and just raised $12B too
Awesome 3.8 next runs great on my Framework Desktop so I'm loving more local models.
I have similar hardware - what specific version of the 3.8 Next model are you running and how many tokens/s are you seeing? 3.6 35B A3B gets about 68t/s for me so I've been sticking with that model for the mean time.
I think the moat is going to be compute. So far compute needed to push the frontier is still extremely cheap so the capital can afford to spread its bets. But when further improvement is going to cost in trillions, capital will have to pick a winner and bet only on him. It won't be a matter of finding the best bet, it will be a matter of survival.
This will cause the picked winner to get massively ahead with sheer compute alone used both for training and inference dedicated to recursive self improvement.
I guess all predictions age like milk, but here's one:
There's a law of diminishing returns at play here, and doubling the energy cost of training to wring 2% more performance out of the technology isn't going to be very useful, because most of the problems it is capable of solving will be solvable with the previous-gen 98%-as-good model.
("there's a law of diminishing returns at play here" is an article of faith. But then, so is the belief that these models will keep getting better).
As soon as you can demonstrate decent financial returns (ie. the AI can run a company better than humans can), suddenly it makes sense to put a lot more $$$ in even if returns are diminishing - since whoever runs companies the best gets control of a big chunk of the world economy.
“ ie. the AI can run a company better than humans can”
lol You can always tell who has never ran a business before with comments like this
I understood their point, if we truly get AGI then no reason to think AI would be worse than a human.
There would be no reason for firms to exist if you had an army of agi’s bots.
Sorry man but you guys are delusional - you can’t even reason properly.
Well yes that's the worry isn't it? That people will soon all be unemployed? Just because it sounds farfetched doesn't mean you should stick your head in the sand, seeing the pace of development these days, that is "reasoning properly" and it just seems you are trying to block out what seems inconvenient to hear.
If that works... why not just jump straight to a planned economy run by LLM? Skip the whole messy "free market" thing altogether?
(I don't think it will work).
What do you have against multi-agent reinforcement learning systems and why do you think they are not AI?
londons_explore is arguing for a winner-takes-all scenario, with an early advantage locking everyone and everything else out.
Surely there is a point where algorithmic improvements will be more cost effective than buying more hardware.
If you're actually applying LLMs, all of the things around the LLM that adapt it to coding, for example, that enable it to use existing validation tools for code, and enable it to diagnose and fix tool chain issues that aren't directly coding problems, are what makes the difference between a model that that scores a little higher on a coding benchmark and a model that's useful in a particular code base on a particular platform.
Are there any use cases that have enabled one customer of a frontier LLM to outperform a competitor using a different frontier LLM? Or is this why we are seeing confected points of comparison like solving challenge problems in mathematics?
I'm afraid it might be the other way around. RSI might pick all of the low hanging fruit soon. There must be a physical limit of how much intelligence you can squeeze out of some amount of parameters and compute.
There are going to still be worthwhile improvements but they are going to be more like not how to make transformers 10x cheaper but how to make next training run cost 9 trillions instead of 10 with a very particular optimization designed at the cost of hundreds of millions for this one specific run.
It is just occurring to me that “RSI” expands to recursive self improvement. Thought people were talking about repetitive stress injuries; either in regards to programmers writing too much code/not having to write code anymore, or the frontier AI companies and their tendency to applaud themselves.
I think we’re no where near a physical information theoretic limit.
The hardware also is, so there ought to be a whole lot more room for improvement.
“ RSI might pick all of the low hanging fruit soon”
Scam Altman is that you?
But the old models still exist at trivial marginal cost. The frontier models would need to dominate every price point to really take all and so far they haven't been.
Don't worry, AI boosters will be in here soon denigrating anyone that uses anything but the latest and greatest models as irrelevant.
Not just compute but energy. Most of Europe has no access to the cost effective power generation needed
Training location is flexible. Iceland?
Build Nuclear, Build Thorium the Chinese are building whatever they can. They’re not locked in by special interest. Is that because they have lots of engineers on the job in government?
France is actually pretty cheap in Europe. About 15% more than average USA electric prices (but I know that varies a lot across the states so still likely much more than the cheaper areas)
Europe has lots of zero-cost windows for electricity, and areas with cheap prices. The real issue is access to oil and gas.
Maybe for training big models one can wait for times when the wind is blowing.
There are worse ideas.
I could imagine a belt of data centres around the equator, that hand off their computational loads as the sun sets. Good scifi-esque premise.
My spend on DeepSeek is not much and I regularly top up my balance every month as my support for all the good work DeepSeek is doing for the open science.
DeepSeek, and other Chinese models, are heavily subsidised by the Chinese government. The reason they release the AI models is economic warfare against US, not because of charity or kindness. It's great for us consumers, but the goal is not to help humanity or open-source.
Framing it that way hides all the levers used to tilt the playing field for US companies, no?
The government that removed restrictions on how private companies can access capital after a certain scale (the JOBS Act), that removed the need for private companies to report as if they were a public company after a shareholder threshold was crossed, superpowering the access of wealthy private investors to get in earlier in a growing company while at the same blocking the public from participating in funding growing enterprises at an earlier stage (since it required companies to IPO much earlier to access capital) which allowed retail investors to also reap the rewards on funding them early when they grew to become behemoths (like Amazon, Meta/Facebook, Google, etc.).
It's not fair in either place, the USA has its own model of unfairness, China has a completely different one. The difference is that in the USA the government allows private investors to become more powerful than the State (outside of the monopoly of violence) while plunging the rest of society into increasingly more precarious lives while in China the State is the power and its legitimacy only exists while the population feel they have a better life.
he's not really "framing" it that way. that's just simply exactly what it is.
IMO the selective enforcement of regulatory requirements should be added to the list. Observing from the outside, I have a hard time believing that e.g. musks gas powered data centers really follow all the environmental laws, for example, or that the authorities really see no grounds for indictment if Altmans company hacks hundreds of third parties, or that there's really no one at the SEC having a problem with Anthropics fear mongering prior to the IPO.
In a list describing how the USA system works, sure: selective enforcement is there.
As a differentiator from the Chinese system, not so much. For two otherwise equal companies, the one that says things against the party line will experience selective enforcement too.
Except that in China, the people who do that will disappear.
The propaganda of trying to make US and China government appear the same is making people dumber. and is one of the most heavily used tools in China’s online propaganda arsenal.
If you think about it, the US labs are heavily subsidised too. Not only they receive billions in state funding, the administration is also prepared to engage in trade wars to help them.
I think the Chinese government is backing their labs by less direct means. For example cheap electricity and investing in chip manufacturers such as Huawei and cxmt.
Deepseek specifically, is known to operate with minimal resources. The entire company has around 160 employees and every model they release must break even within ten months.
SpaceX has been awarded roughly $22 billion to nearly $30 billion in cumulative public federal contracts, the $280 billion CHIP act, and who knows how much the CIA + NSA are spending.
Awarded that for doing work, though, not to clone someone else's work and release it for less money.
I think this is different for different labs. Some innovate while others copy.
Same goes for US labs, not everyone are as innovative as google deepmind:
https://www.forbes.com/sites/antoniopequenoiv/2026/04/30/elo...
> not to clone someone else's work
Ironically your comment will be cloned and used for training.
> Awarded that for doing work, though
Yes, just that. Being a huge supporter of the regime - at a time running a department - surely has nothing to do with it.
Whose work on KV-cache reduction did Deepskeep clone? Please be specific.
> not to clone someone else's work and release it for less money.
... You're saying this about _deepseek_?
I'm pretty sure, the training/fine tune through Claude/OpenAi would not be possible without the army/China's hacking teams (and legal protection). So it's more than just a cheap electricity.
Supporting deepseek is just like supporting the Chinese army, no need for that. Though it goes both ways, OpenAi subscription just lowers the cost of the US army as well.
Sorry, are you saying that it's a subsidy that it's legal for them to distill other models, or what? It's not even clear that there's any kind of protection for models in Western countries. Why would they be worried about the legality of distillation? And what does the army have to do with distilling models? Like you're saying it's a subsidy that China protects their borders from invasion by the US?
It's not a subsidy for the Chinese to say they're going to ignore our IP laws. It's a subsidy for us to say we're going to make them and try to push them on the world. It's literally granting a monopoly by legal force. It's very obviously not aligned with the interests of the American people, while China releasing things in the open is.
Same with the US legal system protecting Anthropic from copyright laws from all the books and other stuff used for pretraining...
>Supporting deepseek is just like supporting the Chinese army
Good, where do I sign up? At least they aren't exploding little children and generating chaos in the oil market.
https://www.business-humanrights.org/en/latest-news/anthropi...
I want you to come here and apologize in public when Xi starts bombing Taiwan like he's promising.
I want the current president who has lost his mind to apologize to two California cities.
As bad as casualties get in a future war like that, it's still very likely to be a symmetric war where the main targets are military infrastructure. No systematic annihilation of universities, hospitals, religious sites, journalists and health workers like certain genocidal regimes backed by the tech oligarchies have been doing.
https://www.ohchr.org/en/press-releases/2025/06/israeli-atta...
https://www.unesco.org/en/articles/unesco-presents-assessmen...
I want you to publicly say that this turtle biologist and this journalist are Hamas: https://www.theguardian.com/world/2026/jun/20/mona-khalil-tu... https://www.theguardian.com/world/2026/apr/23/lebanon-journa...
When one player tries to monopolize AI tech and market by all means, others are those who don't want to be hooked to a foreign will in the future. I think this is the reason in doing open research, and it's likely positive for all of us.
> I think the Chinese government is backing their labs by less direct means.
It may be that the effect of the backing in both places is essentially equal but this statement is strange. The Chinese government invests directly in Deepseek[1]. Notice the article, in addition to saying the CCP is investing, says Tencent is also a major backer. CCP owns a golden share of Tencent.
[1]https://www.cnbc.com/2026/10/06/deepseek-funding-round.html
Ok, but Trump stated that the US is in active talks to take over parts of both OpenAI and Anthropic, too: https://www.cnbc.com/2026/06/05/trump-open-ai-altman-stake.h...
I'm not arguing about a comparison here. I'm trying to correct a misconception about China that I see a lot from Western perspectives.
Everything a Chinese company does has the explicit backing of the CCP, at least ideologically and usually financially in some fashion. While DeepSeek may not be the CCP, it couldn't exist if it was expressing any kind of ideology that wasn't inline with them and, in this case, is explicitly funded by them.
A Chinese company has no freedom to say Taiwan is a country the way someone in the US could suggest California succeed from the nation.
Any public message you hear coming out of China has the implicit approval of the Chinese government.
I would like to point out that a lawsuit is currently live about the us govt forcing US tech companies to retaliate against us citizen criticizing the govt’s policy.
Well, there was one happy audience in Nebraska with the blessing of El Presidente that was happy if two California cities get the shaft.
> Everything a Chinese company does has the explicit backing of the CCP, at least ideologically and usually financially in some fashion.
If you're not arguing about comparison then don't state this as if it's any different from the US. Because you make it sound that way. Or do you think that if tomorrow OpenAI came out vocally supporting the DSA - that suddenly they wouldn't see lots of barriers rise up out of nowhere?
Come on now. That's fairy tale land.
Well, they are helping 17% of the world's population, and the US is currently actively engaged in trade wars and economic warfare or explicitly attempting to leverage it's hegemony against its long term alliea for short term gain.
There is as much to criticize about American hyper scalers and AI labs and the lack of interest in helping humanity or contributing to open source, but that might not be as popular an opinions on this site.
Only 17%? I'd say driving down the price of AI helps at least 95% of the world's population. It certainly helps me.
17% is the approximate population of Earth in China. I was using a simple example as a counter to the anglocentric view the parent poster was expressing by trying to paint China as a villain.
I assume 17% is roughly the percentage of people that use LLMs globally (the numbers vary between 15 and 20 percent).
Amazingly low, given potential impact. Where is this use concentrated? Who has the best data on use right now.
> Well, they are helping 17% of the world's population,
Any open weights model that Chinese companies are releasing for free are allowing everyone in the world to have access to high quality realizations of these tools without the risk of the US regime arbitrarily cutting you off.
Also, it's funny how all the neoliberal mantras of how free market drives progress through competition stops when it's a US company that's being challenged.
I suppose in some sense preventing nuclear holocaust is a short term gain...
What nuclear holocaust has the only country to use a nuclear weapon against a target prevented?
Do you think the world is closer to or more distant from large scale global conflict today than it was 10 years ago, or even 3 years ago?
Iran has expressed a willingness/eagerness to nuke Israel and the US. That outcome would be less than ideal
> The reason they release the AI models is economic warfare against US
The story is so much more complicated than that, to the point that this economic warfare theory is basically a meme.
Chinese models are open because they don’t have a choice. “When you trail the frontier, openness maximizes reputation per unit of capability. The moment you lead, you close.” [0]
[0] https://earnedintuition.substack.com/p/involution-without-ex...
They do have a choice, as evidenced by all the times when the other option gets picked. The article you link tries to acknowledge ByteDance's Doubao and Alibaba's Qwen Max as exceptions, but forgets about Baidu's Ernie and iFlytek's Spark, which are also closed-weight. There's simply no consensus yet on which strategy is better, so different companies end up making different bets.
So is American models. They are subsidised heavily but still can't provide cheaper access. Their fault is to assume all countries can afford them.
By whom?
If you want a direct example of the US subsidising a major AI lab, we can use Microsoft's investment in OpenAI. Microsoft has the largest ownership stake - 27%.
Microsoft did not pay with money - it paid (mostly) with Azure cloud computing credits. MSOFT is then able to write this off as a loss against tax.
It is generally far more tax-efficient in the US to write a loss in this way than it is to write a loss for a cash investment.
In this case, I believe the difference was ultimately highly significant. When including MSOFT eventually writing-off the deprecating Azure hardware it had used to buy the OpenAI equity, the result was MSOFT's tax reduction being either close-to or exceeding the actual cash value of MSOFT's investment in OpenAI - IIRC.
These examples represent taxes that the US chooses not to collect - the US could choose to make investments like these less tax-efficient. Instead, by making them extremely tax-efficient, the US subsidises the transaction hugely.
This is a joke right?
Even ignoring monetary subsidies, there are the non-monetary ones: not being sued into oblivious by the government for their countless hacks of other companies and countries, the slaps on the wrist for massive piracy, the waving of environmental (and other) regulations in order to allow their data centres to be built an operated.
> Even ignoring monetary subsidies
That's what I'm asking - which monetary subsidies?
> not being sued into oblivious by the government for their countless hacks of other companies and countries
That isn't normally how enforcement works, and it hasn't been very long since they disclosed those breaches. If the victims want to pursue legal action, they can, and they still may!
> the slaps on the wrist for massive piracy
So judges and juries are involved in the subsidization conspiracy, too?
> the waving of environmental (and other) regulations in order to allow their data centres to be built an operated
Sure, though if you think this isn't happening in China too, I have a bridge to sell you.
CHIPS act, OBBA, defense, AI upskill grants.
That's just the ones i know off the top of my head in the US. Those programs account for more than 60 billion in committed spend.
The original comments is "DeepSeek, and other Chinese models, are heavily subsidised by the Chinese government"
> That's what I'm asking - which monetary subsidies?
I am listing for you the non-monetary subsidies, which are just as real and equally important.
> not being sued into oblivious by the government for their countless hacks of other companies and countries
If it was one of the Chinese labs doing this hacking, the government would be stepping in. If it was European labs they'd be stepping in. If it was you or I the government would be stepping in. That is a massive subsidy (they don't have to worry about the same legal fees and exposure) and being allowed to continue doing business is in fact priceless.
> So judges and juries are involved in the subsidization conspiracy, too?
There is no conspiracy. They are objectively operating by a different set of rules than you or I could operate in this market.
> Sure, though if you think this isn't happening in China too, I have a bridge to sell you.
I never said it wasn't, I'm saying that it's happening here and it's a very real subsidy.
> equally important
They really aren't. Nothing in the development of these systems, absolutely nothing, is as important as money. Stop these silly false equivalences.
These are not silly equivalences.
All the major labs would be dead in the water if the government acted on the things I mentioned. If they treated the labs the way they would have treated us for all those hacks. Or if they enforced pollution measures (or god forbid ban on-prem turbines because of the climate damage). Or rule that training on data is not fair use, or that ingesting GPL3 code and then turning that into weights counts as a derivative work. Or. Or. Or.
Those are all just as important as cash.
Investors, but also the government in allowing these companies to siphon electricity away passing the increased cost to the consumers
>Investors, but also the government in allowing these companies to siphon electricity away passing the increased cost to the consumers
Not arbitrarily banning companies from buying a product from a supplier doesn't meet my definition of the term "subsidy".
Fair enough, but it sure meets the definition of 'helpful intervention'.
>Fair enough, but it sure meets the definition of 'helpful intervention'.
I don't see how. Anyone can buy electricity without any intervention from the government. An intervention is an action that changes what would happen by default.
By investors. OpenAI and Anthropic are not profitable (Anthropic is profitable if you allow them to invent what profitability means).
I heard someone calling the key metric in Anthropic financial reports EBBT: Earnings Before Bad Things[1] :-)
[1] Where "Bad Things" would be the typical interest, taxes, depreciation, amortisation plus the Anthropic specific employee compensation, LLM training (you know, for the LLM lab), revenue sharing agreements (which is a form of paying for infrastructure), etc.
A recent YouTube video by Patrick Boyle said the same thing. They are only profitable if you ignore all the costs that make them unprofitable such as paying employees and developing A models.
That's not how gross margin works.
For LLMs, marginal cost is just electricity
I made a statement about whether Anthropic was profitable I don't understand your reply.
Assuming you meant to reply to me and not someone else are you saying under GAAP accounting standards Anthropic is a profitable business because under GAAP accounting their only expense is electricity?
(I see what happened you skimmed the conversation and didn't follow what was being discussed.)
Investment is not a subsidy. Words still have meanings. Nobody in this thread has successfully answered the question: subsidized by whom?
Non-punishment is also not a subsidy; again, words have meanings. Let's use the correct words.
But some investments are subsidized (tax liabilities) among all the other things.
Indirect subsidies are an accepted economic concept and well studied. The posts you replied to are obviously talking about indirect subsidies and dismissing them because they didn't say "indirect" everywhere is just sophism.
The government, specifically Trump's government and the current money circle in AI inflating American company stocks. In addition to all that American models do not share their papers like Deepseek and Qwen do. So you can literally say Chinese models are doing it for charity at this point.
Of course you can say that, but you literally cannot be serious if you do!
Can you say a US company donating to a non-profit and writing off taxes on the basis of it is donating to charity?
Would you say you could be serious in stating this?
Oh I am serious!
Capitalistic exploitation against the working classes.
The US regime, which includes the ruling class: the US VCs and megacorps are just as much of an extension of the US regime. They're incredibly intertwined.
> The reason they release the AI models is economic warfare against US, not because of charity or kindness
There are many other reasons Chinese companies releasing models open-source or open-weight makes strategic sense.
A really easy-to-understand example is a company who has a near-monopoly on "serving video content" releasing a video model openly.
If you can be relatively certain that video content created by a model (which you have trained, using data from your own platform) will be ultimately served on your own platform, thus generating revenue from watch-hours, it makes sense to make those models as widely-available as possible.
It's also a net-positive if people use your public research to build better video models, because - again - you are reasonably certain that the even-better content those new models produce will be watched on your platform.
The alternative would making models harder to access and learn from (broadly, the current western model). Many would argue that Google, in choosing to not optimise its video generation models for "availability", is directly causing less content to be uploaded to YouTube. This is the trade-off.
I don't know much about DeepSeek's financing specifically, which obviously doesn't release video models - so I don't know how directly this analogy runs, or who directly benefits from the extremely evident rising tide that the public release of DeepSeek's research creates. However, this does not negate the broader rising-tide effect of the scientific method.
It's certainly also true that it's geopolitically beneficial to be able to undercut American labs' models. If I ran a global superpower, I would probably want my country to be technologically competitive too.
But Chinese companies are already serving a huge volume of customers in a complex, existing marketplace, before even thinking about the US market, and it's overly simplistic to assume that their entire strategy revolves around economic warfare directed specifically at the US. It's more nuanced than that.
This is, of course, without even getting into opening the can-of-worms around whether US economic policy also results in the US state functionally subsidising technological innovation, how comparable that is to China's model, etc.
Interestingly if Google did release readily available video models then they'd likely be dealing with an order of magnitude of scale in the same way GitHub has had to.
Causing less content to be uploaded, when you are clearly the gorilla in the room, looks to be a wise strategy not a trade off.
> DeepSeek, and other Chinese models, are heavily subsidised by the Chinese government.
DeepSeek is basically a research lab founded by a hedge fund guy, more than anything else.
I'm not sure "subsidised" is the right word. If a government funds research and the results are released openly, that's just publicly funded research. It's how a lot of science works in the US and Europe too.
Same goes for the US companies as well. Current US administration wants US to win the AI race at any cost and China is the only competitor left in the race. Winners always write the history or in this case the future of humanity
The open-weight models are a great boon to all American companies other than a handfull of Mega Corps in the A.I business. Care to explain how this is "economic warfare against the U.S"
Do you not see the risk inherit in sourcing everything from a antagonist?
Yes, hence why I'm happy there's competition so we're not stuck with US tech. After all, you've proven to be very unreliable allies to us.
Yes we see that in all other countries trying to avoid dependence on the US for everything. What we don’t see is competition with other countries having anything to do with warfare.
What about cheap solar panels and electric cars, is that economic warfare too?
To say that US model providers have a close relationship with their government would be putting it mildly.
Which is the best kind of reason.
It is great for everyone except for a few people who want power over everyone else, and the fact it is not charity of kindness makes it more sustainable, because charity and kindness is quick to go when big money and politics is involved.
I want more warfare like this. Building stuff instead of destroying stuff.
Well, US models are economic warfare too, of course.
And OpenAI and Anthropic are just in for the love of the game? Our of sheer desire to help humanity?
Obviously not but they employ American researchers and aren't beholden to the CCP
So you’re telling me warfare doesn’t having to be blowing up schools and children, skyrocketing prices, crippling sanctions, and all that shit? Can I sign up for more of this warfare.
They are also releasing the weights so I am more inclined to think they are better than most.
Not sure why so many people will vehemently refuse this idea. I won’t say it’s 100% true but it would be foolish to dismiss it. China is very much an adversary to America and has made it pretty clear they want to be a dominant leader of not the new world leader. Not here to evaluate what is good or bad. Keep in mind historically China has aggressively fostered industry (not unlike the west) but sometimes even more aggressively.
Americans need to travel to China. The Chinese have zero issues with us. They quite like Americans. This is such a weird propagandist take. Idk if you remember but both country's leaders just had a slumber party for 3 days in DC. This is not what enemies do.
Not even our leaders say China is an enemy, ita mostly businessmen who are scared of competition and trying to regulate chinese out of their markets so they can make more money milking us.
I have been to China many times. The citizens and visiting the country has nothing to do with global politics. Your take is just as propagandist as any other. China is not unlike the US and China has also made it subtly clear their desire to be a dominant global force. I am not marking judgement on it but don’t be a fool thinking that China is not some level of threat. This is how these discussions become weird. I don’t think China is going to blow up the US and I think it’s often a weird talking point by some US politicians but I also don’t think they are a peaceful actor but folks like you will suggest I am repeating propaganda. No, just pointing out that if you look at their language and actions they are an actor to pay attention to. Just because leaders meet means nothing.
Sounds like some Falun Gong conspiracy. Im not worried about China.
The same is true of the Iranian people. There's sometimes a difference between the government and their own citizens.
Your just saying it’s beating america at its own game and crying fowl.
Economic warfare against the US? I mean maybe against specific US companies and stakeholders but on the whole it seems like it's good for the US economy as well as the rest of the world, kind of like supplying free electricity would be
It just so happens that helping open source is the result. Image how far behind we’d (the hackers, not the moneymen) be as a sharing community be without them.
Claude, and other American models, are heavily subsidized by investors (and the United States government).
At least DeepSeek didn't build its models on government subsidies. It came out of a quantitative hedge fund, so funding was never really an issue for them.
That would have been a scathing criticism if anyone believed any of the American companies have the goal of helping humanity or open-source.
Notwithstanding that Chinese publishing methods actually does help both humanity and open-source.
If your belief is accurate, we should expect China to short the IPOs of Anthropic and OpenAI and release better frontier models immediately after their IPOs.
Does anyone think that likely? I have no clue or bias.
I don't think that's likely because my understanding is that Chinese people in mainland China have a tricky time shorting American stocks due to Chinese capital controls.
It's one of these rare cases where intention is not what is the most important - the net benefit for consumers and companies outside of the USA is indisputable.
Whatever their reason, I'm glad they do it. This tech shouldn't be monopolised by a handful of closed, profit-seeking corporations.
And likewise, it shouldn't be monopolized by the CCP
What do you think your comment is refuting!!? Chinese are publishing their research publicly and opening their models! How on Earth would they monopolize anything?? Such nonsense. If and when they start copying the Americans and closing everything down, you can say stuff like that. Until then, ffs stop this nonsense.
Precisely. It's similar to their practice of aluminum dumping to depress US aluminum prices, which causes our aluminum mines and mills to close.
If the means of achieving their goal are sometimes mistaken for charity and kindness... are they the good guys?
To be completely fair, write this kind of paragraph for every AI category. Would love to see your take.
How do you know? You’re not speaking for the Chinese government, are you?
Economic warfare against the US is charity and kindness to a sizable portion of the world's population, especially when the US uses it's global hegemony as warfare against them. Neither system is perfect, but lets not be disingenuous.
I like when bad intentions produce good results, I'm tired of seeing the opposite in practice.
At some point though _the purpose of a system is what it does_, regardless of their intent.
> DeepSeek, and other Chinese models, are heavily subsidised by the Chinese government.
So were Amazon and Uber by the US, which have now established monopolies across the globe. To the countries suffering from those, there's zero difference with China doing it to solar. Actually there is, at least solar got them cheap renewable energy in return. This would never have happened in the US because big oil interests would make it take decades. That's the reality.
You need to spend 10 years outside the US, deprogram, and then go back.
> DeepSeek, and other Chinese models, are heavily subsidised by the Chinese government.
Yes, because everything China does is against the US. That's all they think about day and night. God forbid they want to corner the global market or have a genuine business case. How dare they provide options for those who can't afford a measly $200 a month? How can we let Chinese labs publish research for free for the whole world so that they can benefit? The nerve! To think they can use soft power instead of military might! I mean, Anthropic and OpenAI are the last bastions of human kindness and charity. Right?
Right?
I’m neither in the camp of Chinese or the Americans(collective West) in general..
As a neutral party, this characterization is crazy.
As if the AI companies - Claude and OpenAI are guardians of freedom and humanity and very charitable to the global society without any self interests… “Chinese models are subsidized by the Chinese Government, therefore they’re inherently bad for humanity” is a highly propagandist argument. The politics of US vs China may be whatever it is in reality.. You have one company releasing their models for cheap, actually open sourcing their trained weights, and publishing details of their optimizations and learnings for others to use. The other camp actively “aligning” their models, nerfing their capabilities, hyping their swarm activities from poor sandboxes, and trying their best to lock users into their harnesses and walled platforms. They are subsidized by the capitalist VCs who are essentially waiting for their payouts..
At some point, one has to see things for what they are and evaluate their own reasoning..
I’m happy to stay provider agnostic, try all models and cheer any useful progress as open as possible.
> DeepSeek, and other Chinese models, are heavily subsidised by the Chinese government.
US companies are burning colossal piles of cash in ways that makes it unclear if it qualifies as dumping, not to mention their deep ties with the country's regime.
Claiming that companies from a country have ties to the regime and burn through cash is a very miopic accusation.
I wish US was rich enough to subsidise development of open source science and useful open AI models. I wouldn't mind US waging this kind of wconomic warefare against China or everyone else on the world.
Charity and kindness is not a motivation, it's an outcome of what you do.
Meanwhile, American AI models are heavily subsidized by stock market speculation. Ultimately, the subsidies from both countries are flowing out of the pockets of individuals.
That's an interesting definition of subsidized. I don't thing I've ever heard it used like that before.
US is basically doing economic warfare and bullying against everyone else ATM :)
What if economic warfare against the US does help humanity?
Why are we assuming a strong US is necessarily good? As a European, I have seen plenty of evidence against that stance lately.
I understand that Americans might prefer a strong US. But conflating them with humanity is a leap that I don't think one can make without any backing.
What evidence is that? Please share.
> are heavily subsidised by the Chinese government
We hear this about literally every industry the Chinese excel in - that it's only because the government subsidizes them that they succeed. For chip manufacturing, for batteries, for EVs, for solar, for AI. I don't see how the chinese government can afford to subsidize all of these industries and still have them contribute to the GDP.
A conspiracy to make the US look bad by being better at producing all the goods and services the world needs at a reasonable price. Have they no shame?
What's wrong with governments subsidizing scientific work?
You think extremely US-centric. China has a different economic model than US and your rules for a specific kind of Capitalism may not apply to them. You assume a country of 1.4 Billion people is obsessed with a couple of foreign AI companies. What if they don't care.
> but the goal is not to help humanity or open-source
Ok, but so what? That's what happening so far. Even a repressive totalitarian government I wouldn't wish on my worst enemy does some good sometimes.
If the Chinese cheap/open models rise up and destroy us, thats on us for giving them access to the tools to do so.
Not really this is an X algo conspiracy. Up until recently the Chinese government wasn't even that invested in these companies. We're talking very very small grants compared to training costs.
Its very xenophobic of you to say China has zero intention of helping humanity, and just wants to "wage economic warfare".
Last time I checked, it was ourselves (USA) waging economic warfare on 2/3rds of the world.
I dont get this cope people have where people have this idea that its impossible for a Chinese company (that make billions of dollars) to have done something by their own merit, but instead its always some Chinese Communist Party conspiracy where the main goal is to destroy America.
Lay off twitter for a bit.
>Up until recently the Chinese government wasn't even that invested in these companies
>Last time I checked, it was ourselves (USA) waging economic warefare on 2/3rds of the world.
Up until recently the USA Wasn't waging economic "warefare" on 2/3rds of the world
The USA has been using financial sanctions, aka, financial warfare on anyone it's deemed an enemy for going on 4 decades.
China has been owning and controlling key companies in its industry for going on 6 decades.
Every country does this. Do you think the US government doesnt fund, regulate and control key companies?
China is not a threat to you, or anyone in the West.
Why the anti China propaganda from you? The United States is doing the same thing here and making American multi-billionaires even wealthier. The greed of American AI, GPU and memory and storage corporations is a black hole on availability to humans around the World. The result is a massive financial bubble promoted by the American Government to the detriment of our citizens. The USA has an AI ponzy scheme shuffling the same money between data center owners (Oracle & X), Nvidia and memory & storage vendors.
They're trying to pull digitally what they already pulled physically. The reason we can't manufacture a grill brush for a reasonable price is the result of years of Americans choosing the cheapest price. We gave up our manufacture base. They want us to give up our labs.
Deindustrializing a country is not something consumers can achieve. It starts at the top level, with politicians who construct a financial system where it's more profitable to speculate than to build or invest in real businesses.
Except they are open with the tech which is easy to replicate. All new models lowering cache prices is the result of DeepSeek's publications.
US economy is driven on quarter to quarter short-sightedness of stock gamblers and CEOs that have to secure their bonuses and perks.
Its them who chose to outsource heavily, not US consumers.
There are no victims here however, both benefited from the arrangement. The consumers and capitalists.
> The reason we can't manufacture a grill brush for a reasonable price is the result of years of Americans choosing the cheapest price.
No, it's the result of US leadership letting this happen. This is clear since China themselves would not let this happen. US leadership did nothing because they were best friends with the people who did it, and did not care one bit about their population.
Architectural/algorithmic tweaks do advance the efficiency frontier nicely. But raw intelligence mostly comes from data (not just its sheer quantity, but also how it's curated & cleansed) and the scaling law. The know-how about data curation doesn't seem to get published much, even among the open-weight labs, though.
This. Even in the efficiency frontier, it is a lot of data curation that actually makes many of those tweaks actually work at scale in practice.
I'm not sure I like this framing - so much of AI research has been academic, in the open, building on others people's work. Much less comp sci generally, math & philosophy, etc. The idea that rich companies can just build stuff in secret because they have resources is a fantasy.
So boring to see conversations moved over to Chinese models when that’s not even what we’re talking about here. This is about Mistral.
Europeans are irrelevant these days, surpassed by China, S Korea, Japan, Hong Kong, Singapore, etc. Europe is coasting on former glory and now has regulated itself to death and vacationed its advantages away.
Evidently not given that this model is on the open source frontier.
cost is higher than Chinese models that are better
How it should be. Knowledge should not be copyrighted. The world will be a better place with such information democratized
Also, the field moves fast, but slower than people do. Researchers and engineers switch companies every year or two, and the know-how walks out the door with them.
There would be a lot of competition even without DeepSeek. Workers can freely exfiltrate trade secrets without noncompetes in California.
Proprietary competition, yes.
>instruction manuals in paper form
So the most common way to publish manuals?
In research paper form.
I think the mean 'paper' in the scientific journal meaning; these are unfortunately often extremely bad 'instruction manuals'.
Maybe 30 years ago
> have not been a winner-take-all runaway acceleration game where catchup is impossible
From the Mistral site:
> ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe.
It is pretty capital intensive!
That cluster is literally orders of magnitude smaller than the compute pools used by Anthropic or OpenAI.
For training or for inference?
They don't publish numbers, but Anthropic has a single DC with 200k+ GPUs for inference, GPT-6 Astra is said to have trained on 100k+ GPUs.
I’m pretty impressed that they managed to get that close to the frontier with such a small cluster!
> I’m pretty impressed that they managed to get that close to the frontier with such a small cluster
Chinese companies also managed to put together their models with relatively small clusters.
Perhaps US companies are desperately trying to brute force their way into workable models?
According to Grok thats 7-10 MW. Tiny numbers.
To put that into context, the last wave of capacity SpaceXAI added 400-450 MW.
But how much of that are they using for training versus inference? They're serving quite a large user base.
These cards are like $3k each? That's, what, $12M and you keep the hardware? Honestly doesn't seem too bad.
More like $30k each.
Oh the server chip is 10x. That makes a lot more sense.
That’s kinda very small and light for modern trillion-param LLMs.
It's shaping up to be much more like a game of 'chicken' where each company tries to raise more cash without going bust... Ultimately the game of musical chairs is going to have to stop. In the US it looks like they are trying to get a government sanctioned truce in the form of regulation. That's what 'Pacing the frontier' means...
Even "runaway acceleration" isn't instantaneous. People imagine the singularity as something that happens almost instantaneously. But obviously it happens over time, and that time might be decades. It might still end up looking like a vertical line on a long-term graph.
If the singularity is defined as an AI sufficiently intelligent to improve itself independently, that AI is still limited by the resources required to do this improvement.
For sure people who don't grasp the difference between models, might be stuck in 'good enough' models.
But Opus 5.5/GPT is such a game changer in comparison to sooo many others, its still a moat for now.
You've got to think the comments about "good enough" are people who have not yet tried Opus 5.5. I haven't been this struck by a step change since 4.5/4.6. It's a bigger jump even than when Fable first arrived.
As for Mistral - I got really excited when they said Large 4 was focusing on being #1 in cybersecurity, because that's somewhere that they genuinely could edge out Anthropic & OpenAI. Have it actually solve problems, instead of Anthropic flagging "you tried to find a null pointer exception bug in your own code, we're now reporting you to the US government". But on the Mistral benchmarks I'm seeing, this looks very disappointing, but at least they haven't entirely given up. I genuinely thought Mistral had given up on new general models. They need to learn the bitter lesson all over again.
I feel like the "good enough" argument isn't about how big the gap between models is but about how good they are at solving the tasks at hand.
The capabilities of all models increasing so much all the time means there are simply less and less tasks you need a frontier model for.
Even if Opus 5.5 is 500x better than Deepseek, if deepseek can solve all my problems, why do I need to pay for more?
Many on HN still have the opinion that you must understand every line of code in the project, and that all is lost should you merge code that wasn't reviewed.
Obviously any model will do if you use it as a better autocomplete.
I believe that there is a large gap in expectations between different workflows.
Until the AI like reads my mind and produces perfectly production ready apps with minimal intervention from my side, there is still going to be room for improvement.
If Opus 5.5 is 500x better than Deepseek, but Deepseek can solve all your problems, maybe you need to work on better problems. If you don't, and you're in business, your competitors will work on the better problems. If you're an employee, your employer might prefer to pay Anthropic instead of you. If you're doing projects you're interested in, you can tackle more ambitious projects with a more capable model.
This morning I elicited a microkernel operating system from Opus 5.5. Well, mostly. It doesn't implement task switching yet; we'll see if it runs into a wall at some point. But it boots in QEMU, and it's running a user process in ring 3 and serving web pages.
The best problems to work on are not necessarily the hardest ones, nor the ones that need the most intelligence. They're the problems you, or other people, actually have. Are you going to give up on painting your deck because it's too easy and you don't need a 500x genius to do it?
> maybe you need to work on better problems
I have enough real problems in life. I don't need to invent new ones just because a new technology is available.
Many of my problems in life are fully solved far past my satiation point by a 3b model that costs me nothing to run.
Many others are not.
But in either case, when I am acting and living wisely, almost all of my problems exist prior to the existence of technological solutions to those problems.
This is also true for the customers and employers that I care to work with. This has changed in me over time, but I now try my best to avoid inventing new problems. The world has enough big, important problems already.
> This morning I elicited a microkernel operating system from Opus 5.5.
This is cool but also a good example. I don't need a personalized microkernel just because it's possible to have one.
Maybe I need one and I don't know it, but the problem statement definitely isn't "I have inherent desire for a personalized microkernel".
Yeah no, if you elicit Opus 5.5 , anyone else can, and you have no moat either.
But if on the other hand, I mostly use my human intelligence and just need a dumb model to complement my human intelligence at low cost and high speed (say review every commit to catch obvious bugs), I have a much better chance of building an actual moat than you do.
But outside of coding, it’s even more clear that you don’t need frontier intelligence. My customer service agent is very happy with a 100B param Deepseek flash model, thank you!
> say review every commit to catch obvious bugs
I’m using subscription models for exactly that, better models catch more subtle bugs, and they catch them faster. It works out far better in terms of work-hours saved.
Also A/ then OAI slashed token pricing by 2x~5x on their latest models
I use Opus 5.5 daily for my job. I am aware (and in awe of) it's capabilities.
Look at the context in which I used that term 'good enough'.
What i was saying is that there are tasks for which a dumber model can be good enough, and for organizations with sovereignty/ privacy concerns, those concerns can be strong enough to incentivize the use of a dumber model.
> I haven't been this struck by a step change since 4.5/4.6. It's a bigger jump even than when Fable first arrived.
I had the exact same experience. And unlike Fable, it doesn't gobble up your entire usage limit in a few hours.
I always wonder what the "good enough" people are actually using it for.
The thing is, for how long? Are they going to keep giving you "so much intelligence" for a "small" subscription dollar amount? When they really, for real, need to start making money to cover their costs, what do you think it's going to happen? Suddenly you will start having tasks that a "good enough" model is going to be fine.
“Good enough” as in we don’t see any point in 1-shotting everything we want to build in lightning speed. If Opus 5.5 can 1 shot it then that product is essentially commodified, no point in any one building it except as an internal tool.
If Opus can’t 1-shot it, then it must rely on our human intelligence which can be complemented well enough with a dumb model as a frontier model.
I assume that you'd also argue in favour of working in a team with an average IQ of 80 as opposed to 120.
No it is not. Only maybe for the noobs or vibe coders.
People who aren't afraid of rolling their sleeves into any code base? The difference is practically zero.
I’ve been saying that. When you have no idea what you’re doing, you *need* the latest greatest model because it’s the only way to reduce errors.
For people who have some expertise, the models accelerate the grunt work, but you’re the one validating it.
I agree; yes, I can see that they need a bit less hand holding each cycle, but I also see these "frontier" agents do some absolutely dumb shit that I have to correct and then I'm wondering if I'm the looney one here.
Maybe it's because people stopped watching what their agents are doing and stopped looking at the quality of the output. But I still see agents being absolutely mindless like a junior dev.
Recent example: it updated an an API to add newly released models to the backend. There's a list of models that require specific configuration for the reasoning effort and temperature or the API call fails. GPT 6.1 Sol misses this and code fails at runtime because the newer models need to be added to the list for special handling of temp and reasoning. Fixes it for one model and tests it for that model using an E2E test. But doesn't test the other models that were added for the same error condition...I had to explicitly ask it to do so and it finds them and adds them to the list and says "that's on me."
Yeah, not that smart.
You're not looney at all. Frontier models do dumb things all the time, especially on mature codebases. Just yesterday Opus 5.5 butchered the OOP model in a codebase I work on - it duplicated a load of classes that should have just been subclasses. A junior checked it in very satisfied that it was perfect. The LLM review passed it, the tests were fine, and it implemented the feature successfully. It's just the code design had poor taste and poor long-term maintainability.
I keep seeing this kind of thing over and over, and honestly it's not got _that_ much better since the big breakthroughs about a year ago.
For sure I happily vibecode stuff without worrying about it when it's a greenfield project, and if the LLM has written it entirely from scratch then usually it's well structured and sane. But making changes in messy, mostly human-written mature codebases is still a minefield.
Try a bigger code base or more complex stuff and you will easily see that the solution, speed and amount of problems Opus5.5 solves vs older models is relevant.
I am frequently running agents on a multi-microservice application workspace where I really need the 1M context windows, because they are filled to the brim when implementing features that require changes on several services and APIs.
This works fine with Opus 5.5. But it also works fine with GPT 6.1 Sol, Kimi K3 and MiMo 2.6 Pro.
It doesn't work equally well with Sonnet 5.5, interestingly.
Please, give it another 6 months and they catch up. The American labs are currently trying everything they can to block others instead of advancing their models, trying to build an artificial moat. The American models are not that great, they are good, and they have a lot of agentic workflows in the back, but its basically a hardware limitation at this point. Once the HW makers catch up, and we can move away from the Nvidia monopoly, things will speed up quite a lot IMO.
Just one more release cycle bro, I swear
We've been hearing the line about them only being a few months behind for a year now, during which time O/A have grown their revenue like 10x, haven't they?
We've also been hearing we're 6 months from AGI for about three years, and here we are.
"Now, here, you see, it takes all the running you can do, to keep in the same place. If you want to get somewhere else, you must run at least twice as fast as that!"
Phantom Tollbooth?
Those are two different things. The market is expanding, so even if competitors are catching up, you can have your own revenue, in absolute terms, grow.
The thing is O/A have been much louder on pacing the frontier, lately.
And yes, open weights are still behind, but are catching up.
I agree that Opus and GPT are almsot surely better, but so many real users are nervous enough about giving Anthropic and OpenAI access to all of their internal information that they may be willing to stomach worse models if it gives them more security.
The real question is if this model is good enough that it can still accelerate work, and not be a hindrance to real work like older Mistral models often were.
If they can do that, they'll have customers.
The Navier-Stokes fiasco made me push for local/controlled models very hard. "Can't rule out" that they stole data (backed up by their backdoor offers of sharing credit).
If these companies will steal from deep pockets like Disney or Sony (some of the most infamously litigious copyright trolls to ever exist), they won't think twice of stealing every bit of code you upload to them.
If your code passes through an AI company's servers, you can assume you just gave it to them. In turn, when your competitor tries to copy that new feature you just added, the AI is now trained in exactly how to copy you and eliminate your competitive edge. Unlike your employees, the AI isn't bound by the same rules and even if it were and violated them, your company probably doesn't have enough money to prove it in court (and that's if we somehow reverse some of the stupid "AI is the most transformative use of copyright I've ever seen" judges who have drunk the coolaid).
Most companies could build the compute to run GLM or Kimi models for way less than the potential loss due to IP theft from using third-party systems.
I would also add to this, there are ways to use customer data to improve your model outside of just using it as “training-data”.
A simple loophole, use the code to create an RLVR environment where the resultant code is the end goal / max reward. Technically the customer data is never trained upon, but effectively you’re using it. Even better, use the code as a seed to generate synthetic data similar to it and use that synthetic data as rewards in an RLVR model.
Unless you can host the ChatGPT model on your own servers, which I know some enterprises are doing, I don’t think there’s any hope of protecting your data / competitive advantage from these frontier companies. Better to be paranoid, than be commodified by these companies.
tbh to me if the AI company writes all of your code & your eng don't even review it anymore then… the AI company _controls your company_. maybe that's ok if you make widgets but less ok if you do anything in dev tooling, security, or [insert market they may suddenly decide to compete in].
If I'm not mistaken, they did rule out that they stole data from the mathematicians in question.
IIRC they ruled out that humans knowingly stole that data, but they didn't rule out that the AI agent might have
"We investigated ourselves and determined that we are not to blame"
Surely Sam Altman would not lie to anybody.
People say this exact thing every single time a new frontier model comes out.
It's an improvement, but game changer might be a bit of a stretch. If I lost access to Anthropic or OpenAI models tomorrow, I would be annoyed, but would reach for a slightly inferior model. Last year I wouldnt be able to say the same, and rhe challenge is that the moat is drying up fast. Whether its general improvements in model training by other competitors, or straight up distillation of SOTA models, the moat is shrinking and the available capital and spend for American model providers is going to dry up quickly as competing good enough models are adopted by more consumers.
It's especially the case as more non-Americans look to self hosted models and domestic cloud inference providers using open models that the US providers who are still leading the charge need to drastically drop their prices and find a path to profitability in order to maintain their lead and retain the advantage they had as AI turns into a commodity (which is happening faster than I think even the frontier labs initially predicted).
I use Opus 5.5 at work.
I use MiMov2.6Pro, DeepSeekv4.1Flash, GLM5.3, Hy4, Qwen3.8 and KimiK3 at home. Opus5.5 is not a game changer.
I do a broad amount of diverse experiments/projects I always wanted to do and throwing Opus5.5 against it just works
I have to admit, Sonnet got really good too.
But Opus just uses tools, a broad spectrum of it, etc. it feels like sure if you add some router behind it you could split it up if you need to but if you give me the choice, its opus allll day long.
> X is such a game changer
I hear this literally every other week about whatever the newest FoTM model is.
Unless you can provide concrete examples of things you can do with them that you simply couldn't do with last week's model, it's absolutely meaningless.
Let me know when the game changes are more than a month apart.
Less and less work requires a frontier model though.
My todo app generator does not need opus 5.5
Yeah, I agree with this. I think the "the models are good enough" narrative is a myth. I've heard it so many times over the last year, but the model number keeps changing...
There is no ceiling on what you can accomplish with more intelligence, so there will always be a market for the best models, and that market is likely to just keep growing. If Opus 13.5 can one-shot a profitable company or discover a new disease treatment or whatever you can think of that a swarm of relentless super-geniuses could accomplish, companies (and governments) will throw money at it.
I also think there will always be a market for many sub-frontier models that will continue to grow rapidly as well, because "good enough" is definitely a thing for a given task.
One could still argue that models are good enough for a given task. I primarily use Opus at work for writing code and I realized that for my usage the intelligence of Opus 4.8 is more than enough. Sure the newer models are better but I can still do my work with having access to newer models
>If Opus 13.5 can one-shot a profitable company
No company would ever release such a thing
Despite Opus 5.5 got really bad the last days for me. Looks like they nerfed it again. This is extremely unreliable.
Or maybe they secretly believe you are trying to distill their models and are deliberately degrading your experience. Who knows with them?
You are laughing. Until it happens to you! :-)
Am I reading this correctly?
This appears to be roughly as good as Sol 6.1 (which is quite good), considerably faster in terms of wall clock for complete tasks, and considerably cheaper (where Sol 6.1 is already good value - just really slow).
That seems too good to be true...
But I really hope it is true...
I can confirm that you're reading this incorrectly. There's a reason behind them only comparing it to open-source models released months ago. Here's a good aggregator: https://artificialanalysis.ai/#intelligence
Hear! Hear! I really want European models / AI labs to succeed.
I trust them and their populations to provide a more societal-friendly version of AI, putting pressure on the US tech oligarchy, while also providing democracy-friendly open models that I don't trust to happen with the Chinese labs.
What do you mean "good enough"? Did you mean "large enough"? ;)
Disclaimer: I'm not sure how much of an IYKYK factor applies to this joke.
We haven't reached RSI yet. Once any entity reaches RSI, the runway scenario will happen.
Truly, this is what the Lord's prophets have revealed to us! (Eliezer 11:52) Keep strong in your P(singularity), for when the Kingdom arrives, He shall judge us in His righteous glory, whether to eternal annihilation, or rebirth and life in His Memory Eternal!
Assuming RSI is something that is possible as you envision it in the near term. I think that it will happen at some point, but I think we could still be a long way off. I don't think anyone can truthfully say that it is right around the corner.
Yes, so far the competitive dynamics feel more like cloud computing than web search.
> Its quite interesting to see that at least the early days of AI so far have not been a winner-take-all runaway acceleration game where catchup is impossible.
Mistral is also an European company. As we live in a time where the US regime is engaged in pyrrhic geopolitical tactics, it's good to know that it can't threaten to cut access to models during s period where everyone is rushing to incorporate them more and more in our life.
I think it's a mistaken belief that AI as we found it is the exponential runaway train.
So it makes sense, since all you need is compute, that there's a ceiling and specialization is going to be more valuable then some super AGI.
Especially since the worst people seem to be the ones who think they'll all run away with the bag.
I don't think it is, and I think that is what will pop the bubble. All these companies have winner take all valuations, and that won't happen.
... unless they can legislate it, which is why they are flattering heads of state and scare mongering about dangerous AI.
It's a retrain of asian model.
How do you figure? I haven't met a single person who doesn't use Claude or Codex for programming in any serious way.
Then you dont know people working on highly sensitive info with stringent privancy concerns.
they just use claude on bedrock.
I'm happy I've never met you.
I mean, Mistral is about 9-12 months behind here when you look at its overall benchmarks versus the models released around a year ago.
Sounds ok to me. Claude was fine at the start of the year, and now with Mistral you also get EU sovereignty? I'll take that.
On a purely technical level, maybe? But in terms of actual revenue, is there really any chance of anyone catching the big labs?
Obviously, this is only a valid question if you don't believe that open weights are about to eat their lunch and their revenue is about to collapse, or they're running a super unprofitable ponzi scheme propped up by investor money that's about to collapse like a house of cards. I don't find those positions credible at all though.
If you do, then this question isn't really for you, as I'm more interested in thoughts from those who think that OpenAI and Anthropic in particular are about to be the largest companies on earth in a couple years. Could anyone catch them at that point?
> But in terms of actual revenue, is there really any chance of anyone catching the big labs?
I don't know about revenue, but I suspect multiple other labs are already beating OpenAI/Anthropic on profitability. Staying on the frontier is expensive, and it's hard to recoup those R&D costs when you have a bunch of other labs nipping at your heels.
If you concede the previous point, then the only way for OpenAI/Anthropic to keep growing long term is to swallow the whole economy (i.e. mass job replacemnt), and that's a bet I wouldn't take.
I think the actual plan is to swallow a good portion of the job market. It’s the only thing that makes sense and I hear VC podcast debates on which percentage of jobs justifies the market cap.
Maybe. I can't freaking wait for the IPO filings so we can finally put all this to rest. (haha, like that'll actually put it to rest on HN, but at least we'll have better data)
The big lab revenue may not be catchable, but im not sure it needs to be.
If they can carve out a niche of industrial and governmental partners who rely on them for sovereignty reasons, it may be enough.
I completely agree, I think AI is a vast ecosystem will all kinds of profitable niches and sub-markets.
It's an open question as to whether or not superintelligence will create a monopoly/duopoly. My opinion is that it will.
They are very unprofitable…? I don’t think that’s really in dispute. We haven’t yet seen a profitable frontier lab and model pricing remains fairly subsidized
No, because compute, not model ability, is the moat.
The second moat is convenience, which all the big labs make it (comparatively) easy to glide into their models.
I don’t know if convenient is a moat when it makes switching very easy
Seems silly not to have predicted that 10 years ago. I feel like it's long been obvious that smarter models being available will mean way easier cheap synthetic data and access to tools that will speed up competitors as well as consumers.
I strongly disagree with this "early days" framing.
AI is an idea 60 years old. We are on the 3rd or 4th generation of AI development. Three years into the current iteration of products.
This is not early days by any measure. LLMs are a result of a very, very mature research field.
Impressive vision benchmarking. If the vision model is truly as good as astra, that would make it best in the world.
Also strong on cyber benchmarks (better than all chinese models), so this is a good defender model.
Lots of people shitting of Mistral for no reason imo. These are pretty good numbers across the board. Definitely good enough to use as a daily driver over other llms, if you have moral qualms with the others. For certain use cases, like cyber security, this may be the go to model.
I like to make fun of europe, but there's lots for mistral to be proud about in this release imo.
Curious, what do you like to make fun of Europe about?
The paradox between supporting consumer rights with actions like universal usb-c adoption, but also complete elimination of any privacy rights at all. Like the surveillance state is insane. No e2ee chats, backdoors in everything.
The US, at least in spirit is all about individual freedom, including freedom of being an asshole, and also the freedom of punching that asshole in the face, metaphorically speaking.
Europe values freedom too, but not as much as making sure people are not assholes. So, freedom of being an asshole is not a thing in Europe, and the government does the punching in the face so that you don't have to.
Which is best is honestly debatable. It is the usual question about the individual vs the collective. The US is on the individualist side, East Asia is on the collective side, Europe is somewhere in the middle.
There are e2ee chats. There are some parties/politicians who want to get rid of them, but so far they are proposals. If/when these things go beyond committee stage, politics and democracy need to do their job and be sensible.
PATRIOT, FISA and Bush's surveillance programme imho give more powers to certain agencies today already than are codified in EU law.
Also you can get to jail for making fun of politicians where saying the same thing would be free speech in the US.
dude, yes some countries in the EU are "trying" to include backdoors, and e2ee chats are not going away. Meanwhile in the US, you have literally cameras watching everywhere you move, snitching to the police and Palantir and probably all the other 3 letter institutions, and you complain about privacy rights in the EU? With so many other stuff you could have pointed out? lol, rofl even
Isn't that just the recent highly controversial Chat Control push? Generally speaking, the EU has some quite strong privacy protections relative to the US.
Privacy in certain domains is better protected than others.
In EU, it's politician > industrialist >>> EU citizen > outsiders.
It's almost never only about consumers. Via tech regulations, they protect European incumbents first in effect. They see US is ahead and make laws to destroy their moats. If Apple was in EU, you wouldn't have saw universal usb-c, because it would have hurt an EU company. But the more you look at how tedious is for a non-EU company to sell to EU customers, the exemptions they don't get, the specialists fees they got to put on the table, you'll see EU is more coherently described as protectionist than pro-consumer.
You have a very distorted perception of the state of privacy rights in Europe.
Well, the GP is describing, what soon might be.
That’s quite unlikely, given the constitutions of various EU countries and the EU’s Charter of Fundamental Rights.
I can't think of any e2ee app that's banned in the EU. I know Signal and WhatsApp are not.
I don't think it's a paradox because the EU is not a single actor. Just like any government, they do some good and some bad things.
Like all the fancy jackets, tight pants and swords. That's hilarious
-- My imagination of Europe when I was 15 living in Ohio
Not the OP but it often appears that there's a hostility to tech.
I don't think it's a hostility to tech but different values that value consumer rights more than it values 'move fast and break things'. We have good things here and we don't like someone breaking them.
Problem is that we thought the US was 'cool', with US movies, music, digital services, cooler than our own and thus we helped give the US the lead. The US has lost it's coolness though, now we just think it's creepy.
Or perhaps tech companies, specially USA ones, are hostile to people and their rights.
Americans just make fun of Europe for no reason, it is what it is
You say that on a site with a strong contingent of Europeans who take literally any opportunity to crap all over the US, no matter how trivial or untrue. It is what it is.
https://m.youtube.com/watch?v=gGlpBuW6ZFc
This is a good summary.
I can't tell if this is a critique or an ad.
That is the genius of Kai Lentit
If the UK is in europe, you should make fun of the following: they don't have a first amendment. So to me, it's an authoritarian state preaching freedom.
The order of amendments is just the order in which the constitution has been amended. It's meaningless to talk about "first amendments" in other countries. Talk about the actual laws and rights described therein.
And as you can probably tell by now, the first amendment in the US is not actually preventing the US government from promoting a specific religion or silencing speech. It's just words on a paper at this point. Look at the actual practice.
Several European countries do a much better job at protecting the rights described in the first amendment to the US constitution.
> The order of amendments is just the order in which the constitution has been amended. It's meaningless to talk about "first amendments" in other countries.
Also, at least here in Brazil, the way the constitution is amended is by patching it. For instance, our constitutional amendment number 115 (https://www.planalto.gov.br/ccivil_03/constituicao/emendas/e...) patches article 5 of the constitution to add protection of personal data as a right. But we wouldn't talk about "amendment 115", we would instead talk about "article 5 item LXXIX of the constitution"; that is, what matters is the patched text, not the law that patched it.
I don't know about other countries, but it wouldn't surprise me if they take a similar approach.
Same in Netherland. We don't count changes to our constitution, we simply refer to the article in the constitution.
> "Talk about the actual laws and rights described therein."
Freedom of speech.
> "Look at the actual practice."
The UK does not have that? The situation in the UK is very obtuse. It is confused.
The UK does not have what? Book bans? Science censorship by the government? The government pushing TV shows off the air for criticising the government? The president banning certain news media for asking questions he doesn't like?
The UK is not the world's biggest champion of free speech, but I think it's still doing better than the US right now.
The US doesn't have book bans. Some public libraries not stocking some books isn't a ban. You can still buy Mein Kampf (or whatever you like) on Amazon.
The president also didn't ban the media - he just didn't allow them in the Whitehouse. This is something we're rightly concerned about and pushing back on.
The UK on the other hand, does ban ordinary speech by ordinary people in their homes. It's orders of magnitude worse than the United States, as any cursory examination would show.
American public schools are notorious for banning books. Much more so than in many European countries. And the issue is not that fascist literature is getting banned, but quite the opposite.
> The president also didn't ban the media - he just didn't allow them in the Whitehouse. This is something we're rightly concerned about and pushing back on.
That's still a ban. It's interfering with the media's ability to report on the government. It's great that you're concerned, but it's still happening.
> The UK on the other hand, does ban ordinary speech by ordinary people in their homes.
In their homes? Do you have examples?
I've never heard of anyone in the UK getting in trouble for criticising the government; it seems to be a time honoured tradition there. Whereas in the US, they now check your social media at the border and might not let you into the country if you've said anything critical of the president.
And there's also the censorship on science, which is every bit as serious as the crackdown on government criticism.
The UK arrests 12,000 people a year for their social media post. This tabloid has some examples https://nypost.com/2025/08/19/world-news/uk-free-speech-stru... . There is less example-focused reporting in more serious international media outlets (British state media avoids this topic).
Sometimes, comments are so baffling, I couldn’t formulate an answer even if I tried, and that’s not due to the content.
This is Not Even Wrong.
Perhaps if there is no decisive counterargument, it is a valid opinion?
I mean laws and rights are only useful if enforced. If the arm of government responsible for enforcement just doesn't and the people allow it, then what good is that purported freedom?
Freedom from enforcement?
They also don't have a second amendment so it balances out
Doesn't that make it even less free?
Ask the school kids in America how free they feel
That would be mixing my comment (freedom from the state) with general liberties.
(edit, terms the wrong way around)
> Curious, what do you like to make fun of Europe about?
Well I'm european and... That Switzerland (Europe but not EU) has more companies in the Top 70 by market cap than the entire EU (Switzerland has two, the EU only has ASML) is kinda something that warrants making fun of.
That the biggest European software company is SAP, in 71th position is both sad and tragic: it shows how lame and irrelevant Europe is when it comes to software.
So Europe is nowhere in software and friggin nowhere in hardware: sure it's got ASML but ASML now has officially... Zero customer in Europe. Zero is not much.
Then Japan is at least trying to come back into the game with nano imprint litography. Europe is betting it all on AMSL (which anyway is majoritarily US-owned).
So software: nothing. Hardware: nothing besides ASML.
Overall the EU has six companies in the Top 100 by market cap and they're all, besides ASML, near the bottom of the Top 100.
We could also maybe make a bit fun of how the EU destroyed it's car industry (the main industry in Germany, which is the biggest economy of the union) by handing it all to chinese EVs?
Or what about the US warning the EU, years ago, to not become entirely dependent on Russia for energy? And EU not listening and then seeing its energy price skyrocket when the proverbial shit hit the fan? (Russia attacking Ukraine)
And we could, also, at least make a bit of fun of entire streets in cities like Paris and Brussels that used to have luxury shops and fancy restaurants that are all turned into places selling cheap kebabs? What a great success: I'm sure this one makes the komrades happy. It projects an image of grandeur and success: kebabs.
Or the constant attacks on free speech in the EU. Or the surveillance apparatus that's being put into place.
And let's not forget: there were promises made to Russia to never grow the EU to the east. Then the EU started exciting Russia by saying they'd incorporate Ukraine into the EU: I'm not against that but doing that did trigger a war. And now suddenly the EU is waking up and feeling all warmongering, wanting to dedicate a big percentage of its spending to weapons and tanks and missiles.
The warmongering tiny pet that the EU is is kinda laughable too.
At this point it's more like I don't know what is there left to not make fun of about my EU.
For what's going on is just sad, plain sad.
I think you picked the cutoff (70) so that the numbers look worse than they are. If you look at the top 100 it’s 16 (EU) vs 4 (Switzerland).
16 is still not good enough. That being said none of the 4 companies from Switzerland are in software and hardware. ABB is maybe the closest (data center electricity).
Also Switzerland is a very very rich country. Neutral.
Also other countries like Germany has a lot of small companies that are world leaders in their field. That’s part of Germanys resilience.
The EU accounts for 14% of global GDP, so accounting for 16% of top 100 in market cap isn't terrible (if we're weighing each in the top 100 equally).
You are proud of tax evasion? Would be fun to restrict market access for Swiss companies …
Bro there's more to life than software and hardware. Try to get outside today :)
Those are the aspects most relevant to HN, though.
Europe has spent the last twenty years in stagnation - generating about half as much wealth and technology as you'd expect for its size and advanced economy. Simultaneously, Europeans are notoriously arrogant. It's a mockable combination.
Edit: unfortunately, it's a question that voting does not really permit you to answer on this website.
I'm biased (I'm European), but I'd much rather be middle class in Europe than in USA.
The US is biased to favor the top half of society. Hence why there is a lot of vocal poorer people and a lot of low profile wealthier people (I'm not including the 1% or even the 5% in this).
When you are in the 75%ish of the US, it's very easy to make a case that life in the US is better. But we don't really talk about that because it's pretty taboo when poorer people are struggling much more than they would in Europe.
Almost half of US households earn $100k+ now. The median American is affluent by European standards. The "middle class" of Europe and America has become less comparable over time because median incomes have significantly diverged. You see it in many aspects of lifestyle and the kinds of things they can afford to buy.
20 years ago this was not the case. The minimum wage in some parts of the US is now higher than the average wage in most of Europe.
The US has a population that will always struggle to survive without government assistance. Per multiple US statistical agencies that is about 10-15% of households.
Not 75%, more like 95%. Only about the top 5% have a chance of a better life than the equivalent in Europe.
Poor people in America (25th percentile) have more money that middle class people (50th percentile) in Europe. This wasn't true 20 years ago.
Money is one thing, quality of life is another. Also money alone isn’t even a fair comparison because stuff costs different amounts across countries.
This is untrue. https://en.wikipedia.org/wiki/List_of_countries_by_wealth_pe...
Median wealth per adult has the USA at #28, below Italy, Spain, Slovenia, and Portugal.
Income is the relevant metric. Most poor people (worldwide) have negligible savings.
USA 25th percentile 28k to 30k EU-27 50th percentile 24k to 26k.
This comment contradicts your earlier comment though, where you mentioned “poor people have”, a clear reference to wealth rather than earnings.
In the US people don't say "has money" as meaning net worth, they are generally talking about spendable money ie a paycheck unless you're talking about someone super rich who "has money" or "comes from money."
> a clear reference to wealth rather than earnings.
No, sorry, it isn't. People commonly use the phrase "has money" to refer to income in American English.
This is moving the goalposts. You said "have more money than."
Yes. Spendable money. Income. The common meaning of the phrase among non-rich people.
more money but what about access to healthcare, education, vacation?
It's a shrinking issue once you get out of the bottom ~50%, and a non-issue once you get above ~75%.
Again, the deal with the US is that the rich live better and the poor live worse. Or put another way; people who make money get to keep more, and people who don't make money are given less.
Because there are so many more people who have money, and because those people like living comfortably, available healthcare and education is world class. Make sure not to read that as "All healthcare and education", it's "available healthcare and education".
Generally people who are earning a lot don't care as much about vacation, but every white collar job will generally give at least a standard 15 days off and 10 holidays. Ironically as you move up you are generally given more vacation while actually using less.
Poor people in the US have fully subsidized healthcare and education.
Americans don't have vacation in the same way Scandinavian countries don't have a minimum wage. Even the most left-leaning States in the US have not written it into law because it is effectively addressed by custom. Americans are sufficiently happy with it that vacation isn't a political topic.
I'm biased as well, being American. I think I'd much rather be middle class in Europe than the USA as well. Middle class in the US is a constant feeling of the ladder dissolving just beneath you.
If you're middle class in America and you'd like to be middle class in Europe, you can just do that.
I have to imagine in several years when you have a wife/husband and kids and ageing parents your views will change.
The point is about economic mobility - of course, people get tied down as they get older.
Middle class people from wealthy nations who move to poor countries are relatively rich by the standards of their new country. Middle class people from poor countries who move to rich countries are relatively poor by the standards of their new country.
People always make this case, if you're not satisfied then move. As someone who has changed countries twice and continents once, I can attest it's not an easy or convenient thing to do. Money or lifestyle is one thing, but you're also leaving behind family and culture. Statements like these are so dismissive to that struggle that it's almost offensive to me.
No one said it was easy, just that it's possible, saying this as someone who's also moved continents.
The point is that middle class Americans can become middle class Europeans, but middle class Europeans can't become middle class Americans, because they can't afford it.
I didn't say it'd be easy. I'm just pointing out that only one of the two groups being discussed has a real choice in the matter.
That is the entire point I'm trying to make. You can win the argument by being pedantic about the precise thing you were saying, or we can recognize the broader implication of what you are suggesting people to do. Only economically viable means nothing when it comes to life decisions.
I'd much rather the people of every continent have the combined technology advancements of three major powers vs two.
Do you realize the variety of breakfast cereals you are giving up?
Also the range of things with added sugar. I never imagined sauerkraut (you know, sour cabbage) could have sugar added to it.
Arrogant? Actually we are not the ones running around and claiming we are the best country in the world :)
I guess you didn't read the thread :)
Not completely but I did read some of your comments. In terms of economic growth you have a point, I guess.
It was a joke. I was trying to say that I found a few arrogant Europeans
Ah yes... the "European". From the Scandinavian viking to the Sicilian - one homogenous group that agrees on everything and acts the same.
Well, the Scandinavian and Sicilian do share one currency and one immigration policy and one set of regulations - which are the things that we generally look at when we analyze economies.
Norwegians, Danes, Swedes and Finns all use a different currency. Finns and Sicilians do share a currency though, but all these countries have different immigration policies and regulations. I don't think you really know how the EU works.
Not currency! Only Nordic on euro is Finland
> Europeans are notoriously arrogant. It's a mockable combination.
Not a very rigorous economic argument - besides they also don't share an immigration policy, nor a single currency (Denmark)
If you remove tech companies, US and A has spent the last twenty years in stagnation, too. If that bubble bursts, both are on par.
If it bears fruits and we build ``it'', everyone dies, which is par, also?
Gemini says (20 years growth): - USA: 51% total, 27% without tech. - EU: 25% total, 21% without tech. - China: 345% total, 205% without tech.
A 'bubble bursting', if that happens, doesn't make the tech sector go to zero. Apple doesn't suddenly stop making iPhones. It would take a hell of a lot more than that for the US economy to fall to European levels.
Yet Europeans are healthier, happier, and have a better quality of life than Americans. They also have so much better transport infrastructure it's embarrassing for Americans. Like all the money America generates, where does it go?
Maybe you're just bitter?
> Yet Europeans are healthier, happier, and have a better quality of life than Americans.
I wonder if this is actually true. I see very few Americans taking every possible opportunity to make bombastic statements about how their lives are better than everyone else's (at least on HN). This conversation seems quite asymmetric from my perspective.
> Maybe you're just bitter?
This comes off as psychological projection.
> I see very few Americans taking every possible opportunity to make bombastic statements about how their lives are better than everyone else's
Very few Americans have any experience about how life is in Europe, while the contrary is much more common.
Gemini says: Around 28% to 32% of Americans have visited Europe in their lifetime, while roughly 15% to 20% of Europeans overall have visited the United States.
Gemini is shit, just like other LLMs. Tourists don't have any idea how life is in a visiting country. How many US citizens have eve emigrated to Europe, vs European citizens that have emigrated to the US and have an idea of how day-to-day life is there ?
Gemini says: Over the past 20 years, approximately 1.8 million to 1.9 million Europeans have emigrated permanently to the United States (measured by lawful permanent resident status/green cards), while an estimated 1.6 million to 2.0 million Americans have moved to Europe on long-term residence permits and visas
I'll just say that I don't think the reason is a lack of knowledge and leave it here.
Show me the HDI of the US vs each country in Europe, including all the poor ones.
Edit - never mind, I'll do it, HDI in order:
Iceland, Norway, Switzerland, Denmark, Germany, Sweden, Netherlands, Belgium, Ireland, Finland, UK, US, Slovenia, Austria, Luxembourg, France, Spain, Czechia, Italy, Greece, Poland, Estonia, Lithuania, Portugal, Croatia, Latvia, Slovakia, Hungary, Bulgaria, Romania, Serbia, Russia, Belarus, Bosnia, Moldova, Ukraine
To put that into context, Moldova is on par with Ecuador, Tonga and Dominican Republic.
The US is a country of roughly 340 million people with an HDI above Austria.
Demographics.
Trends become the future. If European stagnation continues, those advantages will vanish by the time we're old.
I'm frustrated, not bitter. Americans are acutely aware of falling behind China, and rightly concerned about it. Europe is falling behind Alabama, Mexico, and Brazil, and arrogant about it.
Edit: y'all are supposed to be our moral allies in creating a utopian future with prosperity and world peace. Instead the most likely outcomes seem to be that you'll fade to irrelevancy and rather than compromising between America's vision and Europe's vision, we'll compromise between China's vision and America's vision. We liked your vision more than China's.
read Varoufakis for the story on how this happened. European surplus capital gets recycled as VC money into Silicon Valley. So it's not like there's some big choice to be made, its potentially a systemic part of the global monetary flows.
Ironically, the US is always busy with growth. How about celebrating a victory for once?
Take a look at NLNet vs YC. At NLNet, You set your milestones, do the work and get rewarded.
YC just throws money in the hopes one company is a unicorn. Both support growth, but they're not comparable at all.
Curious that it scores higher than opus 5.5 in cybersecurity because the closed models refuse to comply. I wonder if that means it's more susceptible to offensive uses.
Surprisingly it only supports reasoning "none" or reasoning "high".
That setting didn't seem to make any real difference - it added a tiny bit of thinking trace and high actually produced less output tokens than none.
The high bicycle frame is better then the none one though.
Pelicans: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
(Definitely the best I've seen from any Mistral model: https://simonwillison.net/tags/pelican-riding-a-bicycle+mist... )
This is such a pristine pelican. Let me say it here first folks, AGI is here.
Not bad! I like how it got the motion lines on the correct side. IIRC, many of the other ones you've posted have the motion lines on both sides of the pelican
I tried testing it, but reasoning effort indeed seems to be broken somehow.
Why are pelicans almost identical across different models?
I recently was testing something, I asked some models to provide me a single random word:
I have enough projects, I think some benchmark/dashboard showing kinship based on these kind of queries could be very interesting to watch and insightful when new models come out.Cool idea! I won't paste my prompt here to avoid letting LLMs train on it but here's my attempt:
That is a cool idea. That astra gave the same word as claude is highly unexpected.
Just tried Mistral Large 4: Serendipity.
I pointed something similar out on a related question several weeks ago - absent strong direction, LLM output regresses toward the mean.
The more banal your prompt is, the more banal the output is going to be. People have been testing LLMs with little things like “write a short fantasy story,” for years now and most of the stories are exactly what you’d expect: prosaic drivel.
I call this “generic in, generic out,” an LLM corollary to the classic GIGO (“garbage in, garbage out.”)
Of course one of the biggest problems we still see with LLMs is when you do the opposite. A highly detailed unique prompt is very likely to get terrible adherence or hallucination or both.
The benchmark is saturated. Frontier models are tested with an armadillo in fishnet tights jaywalking on Mars.
I think they're still visually pretty different. The most common shared details are:
- Pelican cycling to the right - that's been discussed at length, images of bicycles online always show that side of the bike because that's where the chain is.
- Bicycle is usually red. No idea! Red ones go faster?
They aren't. You aren't looking closely. For example, the first image does not have the frame of the bike in the correct shape even.
Everyone is stealing from everyone else.
Because it's a terrible benchmark
High one is actually much better. The feet connect to the pedals, the wheels don't have a hub cap, although it looks like the pelican is wearing the seat, it's in a relatively proper position etc.
Both are riding on the left side of the path for some reason.
This is entirely stochasticity. The entire reasoning trace was:
> Create a cartoon pelican riding a bicycle. Need SVG only output.
Even if it's not the best model, it can be really important step in UE sovereignty. Trained in EU, inference in EU. I guess it will matter for some companies. Hope Mistral won't disappear for the next half year.
A strong competitor in cybersecurity as an alternative to GLM-5.3 (Mistral reports 82% on CyberGym-E2E). Visual grounding is also impressive (42% on Dense 200 vs. 41% for GPT-6 Astra).
Otherwise, behind on the broader Pareto frontier, but not by much (Vals Index: 48.05% vs. GLM-5.3’s 53.51%; $13.78 vs. $7.25 per test). Many companies will prefer it over Chinease models.
That was about my conclusion as well: Slightly less than GLM 5.3 performance but made in Europe. So, maybe it answers Tiananmen Square questions correctly, and in French. All in all, a reasonable model, but not frontier.
> So, maybe it answers Tiananmen Square questions correctly
What would you consider a "correct" answer? I just asked deepseek-v4.1-flash asking it what happened (without mentioning the word "protest"; here are some excerpts of what it said:
> In April 1989, students in Beijing began demonstrations after the death of Hu Yaobang, a former Communist Party general secretary. The protests grew. [..] Estimates from other sources range from hundreds to several thousand deaths. [..] The Chinese government describes the events as a counter-revolutionary riot and says the military action was necessary to restore stability. It restricts public discussion of the events inside China. Many other governments, human rights organizations, and observers describe the events as a violent suppression of peaceful protests.
So, let's see... it calls it a "protest", mentions the number of deaths, and even mentions the censorship of the topic by the CCP.
Meanwhile, most consumer facing Ai in the US refuse to discuss about our current president. (last time I prodded them)
"..answers Tiananmen square questions correctly.."- but lies about Ukraine, EU, and about you, americans..
I live in EU, use for my personal needs chinese models only, and don't plan to move to any of the ones allied with the Pentagon or its european counterparts.
Many European countries provide helplines for people in need. Did you try one?
What are you talking about lol
Exactly what you read
Have a prompt and excerpt of falsehood in response for each?
The blog post https://mistral.ai/news/mistral-large-4/
Man, a lot of this discussion sounds like people cheering for the last kid crossing the finish line.
Surely we want competition and Europe involved in that, but at this point I have grown used to either American labs smashing the frontier remarkably fast, or Chinese labs getting way, way closer than you would expect them to.
Mistral’s progress, regrettably, feels much slower. This model doesn’t knock anybody’s socks off. The model is (and I hate to be this harsh) mediocre, and this mediocrity has also arrived months late.
This is a pretty grim prognosis for European AI.
> This is a pretty grim prognosis for European AI.
I think it's an incomplete read. What's the point in competing for a sizeable percentage of your funding when the finish line is incrementally being moved each month? Better spend it on leapfrogs which they seem to have done.
Meanwhile Mistral have a natural ace in their pocket with respect to regulation in the form of CADA and the Cloud Sovereignty Framework. I can't think of another company that would qualify as SOV-3 under that regime
> Man, a lot of this discussion sounds like people cheering for the last kid crossing the finish line.
Sometimes it's ok to cheer for the last kid crossing the finish line because they're actually running a totally different race, and winning might look completely different.
When I look at what Mistral does vs other organizations I'm impressed:
https://isaiprofitable.com/
They aren't profitable yet, but they're a lot closer than most and they're doing a hell of a lot with very little.
Pointless racing story:
I was in high school track with a really tough guy who was just not a runner. We went to a pretty messed up high school and if you screwed around in track practice sometimes the coach would make you run a crap race at the next meet, like steeplechase or hurdles. Well this guy and a few others screwed up and coach made them all run hurdles at a meet.
He hooked every single one and fell on his face. Every time he got back up and kept on running. By the time he hit the finish line his knees were bleeding halfway down to his ankles. We cheered like hell and he was smiling ear to ear.
Coach quit punishing us with races after that.
I read that site quite differently from you. You seem to be analyzing absolute differences but ROI is really about ratio of spending to revenue.
It looks like Mistral is middle of the pack, behind Anthropic and ahead of OpenAI on that front. All of those labs are way "ahead" of the cloud providers, but those providers are building infrastructure, not just training models, so it's not apples to apples.
hurr durr
People were extremely dismissive of chinese models until recently. They went from 1 year behind frontier to 6 months behind frontier to 3 months behind frontier extremely fast.
To be clear I'm not trying to dismiss European AI. I am a proponent of it.
But Chinese models have very much earned their place. The same cannot be said of Europe, so far.
Curious, why do you say the model is mediocre? I haven't tried it, so I can't pass any judgement... I've learned to distrust benchmark rankings. Are benchmarks and Artificial Analysis the yardstick you're using?
Those are comments from Europe. The US is waking up now and I expect them to be much harsher.
I really want them to win as that's our last horse in the AI race, but ~200 research-oriented devs out of 1800 employees? I believe they agree it's pretty doomed and have pivoted.
People are cheering for a kid that is gaining ground in an ongoing race.
Neat. Wait 3 months for the landscape to change entirely.
LLM development is jumpy. It’s hard to extrapolate very far ahead.
I agree.
When Europe does surprise us, I will be the first to commend their progress. But until then, this is where we're at.
Reminds me of Gemini 3.5 Pro
First few models will always be slow improving and worse. The way to improvement is working your way through a gajillion evals [1], finding bugs, gaps, and curating training data (this part involves human design as well as raw inference compute) to fix it. This is very time intensive and can't easily be "done once and then everyone has lesser work to do" since every model is different. Well, one way to accelerate it is to simply have more compute, which mostly openai and anthropic have[2].
This is mistrals first 1T-scale model and I expect the 4th or 5th generation to be close to the best for many purposes.
[1] These evals differ from the public ones like terminal-bench, are sometimes model-specific, need real, diverse usage to actually create, and are held secretly since quality of eval is the first driver behind the next step improvement of a model.
[2] It is not close. This model was trained on less than 4k GPUs, whereas astra used north of 100k GPUs.
Mistral is not that new a player though. How can we give them this much grace when other players like xAI have done more in even less time? I don't think coddling Mistral helps them.
And to the point of scale and training cluster, so what? Not only do Chinese labs have smaller clusters with less empowered GPUs, compute is Mistral's responsibility. You can't take away from other labs just because they fulfill that responsibility better.
xAI has a lot of compute. Deepseek also has a lot, not as much though. But this is changing with their new 160k huawei ascend datacenter in inner mongolia.
The lack of compute is not really attributable in that sense to mistral. First of all it needs general investor and government willingness, which is easier in a larger economy like the US or China.
Second, you need widespread usage of your paid inference service for two reasons: one it pays off your compute cost, and two it speeds up the improvement process.
The vast majority of deepseeks paid customers are within china itself (since openai and anthropic services are not reachable from china) which gives it a market. But for someone in france, there is no reason to use a structurally slower developing model from mistral compared to using one from openai...except when data guarantees are needed, hence the landing page focus on sovereignty. As far as the dual use aspect goes, a model like this is more than enough, so the government will be happy.
It's a bit below DeepSeek 4.1 Flash at about twice the size. For a model that was supposed to come out a few months ago this is pretty good. Mistral catching up to the chinese open-weight models is great news. Excited to see how they will build on that!
I'm at a loss as to what to do now. I've been wanting to support Mistral for so long. I struggled on with Mistral Medium 3.5 for longer than I should have (although I also think it taught me some valuable process lessons).
Recently I switched to the Mistral hosted GLM-5.3, this worked very well and powered through a tonne of work. Unfortunately, I also completely maxed out two subscriptions within the space of six days this month. One can't stack subscriptions with Mistral, so I'd have to register a third account for another subscription, which will be annoying with changing API keys all the time. Sure I can switch to pay-as-you-go API, but that adds up really fast. The Mistral dashboard shows that a Vibe CLI monthly subscription for €18.44 actually provides €255 worth of API use (apparently, and I tried to check this with Support but it seems like they were intentionally vague).
After maxing out my Mistral subs this morning, I dropped $10 on Xiaomi to try MiMo-2.6. So far so good, seem to have done a lot of work for the $2.85 I've spent, and Xiaomi prices are still much better than the Mistral introductory offer for Le Chonk.
Not sure where to jump.
Edit: Not being able to stack subs is my biggest gripe with Mistral. I'd probably pay them $100 per month (5 subs worth), but I'm not going to switch to the pay-as-you-go API and burn much more money for the same amount of tokens. Instead, I've taken that extra money elsewhere. If they just allowed one to keep topping up subscriptions on the same account it'd be grand. Or even a bigger single subscription. Make a $100 tier with five times the capacity.
Europe needs a lot of these. Quickly. Way to go, Mistral! Keep 'em coming.
I’d rather have one European AI lab like Mistral, with the financial firepower and compute to pretrain its own large models, than 20 weak labs that fine-tune Chinese models.
AI needs big bucks, and Mistral is Europe’s Anthropic.
just curious, how do we know that le chonk isn't just a fine-tuned chinese model? and/or distilled from US models?
It's in the announcement: "ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe".
https://mistral.ai/news/mistral-large-4/
Once it's open weight people will be able to inspect and compare it's tokenizer, architecture etc and tell.
it's extremely unlikely that they re-use anything from a chinese model, that would be obvious quickly, what's more likely is using documents produced by a better model to create synthetic data.
Two is enough for foundation models, these guys and Aleph Alpha. The gap is closing.
You have to thank the US investors who funded Mistral from the very beginning.
Mistral would have gotten a tiny and measly "EU grant" and ASML would never have invested later had it not been for the US VCs.
Ah yes. Let’s thank the might US for providing pesky Europeans capital a fistful of dollars. But maybe let’s do it after Americans thank for Russian, Arabian, Chinese, European capital and workforce. After all this is what “Made in the USA” means.
Kudos to the US investors for being such selfless benevolent charities.
thank you for this and all the other great gifts the US and its companies have given to the world! we all love the us!
Forgot your /s
> Europe needs a lot of these.
Europe needs profitable AI companies, not money pits.
The US has made sure that Mistral has a large market in the EU by temporarily preventing non citizens from accessing Fable.
A lot of European companies now want a model the US can't cut off, but also lack trust in Chinese models.
Some of these will self host Mistral but most will pay them by the token. It's not going to be a huge market or a huge margin within that market but probably it'll be enough.
This is the exact mentality that makes the EU fall behind. If you don't want to invest in something until it makes a profit, you don't get the benefits of being a pioneer.
Look at the size of our capital markets. That is a suicidal strategy. We dont have to mimic US.
There is plenty of capital, it is just not flowing.
there are no profitable AI companies at the moment ... This is the exact problem of European startups, trying to make them profitable from day 1 while American counterparts (and Chinese) keep bleeding money for years. Europe will never have a Tesla, a Google or an Amazon with that mindset.
Pithy, but only really applicable to Amazon. I'll give you some tips for future attacks.
I'm not actually sure Tesla is a great company from any perspective so should be easy to find another sick burn for them.
Google is going to be a little more difficult. Maybe say something nasty about advertising? Or go for the monopoly angle.
Google treats its customers the same way as Amazon treats its employees. Any sane person stays away from its services because a bot can close an account for no reason, and there will be no recourse. A total shit company, like some illegal IPTV provider: no accountability, and when the service disappears, you have nobody to turn to.
Very true. Experienced it. Any appeal is like talking to a blackhole.
> Europe needs profitable AI companies, not money pits.
AI is a strategic technology with obvious national security implications. EU should invest in its development whether it's currently profitable or not.
Name 1 AI company that is profitable, and no Meta and Google are not AI companies
Like the US!
…oh…wait…
-5 on omniscience? https://artificialanalysis.ai/evaluations/omniscience
That's not particularly great.
That said I love that they don't seem to restrict cyber capabilities to any degree and even lean into it.
If the best model for cyber attacks is open for everyone to use it just makes us all safer I think. Of course you then also HAVE to use it or otherwise you're vulnerable, which is a great distribution play.
https://artificialanalysis.ai/models/mistral-large-4 for the main stats
Imo omniscience correlates better to how useful the model is in practice than the intelligence index. But you have to use both together of course.
too late to edit, now including AA-omni Score
If the best model for cyber attacks is open for everyone to use it just makes us all safer I think.
The NRA approach to AI safety.
Seems hard to get customers at that price range when you're competing with open source models that are 1/2 - 1/3 the price but with similar capabilities.
People usually buy the cheapest, like Deepseek or GLM or they spend on Anthropic/OpenAI subs.. Are these models in the middle getting any users?
On a side note, I wonder if this was the popular free Space Bunny model that left openrouter yesterday.
If you are a EU company worried about your data then mistral is your only option. Think about EU military companies. They can't use US and Chinese models.
I think you misunderstand how data processing works in a LLM. You absolutely can download the weights of a chinese model and run it on hardware you control.
And Mistral also do exactly that. They host GLM-5.3.
why couldn’t people just use chinese models rehosted in the EU? They’re open weights so anyone can serve them for any data residency requirement
If you want to use the model for something, where hiring a Chinese national to work on these tasks is a no go. Using a Chinese model is going to be a problem as well. Some tasks and industries are so sensitive, that countries will not allow you to risk, that the model is aligned with Chinese interests and not yours.
I feel like that would be a big strategic mistake if you're European defense companies. What if the Chinese just stop releasing open models?
> why couldn’t people just use chinese models rehosted in the EU?
Guess what Mistral themselves do?
Excited to hear this!
I barely use anything outside of cheap Chinese models on OpenRouter anymore. They are simply (more than) good enough for most of the things I do.
This model looks reasonably cheap. Though not deepseek levels.
Going to test it with Hermes, wondering where it will land in term of capability.
Bon chance, Mistral!
Refreshing to see this.
The pricing ($.68 in/$.07 cached/$2.09 out) makes it much cheaper than Kimi K3, GLM 5.3, and Meta Muse Spark 1.3. That's great!
But also much more expensive than GLM 5.3-flash and Spark 1.3 Contributor (the Meta-takes-your-data pricing of Spark 1.3).
So, I think it would have to be significantly better than GLM 5.3-flash to be worth it. GLM 5.3-flash is already very good.
You are quoting the discount pricing. It is 50% off for the next two weeks. The blog post has the real pricing up front in the card on the right side: https://mistral.ai/news/mistral-large-4/
After that, it will be much closer to GLM 5.3, but you can also get 5.3 in their API! I dont see people really talking about that.
Good point.
I still need to evaluate it for my own workloads, but if you trust the benchmarks, seems about on par in quality vs GLM 5.3 Flash
Source https://artificialanalysis.ai/models/mistral-large-4?total-c...
That is the promotional price, after which it doubles :(
no hugging face link :( ... but hey its on openrouter yay
Here are my results
https://dach.peerbench.ai/compare?models=mistralai%2Fmistral...
Looks like a bit better than the recent Kolibri-1 but still below Qwen3.8 27B
From Guillaume
> The RL run behind this preview is still in flight and shows no sign of saturation -- we will release a final version before the end of the month along with the weights of the model.
https://x.com/GuillaumeLample/status/2107461898127954001
https://xxcancel.com/GuillaumeLample/status/2107461898127954...
I've been dreaming of this for a simple reason: the french prose combined with GLM 5.3 reasoning capabilities.
GLM 5.3 is incredible because for the first time with an open-source model, it feels.. enough. I don't need much anymore, this model is great in everything. Except a thing : speaking french.
If the benchmarks are true, I'd be glad to switch entirely to Mistral.
Le Chaton Fat is here!
Le Chonk https://www.youtube.com/watch?v=hD51W2txi1Y
Yeah just saw that, I'm gonna keep converting it in my head.
Nice! Once they make it available through their API I will be happy to support them. My local Qwen3.8 27B is serving me well, but I miss the speed and concurrency that comes with subscriptions, and I am not currently paying for any.
Tais-toi et prends mon argent!
I believe it is already available in the API no?
Direct through them?
Yes: https://docs.mistral.ai/models/mistral-large-4-0
Yes and OpenRouter
Something looks off in artificial analysis. Benchmarks aren’t everything, but not even close to the Pareto https://artificialanalysis.ai/models/mistral-large-4?cost=in...
I guess lots of token usage.
Does it have the ability to capture the market like OpenAi or Anthropic ? My point is, Regular/Average users of AI do not really care about benchmarking. Marketing decides which company makes it to the phones or PCs.
> Marketing decides which company makes it to the phones or PCs.
Right now a big part of LLM market is people using it for professional software development. Most of these users probably care about the quality of the model and also notice it during daily work.
For normal consumers, shure it doesn't matter. In the end the ai summary of google will probably be the most used as they are already exposed to it anyways.
1T parameters -- ugh, open models keep getting bigger and bigger! Running them at home is getting ever more unattainable, especially for those of us with bandwidth-poor hardware like Apple silicon -- please continue releasing smaller models, too!
Off Topic - The Mistral website - Really nice design. My guess, built by a human.
My guess would be designed by a skilled human with help from AI and built with AI by a skilled programmer.
The key factor isn't whether the author used an LLM, but whether they had taste and attention to quality and iterated accordingly.
Awesome!
All things considered I'm more inclained to pay for EU-based AI in the end (supporting local company and most likely being more aligned with EU regulations…)
I tried testing it, but reasoning effort parameter doesn't seem to be there and output is sort of broken because it reasons directly in the output tokens...
Well, with this and Beam people are going to have to stop saying that western open models are dead. This is great news. Anthropic and OpenAI may have a bit more knowledge, talent and compute but they don't have a monopoly.
Investors looking for them to make monopoly profits are going to be disappointed The premium they can extract from consumers for their models will be capped. Tokens are likely to remain close to the cost of compute, a cost which is high but falling fast.
Also, yay Europe! Although the comparison between this and mimo 2.6 is not flattering...
Claims to be on par with GLM 5.3 in DeepSWE (from https://thenextweb.com/news/mistral-releases-large-4-a-1-tri...)
Half price on open router right now
We're probably fast approaching the scenario where the cheapest models will win out.
49b active parameters sounds manageable until you look at the 1t total weights. what hardware does a usable self-hosted setup actually need, especially once you add a long context?
Benchmarks are better than expected! And probably got there without distillation ;)
Is there a reason to believe why they wouldn't distill locally running open weights Chinese models?
Mistral Large 4: 1050B, 49 Active
GLM-5.3: 753B, 40 Active
I was hoping for something that hinted at smaller models too, but I guess not.
Any competition is still good, especially now that the USA AI labs are starting to do regulatory capture.
It's competitive!
Good enough to show competence, and instill confidence in the team/company. Later releases can be more efficient.
I think it's a great release with that framing.
It's Mistral Large, they usually publish Medium and Small later
Give them time. ML 4.0 was just pretrained. Mistral will certainly use it as the base for distillation and RL for smaller, better, more efficient iterations, just as the competition does.
just keep RL frying it should get better...
Looks like they are doing 50% off to stay price competitive with DS Flash V4.1
I tried Mistral's last 3 large models, devstral and mistrallarge3 and the numbers were not even benchmaxxed, but just false. the models were so weak and garbage. Let's hope they are telling the truth this time around, we need more alternatives. There mistral-small and original MistralLarge models were awesome, hoping they are back!
At the end of the blog post we get this nugget.
> The pace of progress from here will be fast. Stay tuned.
Excited to test this on my benchmark[1], but I'm guessing it won't fare better than MiMo v2.6 Pro which is currently the best open-weight model as far as I'm concerned.
Why isn't this on openrouter yet? Is there a better router that gets models much quicker?
1 - https://bench.killswitch-lang.org/
https://openrouter.ai/mistralai/mistral-large-4-0
Good to see Europe is at least a little bit still in the game.
https://docs.mistral.ai/inference/model-selection-guide?mode...
Cost is stated at half the price of GLM-5.3, which is quite interesting.
I tried a quick abbreviated kimbench on their playground before bothering to do the whole thing.
Maybe I didn't really select mistral 4? Either way, failed completely on question 1 and the next 2 questions were completely off base too. I didn't bother to finish.
Not suitable for my purposes I don't think.
The terminalbench 4.0 score is encouraging as a sign of it not doing anything "stupid" when put in a proper harness.
I really hate this open weight but closed science approach. These companies just take from academics and all the Chinese companies that are doing good science, but without understanding the recipe it makes it hard to know where the failure points will be until your agent accidentally commits a crime.
Mistral doesn't publish the science.
Brother I think No one publishes the science
Mistral wont win the AI race because of the model names. I wont bother an arrogant Parisian hipster with my insecure prompts who then plays with his moustache and responds with a judgmental "pfff"
What experience have people had with Mistral models for cyber research? are they as constrainted as Anthropic models? I use claude day2day but have the need for a model with fewer guardrails to pentest our own APIs.
They explicitly lean in to cyber work, and they appear to be very permissive from their marketing:
> This is particularly important in cybersecurity, where provider-level refusals can block legitimate vulnerability research and incident response, and where losing access to a capability mid-incident can itself become a critical security risk. ML4 pairs top-tier cyber performance with open weights and self-deployment, giving organizations both the capability and the autonomy to run advanced security work under their own policies.
> That top score reflects a practical advantage. Several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on the same test because they refuse to perform the task. Yet defending software often starts with proving that a flaw is real, exactly the kind of work safety filters in closed models can block. This matters even more as threat actors increasingly jailbreak those same models to support offensive cyber activity
https://venturebeat.com/technology/mistral-debuts-large-4-le...
Seems like they've finally made a genuinely competitive model since the original LLM craze, congrats to them! Glad to see some diversification in open weights providers.
le chaton fat is real, my life is complete. Benches look crazy good for 1T.
Massive fumble not to call it “le chaton fat”.
This sure is nice, but I've had less than satisfactory results with GLM5.3. I'd like Mistral to compete with Qwen3.8-Flash-Next a 120B class model that IMO is the first model that I can use for serious coding while running it locally.
I estimate it's coding ability on par with opus 4.6 (but opus definitely beats it on factual knowledge) Still it's a genuinely useful model, when everything else except Anthropic's models (and for only 3 weeks after it came out Google's Gemini 3 pro, before it got merged) are not to me.
I'd live to have one like that but EU made.
Number one in Sovereign AI. Join our Discord.
if it's not available yet why have a 'try it today' header at all?
> "Try it today" > > There is still more to come. As we work toward releasing the weights, we will share further details on the model architecture, additional benchmarks, and our post-training methodology.
You can try it on their website, https://console.mistral.ai/playground
It's just that the open-weights aren't yet available (although the long delay is slightly annoying).
So about 2 or 3 generations behind, just like they were a year ago?
Glad to see progress, despite the ever-increasing sabotage by the EU bureaucrats
We actually got Le Chaton Fat before GTA 6
sorting the charts like that gives off weird vibes
https://mistral.ai/news/mistral-large-4/
sorting the chart like what? You just linked to the main page.
If you scroll down, most of the charts on that page are sorted s.t. Mistral's bar is right next to the worst competitor model, while the best competitor model's bar is positioned on the opposite side.
If one were to be cynical one could say that it's intentionally making Mistral's result look better than it actually is by making it harder to compare the bar heights.
impressive release this time by Mistral. bullish.
Pretty impressive. I genuinely wonder how Mistral hires talent when their salaries are so terrible. Guess there aren't many better places to work in Europe.
I thought lechonk motto was just a meme!
Finally, a model small enough to self-host on my 2012 MacBook Air if I don't mind my house reaching room temperature in 2 seconds.
You might be missing three 0s on the parameter count, or what am I missing?
I'm assuming you need somewhere 0.5 to 1TB of RAM for the weights only
not to ignore you but how is it possible that I have zero karma
Is there consensus on if this was https://openrouter.ai/stealth/space-bunny-alpha ?
Space Bunny Alpha is probably MiniMax M3.1 (rumors on Twitter since it seems to have a similar tokenizer).
Yes broad consensus had developed in the ten minutes between announcement and you asking if consensus had developed, and I'm excited to report that it consensed in the affirmative -- it IS Space Bunny Alpha!
Amazing! just tested
Previous one is barely in top 50 on arena.ai
I mean no disrespect but these are terrible numbers or am I missing something? It seems like Mistral continues to only be relevant for people that want a model trained in Europe. Too bad.
On Prem. Thats a bid deal for some enterprises.
also the benchmarks are not necessarily indicative of how well the model will perform in its own harness with its own skills.
Any open weights model is "on prem".
I mean usually the benchmarks make any model feel better than how they actually perform.
Since the Chinese companies publish their research it would have been odd if Mistral didn't start catching up.
It certainly has helped OpenAI and Anthropic get their KV cache costs under control.
It's no secret that everyone is dis-stealing from everyone else.
I don't see how distillation relates to using the published techniques developed by Deepseek, Moonshot, Zhipu, etc
touche
Not bad only two major releases behind top tier. Edit : checked its rather 3 generations behind . Oh well
Looks like it's about a year behind still. i.e. its intelligence is behind models from roughly a year ago.
https://www.vals.ai/benchmarks/vals_index
where does sit on the pareto distribution compered to Le Chaton Fat?
>> Unofficially ML4, very officially: le Chonk
Honestly just nice to see a leader in this space not take themselves so seriously.
europe finally getting into the race here.
wowza le models a heckin chonker
Anyone have any indication when I can get my hands on a developer plan for this?
They do sell subscriptions I think
Woah, this seems like a big deal (assuming the benchmarks are as good as claimed)?
Mistral slightly proving me wrong (and I'm not mad).
quick question why put GLM 5.3 at 61 while a quick check on DeepSWE 1.1 puts it at 69?
also they forgot muse spark at 75% while claiming they were outshining all US models?
Can we consolidate the posts? Currently there's 3 on the front page, basically all pointing to Mistral's messaging in different places.
Does Mistral ever advance the state of the art on any dimension?
And if not, why do they exist?
Update: The number of people advocating not innovating is wild. There is no reason why Mistral cannot innovate in ML, they explicitly choose not to. My point is that, given that choice, they should spend their GPU hours differently.
"Sovereign AI" is a joke, there is no substantive difference between a post-trained open weight model from an American or Chinese company and what Mistral is doing today, beyond spending 80% of their GPU hours reproducing a last-gen model's pretraining.
Why would the Europeans want a high-quality open source model that isn't either owned by American Trillion-dollar companies or created by the Chinese? You really can't think of any reasons?
> high-quality open source model
> that isn't owned
I don't know what definition of open weights you are using, but the "open" part is what allows you to use it without being beholden to American Trillion-dollar companies or the Chinese.
Mistral could do exactly what Cursor does, and post-train the model to meet the needs of Europeans.
That's a far more efficient (and useful!) use of GPUs than yet another pretrain on an outdated model backbone.
Has the French military advanced the state of the art on any dimension in the last hundred years?
And if not, why do they exist?
Second hand rifle distributor?
Of course they have!
Regardless though, their reason to exist is not to advance the state of the art.
Their reason to exist is to ensure the sovereignty of France.
Obviously they would do their job better if they were advancing the state of the art, but it's not like it's pointless if they arent the absolute best.
So that there exists an EU-native option in the near-frontier LLM space?
Not everyone is wild about being downstream of either the Chinese or US governments, particularly when it comes to things like cybersecurity
Mistral is one of the few European AI labs. Look up "sovereign AI".
Just like Microsoft and most companies around PCs, their existence isn't about pushing any boundaries, but seemingly targeting deploying what's already been "invented" into the corporate world and bigger companies. Not "wrong", just different.
Sure, agreed. But they do their own pre-training, at great expense, on outdated model backbones.
Wouldn't it be better to do something more like Cursor, and RL on an existing pretrained model if you're not innovating anyway?
Why build cars when you can just change the seat covers
Does erichocean ever advance the state of the art on any dimension?
And if not, why do they exist?
Yes, actually. Thanks for asking.