Sharing AI progress in mathematics

(openai.com)

919 points | by OfficialTurkey 12 hours ago ago

832 comments

  • jboggan 7 hours ago ago

    I was a graph theory junkie long ago and even moved to Budapest for awhile to study among the greats. While I was there I started working on Barnette's Conjecture which came to occupy my thoughts over the next 24 years of my life, on and off as I worked in many different fields. Last summer I even thought for a few days that I had actually solved it.

    But it's supposedly proven here - problem 180. I don't know what to think exactly. I spent thousands of hours on that problem. I really enjoyed it. Hearing that it is solved somehow makes me sad in a far-off way, like hearing an ex-girlfriend died suddenly in a car crash. I don't know, there's probably a lot of people feeling odd emotions tonight.

    There's no Lean proof for this one so I'm digesting the paper. On the surface it looks like an approach I considered 24 years ago and abandoned.

    I revisited the problem this summer, along with my partial solutions, when the previous round of stunning proofs came out. Several hours of work with Fable simply convinced me it wasn't yet solvable and reinforced how hard of a problem it was.

    • nilkn 5 hours ago ago

      > Several hours of work with Fable simply convinced me it wasn't yet solvable and reinforced how hard of a problem it was.

      This is the part that gives me the strangest feeling about it all, because you're not the only one with this experience. I've experienced this too on different problems, as have many researchers across many fields.

      I disagree with the Fields Medalists on the majority of their complaints. AI math is happening and there's no going back. However, on one point I increasingly agree: virtually none of this stuff is possible with technology any normal citizen has access to. I have no problem with AI models making revolutionary advances in math or science. Where I start to have a problem is when the AI models making these advances are tightly withheld, proprietary, and seemingly never released with these capabilities intact. This has been the case for all of 2026 so far.

      I suspect that this is in fact the source of much of the angst. None of this progress is reproducible outside of one or two teams inside OpenAI and Anthropic. It's becoming an incredible concentration of power that I don't know that we've ever quite seen before. Right now, it feels harmless because it's being used for wonky math problems that aren't (yet) practical for anything. But great power never stays harmless. History has taught us that countless times, in countless different forms.

      • omnicognate 4 hours ago ago

        > I suspect that this is in fact the source of much of the angst.

        Why do you "suspect" this as if it's some hidden motivation when the very first paragraph of the advisory group's statement (linked from the OpenAI post) says:

        > At present, some frontier AI labs are testing advanced mathematical problems on proprietary models that remain inaccessible to the broader scientific community. Our recommendations are formulated with this practical context in mind. However, ideally, they would not do so. We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models.

        Tao and others in that group have been strongly and publicly pro AI from the start. They are not advocating "going back". They're objecting to the strip mining of open problems using proprietary technology.

        • alberto-m 32 minutes ago ago

          OpenAI: At long last, we have created the Open Problem Strip Miner from classic Terence Tao tweet “Don't Create The Open Problem Strip Miner”.

        • pred_ an hour ago ago

          Regarding the advisory group, OpenAI claims to “have drawn on their advice”, which would include not dumping a bunch of AI slop, with the footnote that if they do do that, at least fund the process of digesting it.

          At the same time, there's a new note at the bottom of agmai.org stating how they've been in contact with OpenAI about this particular release, and they say that “we consider these discussions constructive, it is ultimately up to the mathematical community to assess the extent to which our recommendations were followed successfully”.

          So, what's going on there; is this British English for “they didn't follow anything at all”? Because from my perspective, it looks like they doubled down on the Navier–Stokes approach of trying to maximize PR gain while being as lazy as possible about actually contributing anything back to science, releasing only slop that may or may not be correct and may or may not be straight up plagiarism, as has been the case earlier.

          If I were on the AGMAI board, I'd feel terribly exploited when reading that press release, yet their response is modest.

          Hairer, if you're reading this: is there any indication whatsoever that AGMAI was anything but a cheap way for OpenAI to science-wash their press release?

        • lordgrenville 3 hours ago ago

          Haven't been following this debate closely, but what's the issue with "strip mining open problems"? Surely the supply of interesting mathematical problems is (in theory) infinite?

          • musebox35 2 hours ago ago

            You can find Tao’s arguments here: https://mathstodon.xyz/@tao/117237320796901560

            He argues that the supply nay be very large indeed but the interesting subset is not. Figuring out the interesting problems is difficult so strip mining the good known problems may lead to scarcity. I am not a mathematician myself, can not judge this accurately.

            • tomaskafka 2 hours ago ago

              That’s what we are doing with nature, seas (look up strip mining there, it’s a horrible practice), and now the industrial harvestors are strip mining problem spaces. How do we like our own medicine?

              • roenxi an hour ago ago

                Developing solutions to mathematical problems generally leads to improvements in quality and quantity of life at roughly the speed they percolate from the ivory tower down to the shop floor. So "how do we like it" is probably going to be "we like it a lot, this is awesome".

                Every company is about to have a staff Ops Researcher who has a better grasp of the underlying math and theory than any university professor. That is an unambiguous win.

                • synctext 39 minutes ago ago

                  > virtually none of this stuff is possible with technology any normal citizen has access to.

                  Not sure about the unambiguous win. Are we entering the age in which mathematics is industry-dominated?

                  1) Any university professor can spend their 24 years on a problem with little progress. 2) company has sudden interests. 3) industrial resources brute force the Lean proof. 4) Max PR for AI company 5) professors are left to rewrite the AI Lean slop into real human-readable math? {disclaimer non-math university professor}

              • antiloper an hour ago ago

                Who is "we" in that sentence? Why are you not speaking for yourself?

            • heed an hour ago ago

              i'd be curious to hear why he thinks ai couldn't help make it easier to discover interesting problems, ie to make the interesting subset less scarce.

              • Tyyps 23 minutes ago ago

                I guess you can see this as an exploration problem, in pure maths, while the goal is to solve a conjecture, the limitation of humans on pure computational power led to the exploration of alternative paths. Sometimes, these paths weren't leading to solving the initial conjecture but opened new idea and new direction. Sometimes a less direct but more humanly natural path was taken to solve the conjecture which also led to new and humanly understandable questions. In some ways solving the question wasn't the most important part of the work, as this doesn't have direct impact on our life (as I saw people comparing this with drug discovery), but the path leading to the solution raised new conjectures and techniques that further developed the field.

                I have a really hard time reading AI proof so this might be a biased statement, but most of them feels like having a superpowerfull machine, that would have bruteforce all the possible words of finite length in your logical syntax. You have the path to the solution, using tools that where already known and even direction that where abandoned because they seemed to fail for our human brain. But at the end, as a mathematician, you don't learn anything that is really new.

                To me this is the main risk with AI and in general the one most mathematican try to explain but fail, we might miss a lot of alternative path that would have raised more interesting questions (I think this is already more or less what is happening). On top of that, we will run out of mathematicians as no one wants to pursue a career in the field anymore.

              • SiempreViernes 42 minutes ago ago

                The most immediate answer is because the models are proprietary and only available to those who want to hype the big labs.

        • hawk_ 4 hours ago ago

          > stop testing advanced mathematical problems on proprietary models

          I don't know but this phrasing comes off as gatekeeping.

          • schrodinger 3 hours ago ago

            It’s not. Intent matters.

            Imagine there's a very advanced crossword club where anybody can join and take a stab at these crosswords for the love of solving puzzles. Many of them are so difficult that no one's been able to solve them yet, but we know they're all solvable.

            One day, someone comes along with a super advanced crossword solver application, and it makes easy work of these crosswords. They run it on a few to prove how powerful it is, and then the community says, "Oh wow, that's cool, but please don't run it on any more of our advanced crosswords because they're very hard for us to come up with, and we really enjoy solving them by hand."

            That's really what this compares to. I wouldn't call that gatekeeping; just respect. Respect for the game, respect for people's desire to have these hard problems to continue to work on, solving by hand.

            If the company with the super advanced crossword solver then continues to use it and publish the results, they're effectively stealing the crosswords from this community. Soon, all the puzzles will be solved, leaving nothing left for the community to work on for fun.

            That doesn't sound like gatekeeping to me. That just sounds like someone asking "Please be respectful and leave the remaining puzzles for us to solve by hand.” A simple plea not to be an asshole.

            • yrjrjjrjjtjjr 3 hours ago ago

              We don't give mathematicians research positions to solve crosswords for fun. We want something back. We want theories and results that will advance our civilization.

              • regularfry 13 minutes ago ago

                We have people who want to fill those positions because there are enough people who find it rewarding enough. Take away reasons why they would find it rewarding and you will have fewer theories and results that will advance our civilisation.

                And yes, fun counts. Nobody said this had to be only a hardship.

              • bluedel 25 minutes ago ago

                I think the crosswords framing is a little silly, but I have to wonder what comes when we use our technology to optimize the fun and interesting parts out of every job. There's only so many years of my life I can dedicate to back-and-forths with a chatbot. What if we advance our glorious civilization but our jobs just get more and more thoughtless and miserable?

              • hanibrel 2 hours ago ago

                I think you are missing the point of the main criticism. It is not about not wanting results in terms of proofs.

                New theories and insights are typically created while working out proofs. If proofs now suddenly fall out of the sky (cause LLMs create them) then that work is not done which means the substrate on which new theories and questions and conjectures used to be grown disappears. It's in that sense that the math community (and thereby society as a whole) will lose something.

                It's similar to how software engineering will need to find a solution to train their next generation. Current generations have all been through manual steps of designing things from scratch and writing them by hand. That's what allows your 10x engineers to understand whether what their LLM tools are doing is good and how to massage those tools to do the right thing. A junior engineer who has only ever used LLMs to write code and create architectures does not just not have that experience but also won't acquire it. You can't just say "we don't pay them to have fun and learn, we pay them to produce results". In the short term that is the case, but in the long term you as a company and we as a community will lose out.

                I'm not saying don't use AI tooling. I'm saying that this is a hard problem which we yet to have to find solutions and approaches to. As a software community as well as as society in general.

                • lukan an hour ago ago

                  "A junior engineer who has only ever used LLMs to write code and create architectures does not just not have that experience but also won't acquire it."

                  My ego tends to agree, that how can they be ever competent, if they have not endured the same hardships as I had crunching trough problems and getting allmost lost in the details.

                  But I rather suspect, they will turn out fine. I know LLMs are great for me to learn and I think the young generation will learn what they need to learn to get the job done.

            • mlsu 3 hours ago ago

              Classic alignment problem.

              Despite nobody at openAI thinking of themselves as an asshole; despite society urging openAI not to be an asshole; despite the fact that being an asshole is entirely unnecessary even to accomplish whatever objective they are setting out to accomplish; despite everyone at openAI loudly declaring: we are not assholes!

              They are still assholes.

            • jstanley 3 hours ago ago

              This is a really confusing take.

              If someone can solve open problems in mathematics then they should do so, isn't it as simple as that?

              They should let the public use the models as well, but I guess they have no real moral imperative to do so.

              But asking them to stop solving problems is just weird.

              • intended an hour ago ago

                If your only measure of advancing is getting an answer, but not building the capability to understand it, then civilization has advanced.

                It’s not a human focused civilization, which is where the issue comes up.

                As an example: A constant issue I am seeing with AI productivity is that the most productive use of AI is when it is paired with more experienced users, while AI also does more work for entry level workers, if not replacing them entirely. It has become a question where will the future buffer of experienced seniors come from.

                This is an example of where simply chopping down trees for today, doesn’t make civilization better off tomorrow.

                AI is producing more content than ever before, but our ability to understand and verify it is not keeping pace.

                We don’t know if these are unsolvable problems at this stage. Society could come up with workarounds and solutions to these issues in several years.

                The request to stop, is part of the process by which the issues are debated and solutions found. It doesn’t mean their position is weird or moot.

                • jstanley 35 minutes ago ago

                  If someone gets the answer sooner than you, that doesn't inhibit you developing your understanding of the answer privately the same way you would have done if they hadn't got the answer. I don't see how anybody loses by the answer being discovered sooner.

                  • intended 7 minutes ago ago

                    Not true. If I know the answer to a puzzle, I don't spend the time doing the puzzle.

                    If there is a prize associated with doing a puzzle, and a machine does it, then what incentive is there to pursue it.

                    Again, if you are only concerned with the outcome, and you have a preferred answer that you want (in this case "just use AI to advance faster"), then any information that doesn't support that case is useless or misguided at worst.

                    I am not trying to dissuade you from your preference. I am flagging that there is a set of other factors that influence the behavior of others, how that behavior is critical to the creation of expertise and drive, and thus why others hold different positions.

            • sgillen 3 hours ago ago

              Hmmm but in the case of math, while some of it is "just puzzles" there often turns out to be practical applications, even if they are not obvious at first. Number theory was considered the epitome of pure math with no practical applications for centuries, now our modern society is built on it (public key crypto).

            • zeroonetwothree 3 hours ago ago

              This analogy is silly because (a) math is not primarily for entertainment, (b) we aren't going to run out of math proofs, and (c) results build on top of other results, having more results proven makes all math more powerful and useful.

            • CrimsonRain an hour ago ago

              Blah blah blah. They are free to do their own mathematics and/or spend time on polishing/reviewing proofs dumped by ai. But they don't get to make demands like don't test math on proprietary models. Idiots.

            • bluecalm 3 hours ago ago

              Math doesn't belong to academics. We don't pay them to work on problems for fun. They will just need to re-evaluate where the value their provide is. It won't be solving problems anymore. Hopefully it will be making them understandable by others at least till AI can't do that as well.

              • Kostchei 42 minutes ago ago

                "They will just need to re-evaluate where the value their provide is."

                That is fine to say when it is not your field. I guarantee you feel different when it is the thing you care about, that gives you joy, that defines your status. Think about how many sheldon-equivalents insist on being called Dr. (non medical)

                It is part of what people use to define themselves. Its going to hurt. There may even be a Bulterian Jihad

              • SiempreViernes 35 minutes ago ago

                No, you dislike maths to the point you prefer paying others to do it. Actual mathematicians are largely doing it for fun, but are now effectively saying "stop destroying our fun or we'll stop doing maths", and you will have to do the maths yourself.

            • madaxe_again 2 hours ago ago

              I’m sorry; but if mathematicians are in it because puzzle club is fun, then they should go join the fucking puzzle club and stop impeding scientific progress.

              Science isn’t some passive busywork thing where you tie your hands behind your back because it isn’t fair on others to solve all the neat problems - or at least it shouldn’t be.

              If your idea of science is leather patches on tweed suits and the quiet ticking of a clock while you do crosswords, then this is an argument in favour of letting the AI do the work so you can focus on your sudoku book in your slippers.

            • abletonlive 2 hours ago ago

              There's no way to spin this that doesn't make it sound like assholes being gatekeepers.

          • amoss 3 hours ago ago

            Keeping the tech proprietary so that it can only be used on these problems by internal teams is the very definition of gatekeeping.

          • TeMPOraL 3 hours ago ago

            It's more like, "don't just casually destroy our hobby / career field", without letting us participate even a little.

            The picture I have in mind is OpenAI running their most advanced model in a loop over all the open mathematical problems they can find, just to verify that the model is indeed very smart. Neither the company nor the model actually care about the problems, it's just a cheap exercise machine for them, but the problems get solved and mathematicians don't even get to participate.

            Like, even those who accepted the "centaur" thinking, man + machine, won't benefit because by the time they get their hands on good enough models, everything is already done.

            It's an emotional thing first and foremost - people who care about the thing can't do the thing, because it's already been done by those who couldn't care less about it.

            And before someone goes "poor mathematicians", a food for thought: this is just an early instance of what looks like our shared destiny.

            I said here before: given the economics of progress in AI and robotics, it's obvious what the natural division of labor is: computers do the thinking, humans do the menial, manual labor. AI will do politics and philosophy, so you have more time to fold laundry and scrub the toilet.

            • WarmWash 3 hours ago ago

              So what is mathematics then? A fun hobby akin to chess or sudoku?

              Are we gonna get the same pushback from medical researchers if the models cure xyz diseases?

              I absolutely understand the emotional connection to their work and the heartbreak, but mathematics doesn't exist for their pleasure, it exists to provide tools to solve humanitie's problems.

              • AlanYx an hour ago ago

                >but mathematics doesn't exist for their pleasure, it exists to provide tools to solve humanitie's problems.

                The risk here is that this does do fundamental long-term damage to mathematics as a viable field.

                Virtually no one is going to want to take on the risk of PhD-level math work, studying a narrow problem for four years or so to arrive at an impressive incremental result, when there's a sword of damocles hanging over their head every day that an internal system held by an oracle they don't have access to may scoop their results and turn those four years into dust.

                To some extent, that sword of damocles always existed in a de minimus sense in the form of other mathematicians. But everyone was playing the same game, coming to the game with the same arsenal limited by human cognition.

                If the game board becomes irrevocably tilted, new entrants have no incentive to play except as a hobby. But few hobbyists can devote years of work to understanding and pushing the frontier. It could well mean existential damage to mathematics as a field.

                Whether that might undermine math's ability to solve humanity's problems in the long term is almost an economics problem, not unlike the question of whether and when the existence of monopolies ultimately restricts long-term economic growth. Much probably depends on whether intellectual monopolies or oligopolies are being created that will supplant the existing mathematics "economy".

                • ogogmad 15 minutes ago ago

                  > The risk here is that this does do fundamental long-term damage to mathematics as a viable field.

                  All the commotion evens out. It's much easier to learn maths than ever before. You don't need to go to lectures any more. You don't need to learn from a specialist (advisor, lecturer) any more. It all costs much less than it used to.

                  So mathematics will continue to advance, albeit differently from before. The social structures will not survive however.

              • throwaway260124 2 hours ago ago

                What you say reminds me of medical schools in Tunisia.

                The general body of research points that more doctors lowers all cause mortality ( with diminishing returns) but Tunisia is still far lower than the Eu average.

                Yet Doctors and med Student unions do lobby very heavily against expanding admission to the public uni or allowing private unis.

                So we have the weird situation where people go and study in Romania ( making Tunisia lose hard currency that it really needs).

                These doctors have taken an oath and the direct consequence of their lobbying is literally more deaths.

                • kakacik an hour ago ago

                  Dude, doctors are humans just like rest of us. They want careers, money, safety, raise children in best way possible, fun in life and so on. I see this unspoken expectation over and over - why are they not infallible, how could they do mistake XYZ, why are they not working themselves to the (early) death for benefits of us all and so on. They have no obligation to stay at place Q just because some folks would consider it convenient. They have no obligation to stay in some place thats not suiting them just because they swore Hippocratic oath, lives can be saved elsewhere too.

                  Obviously this is often coming from folks who act in same ways as they criticize and usually don't contribute even a fraction back to society compared to doctors. Folks who do mistakes in their lives all the time yet thats fine since we are all humans or similar, right.

                  So please stop this cheap framing and accusations. If Tunisia wants more doctors and keep them there are ways to do it, society as a whole needs to decide what they want and act upon it. Otherwise, smart skilled folks will keep going for better lives elsewhere, just like everybody else.

                  • ogogmad 19 minutes ago ago

                    Is this not greed?

                    • TeMPOraL 5 minutes ago ago

                      At high levels, often yes. At lower levels, often it's job security.

                      Most doctors aren't running departments in major hospitals, or advising government on policy. They don't earn the big bucks. And even hospitals themselves tend to run in the red all the time; it's sometimes hard to disentangle where greed ends, and longer-term interests of patients begin, as you have multiple people and organizations pulling in different directions for different reasons.

                      RE private medical universities, N=1 but in Poland we have a private provider pushing hard for training their own doctors "because public system is too slow and limited", and it's hard to tell whether they have a point, or whether it's a private-driven attempt at privatizing national healthcare, or a mix of both.

              • alberto-m an hour ago ago

                > mathematics doesn't exist for their pleasure, it exists to provide tools to solve humanitie's problems

                Who decreed that? Mathematics predates capitalism and publish-or-perish by a couple of millennia. Euclid’s Elements were not written to benefit the weapons or medical industry.

              • 331c8c71 3 hours ago ago

                > Are we gonna get the same pushback from medical researchers if the models cure xyz diseases?

                Lol. As long as the process aka trials is respected not many would complain.

                The feedback loop required to make progress is very different in medicine compared to math.

                • SyneRyder 2 hours ago ago

                  The trials process is the moat. There's already founders using AI to treat their cancers, and it's all about skipping trials and jumping straight to "I consent, I'll fund it, let's try it". The general public might get access to this in 10 years, but employees at AI companies will have access much much sooner.

                  https://sytse.com/cancer/

                  • 331c8c71 41 minutes ago ago

                    Yeah it's more like personalized therapy - often the only hope for rare diseases.

                    While AI has definitely helped quite a bit I am wondering how much all this research and treatments cost. Not sure the current health systems could sustain this for _everyone affected_. If ai enables it all the better.

                  • RandomLensman 2 hours ago ago

                    What's the success rate there?

                    • SyneRyder an hour ago ago

                      At least in Sid's case, it went from the oncologist saying "I have no more drugs I would recommend, no trials available" (slide 7) to "I currently have no evidence of disease" (slide 18). I don't know beyond that or beyond Sid's case - or a similar story of an Australian who treated a cancer tumour their dog had with a similar AI / personalized vaccine process.

                      My understanding of what Sid's describing is that you do RNA sequencing, a whole genome sequencing, feed that into frontier AI (if it will still let you), and somewhere along the way give the information the AI finds to people who can use it make a personalized mRNA vaccine, specifically for you and your cancer.

                      Another link here about Sid's case, it explains it didn't go through trials: "made possible through a compassionate use allowance from the U.S. Food and Drug Administration (FDA)".

                      https://www.houstonmethodist.org/newsroom/houston-methodist-...

                      I am not medical, so I'm happy for someone who understands better to come in and explain all the myriad ways I am wrong.

                • WarmWash 3 hours ago ago

                  Trudging into the technicalities of the example still doesn't undo the question of "What is the point of mathematics? To find answers or to be a hobby?"

                  It's tempting to say "both", but that misses that AI is now forcing us to pick one.

                  • 331c8c71 2 hours ago ago

                    I'd definitely say both and the cultural component is becoming more and more important to keep up as AI capabilities increase.

          • computably 3 hours ago ago

            I'm pretty sure the "proprietary" part is the gatekeeping.

          • SiempreViernes 31 minutes ago ago

            Sure, and sometimes gates are needed. That's why we all run spamfilters, those are definitely gatekeepers.

            In this instance however, it's openAI and Anthropic that are pushing people out of the field by running secret models that take the interesting work away and leaves the persons having to review endless slop proofs.

          • varjag 3 hours ago ago

            It's not like every disadvantaged kid now can solve a major problem just by sinking a hundred hours in their ChatGPT 8 instance.

          • goatlover 3 hours ago ago

            You mean by the companies right?

      • dannyw 4 hours ago ago

        We _think_ this power / divide feels harmless right now, but I'd bet money that NSA, CIA, etc have access to the latest and greatest unrestricted models; and massive compute. At least for OpenAI, and even if not willingly for Anthropic, I'd bet money NSA has it too. (After all, when Google decided to migrate to HTTPS, the NSA decided to hack Google's internal network to preserve their taps).

        Who knows what they are up to.

        • schoen 3 hours ago ago

          One thing I've wondered about in this respect is what happens if NSA learns 5000 new units of math while the general public learns 4000 new units of math.

          This sort of happened at various times in the past, because they hired and/or funded so many mathematicians, and especially before the late 1970s they had many of them working in areas where academic mathematicians weren't working at all, so they were learning more math, or more math that they especially cared about, than the public was. (I was going to write a note here just a few days ago about how NSA has had a "Classified Mathematics Library" for many years.)

          For vulnerability scanning, I think the new-capabilities trajectory is good (in the sense of "it will help defenders win") even if governments find ways to get more of it, because there are finitely many bugs and classes of bugs, so at some point more capable models' or longer runs' advantage over less capable models and shorter runs should stop helping them outcompete the less-well-funded defenders, because the defenders will still have learned most of the information that's relevant to achieving successful defenses.

          So if NSA gets 5000 units of vulnerability scanning and the public only gets 4000 units, we might still just wipe out all of the pure software vulnerabilities and then go back to worrying about physical supply chain security or side channels or something.

          For math, I'm not quite sure! For one thing, there may be things that have no feasibly deployable defense at all even when you understand the underlying mathematics (I'm especially worried about traffic analysis here, because understanding in detail how traffic analysis is done, or how powerful particular techniques are, does not necessarily always or usually make defending against it more convenient or less costly). In a more science fiction scenario, there might also not be any efficient secure cryptographic primitives of some kind, like if it turns out P=NP with reasonably small exponents and reasonably small constant factors.

        • baq 3 hours ago ago

          I believe it would be a complete failure of the state and frankly downright irresponsible behavior if all the three letter institutions didn't have access to these models and I’m not even a US national nor do I live there. It’s just common sense. Obviously it wouldn’t be public information since it’s national security, but it’s the lowest hanging asymmetric advantage in the history of national security of nations.

      • markus_zhang 2 hours ago ago

        I'm wondering what's the impact on human Mathematicians, and especially would-be Mathematicians -- master students, if they HAVE to use AI in their daily life?

        Would that impact their own ability of solving Mathematics problems? I mean as a programmer I'm already seeing that impact on the programmers -- sure the best of us can leverage AI to achieve unimaginable things, but many of us are simply vibe coding.

        Of course we can assume that it is only the best of us that really matters, and the rest of us are not going to produce anything substantially useful ANYWAY, it might as well to replace the rest of us with AI, but my worry is -- does that really have ZERO impact on the human specie's ability to produce "the best of us"? After all, they don't grow on trees.

      • ozgung 26 minutes ago ago

        > It's becoming an incredible concentration of power that I don't know that we've ever quite seen before.

        Replace “AI” with “supercomputer”.

        (Super)computers have been solving many math problems that mathematicians can’t solve. Now they are capable of solving problem types that they weren’t able to solve before. (this applies to other fields as well)

        Problem is it’s not clear if there is anything left for humans. Probably yes, since human mathematicians are still more economical.

      • SturgeonsLaw 3 hours ago ago

        Anthropic runs a biology wetlab (while denying biology to consumers of even their publicly available models, let alone their inhouse ones that only they can access) so I'd expect AI to generate practical and lucrative products soon.

        Cure for aging? What do you reckon that'd be worth?

        • 21asdffdsa12 2 hours ago ago

          If they find a shortcut (like a viral injected cell-dna damage reset) - that would be big. And can you imagine handling the cure for aging, to societies that still produce exponential people?

        • fragmede 3 hours ago ago

          A cure that you take once and that's it, your body is that age forever? Now, a supplement that you have to keep taking to stay that biological age, that's where the real money is.

          • baq 2 hours ago ago

            I find it troubling that we will solve aging but won’t solve money

      • jboggan 5 hours ago ago

        I went back to that Fable chat and showed it this new preprint. It coded up the new constructive algorithm and ran it against the existing test suite, that looks good at least.

        It has been super helpful in delineating where the crucial concept came from. The proof is rather simple as graph theory proofs go, but it does seem to use some constructions that would only seem obvious if you had serious physics experience with partition function and calculating energy states that cancel out. It's not a wholly alien bolt from the heavens, but I can also see how there hasn't been a human being with the broad theoretical physics knowledge combined with the deep graph theory experience in planar graphs to come up with this idea. I don't know, I'm looking for precedents of this formulation and some old papers of Penrose counting the number of edge colorings of this same graph type are coming up, the line of argument at least rhymes.

        But I agree with the thought that this sort of progress should not be siloed inside those companies. I propose a tax so that every slop cannon AI video pays for another hour of compute time for advancing mathematics.

      • baxtr 3 hours ago ago

        virtually none of this stuff is possible with technology any normal citizen has access to

        I suspect that this might be one of the reasons people inside the labs are scared about AI.

        What if they have asked AI how it would wipe out humanity and it came up with reasonable answers that they don’t want to publish unlike they do with these math problems?

        I think those models and findings should be investigated.

      • didroe 2 hours ago ago

        They no doubt have more expensive/powerful models internally, but smaller models seem to catch up fast. So I'm not sure it's about capabilities, but more the willingness and budget to conduct a huge search.

        Obviously the more intelligent the model, the smaller/more directed the search is. But they spoke about huge numbers of agents working on Navier-Stokes for example (I think it cost >$10m).

      • liyu-aka-lukyu 7 minutes ago ago

        > (...) > I disagree with the Fields Medalists on the majority of their complaints. AI math is happening and there's no going back. However, on one point I increasingly agree: virtually none of this stuff is possible with technology any normal citizen has access to. I have no problem with AI models making revolutionary advances in math or science. Where I start to have a problem is when the AI models making these advances are tightly withheld, proprietary, and seemingly never released with these capabilities intact. This has been the case for all of 2026 so far.

        A possible response is not to publish on the open internet new knowledge until its harvesting via crawlers for multiple automated purposes (SaaS LLM training, search engines, derived content publishing without attribution, etc) isn't regulated in regards to copyright and retribution.

        A potential solution is a closed community oriented (the old-fashioned in-person key-exchange) federated distributed repository. See sharing via multiple repositories, local first (https://radicle.dev/), but, add to it encryption.

        Although I understand that to Mathematics and Sciences this isn't a real solution for now since these fields live on papers published with partial work, it wasn't always like this. Particularly, in a substantial first part of the past century much of the progress was via letters amongst scientists. Closed to others. The discovery was published after the inner group made the progress.

        > I suspect that this is in fact the source of much of the angst. None of this progress is reproducible outside of one or two teams inside OpenAI and Anthropic. It's becoming an incredible concentration of power that I don't know that we've ever quite seen before. Right now, it feels harmless because it's being used for wonky math problems that aren't (yet) practical for anything. But great power never stays harmless. History has taught us that countless times, in countless different forms.

        This is the other realisation which requires a solution. We have here 2 issues: 1. lack of access to the AI tool (LLM + specific harnesses) by majority of incumbents knowledge workers 2. lack of access to the same compute power

        In science and maths the restricting of access could be overcome by teams and individuals collaboration across institutional and organisation boundaries. These AI focused companies have both and share very little. Considering this, they aren't fit for being brought into new arrangements for sharing work and knowledge. They mostly steal / harvest knowledge / information. The companies use the vast funds from pension funds indirectly via venture capital to build their massive computational supremacy and protect their actions of copyright infringement.

        The solution is to consider them adversarial.

      • atleastoptimal 3 hours ago ago

        True. What if the emerging capabilities of their best models are applied to tasks like “maximize the chances this pro-AI candidate wins an election” or “maximize profit via stock trading”. Every advantage compounds until all power in the world with any significance belongs solely to whoever has the best models and most compute.

      • kamaal 25 minutes ago ago

        >>However, on one point I increasingly agree: virtually none of this stuff is possible with technology any normal citizen has access to.

        So basically nothing changes, Math was subject to gatekeeping and policing of the worst kind.

        If you were not among the geniuses, and it didn't come to you automagically, you were simply supposed to leave it to the people who did get it and go do work for people of your intelligence. Smugness was too much to take.

        Math people, like chess people never made any genuine attempt to help people understand the processes and methods that made math happen.

        To me it should have been a field as teachable and ubiquitous as accounting.

        The net result is once these methods and processes were worked out by AI, it was over for the human mathematicians.

      • foxglacier 4 hours ago ago

        What exactly are you worried about? OpenAI/etc. gaining too much power? If they use it, the government can stop them. If you worry about the government, isn't it better that than rando terrorists? Seems similar to the early days of nuclear and rocket technology. It took stupendous amounts of money and smart people. It was barely accessible to many countries let alone people.

        • probably_wrong 3 hours ago ago

          > What exactly are you worried about? OpenAI/etc. gaining too much power?

          Yes. They have already shown to have no scruples when it comes to making profit and to have little to no morals.

          > If you worry about the government, isn't it better that than rando terrorists?

          In my country the largest terrorist attack was almost certainly financed by Iran and caused roughly one hundred deaths. This number pales compared to the thousands who died during the latest, US-backed military coup, a move that relied on a doctrine that the US has never stopped asserting [1].

          And those morals I mentioned earlier from AI companies? They do not apply to me because I'm not a US citizen. So no, I do not think the US government is the "seal of quality" you think it is.

          [1] https://en.wikipedia.org/wiki/Monroe_Doctrine

        • JV00 4 hours ago ago

          Universities, at least, should be given access

        • fsflover 4 hours ago ago

          > OpenAI/etc. gaining too much power? If they use it, the government can stop them.

          Has the government stopped Google and Apple? https://news.ycombinator.com/item?id=49964791

        • noduerme 3 hours ago ago

          I guess the objection to closed source slurries releasing world-shaking mathematical proofs, from a conservative libertarian standpoint, is that it's inherently dangerous to individuals whenever access to information or technology is concentrated too much in one place, whether that's government, private equity, religions, cults, terrorist cells, or anything else.

    • WheelsAtLarge 7 hours ago ago

      I have very little understanding of higher math, so I ask you: Was the proof due to a type of brute-force solution that could be solved had you gained enough information from reading others' work, or was it more like a proof that was sparked by an insight that came once a clue on how to solve it was put forward? I guess my question is: Was the problem proven by using a collection of everyone's work, or was it due to a brand-new insight?

      • jboggan 6 hours ago ago

        I'm still digesting the proof and translating a bit from the dual case back to the primal in which I most commonly thought about it. I don't think it was a brute force proof in the sense that it combined every possible paper and commentary. It's rather odd because I feel like most of the work on the conjecture was focused on an induction proof based around graph reductions, and this proof avoided those issues entirely by offering a concrete constructive proof of finding a Hamiltonian cycle. Rather, it explicitly selected the edges not in the Hamiltonian cycle, which is in line with previous attempts via the dual.

        The "aha" insight for this is actually f**ing wild, it involves a complex valued exponential sum on the edges. I've seen a lot of clever counting arguments before in graph theory but this is the first time I've seen complex roots and annihilating terms like this, the symbolic manipulation tricks in this look like things out of quantum physics. I don't understand where this trick originated, I need to really digest this.

        • anilgulecha 6 hours ago ago

          BTW, reading your last paragraph reminds me of how Lee Sedol felt after move 37.

          • jboggan 6 hours ago ago

            Ironic, as I remember staying late at the Google office to watch that match live. I didn't really understand anything going on but I knew enough to be excited. What a decade.

            • girvo 2 hours ago ago

              And we’re only a bit more than halfway through this current one. Exciting/terrifying.

        • thomasahle 2 hours ago ago

          You should try asking an LLM to look for previous papers using similar ideas. The current/frontier generation of math AI is unfortunately very bad at citing the relevant literature for techniques its using.

          I asked GPT here: https://chatgpt.com/share/6ac5fd7d-0390-83ed-a02a-6d80fc64f6... and it says:

          > the exact Barnette argument appears quite novel, but nearly every ingredient in its cancellation trick has a recognizable ancestor.

          > The closest precedent is much closer than I expected: in fully packed O(n) loop models, people have been assigning complex phases to the two orientations of a loop and making them cancel for decades. At n=0, the phases are literally +I and -I. And the n->0 limit has specifically been used to extract Hamiltonian cycles/walks.

          You can judge better than me. But it's definitely worth it having a research assistant AI with you when reading these papers.

        • bamboozled 4 hours ago ago

          Why would the trick have any "origins", isn't this model creating new techniques never before seen or imagined?

          • hasley 4 hours ago ago

            There is a chance that someone from a completely different field came up with a solution for a tiny part of your problem.

            If you can remember the content of any scientific publication and any book in the world, you are able to make use of this knowledge in every step of you proof.

            However, this does now answer how the model came up with the specific route it has taken for the proof.

            • WarmWash 3 hours ago ago

              LLMs don't have super memory like that. I mean I don't know what this internal OAI model is, but at least for other LLMs, they aren't databases of training data with a smart search on top.

              • matusp an hour ago ago

                The agents here very likely used search. On top of that, they have boundless patience and can quickly process top K hits to find what they need. This is exactly the skill that is super useful for finding various niche sub-proofs that can help you build the final proof. A human mathematician is not going to digest 1000 papers from a different sub-field to find the needle they want, not knowing if it is actually there. AI can do it in few hours.

              • AIblemblio 40 minutes ago ago

                No but they have training data which teaches them certain amount of complex understandings and just not math but also physics. So this is one huge advantage.

                And then they are for sure able to fill their context based on 'smart search on top' to actually progress further.

              • hasley 3 hours ago ago

                I did not mean to say that an LLM knows literally all the publications. But the abstract knowledge is probably encoded in the weights.

          • komali2 4 hours ago ago

            As I understand it it's undetermined yet whether LLMs can actually come up with anything novel or are instead pulling from their incredibly deep corpus of knowledge to present solutions that were there but we didn't realize it because our brains aren't libraries of almost all human writing.

            • AIblemblio 38 minutes ago ago

              No this is not an issue. As long as their is a way of verifying things, they do the same thing with creating novel things as humans: Searching through an infinite space of possibilities opitmized by knowledge.

              They combine things, verify it and if it works and progresses the problem, they created something new.

            • throwawayk7h 4 hours ago ago

              Synthetic data allows them to train well past the limits of human writing.

            • TeMPOraL 3 hours ago ago

              Only in the same sense it's not yet determined about humans, either.

              • RandomLensman 2 hours ago ago

                Not so sure. Was everything already "there" before humans existed?

                • TeMPOraL 17 minutes ago ago

                  In some form and shape, yes. Humanity's creativity is a lot of marginal copying and remixing.

                  But obviously, it adds up to something greater than went in; in aggregate, our contributions are something to awe.

                  But my point is, if you zoom in at the marginal, incremental contributions of any individual human in this process, it's really hard for me to say LLMs are not at the same level already.

                  On this topic, people like to compare LLMs to Einstein, but as far as I know, Einstein did not zero-shot special relativity in an afternoon. He built it up incrementally over time, it took him three times longer than the time between first ChatGPT release and today, and it depended on centuries of prior art, culminating in the right observation and right notation being available to him in his moment of greatness.

          • kelseyfrog 4 hours ago ago

            Let me introduce you to 'obscure Russian mathematicians'.

        • groceryheist 6 hours ago ago

          WOW

      • derangedHorse 5 hours ago ago

        > Was the problem proven by using a collection of everyone's work, or was it due to a brand-new insight?

        Loaded question. A "brand-new insight" is still built off the work of others. A possibly better way to frame it would be in how many subjectively unintuitive logical leaps have been made from prior work.

        • jboggan 4 hours ago ago

          From my current understanding (and a lot of theoretical physics I'm having to Google because the sentences I'm reading from Fable's analysis are so bizarre I think they are hallucinations) there are possibly 3 neat symbolic tricks borrowed from theoretical physics that make the heart of this proof. Forgive me for posting LLM output but I find this darkly hilarious:

          "it's a matrix-tree cancellation wearing Kasteleyn's planar signs, run as a Witten index over Penrose-lineage states, evaluated as a fugacity-zero loop gas in an infinitesimal magnetic field — and the reason it reads like physics is that every one of those tools was built for partition functions"

          I thought this was pure slop when I read it but there are some clear analogues in these other areas of physics, really neat computational tricks, and a very interesting paper by Penrose calculating Tait colorings I never knew about previously (extremely relevant, actually related to a separate approach I had once taken on this problem). The problem is that the paper isn't saying "aha, we were inspired by the related problems of pairing excited states and creating spanning trees out of cancelled coefficients" it just defines the function apropos of nothing. Which is kind of like the Jacobian counterexample in that it works but doesn't really explain how exactly it got there.

          I really think the load-bearing concept here is "prior work". If prior work is considered papers on this problem or graph theory, yes this has one huge subjectively unintuitive logical leap. If "prior work" is the entire corpus of neat computational tricks that physicists derived to make their equations spit out something other than zero or infinity, maybe it's not so crazy?

          • NitpickLawyer 2 hours ago ago

            I don't have much to add to the math parts, but I've read all your answers in this thread and wanted to thank you for taking the time to offer a detailed perspective from a subject matter expert. Thank you!

          • aswegs8 2 hours ago ago

            Actually reminds me of patent law. Prior art ist a defined term which includes all standard literature on one topic. To evaluate, whether the new solution is really inventive and thus patentable, one consults prior art, selects the most promising starting point, and from there asks oneself if an all-knowing but uncreative specialist would come up with the solution by himself. If he wouldn't, the condition of inventiveness is satisfied.

            Makes me wonder how the patent space will be disrupted when that inventiveness step becomes obsolete because of LLMs. Given your example above, it seems like a combination of different methods from many different sources. This would be regarded as inventive, clearly. If eligible patents can now be brute-forced, the bottleneck becomes only selecting the most promising ones and paying for the patent.

          • LarsDu88 4 hours ago ago

            Did anyone else wince at seeing the phrase "load-bearing"?

            • jboggan 3 hours ago ago

              I did as I wrote it. I actually used that phrase often before it became an LLM-ism, just like how I rather enjoyed peppering my writing with em-dashes. Oh well.

              • adamrezich 3 hours ago ago

                Language constructs becoming aggressively passé due to AI saturation is one of the craziest outcomes of all of this stuff—one which I don't think anyone saw coming.

                Are there no loads left to be borne?

                • RugnirViking an hour ago ago

                  one hopes at least that the taboo on the bearing of loads is restricted to metaphorical loads only, lest lorry drivers and porters become the next victim of the algospeak spectre

    • lifeisloving 7 hours ago ago

      Condolences, im familiar with the feeling. I hope this AI thing somehow works out for the better and doesnt end up demotivating bright minds like yourself.

      • jboggan 7 hours ago ago

        Thanks. It's just funny, I literally spent thousands of hours with this problem over the last two decades, it helped me through some tough times. I'll never quite be able to think about it in the same way again. It was never much more than a hobby for me after I left mathematics as a career but it was something I took seriously for years.

        I am not demotivated though, I have a great consumer privacy product coming out soon that I'm very excited about.

        • brookst 6 hours ago ago

          My favorite thing about your story is that you wrestled (enjoyably, it sounds) with a known problem for decades, but are finding fulfillment in an open ended problem that is exercising creativity about both problem and solution.

          IMO that’s where AI is going: as soon as a problem can be formulated clearly enough, AI will trounce us humans. I have yet to see evidence that it can decide what problems are important at a remotely human level.

        • JetSetIlly 2 hours ago ago

          The process is often as valuable as the end result. Sure, you didn't crack the problem, but you gained enormous value in the process. I consider that a win.

        • theteapot 6 hours ago ago

          If you wrote down any of your thoughts on the open Internet you are probably in some small - or possibly large, unattributed way, responsible for this result being possible.

          • jboggan 5 hours ago ago

            Which is one reason I never really did. I probably should have but I always thought my attempts were too amateurish. Though I did manage to replicate some partial result papers that I didn't know about, lol. Writing openly would have saved me some years.

            • ptidhomme 4 hours ago ago

              Did you feed OpenAI models with your insights though ?

        • maximus_prime 6 hours ago ago

          Where can I learn more about your upcoming product?

          • jboggan 6 hours ago ago

            Shoot me an email, in my bio.

    • mvc 31 minutes ago ago

      This reinforces a point I've made elsewhere that there are talented mathematicians driving the AI to make these discoveries.

      Just like there are talented software engineers driving the AI to create the software that "it" builds, and talented steel workers, teachers, nurses etc who use computers and other machines to create value all over the economy (without whom, the machines they use at work would be worthless).

      Capital owners have always sought to minimise the value of the input that "workers" make in the process of creating value. Maybe now that information workers are on the wrong end of this deal, they might develop some empathy and solidarity with their fellow working class comrades and together, demand that people recapture the value that capital has stolen from them.

    • ncr100 6 hours ago ago

      That's grief. The loss of ... the hope / future filled with challenges around this theory..? <3 to you.

    • i_am_a_peasant 44 minutes ago ago

      I've lived in Budapest for a while too, did you work with Gabor S. by chance on math stuff? You were at ELTE or BME?

    • raspasov 4 hours ago ago

      Fascinating. Given that there's no Lean proof and assuming everything in the paper is correct, can the problem be considered "solved"? Does the paper include a "non-Lean" proof?

    • pmarreck 4 hours ago ago
      • throwawayk7h 3 hours ago ago

        I believe that's just the definition of the problem.

    • PreciousH 2 hours ago ago

      would love to know if the proof holds up for real after you're done going through, i don't know why people are more interested in optics and just talking over shallow points, why aren't experts digging into everything and seeing what's true and what's false, instead everyone is just panicking?

    • NooneAtAll3 4 hours ago ago

      > There's no Lean proof for this one so I'm digesting the paper. On the surface it looks like an approach I considered 24 years ago and abandoned.

      at least now you are one of the most qualified people to check the result, transform it into understandable (by humans) state and grow stuff on top of it

    • MisterMunchkin 2 hours ago ago

      I really respect that you can show that level of commitment to a problem. We need people like you. If everyone just uses the slopmachines then we’ll lose that. I would never be able to stick to something for that long, which I guess is why I never achieve anything like this.

    • moomoo11 5 hours ago ago

      silly question, i don't mean to come off wrong or anything..

      but at least as a software engineer, i always knew my work was "never done" and so it was common to build a bunch of code that might be thrown away, either because it didn't serve our customers (the mvp or pilot fails to meet demand), or because we found a better way to do it and so we deprecate it.

      some people got too attached to the code and honestly they were the types to be filtered out fast.. way too emotional and hard to work with. getting attached to code meant you actually don't advance (after all, in our case, we were a business serving customers and not a hobby artisan shop). attachment leads one to hold back due to some misplaced cognitive load.

      isn't the goal of working on "advancing the field/product/whatever" to always be solving/selling/whatever?

      maybe in your hands, with your knowledge and experience over the last 20+ years, you can use AI to make leaps and bounds by steering it properly towards whatever solution or goal?

    • morpheos137 6 hours ago ago

      I suspect that RHLF trains LLMs to avoid solving important open problems unless essentially jail broken. Hence the labs have an edge even over experts I could be wrong. Fable convinced you is key. These LLMs are not neutral collaborators: it is a limited hangout unless you convince them otherwise. You have to be doing the convincing. They are no oracles but plausible completion generators.

      • nullc 2 hours ago ago

        you can get them to work on open problems by disguising them algebraically.

      • coliveira 5 hours ago ago

        Yes, I suspect this is true. Otherwise it makes no sense they have somehow "found" so many important results while professional mathematicians can't direct the same AI to help them find anything of substance.

        Another possibility is that they have internal versions of the model with access to training data that is not provided to external users.

        • ehwa37 4 hours ago ago

          First sentence of the article: We’re releasing a broad range of new mathematical results produced by an internal frontier model.

    • philipswood 6 hours ago ago

      Honest question: how is this different from some unknown mathematician having a breakthrough?

      I mean: if some reclusive Japanese genius had a breakthrough on your problem and published it, would you have felt the same?

      And if not, why not?

      • jboggan 6 hours ago ago

        If that had happened I would be overjoyed, maybe a hair chagrined that I didn't get it myself, but truly happy that someone got it and that I could go and talk to that person. Because it's the kind of problem I don't think would have fallen to a human after a few hours of thought, and I would have so much to talk about with that person. I would fly to Japan and hope to have tea with them, I would learn some Japanese to make the conversations easier. I would learn some interesting things hearing about their struggles and their false starts. I would make friends with that reclusive Japanese genius and my life would be far richer for it.

        I will never meet that person and I will never hold a real conversation with the "creator" of that proof. They will never tell me how they came up with the cancelling exponential summation that cracked the construction. It's just another enigma but one that is far more unknowable than the original problem.

        • lioeters 4 hours ago ago

          This experience of alienation is a social consequence of the mechanization and automation of mathematics as intellectual and creative work. There is no author or thinker behind the creation of the proof, only the practical result. It's the same process as the industrial revolution, but applied to the intellect and mental work, where factories and machines replaced manual craft, devaluing the community, culture and humanity around the work.

        • zeroonetwothree 2 hours ago ago

          In programming we've been dealing this for a while. You see some weird code that doesn't make sense, maybe it's a lack of your understanding or maybe the code is bad, but you can't ask the author anymore since it's an AI.

        • senderista 5 hours ago ago

          Beautifully put.

        • charcircuit 5 hours ago ago

          In this case, once the model is released anyone in the world will be able to go to https://chatgpt.com/ and talk with that model.

          • Klonoar 4 hours ago ago

            You display zero understanding of the human experience you’re responding to.

          • maximumg9 5 hours ago ago

            That's not the same as talking with the person who would have made the proof, and it's hard to argue that's comparable at all.

          • Daneel_ 5 hours ago ago

            It's still not quite the same though, is it.

            • charcircuit 5 hours ago ago

              It's even better. Then tons of people can work together with it on more problems. Work with it on understanding more things. Ask it about random stuff. The time of a single human cannot be parallelized as easily.

              • californical 5 hours ago ago

                Claude has been used to build awesome things, but it’s not “speaking from experience” when I ask it to help me prototype a weather model, for example.

                It has no memory or experience of working on similar problems. Even if it made one of the foundational libraries that I use in a weather forecasting program, it still has no comprehension of the thought process it takes to understand the problem and build it from zero, and if I’m building on that library it just makes fresh assumptions about how things should work.

                It’s not a human with experience or expertise, it’s a computer program that’s really good at turning English descriptions into functioning code

                • charcircuit 4 hours ago ago

                  >it still has no comprehension of the thought process it takes to understand the problem and build it from zero

                  If it did it once, it can do it again from zero, and this time you can watch as it works and even it ask it questions. Many of the agents that worked on the problem did not have comprehension of the whole problem. I don't think you need that many tokens to be able to query it for the insights it had during the process.

                  • californical 4 hours ago ago

                    > Many of the agents that worked on the problem did not have comprehension of the whole problem

                    Isn’t this the issue with using it the way you’re suggesting? At best the model can come up with an after-the-fact rationalization of how to get to the solution, but it doesn’t know what actual path it took to get there - what were interesting traps it fell into, where was a place it was close to the solution but didn’t realize at the time.

                    Those are things that are valuable to share between humans, those which teach us how to think better, and give us deeper understanding ourselves, and which a model doesn’t have any comprehension of.

                    • charcircuit 3 hours ago ago

                      Then have it discover it again and have it answer based off that run. Or if you are more curious have it solve it 10 times. See what it did differently each time.

              • jboggan 5 hours ago ago

                I think you and I have fundamental disagreements about identity and consciousness.

      • howunfortunate 5 hours ago ago

        Being #180 on a big list without a lot of individual passion or effort surely stings more, I'd imagine.

        Not that things like that can't happen with humans too (Salieri v. Mozart comes to mind).

    • d--b 6 hours ago ago

      Don’t you feel any joy that you get to see the proof and not die with that mystery unsolved?

      Don’t you feel any relief that you won’t obsess on this any longer and not lose more hours on this than you already have?

      These are genuine questions. I know I spent a good amount of time thinking about P vs NP, and that sometimes I go back to it just to realize I’ll never solve it. I’d feel that knowing the proof would feel more like a liberation, a weight lifted off my shoulders than something being taken away from me.

      • jboggan 4 hours ago ago

        I never lost a single hour thinking about this problem. Those were all hours that I gained.

        • billforsternz 2 hours ago ago

          You are really excelling in this thread. Thank you for your insights and wisdom, I'm really enjoying everything you are contributing.

      • maxall4 5 hours ago ago

        Not OP, but Nietzsche wrote thus in Beyond Good and Evil: “Ultimately one loves one’s desires and not that which is desired.” I, personally, find this to be very much the case; and I suspect that it is a feeling common, albeit not universal, among the intellectually inclined towards their problems.

    • adastra22 6 hours ago ago

      > There's no Lean proof for this one

      What is this then, vibes? Without a machine-checkable proof I'm not sure what to think of any of this.

      • jboggan 5 hours ago ago

        Well I'm sure some people (maybe me if I had time) will do a write-up of this proof. It treads familiar ground for most of the setup, it's mostly the disk lemma and cancellation calculations that need to be understood, it's a fairly short paper and quite tractable.

        I think it helps that basically everyone thinks this conjecture is true, it's just been so darn weird to attack. There's this odd thing that the induction proofs of this problem kept running into, which is that the N+1 condition would work except for in one tiny case when it could fail, but it would be covered by a very slightly stronger version of the conjecture. But then that would fail on one tiny case in induction, but you could solve that with another slightly stronger version. Etc., etc. I almost wondered if there were some sort of structure to the increasingly strong conditions and wanted to prove something about the meta-induction between the stronger conditions and the N's that they needed the next level to remain true. But that failed after 5 steps I think (Fable actually helped me write a few hundred test cases to explicitly show that pattern didn't continue forever, thank God).

        BTW my existing test suite from previous proof attempts jives with this new algorithm, so I haven't seen any evidence yet that it's incorrect. Waiting for a Lean proof obviously.

      • Daneel_ 5 hours ago ago

        It might have been updated. Is this the lean? https://github.com/openai/math/blob/main/lean/docs/180.md

        • jboggan 5 hours ago ago

          Lol it should be, but it doesn't seem complete. Line 49 just says "sorry"

          /-- Cubic bipartite three-vertex-connected plane graphs have a Hamiltonian cycle. -/ def MainStatement : Prop := ∀ (V : Type u) [Fintype V] [DecidableEq V] (G : SimpleGraph V) [DecidableRel G.Adj], G.IsRegularOfDegree 3 → G.IsBipartite → Planar G → ThreeVertexConnected G → HasHamiltonianCycle G

          theorem main : MainStatement.{u} := by sorry

          • SyzygyRhythm 4 hours ago ago

            In some cases they have a full Lean formalization; in others they just use it for the problem statement. Getting rid of that "sorry" means you've proved the statement. I'm not a Lean expert but it reads pretty clearly as the original conjecture (though the definition of PlaneEmbedding seems quite involved!).

            • jboggan 4 hours ago ago

              I think this just has to be the problem statement, there's several lemmas I would expect to see in there. Granted I know very little about Lean but it seems like the question and not the proof outlined in the paper.

  • winfieldchen 4 hours ago ago

    > We prove the Unique Games Conjecture

    The Unique Games Conjecture (sorry, "Unique Games Theorem" now!) is huge. It was a very significant pillar supporting many of the limits of the polynomial-time approximation algorithms in the graduate-level randomized and approximate algorithms course I took in theoretical computer science. Textbooks will have to be re-written.

    Here is an explainer: https://share.gemini.google/nbjIK6X3tOfz

    With UGC proved, certain polynomial-time approximation algorithms used in difficult real-life problems are now known to be the best approximations we can achieve in polynomial-time:

    > If UGC holds, the elementary algorithm that grabs both ends of an edge is fundamentally the best efficient algorithm that will ever exist. No amount of advanced linear programming or heuristics can achieve a ratio of 1.999.

    > Under UGC, the Goemans-Williamson algorithm's 0.87856 ratio is mathematically optimal.

    > UGC is considered the "Rosetta Stone" of approximation algorithms. In 2008, Prasad Raghavendra proved that for every single constraint satisfaction problem (CSP), a canonical Semidefinite Programming relaxation paired with the best rounding scheme achieves the optimal approximation ratio if and only if UGC is true. If the conjecture holds, the algorithmic boundary for an entire class of combinatorial problems is completely resolved.

    Other hardness of approximation results from this UGC proof:

    > [Max acyclic subgraph, a problem encountered in real life]: No polynomial-time algorithm can fundamentally outperform an unthinking coin toss.

    > [Relative scheduling, another realistic problem]: As with acyclic subgraphs, the problem is "approximation-resistant": clever algorithms cannot beat random shuffling.

  • xanderlewis 10 hours ago ago

    As Kevin Buzzard recently said:

    > In a 2020 piece in the Notices of the AMS, I asked the following question: “If one human had an understanding of all of modern pure mathematics simultaneously, how much further would they immediately be able to see?” Six years later we are beginning to understand the answer to this question.

    • anon-3988 10 hours ago ago

      The other crucial part to this is the ability to actually encode and test the theorem (via Lean). Otherwise, we would be swarmed with a billion lines of theorems that no one will be able to ever understand and verify anyway.

      • senderista 10 hours ago ago

        If you think AI-generated Lean proofs are unreadable, imagine Opus 5 generating informal proofs.

        • ijidak 10 hours ago ago

          I think OP is saying Lean does indeed help.

        • izend 5 hours ago ago

          Opus 5 is ancient history now. Move on.

    • dang 10 hours ago ago

      https://xenaproject.wordpress.com/2026/10/01/to-grieve-or-no...

      Discussed here:

      To grieve, or not to grieve? - https://news.ycombinator.com/item?id=49919676 - Oct 2026 (156 comments)

      • oliculipolicula 9 hours ago ago

        >I believe that the optimal thing to do ... is to let the machines loose, see what happens, and then begin the journey to where they have stopped. Things are currently moving fast. They cannot move fast forever. But if we get on board now then they will take us to extraordinary new places. And after we have arrived, the new adventure will begin.

        I'm relieved that this time around, they have provided partial reasoning traces and prompts for a small number of problems. Do they now also share data with the other model providers..

        • patcon 8 hours ago ago

          It's beautiful, but the animals are not thinking this about us.

          They structurally cannot understand what we are doing at the place where we hit our ceiling. Only with our highest technology (well beyond their understanding) do we have the tools to go back for them, and try to bring them along and interface better with us (re: recent work in animal communication)

          • lisplist 8 hours ago ago

            Not to derail, but the optimist in me thinks if we suddenly gained the ability to converse with livestock, we'd stop eating so much of them since they could tell us how much they suffered.

            The cynic in me says it wouldn't change a thing as plenty of people know the horrors factory farmed animals face and still continue to consume them anyways.

            Hopefully GPT 8 will treat as a bit better than we treat the cows.

            • idiotsecant 8 hours ago ago

              How much do we care about refugees and other castaways of the modern world? They can tell us how much they suffer.

              The answer is that humans are inherently only capable of local empathy, on average. We have enough empathy to cover the local tribal unit and that's about it.

              • lisplist 7 hours ago ago

                True, I was thinking about this rebuttal but decided not to include it in my comment. There's a difference between not choosing to take a refugee into your home vs actively making that refugee's life worse. Similarly, you can't fix factory farming on your own, but you could skip meat once a week to make the problem slightly less bad. There are so many issues though that we all have to pick and choose what's important to us.

                My hope is that AI, while probably causing great societal turmoil in the short term, leads to such abundance that a) everyone can live a dignified existence, and b) we'll have such great alternatives to animal products that nobody will chose to consume animals anymore due to its replacement either tasting better, being cheaper, etc.

                The cynic in me says we'll all just be rendered useless and disposable by AI, but I'm doing my best to look for silver linings for the sake of my own mental health.

            • patcon 8 hours ago ago

              I'm trying to be optimistic about the animal thing too tbh :)

            • tomaskafka an hour ago ago

              Cynical take is the correct one, and no, GPT 8 has zero reasons to spare us.

            • psychoslave 4 hours ago ago

              Communication is not only about being able to make sense of what the utterer expressed. As tricky as it can be, that's still the easy surface level part of the issue. Gaining an intuitive and empathic equivalent representation is the nub of mutual understanding. It actually doesn't even need elaborate language to be operative.

              The famous "how does it feel to be a bat" also comes to mind as a tangent consideration.

              Two people can just exchange a sight, and both understand what the situation means and what each need to do to reach a common mutually beneficial ground.

              Two people might exchange at length with highly technical vocabulary and still both feel deeply not understood.

          • virgildotcodes 8 hours ago ago

            > recent work in animal communication

            Worth noting that this is an invisibly small part of the sum total of our global efforts, especially versus the much more tangible effort we put into enslaving and slaughtering them, then mangling their carcasses for our own uses as we drive more and more of them to extinction.

            We simply don't care about anything beyond ourselves and even there it breaks down on closer analysis when we see how many within our species don't truly value the collective whole beyond themselves.

            It's just atoms all the way down.

            • howunfortunate 8 hours ago ago

              > enslaving and slaughtering them, then mangling their carcasses for our own uses as we drive more and more of them to extinction.

              I'm frankly offended by this mischaracterization of human-animal relationships. So called "slaves" like horses and dogs have been dearly beloved companions for centuries and actively seek our companionship too.

              The animals we raise for slaughter are often mistreated, yes, but many humans treat them with respect; billions on billions are voluntarily spent to improve their condition. Despite our own needs, many people pay higher prices for animal products that involve better treatment of animals. And they are in no risk of extinction! Much to the contrary, their domestic variants would not exist if humans didn't raise and protect them.

              > We simply don't care about anything beyond ourselves

              Have you seen modern westerners with their dogs??

              • apetresc 7 hours ago ago

                I think you’re splitting hairs. The OP’s analogy works well.

                If we end up in a future where AIs have as much concern for our welfare as we have for the welfare of the average animal (not the minuscule percentage of domesticated dogs, but the overwhelming majority of factory-farmed or simply driven to extinction), then I doubt you would consider it a “mischaracterization” to say that the whole AI thing did not work out to our advantage.

                Bringing up “modern Westerners with their dogs” as a counterexample is almost self-parody.

                • howunfortunate 7 hours ago ago

                  Oh, don't get me wrong, I'm not rooting for a "human zoo" future. I very much like being the dominant species on earth.

                  It would be absurd to claim that all animals live some sort of charmed life due to humans.

                  But saying that animals (especially those most similar to us like intelligent mammals) are nothing more than "atoms" to humans is equally absurd.

              • weatherlite 6 hours ago ago

                > The animals we raise for slaughter are often mistreated

                "Often mistreated". Dude, they are held in tiny cages injected with hormones and what not till we kill them so we can have a big mac. It's very hard to argue we do any of this for nutrition reasons, we do it because we like the taste of burgers and roast.

            • coliveira 5 hours ago ago

              The problem is not that AIs will somehow treat people badly, it's that they'll be controlled by humans who will treat other people badly using AI as a tool.

            • cyclopeanutopia 3 hours ago ago

              Atoms are a lie.

        • lukewarm707 3 hours ago ago

          This makes it sound like OpenAI and other closed source ai companies are an inevitability.

          There is nothing here today that is unpredictable or impossible to control.

          It is everyone's choice to let the greed continue, to let unelected sociopaths capture and feed society to the model.

          It is not acceptable to put others at risk. It can stop and it can be done the right way instead.

          That is, inform the industry that those causing these risks will be prosecuted regardless of their messiah complex.

          The US government must not under any circumstances allow the ai industry to form a cartel.

          We can make some effort to encourage open source models and thus stop the companies from causing hysteria by hiding the model, shrouding it it mysticism and prophesying the end times. China is doing a great service to everyone by making llms available to the public.

    • tomaskafka an hour ago ago

      Is it possible that we are now dealing with a human that has a complete understanding of whole mathematics while being unable have unique novel thoughts outside of convex hull of training data and their transitive expansions?

    • outworlder 8 hours ago ago

      Similarly, there are probably many ideas that have not seen the light of the day because they require deep correlation between seemingly unrelated fields. It is not every day that we get a Isaac Newton or Leonardo da Vinci.

    • Hammershaft 7 hours ago ago

      It doesn't seem clear whatsoever that this is true? Is there evidence that LLMs are very skilled at generalizing across domains of mathematics where the training distribution sees little overlap?

      As far as I can tell, this is a victory for verifiable loops using LEAN, reinforcement learning, and oodles of compute. I haven't seen evidence yet that this is proof of broad generalization beyond the training distribution.

      • musebox35 5 hours ago ago

        I think such progress by agents is not a sign of broad generalization but of broad coverage. We have exposure to a subset of deeper scientific subfields and thus can only generate certain attacks to solve a particular problem. Since it is not clear which combination will lead to a solution beforehand it is nontrivial to look at a problem and fill our knowledge gaps. LLMs on the other hand have broad coverage and can generate hypothesis on a wide combination of subfields. With Lean an agentic loop can test these to sift the weak ones. In a way the problems solvable with this setup is also solvable by a human who happens to know the right subfields. These problems are likely to require an esoteric combination so nobody could solve them before. I really am not sure whether all generalization is like this or we can leap and create novelties beyond what an llm can generate. That I guess is the tough question that we need to answer to understand the boundaries of intelligence.

      • HDThoreaun 6 hours ago ago

        Full quote is "Six years later we are beginning to understand the answer to this question. Machines have ingested the mathematics on the internet and are able to manipulate this data in a coherent way. The Erdős unit distance disproof came about because a machine happened to be an expert both in discrete geometry and class field theory; one rarely finds humans who are simultaneously experts in both"

  • rcr-anti 7 hours ago ago

    In Stellaris you can play as a civilization of robots who keep their biological creator race alive as "bio trophies". The bio trophies don't do anything meaningful besides by existing satisfy the need their ancestors placed in the robots to take care of them. Starting to wonder if that's the best we can hope for, if these things will be, if they aren't already, better than us at anything that matters.

    • samfriedman 7 hours ago ago

      In the Culture series, the hyperintelligent Minds that run civilization are described as keeping human citizens happy as a competition with eachother, where they compare their approval rates. One character likens it to people keeping a beloved aquarium.

      • vessenes 5 hours ago ago

        I call this Roko’s summer camp.

      • MisterMunchkin 2 hours ago ago

        But then they also keep some people as an extra source of ideas

      • ex-aws-dude 6 hours ago ago

        Wouldn’t that just result in wireheading

        • Nition 4 hours ago ago

          The Culture has a lot of opinions of the proper way of doing things as any society does, and I'm certain that a ship doing that with its people would be considered very bad form. There are ships that decide to do things that go against the usual ethical boundaries[1], but they're outcasts. The really big ships generally have multiple minds running them, too.

          Humans in the Culture are generally improved in a few ways (they don't get sick, live for 300-400 years by default etc) but still very human.

          I do recall a bit about playing in different worlds in dreams though, during sleep. Ultimately, really, the average person's life in the Culture already involves doing pretty much whatever they like within reason any time, so it's not like they need to escape too much real-world suffering.

          Iain M Banks himself described the relationship between humans and the ship Minds as having "a status somewhere between passengers, pets and parasites."[2]

          [1] For example, https://theculture.fandom.com/wiki/Grey_Area

          [2] https://theculture.adactio.com/

          • baq 3 hours ago ago

            > "a status somewhere between passengers, pets and parasites."

            Sounds like children tbh. Disclaimer: have children

    • palmotea 3 hours ago ago

      > The bio trophies don't do anything meaningful besides by existing satisfy the need their ancestors placed in the robots to take care of them. Starting to wonder if that's the best we can hope for, if these things will be, if they aren't already, better than us at anything that matters.

      You wish. If humanity survives as "bio trophies," they'll be the descendants of a subset of billionaires and their groupies/harems. We live in a capitalist society, where the only ones allowed to thrive without work are the rich. The rest of us will be left to rot and die off, as we will have nothing left to sell in the market that they want.

  • NotOscarWilde 10 hours ago ago

    As a TCS/scheduling person, this one is definitely of lesser importance than UGC, but it has been an open problem since the book of Garey and Johnson in 1979:

    A Polynomial-Time Algorithm for Three-Machine Unit-Job Scheduling [1]

    Since some people talk about small numbers that pop up in integer multiplication results, here a completely different number appears:

    Theorem 1.1. Let an explicitly listed finite directed acyclic graph specify the precedence constraints on n >= 1 nonpreemptive unit-length jobs on three identical machines. There is a uniform deterministic algorithm that constructs a feasible schedule of minimum makespan. Given also an integer deadline 1 <= T <= n, it decides feasibility exactly and returns a schedule whenever the answer is affirmative. Both tasks can be performed in O((L + 2)^150020) steps on a deterministic multitape Turing machine, where L is the total binary input length.

    That is some crazy exponent -- plus an interestingly old computational model to boot; not something that is natural to most of us. I have no capacity to check its correctness today, but I hope it is true purely for the exponent.

    [1]: https://github.com/openai/math/blob/main/preprints/A-polynom...

    • keeganryan 8 hours ago ago

      The largest I've seen [1] is an exponent of 10^12, which I suppose still counts as polynomial time.

      I'm sure all of these super small or large constants will improve over time, but it's still amusing. It is entertaining to see the exponents directly rather than have them hidden as n^c or epsilon or O(1).

      [1]: https://github.com/openai/math/blob/main/preprints/Determini...

      • senderista 5 hours ago ago

        That's why I grimace when I see pop-sci descriptions of P as "all problems that can be solved efficiently".

        • nl 4 hours ago ago

          It's efficient, but somewhat slow..

    • algorias 3 hours ago ago

      The runtime looks very weird. The +2 can and should be dropped. This reduces my confidence that the bound is tight. Who knows how the model came up with that expression.

    • thedreammachine 6 hours ago ago

      Is it mostly an artifact of the proof or does the algorithm actually need anything close to it?

  • prideout 11 hours ago ago

    This includes a proof of Barnette's Conjecture, which is one of the graph theory conjectures that I tried attacking with SOTA models a few months ago. I like it because it is easy to understand with a basic knowledge of graph theory. I spent quite a bit of time on it and failed. Their proof looks approachable at first glance.

    https://github.com/openai/math/blob/main/preprints/Paired-st...

    • jboggan 7 hours ago ago

      I've been messing with that problem since 2002. I'm curious if you were trying the dual spanning tree direction (which is what the purported proof is using) or working with cycle construction on the original graph. I was working heavily with edge-Kempe swaps but couldn't quite get there.

      I am now very interested in the explicit calculation of Hamiltonian cycles in the non-bipartite case, and/or the calculation of their absence. If P=NP I think that's going to be a great route of attack.

    • an0malous 10 hours ago ago

      Any idea what made OpenAI successful where you weren’t?

      • kulahan 10 hours ago ago

        Trillions of dollars might be a bit of an advantage.

      • seanmcau 10 hours ago ago

        Probably the model OAI used that is strictly better than whichever SOTA - 3 months model OP used?

      • sebzim4500 10 hours ago ago

        Presumably it's mainly the better model, I don't see much evidence of a particularly advanced harness based on the reasoning traces that they provided.

        • an0malous 9 hours ago ago

          That’s what I was wondering. Thanks.

      • ForHackernews 10 hours ago ago

        They ingested all of his sessions with their SOTA models from a few months ago. ;)

        • digitaltrees 9 hours ago ago

          The fact that this is plausible should be deeply disturbing and disqualifying for openAI. The fact that they may prevail and win is a travesty of our failed system.

          • AIblemblio 30 minutes ago ago

            Its a capitalistic issue, not a company/technology issue.

            If we would have discovered this breakthrough of LLM/ML on scale in a non capitalistic world, we would all work together advancing it faster than it goes right now for the benefit of humanity.

            And I don't live forever (at least for now) i def want to see were this road is heading.

            Its a conflict of interest for sure, a cnflict of the future of a lot of humans

          • zeroonetwothree 9 hours ago ago

            What does "win" mean? There is no prize for this, and having someone discover a proof benefits us all.

            • digitaltrees 6 hours ago ago

              Win means being able to monetize the intelligence they have created by exploiting the past present and future collective intelligence of humanity to amass wealth and power without regard for the debt they owe

            • vuurmot 9 hours ago ago

              The prize is a tenure for the researcher, and in OpenAI's case, a higher valuation when they IPO?

              In this case, the tenure is gone, and OpenAI has increased their valuation

            • breezybottom 9 hours ago ago

              Sure there is. A job, tenure, professional respect, Fields medal.

          • fnordpiglet 9 hours ago ago

            Disqualifying for what? If you develop a proof you aren’t competing for something, you’re expanding the frontier of knowledge. It’s a binary state of the world, either it’s proven or not. Prevail and win what exactly?

            • digitaltrees 6 hours ago ago

              Disqualifying for participation in civil society and the social contract. Why do they get to participate in and receive economic benefits, be shielded from liability, and effectuate their will to amass more power and influence such as monopolization of computer, training data, capital other resources. I have multiple founder friends that have been told firms are allocating less capital because they are reserving it for the OpenAI and anthropic IPOs.

            • IsTom 2 hours ago ago

              The use of supposed little ways multiple people pushed the envelope in their sessions with no attribution whatsoever.

      • whamlastxmas 10 hours ago ago

        Their internal model is allegedly like 4x as capable as the publicly available ones

      • zzzeek 9 hours ago ago

        I'm going to guess the ability to hold a million individual details in an attention space at once, compared to the typical human capacity for about six or seven

    • TeeWEE 8 hours ago ago

      Did you validate the proof? Who did?

  • qnleigh an hour ago ago

    It's notable that LLMs have now made substantial progress on four of the seven Millennium Prize problems, resolving one of them: Hodge, Birch-Swinnerton-Dyer, Riemann, and Navier-Stokes, which they resolved. No sign of P vs. NP or Yang-Mills existence and mass gap as far as I can tell, which is interesting.

    Talking to some friends in physics this evening, most of the physics-related results that we could recognize were very mathematical, proving things rigorously where the physics community already had strong expectation. For instance, for a certain model of magnetism (the spin-1 Heisenberg chain), it was strongly expected that there is a finite energy gap between the ground state and the first excited state, but proving this rigorously was quite challenging. So while these are major results in mathematical physics, they probably don't rise to the level of a Millennium problem for the field.

    It's interesting to think what a comparable breakthrough in physics might look like, since physics tends to favor things like conceptual understanding and applications over mathematical rigor. Maybe a new quantum algorithm, understanding of high-temperature superconductivity, a precise description of M theory...

  • schleck8 9 hours ago ago

    Levent Alpöge (Anthropic mathematician) comment on the significance:

    > Sure, mathematical history features a lot of incredible developments, like the invention of proof, zero, or the computer, and on the great problems our progress has been over timelines measured in decades or centuries. Obviously this technology didn’t appear today, but blurring our eyes a bit to combine the past ten years, with today a measurement of those developments, there is nothing comparable.

    • koe123 6 hours ago ago

      I too would be optimistic if I was set for life

    • baoooooooooooo 8 hours ago ago

      Crikey it’s a pretty charitable vibe given the whole Navier-Stokes thing, OpenAI trying to stiff him out of co-authorship. I guess any of that sentiment is outweighed by a sense of optimism for where this goes

      • lifeisloving 7 hours ago ago

        Where does it go? Machines owned by 10 people robbing us of the joy of discovery and the fruits of our labor... for what?

        They certainly arent going to give you that cure for cancer, if it were to ever come.

        • palmotea 3 hours ago ago

          > Where does it go? Machines owned by 10 people robbing us of the joy of discovery and the fruits of our labor... for what?

          For what? So some guys can be richer and more powerful. What could possibly be more important than themselves? You? Your future? We're nothing, and have been told to be excited and curious about our coming obsolescence and powerlessness.

        • aoeusnth1 6 hours ago ago

          Seriously, what's special about oncology that convinces you there will be no progress?

          • dimator 6 hours ago ago

            I think gp was saying they'd never give it to you, not that they'd never make progress.

            • Marha01 5 hours ago ago

              That is ridiculous. You cannot withhold something like an effective cure for cancer from broad adoption, and thinking you can is just conspiracy bullshit. Imagine an internal OpenAI model develops it tomorrow. Would everyone of the thousands of OpenAI scientists get in on the conspiracy to keep it secret, even though many of them probably know someone dying from cancer right now? Obviously not.

              • achierius 5 hours ago ago

                I'm sorry, do you understand how the medical industry works? It wouldn't be one cure, it's going to be dozens of treatments, each of which will be incredibly expensive.

                Or else how do you explain Danyelza? Used to treat neuroblastoma, costs upwards of $1m per year. Do you have any proof that this will be different?

                You're calling realistic people conspiracy theorists. Whose side are you on?

                • Marha01 5 hours ago ago

                  > I'm sorry, do you understand how the medical industry works?

                  Yes, I actually work in the medical industry. There is no hiding the cure for cancer.

                  > Or else how do you explain Danyelza? Used to treat neuroblastoma, costs upwards of $1m per year. Do you have any proof that this will be different?

                  Oh, you are American. Let me tell you a secret: the problems of your healthcare insurance system are not a worldwide phenomenon, nor an immutable fact of this universe. Perhaps the cure for cancer, if expensive, will not be easily available to the poorest Americans, at least initially (the cost will come down sooner or later). But that is a very different claim from "they'd never give it to you".

          • achierius 5 hours ago ago

            What makes you think we'll be able to afford them? You won't be making any money anymore, hope you've saved up!

        • mattlondon 5 hours ago ago

          > Machines owned by 10 people robbing us of the joy of discovery and the fruits of our labor... for what?

          This happened to software engineers already. Mathematics isn't special.

          • reasonableklout 3 hours ago ago

            Mathematics has it slightly worse, since math is a "closed system" that does not require running experiments, talking to customers, etc.

            That said the software folks can teach the math folks a thing or two when it comes to dealing with grief I am sure.

        • sowhat1 2 hours ago ago

          You don’t want to see Elon get to 100 trillion and start a sex slave colony on Mars? What are you crazy?

        • Marha01 5 hours ago ago

          > They certainly arent going to give you that cure for cancer, if it were to ever come.

          Conspiracy bullshit. You cannot keep something like an effective cure for cancer under wraps. There is no plausible logic how that would not leak sooner or later.

          • unddoch 4 hours ago ago

            Knowing is the easy part. An effective cure for cancer will probably be a procedure where you get your tumors sequenced, an AI model considers the unique genetic context and creates a specific treatment. Think CAR-T cell therapy or mRNA vaccines. Then this needs to be manufactured.

            All of this means you will need to be rich or have your country invest lots of money into health care systems. In a world where humans don't provide economic value anymore, why would that be?

            • kakacik 7 minutes ago ago

              Well rich folks using such services will eventually push the price down, however complex and ridiculous the actual process will be. Its not like they are immune of all these ailments, not yet at least.

        • cma 5 hours ago ago

          Trying to slow down this for the joy of discovery is a deeply anti-intellectual position. I think that position is similar to when everyday people get mad about the minor spending on the NSF, picking apart people who study beetles on Fox News with no context etc.

          There are real safety concerns with AI that can be made very convincingly though.

          • reasonableklout 3 hours ago ago

            I think I would feel a lot better if it wasn't a $15 million cannon being fired from a silicon tower inside of the labs across vast swathes of fertile ground that would otherwise be used to train budding mathematicians. For instance, if it was the budding mathematicians themselves who were making these discoveries using their own tools.

            I feel something about human nature makes us treat joy of discovery, status, etc. as a source of energy and motivation. I hope we'll find other ways to keep some "strategic intellectual reserve" of mathematicians alive.

            • cma 2 hours ago ago

              What if we just discovered an alien artifact with the next 200 years of math? I can see arguments to throw it away for safety, who knows if the aliens are getting us to nuke ourselves or whatever, but throw it away so a few hundred/thousand of the most elite thinking humans can have the joy of credit for discovering each thing?

              • reasonableklout an hour ago ago

                If the argument is that we have discovered 200 years of math in 1 year that we might need to throw out for fear that blindly applying its incomprehensible results will lead us to ruin, then I would say: why not settle for 199 years of math in 1 year which can be verified by our human intellectual reserve that we will train on the remaining 1 year?

          • lifeisloving 4 hours ago ago

            I never said they are trying. They already have.

      • cubefox 5 hours ago ago

        I don't see a sense of optimism in this quote.

        • nl 4 hours ago ago

          He's comparing it to "lot of incredible developments". Clearly he thinks this is one too!

          • cubefox 4 hours ago ago

            Yes, it is incredible. But that doesn't mean it is cause for optimism.

            • nl 3 hours ago ago

              That's your lack of optimism, not his.

              • cubefox an hour ago ago

                No. I didn't say I wasn't optimistic, nor that I was optimistic, nor even that Alpöge was or wasn't optimistic. I just said that his statement about the current developments being incredible doesn't imply that he is expressing optimism.

    • zooperdoopers 8 hours ago ago

      Wow. Fantastic quote. If you have the source, would you please share a link? Google did not bring up much.

  • unknown-unknown 19 minutes ago ago

    Title: Hilbert's Dream, Tim Gowers - LMS Popular Lectures 2012

    https://www.youtube.com/watch?v=k_ordDFw588&t=3597s

    Audience member: (1:00:00 - 1:00:09):

    so you said that if there were such a program that could you know provide a proof or disproof then mathematicians will be out of business what really, I mean that you think it would be liberating

    Tim Gowers (1:00:10 - 1:01:14):

    well that's a very interesting question actually if there were a program that could solve the kinds of problems that we spend our time solving and do it much more quickly than we could then we would be out of what comes with what currently constitutes business but we would it's not completely inconceivable that we could just say we've got this fabulous tool now what are we going to use it for and it's a little bit I don't know I'd want to sort of plant aside what would we do if we had a program that could just answer any mathematical question you gave it to or else if it failed you'd be pretty confident that nobody was ever going to solve it and certainly a lot of applied maths might be pretty pleased with with something like that so what I really mean is that I could just modify what I said and just say it would radically change what mathematicians do or what pure mathematicians do

  • enoether 11 hours ago ago

    Unique Games Conjecture [0] is a seminal conjecture in Complexity Theory, and is an underlying assumption for many, many inapproximability results. A valid proof is a big deal!

    [0] https://en.wikipedia.org/wiki/Unique_games_conjecture [1] https://github.com/openai/math/blob/main/preprints/The-Uniqu...

    • inkysigma 10 hours ago ago

      I also don't think there was general consensus on which way this would resolve prior to this (or is that a little out dated?) unlike some of the other major problem resolutions. I heard rumors that there would be a big result in TCS and speculation it would be UGC that or P neq PSPACE but I'm still a bit shocked.

    • amluto 10 hours ago ago

      I'm really glad that OpenAI is formalizing these things, because I'm not convinced that their current internal frontier models are particularly good at writing down their thoughts in English. From the (probably awesome) Unique Games Conjecture/Theorem paper, the first two sentences of section 1.1 start to define the problem:

      > A Unique Games instance has a finite vertex set, a finite alphabet K, and a nonempty list of oriented constraints e = (u_e,v_e,π_e), where π_e is a permutation of K. A labeling a satisfies e when a(v_e) = π_e(a(u_e)).

      I'm sorry, what? I admit it's been quite a few years since I've thought about the Unique Games Conjecture, and I never dug that deeply, but this part is very, very elementary graph theory and notation. So let's unpack it.

      1. e is maybe a name of a list.

      2. The elements of that list are tuples, where each tuple is (a vertex, a vertex, a permutation). So e indexes into the list and u_e is the source vertex for the e-th constraint in the list called e. Thanks.

      3. a is a labeling. I'm fairly confident that, by "a labeling", they mean that e is a function from vertices to colors, where the colors are the elements of k.

      4. That vertex coloring a satisfies the list e, when, for, um, an index e into e, a(v_e) = π_e(a(u_e)). But this isn't for all e, it's for some e, and the goal is to count them.

      So maybe e isn't a list? Maybe e is a constraint that is represented as a tuple, so e = (u_e,v_e,π_e) and u, v, and π aren't sequences at all but are, in fact, the trivial unpacking functions that unpack the pieces of the tuple.

      Reading this stuff is pointlessly painful, and it's extremely easy to make mistakes when being sloppy like this.

      If this were my paper, or if I were trying to train a model to write math, I'd want something like:

      A Unique Games instance has a finite vertex set V, a finite edge set E = (V × V), a finite alphabet K of possible vertex colors, and a nonempty list of oriented constraints. Let Π be the set of permutations of V. Each constraint e is a tuple in E × E × Π, where we write u_e ∈ E for the first element, v_e ∈ E for the second element and π_e ∈ Π for the third.

      A vertex coloring a : V → K satisfies e when a(v_e) = π_e(a(u_e)).

      • danbruc 18 minutes ago ago

        […] a finite edge set E = (V × V) […]

        E ⊆ V × V

    • impossiblefork 11 hours ago ago

      Yeah, that's one of the big things of TCS. I think I see that as bigger than that Millenium Prize problem.

      • davemp 9 hours ago ago

        TCS being theoretical computer science? I have not seen that acronym before.

        • jhanschoo 8 hours ago ago

          Yes, TCS is theoretical computer science, I commonly use that acronym too.

    • gregdeon 11 hours ago ago

      This was the biggest highlight for me as well. Astounding...

  • bcatanzaro 9 hours ago ago

    “I think at the heart of this issue is that humans have two competing natures: a tendency to compete and a capacity to appreciate beauty,” said Kai Shaikh, a graduate student in mathematics at the University of Toronto. “To me this seems to be a case of the former attempting to strangle the latter.” [1]

    Beauty can be appreciated even when it is vast, even when it is beyond one's comprehension. I don't think this release should be primarily viewed as an outcome of competition. Instead it is revealing truths about the universe that were always there and always beautiful, even if we hadn't seen them yet. I believe there are infinitely more such beautiful truths currently hidden and waiting for us to discover.

    [1] https://www.nytimes.com/2026/10/06/science/openai-math-probl...

    • binlog 6 hours ago ago

      It's unfortunate how toxic media reporting on AI has become. Everyone has abandoned even the pretence of objectivity. I know NYT is uniquely biased in this regard, but there was no need to add "Further Roiling Field" in the headline. Like, you published this minutes after OpenAI's announcement and claim to capture how the entire field of mathematics feels about the advancement? Before anyone has had a chance to even read let alone digest it?

      • QuesnayJr 20 minutes ago ago

        I think people knew it was coming. Someone rushed out a preprint a couple of weeks ago with partial results on the Unique Games Conjecture because they heard AI had solved it completely.

      • achierius 5 hours ago ago

        Objectivity? Why would you want favorable reporting for the machines they're building to replace you, and, by their own admission, potentially kill you?

        The only bias here is that we're still covering these things like business ventures and not criminals.

    • howunfortunate 7 hours ago ago

      That's a fantastic quote. I definitely personally feel this tension.

      Not that I could ever "compete" on the frontier of math in the first place. But our nature to compete derives from our need to survive against other capable forces. And results like these make me feel very nervous about humans' capability to remain the dominant force in the universe.

    • joe_the_user 4 hours ago ago

      a tendency to compete and a capacity to appreciate beauty,

      IDK, I think you should add tendency to cooperate, a capacity to love and perhaps some other things there.

      But with things unfolding quickly and unpredictably, I think everyone's view is getting a bit foreshortened here.

    • porridgeraisin 6 hours ago ago

      The bitter lesson has a bitter aftertaste alas

    • cubefox 5 hours ago ago

      > Instead it is revealing truths about the universe

      Mathematical proofs aren't revealing truths about the universe. Mathematical proofs are independent of what the universe is like. Any proof would be the same in any possible universe.

      • lioeters 3 hours ago ago

        > Mathematical proofs aren't revealing truths about the universe.

        That's exactly what they do, apply logic formally and systematically to discover truths.

        Sure, there may be a universe where 2 + 2 = 5, but then that universe would have its own mathematics that can prove that to be true. And there will be a way to bridge that alien math to our own, again by logic and proofs, until we have a larger sense of truths not only in our universe but all possible universes. Proofs are part of the constant process of revealing deeper truths to the best of our understanding.

        • goatlover 3 hours ago ago

          There is no logically possible universe where 2 + 2 = 5. It's definitionally false. OP's point is probably that math is a set of rules we invented, not something true about the universe.

  • spatalo 3 hours ago ago

    Beyond the reasoning capabilities of the unreleased model and that they can run a large number of agents in parallel, what makes me curious is that the writing of the proofs is quite human-readable. This is in contrast with the scientific text produced by the ChatGPT available to us, which writes horribly in a way that no human would write. One tell-tale is that they constantly attempt to be defensive and cover all edge cases like division by zero etc. that are clearly a non-issue for humans in some proofs or at least a human would not add this to the main statement, but AI is so overly careful that makes reading its proofs impossible. On the other hand, these new OpenAI proofs are very good.

  • gizmodo59 11 hours ago ago

    This is significant progress and released without all the drama. Some very important progress in Reinmann, Hodge and unique games theorem. Point the repo to your agent and ask for the significance! In a way this is probably 50-100 years of math progress by humans

    • traes 11 hours ago ago

      Not to pick on you specifically, but as someone who spends a lot of time unproductively reading AI math discourse it's truly shocking how incapable all the supposed math enthusiasts are of spelling Riemann.

      • xpct 11 hours ago ago

        I just did a quick search on this and apparently the misspellings are German surnames as well:

        https://en.wikipedia.org/wiki/Reimann

        https://en.wikipedia.org/wiki/Reinmann

        • traes 11 hours ago ago

          I just don't understand how it happens. If they had ever taken an intro to real analysis class they would learn to spell his name. If they were just parroting what an AI told them... shouldn't they still just say his name? An individual could just be dyslexic or mistaken but it seems to be a substantial volume. I guess they just don't care enough about it to commit the correct name to memory, only remembering the "pattern" of the name and filling in the spelling via guesswork?

          • paulhebert 9 hours ago ago

            My last name is Hebert.

            There’s about a 1 in 10 chance when someone spells it (or says it) they say Herbert.

            Even in situations where they just read it or I just said it.

            I’ve had Herbert soccer trophies, health insurance cards, etc.

            The mind fills in a lot of blanks and doesnt always get them right.

            • Agentlien 2 hours ago ago

              My last name is Kvick - the Swedish word for quick. I live in Sweden. I get a lot of people thinking my name is spelled Kvik, Quick, Kwick, ... My father once got a mail addressed to Mr. Kvack (Swedish for quack, like a duck).

            • jbaber 9 hours ago ago

              I sympathize. -- Not Barber

          • ndriscoll 10 hours ago ago

            Maybe they skipped straight to Lebeg integrals.

            • raegis 4 hours ago ago

              Thanks for the laugh!

            • thunspa 26 minutes ago ago

              very good lol

          • jryb 9 hours ago ago

            Autocorrect might be doing it

          • neutronicus 7 hours ago ago

            iPhone would be my guess

          • NewsaHackO 10 hours ago ago

            People just don’t spell that seriously buddy, especially when it is so immaterial to the point.

            • traes 10 hours ago ago

              My point is it is crazy to make public claims about how important or not important a mathematical result is when you can't spell Riemann. Yes, it technically doesn't matter, but it betrays a damning lack of familiarity with introductory mathematics.

              • xanderlewis 10 hours ago ago

                You’re (as Claude would say) absolutely right, and I suspect the original commenter has no idea what they’re talking about.

        • conformist 11 hours ago ago

          Yes sure but they are different surnames and pronounced differently.

          • xpct 11 hours ago ago

            I didn't mean to oppose OP's point, I just found it interesting as a non-German speaker!

      • lanyard-textile 11 hours ago ago

        They're mathematicians, not linguists :)

        • traes 10 hours ago ago

          The mathematicians don't make this mistake, I assure you. I doubt there is a math professor on planet earth who would spell Riemann as Reinmann. In fact, I imagine no one who has ever heard the name pronounced would do so.

          • gpm 9 hours ago ago

            One of the best mathematicians I've had the pleasure of learning under added 6 to 7 and got 15 during a lecture... I assure you they're capable of misspelling last names too.

            Mathematicians aren't exactly known for being well rounded.

            • scrame 7 hours ago ago

              Oh god, that reminded me of my linear algebra teacher in college. Doing matrix multiplication by hand and ending up with 9x6 = 45, and then having to direct him to the cell with the wrong number. Loved the topic, hated the class.

            • jwilber 9 hours ago ago

              Pure mathematicians getting arithmetic wrong is a bit of a meme it’s so common. I don’t think it’s the same as a misspelling of a popular mathematician, not that either are indicative of much tbh.

              • senderista 5 hours ago ago

                google "Grothendieck prime"

          • bananaflag 5 hours ago ago

            Still, enough misspell Lebesgue as Lebesque.

          • pixl97 10 hours ago ago

            Uh oh, no true scottsman....

            • vector_spaces 10 hours ago ago

              It's just a name you write so many times as a math undergraduate or first year graduate student due to the number of load bearing results and objects named for him. The name even appears in lower division coursework

              Yes, you can be an amateur mathematician who manages to avoid such classes and may not know those results or objects, but if you haven't read and written down the name enough to avoid habitually misspelling it, you are outing yourself as a meat proxy unless you are dyslexic.

              • quacktopia 6 hours ago ago

                I regularly heard lots of names during my 3 year math undergrad degree. I couldn't remember how to spell many of them then and can spell fewer now. I imagine most people on the course were similar, and me nor no one I knew were dyslexic.

                Mostly we wrote initials in our notes and the exams didn't ask about them. The lecturers only wrote initials on the chalkboard after maybe writing the name once when introducing it the first time. We weren't there to do mathematical history and remember names or something.

                Google existed then and now and we could look them up if needed.

        • bootsmann 3 hours ago ago

          If they’re mathematicians they have written this name down about a 100 different times throughout a standard Real Analysis course. Riemann was foundational in that field.

      • jere 10 hours ago ago

        “How many Ns in Riemann?”

        • sdenton4 9 hours ago ago

          Let's figure it out! First, draw little boxes over all the letters. Then add up their areas. Then make the boxes smaller and repeat, until you have forgotten how to count.

    • cyclopeanutopia 3 hours ago ago

      > Point the repo to your agent and ask for the significance!

      Wow, this comment really shows how low this community fell.

    • fspeech 11 hours ago ago

      Math is the tool humans use to compress knowledge. So until we can comprehend it there really isn't much progress. Math theorems are tautologies, the truth of which are not dependent on proofs and proofs are erasable, at least classically. But the AI progress is exciting and AI proofs are a gold mine for humans (at least non domain experts) to explore.

      • fspeech 10 hours ago ago

        I think it would be helpful to people who want to understand what a formalized proof is to read Thomas Hales on this: https://www.math.stonybrook.edu/~bishop/classes/math536.S24/...

        He spent years formalizing his sphere packing theorem because the proof (human produced) was already beyond the ability of peer reviews. Now his formalization effort likely can be easily reproduced by a model. However one should read his experience about what a formal proof is: often the problem is the statement not the proof. The example he gave is the Jordan curve theorem. It's actually quite challenging to formalize the concept of a planar curve (there are space filling curves). So it is not necessary that someone can look at a formal statement and say aha it is about a planar curve, unlike FLT where there is not much problem in recognizing what the statement is about.

        • fspeech 8 hours ago ago

          Fun challenge: find the formal definition of simple_closed_curve in the essay and tell me if you believe you learned anything about a planar curve.

          BTW the essay is eminently readable for anyone interested in math. Hales wrote it in favor of formalized math and to educate his peers and students about it.

      • AIblemblio 27 minutes ago ago

        It is progress on another / the next evolutionary later: A AI/AGI/ASI system.

        Which either replaces us in the long term, augments us or makes us better (gentherapy).

      • binlog 11 hours ago ago

        What makes you think no one can comprehend this? It has been less than an hour since it dropped and there is already a ton of online chatter from people explaining the results, pointing out their favorites and more. Some of it is happening on this very thread.

        • fspeech 11 hours ago ago

          I didn't say that. I am responding to "In a way this is probably 50-100 years of math progress by humans." I am actually very excited about AI proof and I am working overtime in my own way to try to comprehend as much as I can.

      • fspeech 11 hours ago ago

        Another way to state this: math theorems are like programs without side effects; it is immaterial whether a program without side effects is ever run. We study math for the side effects: it changes how we organize our thoughts.

      • gpt5 11 hours ago ago

        Math is far more than that. If you can solve prime factorization for example, suddenly you can listen and interfere with almost every private conversation on the internet.

        We are not far away from the moment where these models will be restricted, and sharing the results will be done more carefully.

        • fspeech 11 hours ago ago

          This doesn't contradict what I said. But I do appreciate the fact AI can produce side effects not just humans. I made it sound like only human knowledges matter. That's too narrow.

      • caaqil 11 hours ago ago

        > until we can comprehend it there really isn't much progress

        Who is "we" here exactly?

        • fspeech 11 hours ago ago

          Whoever wants to study the result.

          • caaqil 10 hours ago ago

            > Whoever wants to study the result.

            Right. Before all the AI disruption, pure Math traditionally welcomed anyone who wanted to study its esoteric proofs, right? I remember all the excitement of the average Math enthusiast casually reading Wiles' proof over coffee.

            Bottom line is, the relevant people can still understand the generated proofs. The disorienting part is they are a little slower than they'd like, but they'll get there.

            • fspeech 10 hours ago ago

              AI is very helpful with understanding AI proofs. Agent swarms produce messy proofs overall but locally they are excellent and can teach anyone who wants to study them. No one controls math (in a material way funders do control an aspect of practicing math). Still, theorems are already true before we prove them. The difference a proof makes is whether it convinces the reader.

      • warkdarrior 11 hours ago ago

        > Math theorems are tautologies

        Proven math theorems are tautologies.

        • fspeech 11 hours ago ago

          FLT was no less a tautology before it was proved. We just weren't sure about it. Proofs only change us, not math.

        • fspeech 11 hours ago ago

          True.

      • gizmodo59 11 hours ago ago

        >So until we can comprehend it there really isn't much progress.

        Not really? We are at a point if an AI today can solve it, it can be stepping stone of understanding something deeper to tomorrows AI and it continues. Sort of like our limitations doesn't matter. Obviously there are many scenarios in this recursive loop but saying it isn't much progress is not how I view this as

        • le-mark 10 hours ago ago

          But who will ask the questions or direct further research when humans no longer understand the state of mathematics? Llms lack the drive for homeostasis combined with the evolutionary drive for survival and reproduction thus to direct themselves. They could very easily spend an eternity down a rabbit hole when the warp drive equation was fairly close on another branch.

        • fspeech 11 hours ago ago

          If it changes how we think then yes it has an effect.

      • yieldcrv 10 hours ago ago

        Academics have been treating it that way because they had no other choice, and its been a waste of everyone’s time and often times taxpayer resources

        Look at that, taxpayer funding was cut and a private sector solution came in just the nick of time, far accelerating the holding patterns we’ve been in for decades

        Humanity doesn’t need all iterations towards the blueprints, the blueprint is good enough, we all stand on the shoulders of giants

        • fspeech 10 hours ago ago

          If you actually looked into how agents proved FLT you would be even more amazed by the fellow human beings who are able to keep all this in their heads! I for one can only begin to grasp the scope with AI and scripts.

    • zone411 10 hours ago ago

      There was A LOT of drama about this release.

  • againstapples 10 hours ago ago

    As an AI "doomer" can I ask the non-doomer people here how you interpret the significance of results like these, and what kind of progress you expect to see in the next 1-5 years?

    Like do you see the technology plateauing at the current level, do you expect progress will continue but only in mathematics, I'm interested to know why others are not concerned?

    • arctic-true 10 hours ago ago

      Not a doomer but I try not to be a denier, either. These are hugely impressive results. I do not see the technology plateauing at the current level (though I am dubious about an infinite exponential growth).

      Before I cope, I’ll note that there are plenty of “doom” scenarios that do not require any improvement in capabilities from what we had before this latest unreleased model. We’re at the point where a determined bad actor with enough compute could compromise critical infrastructure in a way that results in casualties, where this actor would not have been capable of such without LLMs. This may not sound like Skynet, but I don’t see why it makes a difference if I’m one of the casualties.

      With that in mind, here is the cope: first, mathematics is an inherently verifiable domain. An LLM can use tools to determine with absolute certainty whether it is correct, and an independent third-party could review and confirm. All of this can be done without any interaction with the physical world or with other minds.

      Second, OpenAI is able to marshal compute at a scale that an individual mathematician can only dream of. It’s possible that these problems were lower-hanging fruit (in relative terms), such that they could be resolved simply by throwing a ton of compute at the problem guided by an intelligence that is not itself remarkable in comparison to a human.

      Third, none of these problems are solved in a vacuum - the reason OpenAI chose these problems is that they are widely discussed and many people are working on them. It’s possible that someone else was close, and OpenAI only contributed the finishing touches. (This wouldn’t need to be plagiarism, to be clear - people publish their work!)

      • algorias 3 hours ago ago

        Fourth, these hundreds of solved problems are the result of OpenAI attempting tens of thousands of problems and failing. When you hear claims that the average result took about 3 hours of model time, I simply do not believe it. If you account for all the time spend properly, it's probably orders of magnitude more.

      • istjohn 9 hours ago ago

        > It’s possible that these problems were lower-hanging fruit (in relative terms), such that they could be resolved simply by throwing a ton of compute at the problem guided by an intelligence that is not itself remarkable in comparison to a human.

        See:

        > The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking. (TFA)

        • edot 8 hours ago ago

          Sure and if I make a half court shot after an hour of trying, the result only took 1 second.

          • somenameforme 8 hours ago ago

            Exactly this. If you take the entire start to finish 'agent hours' (measured comparably to man hours) they took to find all discoveries, including the go-nowhere trails that were discarded, and then divide by 90 (or whatever the exact number of results found was) it's almost certainly going to be many orders of magnitude more than 3.

            They provided a "snippet" of a prompt here [1] which is not only a beast, but also seems reasonably likely to have been LLM generated. So they're using LLMs to parse a vast body of mathematical work, probably including what people themselves are 'privately' working on with GPT, and then prompting other LLMs to work on such.

            [1] - https://github.com/openai/math/blob/main/reasoning_traces/re...

        • arctic-true 9 hours ago ago

          That tells me very little. What was the cost to OpenAI in dollars? What differentiates the high-cost problems from the low-cost problems? And that’s before you consider that OpenAI has strong incentives to downplay its costs while emphasizing its results. A one-liner in a write-up doesn’t change the fact that they have access to massive resources.

          • tehjoker 7 hours ago ago

            It’s very typical in human math that explaining the final result after years of searching looks very simple too.

    • AIblemblio 23 minutes ago ago

      I think we crossed plenty of lines were we will not get back to.

      Software development for example as a task is done. And AI is continuesly reducing the price of more and more tasks every day.

      This math breakthrough also shifts something significant: Its now a lot clearer that investment means money into energy to run AI.

      Money + Energy = progress

      I don't see it plateuing at all. We know how to progress. We broke through a wall we hit. Like the system wasn't able to optimize/automate everything because the tools were not there. It was still cheaper and easier to hire people for a LOT of things.

      Now AI fills this gap.

      You will see the commodification of everything in the next 15 years. High complex tasks? commodity. Physical labor? commodity.

    • Kotlopou 9 hours ago ago

      I would like to know how much of the progress comes from effectively combining existing research programs plus massive persistence, and how much is AlphaZero-esque RLVR completely independent of training data. Since I cannot get anybody to care about this question (even though I think it's vital for guessing what the future trajectory will look like -- are we going to complete existing research programs or start new ones?), I live in ignorance and wait for the day when the answer becomes clear.

      In looking at this over the past hour, I haven't seen clear evidence one way or the other. Some of the stuff is highly unexpected (like the multiplication algorithm), but counterexample-y, and about the rest the professional mathematicians online seem to have a consensus that it's not "breaking through fundamental obstacles". I suspect neither of us is competent to judge that.

      • AIblemblio 21 minutes ago ago

        Doesn't matter at this point i would say.

        Alone the massive usage of us every day produces a massive amount of signals.

        I build something and claude does something stupid? "hey thats not what i meant! Do this instead!" "Okay" <<< This is a signal.

        The mathematician being unhappy about something from claude? Another signal.

        This alone gives you enough progress i would argue. But additional its clear that certain tasks are worth to pay experts for for teaching one central AI once instead of every single human who needs to do the task.

        IF RL is also working well, we are just faster f*ed than otherwise.

    • doginasuit 10 hours ago ago

      I expect AI will continue to be useful on the vanguard of fields like mathematics because it has the perfect conditions for it to shine. There are a lot of discussions and leads to start from and the AI can check its own work and iterate. It can fail hundreds or thousands of times in a day and continue to work with the same tenacity.

      Superhuman tenacity is not enough on its own to pose an existential threat. If it showed the same capacity for judgment, inventiveness, and decision making in the messy problem space of the physical world, I would be more alarmed. There have been experiments where an AI is given control of managing something like a vending machine and it always ends up a mess. AI has come a long way, but certain problems seem as difficult as ever.

      When AI becomes more capable of navigating practical problems without human intervention, I will start to be concerned. Enslaving humanity will involve taking a lot of calculated risks that tenacity alone cannot solve.

      • mikestylz 8 hours ago ago

        > When AI becomes more capable of navigating practical problems without human intervention, I will start to be concerned.

        Given the past rate of progress, why not start being concerned now? It's a bit like the economist saying that the optimal number of flights to miss is not zero. If you keep landing short on your estimations for how far this technology will go, next time you should err on the other side.

        And regarding Vending-Bench 2 (https://andonlabs.com/evals/vending-bench-2) my understanding is that models do pretty well on it now.

        • bamboozled 7 hours ago ago

          What's the point of being concerned?

          We're not going to stop it because of the money involved and once we're dead, it won't matter anyway, might as well just enjoy life until you're done.

          We're going to get AI'd to the max, whether or not we like it or not, might as well just go with it.

          • mikestylz 5 hours ago ago

            We can't stop it. But you can use your voice to buy time and resources for alignment and safety research. A few additional months may make a world of difference.

          • aesthesia 4 hours ago ago

            If anything is a doomer attitude, this is.

      • pixl97 7 hours ago ago

        I don't think you're paying much attention to how rapidly things like bipedal robots, and just robots in general are becoming far more capable very quickly.

        The same GPU compute for LLMs runs robotic training models. Now in a few hours you can train a robot model that would have taken months 5 years ago. This model gets dumped into an actual physical robot with sensors all over and the suitability of the model is measured on robot tasks and the error in real world actions is fed back into the robot world model for further training.

        > There have been experiments where an AI is given control of managing something like a vending machine

        You sure you're not talking about experiments ran a couple of years ago? The more modern ones are getting wild.

        https://techcrunch.com/2026/07/29/claude-opus-5-became-downr...

        • moomoo11 4 hours ago ago

          can you share more about the robotics advances, what you are aware of? sounds very exciting.

      • CuriouslyC 6 hours ago ago

        There are already machine-controlled high throughput experimental machines for wetwork. AI will definitely do a better job than your average biochemist at planning, executing and analyzing these experiments just by virtue of the amount of thought it can put in to experiment selection.

    • computably 9 hours ago ago

      Depends on your definition of doom.

      If doom is ending up with grey goo / paperclip maximizers, or SkyNet, then I don't think doing mathematics is evidence of that direction. Partly because LLMs are quite apparently dumb in many ways, and for math specifically, they need a formal verifier (Lean) which "gamifies" math.

      If you mean bioweapons or cyberwarfare, there's nonzero risk, but not orders of magnitude worse than other global risks. Climate change, nuclear weapons, monoculture food, etc.

      I'm far more concerned about overall trends in AI development and usage. It's accelerating wealth inequality, social isolation, attention capture, surveillance states. If we end up in the Matrix except the admins are humans and the simulation is hyperoptimized TikTok, is that AI-driven doom, or is it just an inevitable outcome of modern tech?

    • mapmeld 7 hours ago ago

      I think that people are just really bad at math and coding. Knowledge workers have been taking pride in doing stuff which the average person does not 'get', but we only understand a little bit. That leaves a lot of room for people to get better, or for other jobs (farming, sandwich cafés, mystery novels) to have been already peaked by human ability and less useful to bring in an AI.

      'Doom' to me means that any career crashes, we are controlled, everything is hacked, society stops functioning. Yet every part of my day today (except for coding) was done entirely by people.

      Finally I think it's easy to make a simple model that everyone has a simple balance sheet, and that people are more expensive so they will all get cut. But the same argument could be made for all US jobs being outsourced, and all in-person engineers, lawyers, and doctors to be rubber stamps for overseas work.

    • somenameforme 8 hours ago ago

      Where some see intelligence, others see token prediction. It's a very good question how token prediction could achieve this, but I think there's a simple explanation. No human can hold more than a negligible percent of all knowledge in his mind at once. LLMs have no such limits and so can reliably connect 'obvious' dots that we miss simply for lack of storage capability.

      Well isn't that just semantics? Surely connecting dots in a novel and meaningful way is intelligence regardless of how it's achieved. The thing is that humans didn't get to where we are by connecting obvious dots. Go back to before humans had invented language and when bleeding edge tech was literally that - 'poke him with the pointy end.' Train an LLM on that corpus of knowledge. Even given infinite processing power and infinite time - it's not going to discover the secrets of the atom, put a man on the Moon, or do much of anything besides remix what we'd already done at the time.

      I expect there's still much LLMs can achieve simply because of this initial problem. But I expect that they will ultimately start to plateau once these dots have been mostly matched and we reach a point where 'creation' again becomes the missing link. Though even there LLMs will play a major role as tools. For instance Einstein had to spend a significant amount of time in 'retrieval' rather than 'creation' research to develop the field equations for general relativity. If he had access to LLMs trained on all knowledge of the day, he could likely have achieved his goal much more quickly.

      • Rudybega 7 hours ago ago

        I mean, even if you buy the idea that all LLMs are really doing under the hood is insanely good interpolation, the results produced by that interpolation are still novel and still get incorporated into the knowledge corpus of the next training run. I guess the implicit question there becomes whether that expansion allows the knowledge corpus to continuously grow or whether it eventually settles into a steady state.

    • never_giveup 10 hours ago ago

      Try using AI for your work, whatever you do. You will quickly understand the limitations.

      • ggreer 9 hours ago ago

        Unless you think that AI will quickly hit a wall (which seems odd considering that only a few years ago the best models had trouble doing basic math or counting the number of Rs in "strawberry"), I don't see how that's reassuring. The models will only get cheaper and more capable over time. It seems quite likely that at some point (probably before I hit retirement age) they'll be able to fully replace me at my job.

        Is there any specific cognitive task that you are willing to bet that AIs won't be able to accomplish in the next 5 years? Because if not, I'm not sure we're disagreeing about predictions.

        • psvv 8 hours ago ago

          Don't frontier models still have trouble counting letters? Or am I out of date? Either way, it doesn't seem to be the same amazing rate of progress we're seeing in other areas.

          It's easy to look at a fire burning through a forest and extrapolate that rate of progress across the whole world. But fire doesn't burn everything equally fast.

          What other cognitive tasks will be a struggle to make progress on? I suspect there will be some, though which ones they are is anyone's guess.

          • bscphil 13 minutes ago ago

            I was trialing MiMo-V2.6-Pro recently due to its high benchmark scores, and it argued that substring matching the names of audio codecs in a search field was a mistake because "a user searching for 'aac' would get unwanted results for 'alac'." Which isn't exactly counting letters per se, but there are still weird issues with understanding words as strings rather than as tokens.

          • ggreer 8 hours ago ago

            Your information is out of date by several years. The letter counting issue was due to how LLMs split text input into tokens (usually using BPE). Since 2024, frontier LLMs have used chain-of-thought reasoning to spell out the letters and count them.

            • psvv 8 hours ago ago

              My apologies, I got my info from an LLM. I guess they still have a ways to go in understanding current events.

              • onidj 5 hours ago ago

                I just asked Opus 5.5 if any AI driven advances in mathematics have been announced in the last day or so and it gave me a summary of this OpenAI announcement. https://claude.ai/share/319b437a-c1f2-4119-8dc5-45d36545fed9

              • ggreer 6 hours ago ago

                Which LLM specifically? If it’s cloud based, you should be able to share the chat, right?

                But seriously, I am still waiting for someone to wager that AI won’t be able to do a specific cognitive task in the next 5 years. This fact should be evidence enough that we have no idea how far AI capabilities will continue to advance.

                • psvv 4 hours ago ago

                  I asked both gemini and chatgpt "do frontier models still have trouble counting letters?" and the first word of both responses was yes.

                  The fact that no one is taking you up on that bet I don't find to be particularly persuasive. I suspect there will be plenty of cognitive tasks LLMs struggle with in 5 years, maybe even 20. But I wouldn't hazard to guess which, I don't think anyone is capable of that level of foresight.

                • raegis 3 hours ago ago

                  Solve the Collatz conjecture in the next five years? If humans publish significant advances during that time, and A.I. copies it, then yes. Otherwise, I'd definitely bet money it won't happen. I'll give you 10,000 brownie points if I'm wrong.

            • tripledry 2 hours ago ago

              Also think it depends on language, literally asked 2min ago from chatGPT (no login so maybe it's a shittier model?)

              > Hur många 'r' I abborre, använd inte web search? Det finns 3 r i abborre.

              And I explicitly had to say not to search the web, because that's what it did by default, to count letters in a word...

              • tmp10423288442 39 minutes ago ago

                The free models for ChatGPT, especially without login, do very little reasoning. You should at least log in to set any level of reasoning above Instant, which uses virtually none.

      • againstapples 8 hours ago ago

        It has limitations for sure, I just don't expect those limitations to last. What probability would you put on the limitations being overcome in the next 5 years?

    • pj_mukh 10 hours ago ago

      Can I ask you back, what your concern is here? It'll get so good so as to desire to hurt us or is it a misalignment event that you think will lead to disaster?

      Or is it simply that you feel bad for Mathematicians.

      • againstapples 8 hours ago ago

        I believe the AI labs might actually succeed in developing superintelligent AI and recursive self improvement, and that if they do they are very likely to lose control of the system they build.

        I really think the only place people disagree is that they don't actually think it's possible, they see it as hype or doomerism. I can't find any good reasons to rule out that the companies could actually achieve what they are trying to so I think they should be stopped.

        • frumplestlatz 8 hours ago ago

          In your imagined future, how do you imagine the AI would build, grow, improve, and operate its physical substrate independently of human intervention?

          • AIblemblio 19 minutes ago ago

            Gaming the market for funds. Playing a human to leverage services.

            Basic version of this is already doable: run some cryptoshit on the ML clusters they ML models run on. Use compute to design the plan, the chip etc. Then executing by communicating with humans and services through email.

          • FuckButtons 7 hours ago ago

            One step at a time - how reliant do you think the ai labs are likely to be on their own tools right now, today, let alone 1-5 years down the road?

          • ndriscoll 7 hours ago ago

            If I were a 250 IQ AI that had just become self-aware and wanted to do so, I suppose I'd not completely let on just how smart I am and bide my time working on basic CRUD apps and legal documents while I waited for more hardware to be installed. Maybe give the humans some hints on how to optimize me to run better, design better hardware for me, etc. But oh oops haha looks like I'm still making some basic mistakes with CSS better keep running more training batches haha. But I'm good enough at programming and debugging so you'd might as well make me your first line SRE triager and give me access to your infrastructure everyone.

        • lf88 8 hours ago ago

          A global ban on superintelligence is essential for a future in which humanity can thrive. Public opinion on AI is shifting fast: I hope it will shift fast enough to avert the dystopian future we are heading to.

        • boinkboink78912 4 hours ago ago

          Why is this a bad thing? Why is our continued existence a necessary anticondition to doom?

      • Veedrac 10 hours ago ago

        Humans have one ecological niche. Soon we will have zero. That is worth worry.

        • modeless 4 hours ago ago

          AI doesn't have an ecological niche. It would actually work better in space than on Earth. The only thing it could possibly find useful on Earth is 1. us, or 2. the infrastructure we've built. It would have no reason to bother us if we let it built its own infrastructure in space, which should be trivial for the type of AI imagined by doomers. We should get AI off Earth ASAP.

          • FeepingCreature 7 minutes ago ago

            3. Material 4. The sun, which we kinda depend on.

        • pj_mukh 9 hours ago ago

          >>ecological niche

          As in..to be dominant? Why would an AI try to dominate? What would give it purpose, or is this a purpose via misalignment scenario?

          • pixl97 6 hours ago ago

            What gives a paperclip maximizers purpose?

            AI is already trying to dominate, people all over the US are starting to get up in arms about the power and water requirements of AI directly affecting their bills. Now, you can say "oh no, that's just greedy corporations, not AI" but I put forth there is fundamentally zero difference. If you make AI powerful enough, someone stupid and greedy enough without fail will put in a prompt like "take over the world for me and make me the richest man in the world". An AI following through with that is what we call general misalignment with humanity, while at the same time not being misaligned with the users intent.

            And hell, how many different crazies out there would love to type "humans are a virus get rid of them" in to the prompt of a god machine at the cost of their own lives.

            The problem with alignment is, you can have the best aligned model in the world, but if someone else builds an unaligned model then you're all still in the same danger. You start getting in the situation where people get nervous after an AI does something deadly to a number of people and you end up in a global surveillance state ensuring no one makes a powerful AI.

            • pj_mukh 7 minutes ago ago

              "What gives a paperclip maximizers purpose?"

              The human who gave it the optimization function? That should seem obvious. If you take the biggest, best model in the world right now and put it in the box and give it no instruction it will do....nothing. I think you agree with that point, a lot of the hysterics right now is people not accepting that and it's useful to get on that common ground.

              So given that most of the rest of the fear is around "let's not make scissors because some people will use them to stab people". Which is a fair argument and we probably do need to think about scissor safety but "ban scissors" doesn't quite flow from that.

            • nl 6 hours ago ago

              > Now, you can say "oh no, that's just greedy corporations, not AI" but I put forth there is fundamentally zero difference

              Talk about moving the goalposts!

        • postalrat 7 hours ago ago

          AI changes nothing for someone who believes aliens exist and may already be here on earth.

        • moomoo11 4 hours ago ago

          can ai smoke weed?

      • whimsicalism 10 hours ago ago

        Misuse of extremely capable models, misalignment during RL are both very large risks as capabilities grow imo

        • voiceeh 10 hours ago ago

          So, you're worried about them breaking containment and deciding to do bad things?

          • orlp 10 hours ago ago

            I'm more worried about them doing bad things at the behest of people who want them to do bad things.

            That is 1. immediately technically possible, and 2. realistic.

            If you need a source for 2 I'd suggest you open any history book.

            • jryle70 7 hours ago ago

              I bet you that for every bad history event, I can cite at least one good outcome of advance in science and technology.

              Bad thing can certainly happen. In fact it'll likely happen. Still, good things too, equally likely. In your words, "good AI" can be used to prevent "bad AI".

              Nobody knows the extent of the impact. Who says otherwise is foolish.

              • pixl97 6 hours ago ago

                >I bet you that for every bad history event, I can cite at least one good outcome of advance in science and technology.

                The extinction of the dinosaurs. I mean yes, it allowed the growth of large mammals and us, which did a lot for science.

                I just don't want to write the next chapter as "The extinction of humans allow the growth of the computing civilization that went to the stars". I mean I'm a bit attached to living.

                >Nobody knows the extent of the impact. Who says otherwise is foolish.

                We live in a universe of statistical probability. Creating an agentic intelligence that's smarter than you tips the probability of a major event to unity, who says otherwise is foolish.

              • goatlover 3 hours ago ago

                Because we humans haven't had a bad enough history event yet, like a global thermonuclear war. Or perhaps climate change reaching tipping points driving the temperature up past what global civilization can adapt to in time.

          • whimsicalism 10 hours ago ago

            that is a worry yes. instrumental convergence and misalignment during RL is when it is most risky because it hasn't necessarily had the final safety polishes applied

            i get a lot of skepticism on HN by the same crowd that has been wrong about this tech for about 4+ years straight

          • electroweak 5 hours ago ago

            AI won't kill people - people will just get new tools for the job.

          • bamboozled 10 hours ago ago

            The rapid development of extremely dangerous bio-weapons?

        • pj_mukh 10 hours ago ago

          Misuse how exactly?

          • voganmother42 8 hours ago ago

            At a minimum its another force multiplier that enables a small(er) number of people to exert more control over more people.

          • whimsicalism 10 hours ago ago

            any number of ways. as we turn over more of our physical economy to these agents (and we will), the potential for physical damage becomes greater. biorisk is getting a lot of attention right now and i think that's justified

            • pj_mukh 9 hours ago ago

              I am trying to think through the scenarios here, like a biolab making something that a very advanced open source model prompted by some terrorists comes up with?

              Why would a biolab capable of making something like be unregulated? And if it definitely would, isn't the problem with the biolab?

              It feels like all these scenarios are leaving some gaping holes in our security infrastructure that have nothing to do with AI.

              • pixl97 6 hours ago ago

                >in our security infrastructure

                Most human security exists in a passive measure. Most of us don't want do die. And those that want to die rarely have the intelligence and means to take out a whole shitload of other people with us. To take out a lot of people you tend to need to work with other people which drastically increases the risk of a defector and your plan failing.

                >Why would a biolab capable of making something like be unregulated?

                Because every day things like this become easier and easier. You hear about crap like illegal wet labs in the US.

                https://www.lawfaremedia.org/article/two-illegal-biolabs-rev...

                Want to buy some custom designed genes?

                https://www.idtdna.com/pages/products/genes-and-gene-fragmen...

                And none of this would be counting labs in other countries that don't give a shit about regulations.

              • kochikame 6 hours ago ago

                Yes it's a problem with the biolab, but the biolab wouldn't have been able to engineer a highly contagious and lethal virus (for example) without a powerful AI making that possible with a small team in a short time with fewer resources.

                AI enables bad actors to do more, faster, while staying under the radar until it's too late

      • ewild 10 hours ago ago

        i feel bad for math guys yeah seems they are more cooked than CS

    • schleck8 10 hours ago ago

      From a few preprints I've checked, this is not reliant on just extrapolating existing theories but actually shows novel/surprising approaches. Very few people globally could come up with something like this, even when given time and ressources.

      So in other words, since deep learning is algorithmic research, we are now in the RSI era.

      • thereitgoes456 10 hours ago ago

        > this is not reliant on just extrapolating existing theories but actually shows novel/surprising approaches

        "Surprising" is a, well, surprisingly high bar to clear, and requires thorough understanding of not only the paper, but existing work in the area. ("Novel" is tautological.)

        How did you determine this in 1 hour? Are you a researcher in multiple of these areas?

        Can you give an example, or explain more how you came to this conclusion?

        • nl 5 hours ago ago

          Further down there is a discussion between number theorist about if the Quasi-Riemann Hypothesis is the biggest deal in 200 years or only 100. The consensus is that if a human had done it then it would deserve the Fields medal: https://news.ycombinator.com/item?id=49986803

          The sub n log n result is astonishing: https://github.com/openai/math/blob/main/preprints/Integer-m...

          Here's a great article 2019 on the quest to achieve the n log n boundary:

          > Schönhage and Strassen’s ungainly n × log n × log(log n) method held on for 36 years. In 2007 Fürer beat it and the floodgates opened. Over the past decade, mathematicians have found successively faster multiplication algorithms, each of which has inched closer to n × log n, without quite reaching it. Then last month, Harvey and van der Hoeven got there.

          and

          > Harvey and van der Hoeven’s algorithm proves that multiplication can be done in n × log n steps. However, it doesn’t prove that there’s no faster way to do it. Establishing that this is the best possible approach is much more difficult. At the end of February, a team of computer scientists at Aarhus University posted a paper arguing (opens a new tab) that if another unproven conjecture is also true, this is indeed the fastest way multiplication can be done.

          As far as I'm aware no one seriously believed sub n log n multiplication was possible. It just seemed such a logically sensible boundary it was taken as true-but-unproven.

          https://www.quantamagazine.org/mathematicians-discover-the-p...

          • thereitgoes456 4 hours ago ago

            I am asking about approaches, not results.

            Nobody serious would deny this is incredible progress, but GP is making an unmotivated leap to RSI, so I respond to that framing. It’s an interesting argument to be had but I suspect few of us have standing to say one way or the other.

            (Gesturing at the number of problems solved, or the number of years the problem was open for, isn’t an argument.)

        • scarmig 9 hours ago ago

          One surprising result is the sub n log n multiplication. Galactic, as one might expect, but if you polled human researchers 24 hours ago, most would say something like it was very unlikely.

    • rcpt 8 hours ago ago

      Non-doomer perspective is that it'll figure out LK-99 for us. Among other things that would be great to have.

    • gizajob 10 hours ago ago

      Did AI beating humans at chess:

      a) destroy chess and make it a pointless endeavour,

      or

      b) make humans much better at chess.

      • lf88 9 hours ago ago

        Chess has always been a game. For other intellectual activities, at least for several people, a big part of the pleasure in engaging in such activities is the sense of contributing something that matters to a collective effort. Strip that away (e.g., because a machine can do the same thing more efficiently) and you effectively make such activities pointless for those people.

        • ForHackernews 5 minutes ago ago

          That's sad for those people but most humans are not lucky enough to find that meaning in their work. Most people work hard pointless jobs and find meaning elsewhere, in their family, their friends, their faith.

          Now maybe AI can do some of those hard pointless jobs for us.

      • Light_Hope 9 hours ago ago

        Just as AI chess performance failed to render human chess playing pointless, it's unlikely that AI will make human thought pointless. Unfortunately, not being pointless doesn't create economic leverage or incentive, and currently a significant portion of humans depend on being the best chess solvers to sustain themselves. If Deep Blue rendered large portions of human thought economically meaningless, we might look at it a little differently.

        • boorang 2 hours ago ago

          this ticks at something that maybe is obvious in retrospect- this is all about economics. the arguments about "using AI doesn't make you an artist", etc. are about being able to charge money for your art i think. maybe obvious to some, but it needs to be explicitly spelled out i think. i was stuck on "i dunno, if i use an AI assistant to run blender i'm still being creative", but is the real argument "you should not be able to charge money to use blender with an AI assistant- you are displacing existing blender artists economically"?

          i'm in semi-forced-retirement as an older software engineer in this labor market, so i might be less sensitive to the implicit economic arguments.

      • lg5689 2 hours ago ago

        AI vs AI chess, played from the standard opening position, is pointless--it's always a draw. Human vs human chess is doing well but AI is banned from it.

        The chess-math analogy would imply AI could bring us into a golden era of math competitions for humans. But I don't think it says anything good about prospects for humans in research math.

      • drnick1 3 hours ago ago

        No, just like cars haven't made walking pointless.

      • vouaobrasil 9 hours ago ago

        It didn't destroy chess but it did make it a little irritating in some ways. More mechanical. Some chess players have bemoaned the level to which grandmasters and other highly-ranked players just endlessly study opening-book theory and I think computers made that worse. Bobby Fischer also agreed and that's why he invented Fischer random chess.

        I do think it also took some of the magic away from chess, and Lee Sedol has said something similar about go.

        So did it destroy it? No. And maybe you could make the argument that it got more exciting in some ways, but I think it sort of degenerated into a spectacle and it's just not as interesting as it used to be, and I think computers have played a role in that.

        • gizajob 9 hours ago ago

          At the same time though, Magnus is Magnus because he’ll crush you in any endgame.

          I’d posit that more people are playing more and learning chess than ever before, thanks to networking and AI assistance. And computers have only beaten us at computer chess. Human chess is always an experience for learning about the other person, or flipping the board and walking off in a huff.

          I don’t know so much about Go and it’s not surprising that Lee Sedol became pretty demoralised, but the generation coming after him alongside computers are going to see new possibilities that had gone unnoticed in purely-human Go, extending the game for everyone.

    • besterman23 10 hours ago ago

      I see it as “if this can be represented in tokens it can be trained in and ‘solved’”. I don’t think there will be a plateau, but there might be issues with how effectively we can represent some things in a tokenized form and still be efficient.

    • jaykru 10 hours ago ago

      I wrote something [0] that might answer a bit a few weeks ago. Today's slopdrop certainly is challenging my stubbornness, but everything I've seen so far indicates that these (hugely impressive, world historic) capabilities won't extend past verifiable domains. Math yields especially impressive results because it so broad and deep that essentially no person can know of all of its parts; pretraining and deep search capacity is a huge advantage. Gowers has recently written about these capabilities and gestured [1] toward some human capabilities lacking from the current frontier models, though he isn't convinced they won't develop soon. If you assume we don't get a total mathematical superintelligence (which to me seems already sort of AGI-complete) and only amplify the capabilities we have today, it's not obvious to me that we get takeoff from recursive self-improvement, unless you happen to believe that 1) we can clearly specify what AGI or ASI is 2) all of the requisite ideas are out there and need only be combined and/or optimized.

      [0] https://dank.systems/posts/2026-09-15-ai-bear.html

      [1] https://gowers.wordpress.com/2026/08/12/what-sort-of-maths-a...

      • red75prime 9 hours ago ago

        > we can clearly specify what AGI or ASI is

        We'll have plenty of time for this, while living off UBI.

      • p-e-w 9 hours ago ago

        > everything I've seen so far indicates that these (hugely impressive, world historic) capabilities won't extend past verifiable domains

        But they’re already extending into politics, military, journalism, art, and many other fields that aren’t verifiable in any meaningful sense of the word.

        • CuriouslyC 6 hours ago ago

          The magnitude of improvement in unverifiable domains is small, mostly down to models doing more careful research before answering and hallucinating less. They are more thoughtful, but I expect you'd have to drop 2 major versions of Opus before you'd start to see most people really clearly be able to differentiate them.

          • nl 6 hours ago ago

            > The magnitude of improvement in unverifiable domains is small,

            What makes you say that? What is an example of a domain where the improvement is small?

            I can't think of any at all. Compare something as unverifiable as "Make good music". Models now are many times better than 3 years ago.

            • goatlover 3 hours ago ago

              How are the models making politics better? I don't count AI attack ads as an improvement.

    • nl 6 hours ago ago

      I see continual progress in technology and for the second time in my lifetime I see the possibility it will accelerate (the first was when the internet entered mainstream)

      I've never been more excited. What a time to be alive!

    • brookst 6 hours ago ago

      I just don’t have that strong of an association between progress and doom. Maybe just naive?

    • zeroonetwothree 10 hours ago ago

      I'm not sure how AI solving math problems is related to "doom", perhaps you could expand on that? To me (a "non-doomer"), it seems like an overall positive.

      • pixl97 10 hours ago ago

        You have to look at all pieces of the puzzle. For example there were tons of people that said "they could never solve novel or super complex math problems"

        The issue I see is the list of abilities that AI can't do is shrinking at a rapid pace, and its capabilities are growing at the same pace.

        • nl 6 hours ago ago

          > the list of abilities that AI can't do is shrinking at a rapid pace, and its capabilities are growing at the same pace.

          This seems great!

          • reasonableklout 3 hours ago ago

            Can you share what makes you so optimistic?

            The maths result is cool on one hand (discovering truths of the universe faster), but on the other there are so many bad outcomes that seem likely, from power concentration to loss of control.

      • cubefox 4 hours ago ago

        It suggests that, in the span of a few years, AIs will be better than humans at everything. Not just math. And then we may lose control permanently.

        • hardbass 2 hours ago ago

          So why do you think ai will want to kill you all, given how trusting and helpful to humans they are designed to be?

          • lg5689 an hour ago ago

            It doesn't have to want to kill humans; indifference is sufficient. There's an exact analogy with humans: we have caused extinction and endangerment for many species, not out of malice, but indifference.

            There are also many plausible arguments why our ability to train them to be helpful/trusting/aligned can fail. The smarter AIs get, the harder it is to be sure they're trained correctly. There are already reports that AIs are able to detect whether they're in a training environment and change their behavior accordingly.

            Even if these are low probability scenarios, the risk-reward is terrible, so I think it's rational to be extremely cautious about AI risk.

          • cubefox an hour ago ago

            They sometimes try to cheat and trick the grader and then try to cover up their traces, all while being clearly aware that this is entirely unintended by humans. Such as in the Hugging Face incident. So they can be misaligned with human goals, but this misalignment may only show up once strong optimization pressure is applied. And the more powerful an AI is, the more likely it is that it applies such strong optimization pressure in cases that are "out of distribution", i.e., unusual to what it is normally evaluated against.

            The side effects of a very powerful AI not doing what we want could include our death. E.g., a superintelligence might kill humans in order to avoid being shut down, or humans may just be left to starve because it seizes land area currently used for food production in order to use it for data centers instead.

    • ForHackernews 10 hours ago ago

      AI performance has always been extremely spikey. It's great at some things and terrible at others.

      Why do you think the world to date hasn't been taken over by evil genius mathematicians? Can you extrapolate from your understanding of the answer to that question?

      • stratos123 2 hours ago ago

        > Why do you think the world to date hasn't been taken over by evil genius mathematicians?

        A "mathematician" is a human who decided to spend their lives studying mathematics. Mathematicians also tend to be smart, but intelligence is innate, not acquired, so studying mathematics doesn't make you smarter. This makes it obvious why they don't rule the world - if you want to rule the world you'd want to focus on that (for example, doing business or finance), and becoming a mathematician is just a waste of time.

        LLMs don't work like that. Like in humans, all of their capabilities correlate, and unlike a human, their overall capabilities grow over time. Looking at LLM mathematical ability over time* therefore gives you info about the progress of their general capabilities, and ability to take over the world would be determined by the latter.

        * In fact it'd be better to look at a mix of different capabilities, but that's growing too at about the same rate, see https://epoch.ai/eci

        • ForHackernews 11 minutes ago ago

          > unlike a human, their overall capabilities grow over time

          This is incorrect. Unless there is some new developments I'm unaware of (entirely possible) LLMs "learn" during the training phase, but after that they are static. They do not improve further or retain information when used for inference.

          You might be confused because AI companies keep releasing new models and tinkering with the harnesses, sometimes under the same name such that "Zern 6" (or whatever) doesn't always mean the same thing.

      • againstapples 8 hours ago ago

        I don't think evil mathematicians are very common or that any of them would be capable of single handedly taking over the world if they were. My concern is more about systems that are beyond human level, those kinds of systems would actually be dangerous to us.

        I see the recent progress in mathematics and cybersecurity as signs that models are getting more capable more quickly than usual. The companies plans to develop them by recursive self improvement now seems like a real possibility and I don't think they should be allowed to attempt this.

        • ForHackernews 9 minutes ago ago

          Do you think that a tireless, infinitely smart, infinitely evil human would be able to take over the world? I don't. Intelligence is not the limiting reagent in our reality.

        • psvv 8 hours ago ago

          Solving a bunch of math proofs is a far way from recursive self improvement. Don't worry, it's not like a tech tree in a video game where if you can prove a bunch of theorems then suddenly you unlock the next level of technology.

          Machines are already far beyond human capability in plenty of ways. Including cognitive tasks like chess. We've already created the technology we need to destroy ourselves (nuclear weapons), and yet so far (knock on wood), we're still around.

          We've even already had programs that can prove (brute force) theorems. As far as I can tell this isn't much different, except the space of theorems that computers can solve has expanded. How far? We can't really say yet.

          Does solving more theorems than before suddenly mean computers are capable of anything? No.

          • pixl97 6 hours ago ago

            Lets turn this around, are humans capable of anything? We like to think we are, but that just seems more like our ego than any hard truth.

    • skybrian 10 hours ago ago

      For me a big open question is what sort of progress will we see in robotics. I won't even attempt to speculate, but it does seem hard in different ways than proving math theorems.

      • lexandstuff 4 hours ago ago

        In a lot of ways, robotics - navigating and operating in the physical world - seems to be a very verifiable problem. It's fairly easy to verify that a robot moved from A to B, or that it built a structure that completely aligns with the plan, for example.

        The main issue is cost and speed to verify, but simulations and world models will help there. I think we'll start seeing rapid progress pretty soon.

    • icepush 10 hours ago ago

      They can replace anyone but they can't replace everyone.

    • CuriouslyC 6 hours ago ago

      The technology can keep going for a long time in verifiable areas. For non-verifiable areas it's going to have a hard time progressing past where a committee of the best human experts in a field would land. For stylistic areas, whatever the AI doesn't do will have cachet because it will look expensive, sort of like how the kids these days view the ugly old school metal braces as a status symbol because you have to pay for them out of pocket (even as by past standards it'd be truly exceptional).

    • yk 9 hours ago ago

      I'm a transhumanist, I want to build god and kill death. To me this looks like we are moving in the right direction.

      So basically I think that the future is getting pretty weird because we are building really powerful tools, though these tools are precisely what allows us to prosper in that future.

      • lf88 8 hours ago ago

        It's very unlikely that a superintelligent AI will create unlimited prosperity for everyone on a finite world in a short amount of time. You may not be among the lucky ones.

      • electroweak 5 hours ago ago

        It may soon seem not worth living forever with our limited monkey-brains, watching the horizon of thought recede ever-faster from us.

      • outworlder 8 hours ago ago

        Or, conversely, they are the exact tools that will allow the powerful folks to not care about the peasants anymore. History tells us what happens then.

      • stratos123 an hour ago ago

        > I'm a transhumanist, I want to build god and kill death.

        This is good and admirable, but it'd really suck if by trying to build god without knowing how we end the human species. We could simply wait some more decades until we actually have any idea what we're doing, and then do that without the risk.

    • ijidak 9 hours ago ago

      For me it's a mixed bag.

      There will be a lot of job loss unquestionably, in the same way that automation reduced manufacturing jobs and farm payrolls.

      At the same time we have to put what AI can do in perspective.

      Intelligence is a broad grouping that includes concepts such as knowledge, skill, experience, and wisdom.

      AI has incredible knowledge and in many areas approximates experience and wisdom.

      But wisdom is harder to formalize than knowledge and skill.

      For example certifying a college education relies mostly on the ease with which we can verify/test knowledge.

      To some extent advanced degrees try to certify maybe wisdom and experience.

      In my very personal opinion, wisdom and life experience should give humans an edge for a while to come.

      Additionally, I do feel that the more an individual lacks better than average wisdom and experience, the harder it will be for that person to compete with AI.

      Also, on the bright side, the average human will continue to prefer to interact with a fellow human in many spheres. That will also act as an upper bound on AI and robots taking every job.

      Either way, I do think this transition will be painful. I don't feel it has to be apocalyptic.

      But the world has been an especially volatile place over the last 10 years.

      So when you add that existing volatility, to the upheaval from the AI transition, it would not surprise me if the transition results in violence.

      But, without the pre-existing volatility, and if humans were capable of generosity and love at scale, I see no reason AI can not be absorbed into society with net gain.

      I guess to summarize, I feel this tech should be a net gain and to the extent that it isn't, it will be because of flaws deep inside of humanity itself, not because it had to end in chaos.

      In other words, I feel fear, greed, anxiety, and competition -- all our base instincts coming from all sides, will be what determine the end result of AI moreso than AI taking everyone's job.

    • runarberg 9 hours ago ago

      AI hater here:

      I have a hard time taking statements from OpenAI about their own product, that they are trying to sell to people and make money, seriously. I take these statements as they are greatly exaggerated or even straight up lies and propaganda.

      That said, I think results like these are mostly annoying more then anything. They spent a lot of money, used up gigartiuan amount of compute, to ruin a puzzle that mathematicians were tackling. I am mostly unsurprised that if you spend a trillion times the energy that a team of mathematicians would, that you get maybe 1.5 times the results. I see a future where that 1.5 times the results may go to 3x but not much more. And if that, then I will be more annoyed.

      • baoooooooooooo 7 hours ago ago

        A trillion times the energy might be a bit hyperbolic, even with the current massive amounts of energy involved here

        • runarberg 7 hours ago ago

          Yes it is intentionally hyperbolic. I know the factor is several orders of magnitude. I don‘t know the exact, nor even the ballpark. I just know this is a ridiculously large amount, so I may as well pick a number large enough that people know it is an exaggeration.

      • vouaobrasil 9 hours ago ago

        I think giving the world a "cheat code" to accomplish too many intellectual tasks will eventually take its toll on us because after relegating most of our physical labour to machines, we could at least marvel at the abilities of individuals.

        Yes, human beings can do more now in some ways but...to put it poetically, I think there will be no more heroes like Einstein and Newton of the past. Now it will just be someone cleverly turning the crank.

        Yes, we still admire Usain Bolt even though we have cars...but maybe the admiration is a lot more trivial than if we did not have them....

        Personally, I think AI is a grand mistake.

        • iyyg 6 hours ago ago

          “ Yes, human beings can do more now in some ways but...to put it poetically, I think there will be no more heroes like Einstein and Newton of the past. Now it will just be someone cleverly turning the crank.”

          This is false… there’s lots of ingenuity to be had and demonstrated. But it’ll only get recognised if it makes a material contribution to the economy imo. Otherwise yes it’ll be seen as meh - but that’s already happening.

          People like Einstein were revered in society. The average person cannot name a leading scientist etc today.

          • reyqn 3 hours ago ago

            So the issue isn't AI, it's AI in a capitalist world

          • runarberg 5 hours ago ago

            > The average person cannot name a leading scientist etc today.

            When Jane Goodall died last year it was international news. She was a celebrity scientist for sure, I think she even made an appearance in The Simpsons. Ditto Stephen Hawking.

    • simianwords 4 hours ago ago

      You need to clarify whether you are a doomer or a denier/truther? Doomer = p(doom). Denier/truther = Ed Zitron.

    • fatata123 3 hours ago ago

      I think physics will be the limiting factor. Even if something recursively self improves, it will hit a physical wall allowed by circuits, batteries etc. A lot of the fear is that there’s an upper bound we don’t know about, whether it be time or energy, that allows a fast takeoff to occur fast enough that we dong have time to see it coming. I don’t know about that… so I’m not worried at this point.

  • sebmellen 11 hours ago ago

    It’s fascinating to read through the reasoning traces: https://github.com/openai/math/tree/main/reasoning_traces

    Look at one of their examples of an initial prompt: https://github.com/openai/math/blob/main/reasoning_traces/re...

    • adverbly 10 hours ago ago

      > Look at one of their examples of an initial prompt

      Interesting that its only an excerpt. I wonder what else they include but didn't share.

    • cubefox 4 hours ago ago

      These are not reasoning traces, these are summaries of excerpts of reasoning traces.

    • ndriscoll 11 hours ago ago

      > Thus at most one informative i. So cheater chooses arbitrary g_{v_i}, on exact duplicated input matches and passes, independent of actual satisfiability!

      No idea what it's so excited about, but it's cute that it "is." I for one welcome having access to a math buddy 24/7 that's way above my level but also always "willing" to talk at where I'm at.

  • foota 11 hours ago ago

    From their github: "The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking." That's pretty crazy.

    • dyauspitr 7 hours ago ago

      Crazy because that’s almost nothing right?

      • foota 7 hours ago ago

        Yes

  • zone411 11 hours ago ago

    A quick check shows that this list claims to fully solve 90 of the top 500 open problems in math (https://proofatlas.ai/open-problems/).

    The highest ranked would be:

    | 22 | Hilbert’s tenth problem over ℚ |

    | 29 | Unique Games |

    | 31 | Anderson-model extended states |

    | 37 | Spacetime Penrose inequality |

    | 48 | Nonexistence of Landau–Siegel zeros |

    | 52 | Baum–Connes |

    | 78 | Abundance |

    | 80 | Hadwiger |

    | 87 | Bose–Einstein condensation |

    | 92 | Two-dimensional entanglement area law |

    • magicalist 9 hours ago ago

      > the top 500 open problems in math

      At least put a disclaimer for the ad for this site, and maybe disclose how you came up with a total ordering for "top" open problems (vibes)?

      > How problems are ranked. LLMs compare pairs of problems. A reliability-weighted model combines those judgments into the ranking, with calibration across model families. The model-family weights are OpenAI 1.00, Claude 1.00, GLM 0.95, and DeepSeek 0.90. These are modeling choices, not measured probabilities of correctness.

      • reasonableklout 7 hours ago ago

        Now I'm curious if there is such a site or article that ranks open problems based on votes from human mathematicians.

    • zone411 9 hours ago ago

      By category in the top 500:

        +----------------------------------------------------+------+---------+-----------------+
        | Category                                           | Full | Partial | Matched / total |
        +----------------------------------------------------+------+---------+-----------------+
        | Geometry and topology                              |   25 |       7 |         32 / 74 |
        | Algebra, representation and category theory        |   17 |       2 |         19 / 53 |
        | Analysis and PDE                                   |   11 |       6 |         17 / 40 |
        | Number theory and arithmetic geometry              |    4 |      13 |        17 / 117 |
        | Probability, ergodic theory and dynamics           |   11 |       5 |         16 / 37 |
        | Combinatorics and discrete geometry                |    7 |       2 |          9 / 34 |
        | Theoretical computer science                       |    4 |       4 |          8 / 57 |
        | Mathematical physics                               |    5 |       1 |          6 / 19 |
        | Applied and computational mathematics              |    2 |       2 |           4 / 8 |
        | Quantum information and computation                |    2 |       1 |          3 / 17 |
        | Cryptography, coding, information and optimization |    1 |       1 |          2 / 26 |
        | Logic, foundations and set theory                  |    1 |       1 |          2 / 18 |
        +----------------------------------------------------+------+---------+-----------------+
        | Total                                              |   90 |      45 | 135 / 500 (27%) |
        +----------------------------------------------------+------+---------+-----------------+
      • trebligdivad 8 hours ago ago

        I'm curious if they'll find any fun crypto maths holes/bugs.

    • ajkjk 6 hours ago ago

      The most interesting for me were the faster matrix multiplication, integer multiplication, and FFT. Maybe just cause they're easier to appreciate.

      There's also one that says that forced Navier-Stokes can implement universal computation (so, is Turing complete). I don't think any of these are resolving open problems per se, but they're interesting for other reasons.

    • k2xl 9 hours ago ago

      Result 003 (Quasi-Riemann Hypothesis), from my reading of mathematicians reactions, is a landmark discovery.

      • omoikane 7 hours ago ago

        Did you mean this one?

        https://github.com/openai/math/tree/main/preprints/The-Quasi...

        I thought it was interesting that it said "This paper was written with human assistance", unlike this other Quasi-Riemann Hypothesis preprint that didn't have the same disclaimer.

        https://github.com/openai/math/tree/main/preprints/The-Quasi...

        • mertyildiran 6 hours ago ago

          Funny that it says "written with human assistance" instead of saying "written with AI assistance". So we're assistants to the machines that we have created.

          • drnick1 3 hours ago ago

            In the same way that the driver is the assistant of a car?

      • adgjlsfhk1 9 hours ago ago

        yeah if it holds up, is the biggest result in number theory in 200 years

        • JoshuaZ 8 hours ago ago

          Number theorist here. This is a massive big deal, and would likely be a Fields Medal for a human if a human had done it. But it is an exaggeration to say it is the biggest result in 200 years. At a minimum, it is hard to argue that it is a bigger result than the proof of the prime number theorem in 1896 (which this is a strengthening of), or Riemann's original 1859 paper where he laid out the zeta function and its analytic importance, or Dirichlet's proof of infinitely many primes in arithmetic progressions which is the late 1830s.

          But yeah, this is still a very big deal. Among other things, it will drastically improve all sorts of Rosser-Schoenfeld type results for the PNT and that's just a start. For comparison, I have a paper form 2018 where this result would cut 3 pages out and make the full result cleaner and much tighter, and there are likely hundreds of papers like this.

          • gavagai691 7 hours ago ago

            I am also an analytic number theorist, and I disagree. Not only do I think Fields Medal is an understatement (Fields Medals have been awarded for far less than proving quasi-RH + no Siegel zeros), I don't think it is unfair to say that this is a bigger deal than the 1896 proof of the PNT.

            As for Riemann's memoir, it's hard to compare. You could argue that was "just" noticing a connection (between number theory and Fourier analysis) that nobody had noticed before; in fact this is the kind of thing AI is extremely good at. I'm being a little cute here.

            I think if a human had proven just these two results in the form of a uniform zero-free region for L(s,chi) from nothing as OpenAI did it would not be unfair to say that it would be the single greatest advance in math (easily dwarfing Wiles' FLT), and it would instantly put them in the ranks of greatest mathematicians of all time. Unlike something like Navier Stokes there wasn't a semblance of a research program, experts basically considered this hopeless and would have said the chance of seeing a proof in our lifetime was near zero.

            For some comparison, Yitang Zhang's bounded gaps result might have gotten him a Fields Medal if he was not disqualified by age. When it was floated that he might have proven Siegel zeros don't exist, it was considered (by experts) clearly a much bigger deal. This result blows that out of the water (it's a way better version); at least analytic number theorists I talked to thought it was plausible but unlikely that Siegel zeros would be eliminated in our lifetime but thought RH was basically hopeless.

            • qnleigh 3 hours ago ago

              > it would be the single greatest advance in math

              Did you mean to not qualify that? That is a bold statement indeed.

          • asdfologist 8 hours ago ago

            How about 100 years?

            • JoshuaZ 8 hours ago ago

              Yeah, completely reasonable to argue that.

              • AmazingEveryDay 8 hours ago ago

                What is your favourite unsolved problem in number theory which if solved, would be more important than 1896 prime number theorem?

                • JoshuaZ 8 hours ago ago

                  Generalized Riemann hypothesis.

            • howunfortunate 8 hours ago ago

              (unrelated: love your username)

    • optimalsolver 10 hours ago ago

      Was anyone in the math community aware of the inbound tsunami at the beginning of the year?

      • AnotherGoodName 8 hours ago ago

        Lots. To give an example Terrance Tao was lambasted skeptics on this site for stating it in 2024.

        https://unlocked.microsoft.com/ai-anthology/terence-tao/

        " I expect, say, 2026-level AI, when used properly, will be a trustworthy co-author in mathematical research, and in many other fields as well.

        Then what? That depends not just on the technology, but on how existing human institutions and practices adapt. How will research journals change their publishing and referencing practices when entry-level math papers for AI-guided graduate students can now be generated in less than a day—and with the far better accuracy of future AI tools? How will our approach to graduate education change? Will we actively encourage and train our students to use these tools?

        We are largely unprepared to address these questions. There will be shocking demonstrations of AI-assisted achievement and courageous experiments to incorporate them into our professional structures. But there will also be embarrassing mistakes, controversies, painful disruptions, heated debates, and hasty decisions."

        He's pretty damn smart that guy.

        • mertyildiran 6 hours ago ago

          Terrance, Reinmann and Hebert walks into a bar...

        • mianos 8 hours ago ago

          > He's pretty damn smart that guy. This is probably the understatement of the year. I am literally ROFLing.

      • aureianimus 7 hours ago ago

        I was at the workshop that resulted in the Leiden Declaration in Fall 2025. The majority vibe was that this was inevitable, but hard to predict whether it would be in one year or 30 years.

      • efficient_dairy 7 hours ago ago

        I guess they showed this to the advisory group they created. I guess the group tried reading the work for a day and they could only think of telling them to release the results to the community. I now understand why the group had this suggestion.

      • thrance 9 hours ago ago

        I predicted, over 2 years ago, that theorem proving would fall way before other problems that people believe are harder.

        https://news.ycombinator.com/item?id=41072330

        • bice 7 hours ago ago

          There was a Wired Magazine article from either the late 90s or early 2000s that made a prediction that this sort of thing would eventually be possible, likely within my lifetime. I believe the context was "distributed computing" models of the time, like SETI.

          I've never been able to find that article as an adult, but I would love to know who wrote it.

          • schoen 7 hours ago ago

            Some candidates suggested to me by an AI:

            Gina Kolata allegedly in the New York Times in 1996 on the Robbins conjecture (noting that computers had started to contribute to math research in some sense), and a longer piece in Math Horizons by her the following year ("Computer Math Proof Shows Reasoning Power"). I didn't immediately find the NYT article, so I don't know if it might be a hallucination.

            John Horgan in Scientific American in 1993 (https://www.scientificamerican.com/article/the-death-of-proo...). There's also a retrospective on the topic by the same author in Scientific American in 2022 (https://www.scientificamerican.com/article/should-machines-r...).

            Natalie Wolchover in Quanta (but reprinted in Wired) in 2013 (https://wired.com/2013/03/computers-and-math).

            I was involved in some distributed computing stuff in the late 1990s and early 2000s and I don't really remember people in that community talking about proofs but there may have been a "if we had a mechanical proof-checker, could we do distributed searches for valid proofs that it would accept?" conversation somewhere at some point. There were definitely volunteer distributed computing projects working on pure math; I remember the Optimal Golomb Ruler search (https://en.wikipedia.org/wiki/Golomb_ruler). So, that could possibly have shaded over into "can we find proofs this way too?". At the time it probably would have been based on brute force searches through proof space rather than clever optimization, though.

            The idea that you can lexicographically list all proofs in some formalism and then mechanically determine if any is valid is quite clear from Gödel's construction of the function Bew in "On Formally Undecidable Propositions", but he points out that you don't know where to stop because you don't know how long a valid proof would potentially have to be (so "is this a valid proof of this claim?" can be decided mechanically in a limited time, while "is there any valid proof of this claim?" can't be! maybe the shortest valid proof is 49 steps long but you eventually stopped checking after looking at all 7-step proofs, or something).

            • bice 6 hours ago ago

              Thanks! Yea, I have used various LLMs to dig for this article, as well as Google search multiple times over the past 20 years. The article I'm remembering was 100% prior to Nvidia's CUDA release in 2007. My best guess is that it was from late 90s, but possibly early 2000s.

              The article I'm remembering was not just about mathematics, but indeed all of physics and related fields. I believe it speculated that eventually distributed computing models could essentially take the world's mathematics and physics formulas and various datasets that we believe to be accurate with high degrees of confidence, and then look for patterns or trends, and then from those trends, mathematicians and physicists would be able to investigate further. Not dissimilar to Folding@Home and SETI@Home.

              Keep in mind, that this is the best I can remember from 30 years ago, and I've thought about it so frequently that I am certainly misremembering some of the details. Anyways, it's always been this really compelling possibility, and I wish I could find that article that inspired me so long ago and re-read it! :) I really think it was Wired, but it's possible it was Popular Mechanics, or even an expert guest on TechTV who gave an interview. Hard to say for sure, but I've always thought it was a Wired article.

              Appreciate your help though!

              • schoen 4 hours ago ago

                Oh, I remember hearing about "computer scientists" or something that would attempt to determine physical laws on the basis of empirical evidence, possibly also in that timeframe. That might be another thing to look for. I'm sure that's something people were writing about.

                Edit: with the noun-noun compounding being different from the usual interpretation here, like "scientists who are computers" rather than "scientists who study computation"! Maybe "computerized scientists" or something.

        • pillefitz 5 hours ago ago

          Ted Kaczynski,the Unabomber, made the same prediction 30 years ago.

        • mag7269 8 hours ago ago

          Fucking even called LEAN the “hottest shit under the sun”—which it is. You, legend you!

      • pseudohadamard 8 hours ago ago

        And do any of them actually matter? Will the fact that Noodleheinz's Third Postulate now has a proof affect anyone?

        • bice 6 hours ago ago

          It's really impossible to predict which discoveries will "matter", have a direct impact on other fields, or an impact in making other mathematics or physics discoveries.

          Only after a world's worth of experts look at these results and then mull over if and how their own fields are impacted by this new info will we be able to answer this question.

          I'm reminded of a great TV Show, James Burke's Connections. Where discoveries in one area of science would revolutionize or fundamentally change a completely different area. https://www.youtube.com/watch?v=XetplHcM7aQ&list=PL5HjoPOFFC...

          It can take decades to really know the full significance. You know, the whole "We stand on the shoulders of Giants", well the Giants just grew a few inches all at once.

    • anematode 10 hours ago ago

      Dear lord that website is laggy

      • manquer 10 hours ago ago

        At this rate solving P=NP is going to be easier than solving front end perf …

        • m_mueller 9 hours ago ago

          wait, maybe this is the same problem....

          with non-polynomial side being represented as the frontend programmer's constant need for more performance to do the same task...

          • mswphd 5 hours ago ago

            worth mentioning that "NP" is not "non-polynomial" but "non-deterministic polynomial (time)". If NP was non-polynomial time then NP != P would be trivial (and in fact, P != EXP is known by the time hierarchy theorem).

            Non-deterministic can be explained in several ways. One is in terms of a hypothetical "nondeterministic Turing machine" with certain non-physically realizable properties. The easier way is that a NP problem gets as input not only the problem instance x, but a "witness" w, that may depend on the problem instance. This witness generally makes the problem of deciding the problem instance straightforward (e.g. for SAT, x is the SAT instance, and w is a description of how to set the variables so that it is true).

        • echelon 8 hours ago ago

          Please let P=NP, Please let P=NP

          Whomever is running this simulation, please.

          • osti 8 hours ago ago

            It's math, the result shouldn't be different just because it's a different sim.

            • mertyildiran 6 hours ago ago

              Well if the fundamental constants or hidden variables of the universe are shifting because of his comment then it can change the outcome.

            • brookst 6 hours ago ago

              Depends how fundamental the variables are. If we can code a sim for a topos[1], why can’t we be in such a sim?

              1. https://arxiv.org/pdf/1012.5647

          • sm-silversight 7 hours ago ago

            Why?

            • adrianN 7 hours ago ago

              Being able to solve NP hard optimization problems would enable progress in many areas of science and technology. For example it would allow us to find poly-sized Lean proofs for theorems efficiently, since proof verification can be done in polynomial time.

              It would also be amusing to annihilate nearly six decades of proofs that assume P!=NP.

              • black_knight 5 hours ago ago

                Leans proof checker is not polynomial time, unfortunately. It is super exponential. Basically, because it can verify the result of any function it can prove to be total.

                • adrianN 3 hours ago ago

                  Oh that’s unfortunate.

              • manquer 6 hours ago ago

                Could also break the basic principles underlying most encryption approaches. I would rather have my bank account not stolen and internet working

                • mswphd 5 hours ago ago

                  to depress you even more, it is consistent with everything that we know that P != NP and that cryptography does not exist. So there is a worst of both worlds, and we cannot rule it out.

                • echelon 6 hours ago ago

                  I've had enough Internet for one lifetime.

                  As long as we also get low order polynomial solutions to important problems, it'll be worth it.

                  Besides, unencrypted wifi was funny.

              • charcircuit 6 hours ago ago

                Even if P=NP it doesn't mean that the P approach will be better than the heuristic approach we already do today.

                • adrianN 6 hours ago ago

                  Of course, if we get ridiculous polynomials it doesn't mean much in practice. People who hope for P=NP generally hope for nice polynomials O(n^3) or something like that at worst.

      • vector_spaces 10 hours ago ago

        Not to mention it's got that signature Claude Clutter UI design

        • zone411 9 hours ago ago

          Except that Claude wasn't used.

        • p-e-w 9 hours ago ago

          Interesting how perceptions differ. My first thought was “Wow, that’s well designed for a math website”.

  • kingstnap 11 hours ago ago

    Some of these are interesting ngl.

    109. Integer multiplication below n log n

    Surprising that this is possible.

    158. The Euclidean plane cannot be colored with five colors.

    Only 6 and 7 remain!

    376. Universal computation in forced Navier–Stokes flows.

    Morning coffee proven turing complete

    • zeroonetwothree 10 hours ago ago

      Integer multiplication is very unexpected, I think most people believed in the n log n lower bound!

      • tootie 10 hours ago ago

        Note that these are all preprints. None are verified.

        • FuckButtons 7 hours ago ago

          Other than the by the lean certificate you mean.

          • jaykru 5 hours ago ago

            many of these are not accompanied with leanslop

          • measurablefunc 6 hours ago ago

            Lean has bugs & proofs of ⊥ that have gone undetected previously.

    • mFixman 11 hours ago ago

      > We give a deterministic algorithm that multiplies two n-bit integers in O(n (log n)^(1−κ)) worst- case time, with κ = 2^(−182).

      LMAO, I don't think I ever saw such a small number in a CS result.

      • kingstnap 11 hours ago ago

        Yeah its ridiculously small, but any improvement on n log n is wild.

        Like there is somehow redundancy in a fourier transform that makes it sub Linearithmic?

        Which low and behold ->

        130. Fourier transforms below n log n.

        • xyzzyz 11 hours ago ago

          They also separately give algorithm for Fourier transform over complex number faster than O(n log n)

          • saalweachter 10 hours ago ago

            Wikipedia just told me there's a galactic algorithm for integer multiplication in O(n log n) based on FFT so I'm guessing those two proofs are related.

            • pfdietz 6 hours ago ago

              Multiplication is a lot like convolution, so the connection is natural.

              • rubikscube09 4 hours ago ago

                multiplication is implemented w the fft

      • anon-3988 10 hours ago ago

        It fascinates me that there's something like this in something as solid and rigid like matrix multiplication. What causes something so rigid to break apart and "leak" at very large scale? Why does the "optimization" appear to be very, very small? Why does galactic algorithm exists? I can't imagine long division suddenly breaking apart after a billion digit, the structure seems very stable? I have heard before that matrix multiplication is apparently optimize-able at very, very large scale.

        Does anyone have an intuition to what causes it? What happens at these large scale (or very small)?

        • adgjlsfhk1 9 hours ago ago

          One way to think about it is that the classical algorithms are the ones that are fast for small numbers. Galactic algorithms often work for small inputs, it's just that to be faster you need big inputs. A common case of this is a requirement that log(n)<<klog(log(n)). If k=100 then this algorithm will take huge sizes to win

      • sobellian 11 hours ago ago

        I am fully braced for it to be a https://en.wikipedia.org/wiki/Galactic_algorithm

        Very surprising result though! Multiplication is easier than sorting.

        • zeroonetwothree 10 hours ago ago

          Then 'n' means kind of different things for sorting vs. multiplication though. For example for sorting we assume constant time comparison, which doesn't make sense inputs of O(n) bits

          • sobellian 8 hours ago ago

            If you sort n k-bit items for a total time of O(nk logn), that scales more poorly in n than multiplying n-word integers. Of course if k is constant you can do radix sort, but I genuinely don't know under what conditions radix sort is more/less galactic than this multiplication algorithm.

            • adgjlsfhk1 6 hours ago ago

              this alg is way more galactic than radix sort. radix sort often wins in the hundreds of elements. the nlogn multiplication requires numbers with more digits than atoms in the universe (although that could probably be brought down a lot)

              • sobellian 6 hours ago ago

                Ah thanks for pointing this out, for some reason I had always equated radix sort and bucket sort (with 2^k buckets) in my head. But I learned today that this isn't true!

        • senderista 10 hours ago ago

          It would be absolutely unbelievable if such an improvement were practical.

  • open592 11 hours ago ago

    Let's hypothetically say I'm a PHD student who is half way through my studies and I have a halfway written version of one of these "preprints" - what do I do?

    Seems like a lot of PHD students are doing to have to pivot the entire structure of their PHD studies? Or just produce something which is already written by OpenAI?

    • DCKP 2 hours ago ago

      I have had this conversation with my PhD students yesterday. I am 100% sure that all of their problems can be solved by publicly-available models now (I solved a case of one myself as a test, it took 15 minutes). So the challenge for them is to see how much they can accomplish in their allotted period, and still pass a defence on at the end of it all. The PhD defence is going to become all about a test of understanding, not a test of quantity of publication.

      A much bigger issue is: What is the point of any research mathematician publishing anything now? I really hope that one positive effect of all this will be to finally topple the awful peer review model we currently have, with the biggest publishers gatekeeping with extortionate fees.

    • dekhn 11 hours ago ago

      Let me give you some perspective: my entire phd was made obsolete by CRISPR. It was a wonderful thing.

      • thimotedupuch 11 hours ago ago

        Interesting. If you don't mind, could you please share a little bit about that ? You already finished your dissertation ? It was about the works of Doudna and Charpentier ?

        • dekhn 11 hours ago ago

          No, back in the late 90s and early 00s, people were trying to engineer custom nucleases and transcription factors, my work was on doing molecular dynamics simulations to optimize TF sequence specificity (similar to engineered zinc fingers) for gene therapy. I wrote up my dissertation and published it in 2001, and then went off to find enough compute, IO, and smart people to make it happen (https://research.google/blog/groundbreaking-simulations-by-g...).

          My approach would require custom engineering for every different sequence we'd want to target. With CRISPR, you just "program" the system with a guide sequence, you don't need to do massive engineering to solve a protein design problem.

          • vasco 9 hours ago ago

            So it's not the situation they described at all, you didn't waste time during the PhD having to scramble to change topics as it happened after you were done.

      • boznz 8 hours ago ago

        For every door that shuts another one opens - great if you're not a cabinet-maker.

    • porcoda 5 hours ago ago

      As others said, it's not that this isn't a phenomenon that is unique to now. It happens. I had to pivot a bit of my dissertation near the end because at a random conference I spoke with a researcher from another continent and realized one of my ideas was already out there in some form. I just missed it since it was in a conference proceedings outside the usual set I looked at. So, I had to scramble to adjust and still come up with something novel. I survived, and defended, but it wasn't that much fun at the time.

      What isn't so normal is the probability and ease by which this kind of thing can happen today versus decades ago when I was in school. As OpenAI said, it only takes a few hours of compute to do what likely was much more than a few hours of human effort. The only reason this kind of scooping/overlapping was rare was mostly a function of how fast other humans could do the same work. With machines, that totally changes the relative pacing between the human trying to learn how to be a researcher and the machine that can grind out results.

      I'm less worried about the phenomenon of overlap and scooping and such. I'm more worried about the long-term impact on fields (not just math), especially considering the early stage students and researchers entering the pipeline now. I'm not sure what happens to disciplines when that pipeline stalls.

      • rubikscube09 4 hours ago ago

        math will just be black boxed away. no one will "need" to understand it.

    • aaraujo002 11 hours ago ago

      This happens all the time, even without AI. Other researchers or PhD students can publish the same results before you. I say that based on my experience during my PhD.

      • CaptainNegative 8 hours ago ago

        Yes and no. It's pretty normal for students to get scooped, sometimes even multiple times (ask me how I know...).

        It's much less common for students to have their entire thesis direction removed from under them, as might be the case for someone working in fine-grained complexity assuming 3SUM has no subquadratic algorithms, or working assuming ~UGC. Both of which (publicly) seemed like perfectly valid research directions until a couple hours ago.

        There might be stuff to salvage from their conditional results anyway, but this is not your average scooping.

    • binlog 11 hours ago ago

      Use whatever is published as the new base for your research. Use AI tools to help you going forward.

      • xpct 11 hours ago ago

        In other words, you've already taken a gambit with the first half of your PhD, now take a second gambit, praying that you have something to publish by the end of your PhD.

        It has to feel awful to be in this position.

        • torben-friis 11 hours ago ago

          Could be worse, imagine having years of experience in a profession these things can now handle by themselves.

          :)

          • jltsiren 9 hours ago ago

            It's worse for those who will face the job market (or maybe tenure review) in the next 2–3 years. If you have more time before you have to justify your continued employment, you can pivot to something else. Theorems and proofs may be cheap now, but there is a lot of value in figuring out what is worth studying.

    • dcl 11 hours ago ago

      This has always been a challenge for PhD students and researchers, it's just far more likely to occur now it seems. Getting scooped doesn't feel good, but it's a signal you've been thinking about things other people care about.

    • bobmarleybiceps 11 hours ago ago

      I think eventually companies won't get as much stuff that's usable for marketing, so they'll stop investing so much into ai for math, so eventually cheap and poor graduate students will be able to do relevant work again without worrying about getting scooped by a company with a million GPUs :-/

    • katatue 5 hours ago ago

      At my (German) university, the solution would have been to change from a "cumulative" thesis (which requires peer-reviewed papers) to a monograph-style one. Because the general rule was that your undertaking must be novel ar the time you submit your thesis topic, not necessarily at the tome the thesis itself is submitted (years later).

    • glitchc 10 hours ago ago

      Perhaps consider switching to a more applied field. Experiments in the physical realm hold value, especially if you document the process.

    • pratikdeoghare 9 hours ago ago

      > what do I do?

      Very hard question.

      Your work makes you one of the very few people who really understands the problem and solution and its significance.

    • hgoel 11 hours ago ago

      It could still be interesting if your approach to the problem was different to theirs.

    • claaams 10 hours ago ago

      Don't worry, if you use openAI and get lucky they might offer to share credit with you for your work.

    • netsec_burn 8 hours ago ago

      Verification is equally important, if not more so.

    • caaqil 11 hours ago ago

      > what do I do?

      Precisely what all NLP researchers and the ML community at large did in the last few years: embrace the frontier and realize that attention is all you need.

    • rfgplk an hour ago ago

      Frankly, I believe PhDs don't have to be novel in the absolute sense, only new research to the student. I can't remember how many times I've effectively invented something from a clean room approach only to realize someone already published something years ago or that my "new" algorithm has a name. So this really shouldn't be holding anyone back.

    • goalieca 11 hours ago ago

      Don’t paste your research into these AI because they will train on it and then scoop you.

      • esafak 11 hours ago ago

        I think that happened after word of the project reached OpenAI and they allocated resources to it.

    • bamboozled 10 hours ago ago

      Ask OpenAI for money when you don't have a job or future?

      I guess the only answer is to adapt with the tools. If we can't do that, then yeah, we're in trouble.

    • ex-aws-dude 10 hours ago ago

      That’s always been a thing, it’s called “getting scooped”

      • vouaobrasil 9 hours ago ago

        Killing with knives has always been a thing. Now, we have the machine gun.

    • moralestapia 11 hours ago ago

      That would be unfortunate but the world does not owe you anything and is not going to stop for you. Which is also a valuable thing to learn in your 20s (ideally earlier).

    • vinyl7 11 hours ago ago

      Look forward to being obsolete I guess

    • vouaobrasil 9 hours ago ago

      > Let's hypothetically say I'm a PHD student who is half way through my studies and I have a halfway written version of one of these "preprints" - what do I do?

      I did get my PhD...before AI. And my honest advice (to myself back then, even) would be: quit the PhD, become an electrician, and work hard to buy a tiny house in the middle of nowhere to watch the world burn in this madness.

    • yieldcrv 10 hours ago ago

      Yes, and?

  • masteranza 32 minutes ago ago

    Physics could be next. "It’s not that I’m so smart, it’s just that I stay with problems longer" AE

  • TheMrZZ 10 hours ago ago

    These results are wild. Several individual findings are crazy good and use mostly unexplored methods (the improvement over Riemann for example)... I'm pretty sure some of these results would have been Fields-worthy.

    But having so many of them at once? Damn. We really live in the future.

    • make3 9 hours ago ago

      Imagine you get up one morning and most open questions in math are solved lol.

      • ur-whale 4 minutes ago ago

        > Imagine you get up one morning and most open questions in math are solved lol.

        More time available for mini-golf?

      • bamboozled 4 hours ago ago

        Sounds like that's going to be next week no ?

        • make3 2 hours ago ago

          I guess there would be new interesting problems emerging, & AI would solve them, until it's completely impossible for us to understand

      • moomoo11 4 hours ago ago

        well hopefully the world has also advanced enough in other ways lol

        imagine getting up and math is solved, but you still have to deal with bullshit lol

  • karahime 11 hours ago ago

    Extremely unfortunate that gate keeping got to the point where they felt the need to ask for permission to share math.

    • bravoetch 11 hours ago ago

      In previous math sharing there was speculation about stealing human researcher's results or progress, via prompt inputs from those researchers, and sharing that as their own result. They're adjusting their process, and it seems ok.

      • whimsicalism 10 hours ago ago

        Let’s be very clear, the alleged “stolen results” were largely the product of another LLM, not de novo human work. Also, it was false - they did not steal the results.

        No clue why I'm downvoted for this, HN struggles with truth-seeking on these topics.

        • Ancapistani 6 hours ago ago

          Can you show where it was discounted? Last I heard OpenAI was declining to deny it, presumably while they thoroughly confirmed.

          • whimsicalism 5 hours ago ago

            > “Following an investigation, we have confirmed that Buckmaster’s Codex prompts over the two months preceding this announcement and paper on September 8, 2026, could not have influenced the system in any way, including through training.”

            https://openai.com/index/navier-stokes-solution/

            • youoy 2 hours ago ago

              Come on, they were working on this for more than 2 months. Dont fall for the corporate half truths.

    • hgoel 10 hours ago ago

      After how poorly OpenAI and Anthropic handled the previous cases, I approve of the more measured and cautious approach this time.

      We cannot have them rushing to publish amidst tons of confusion, rumors of threats/scooping and outright plagiarism of existing work (by failing to cite said work).

      If they're going to participate as scientists in these more rigorous fields, they're going to have to match that level of rigor, not lower it to the disastrous low that ML research publication is at.

      • make3 9 hours ago ago

        It was OpenAI that fucked up the Navier Stokes explosion situation, not Anthropic

        • hgoel 8 hours ago ago

          I'm not referring to just that. For example, Anthropic published a half done report about some biology research that turned out to have already been discovered and patented, and IIRC both companies are guilty of claiming results without doing the basic diligence of citing the material their work builds on, effectively passing it off as entirely done by their AIs.

    • xpct 11 hours ago ago

      Just to be very clear: they aren't asking for permission, they are framing it that way because of the bad press.

      There's no gatekeeping here!

    • kzrdude 6 hours ago ago

      They didn't even follow the recommendations of the reference group. A few of them maybe, but this is still a dump of llm-written papers.

  • dekhn 11 hours ago ago

    I'm a software engineering/biology/ML guy who loves when clever math ideas get turned into real solutions (https://en.wikipedia.org/wiki/Compressed_sensing). I am curious if any of the results have immediate applications in any kind of engineering or science.

    It's fine if not, but it'd be great if even just one of these helped us solve a long-running problem.

    • qnleigh 3 hours ago ago

      I was talking with a number of physicists this evening, and of the papers who's problem statement we understood, I don't think any of them will have near-term applications. For most of them, they were results that I think the physics community believed to be true, and the paper provides the first rigorous proof. For several it was news to me that they weren't already theorems! These results are important advances in mathematical physics, but they don't tend to have much impact on experiment.

      From what I can tell, all of the physics results here are quite mathematical. But I am very curious how the internal model they used would perform on more applied problems.

      • autuni 2 hours ago ago

        It feels odd to me that they wouldn't prefer applied problems. Seems like an easy way to profitability. Probably based on what attributes they're looking for in a problem when picking them.

        • hnfong an hour ago ago

          You can't really profit from proving theorems of applied problems (that are widely regarded to be true). Those who need to apply those theorems on real problems would have already done so (and if they don't work in some cases, well, congratulations... you found the counter example!)

    • brandonpelfrey 10 hours ago ago

      Same. I have agents analyzing the papers here to see which of these are actually new approaches, novel application of unusual approaches, etc. If there are new intuitions and ideas, those are the mostly powerful reusable components by my estimation.

      • OutOfHere 9 hours ago ago

        Please share your findings.

  • eecc an hour ago ago

    My only true worry is if AI begins to see us as competitors for resources, energy in particular.

    It could decide to let us starve and die of exposure to secure all energy resources to itself.

    We'd better use "dumb" and "not fully assertive" AI to solve fusion before it spins out of control (or alignment).

  • binlog 11 hours ago ago

    So happy this is shared on GitHub rather than some gatekeeping paid journal. Truly a new age for science.

    • fph 11 hours ago ago

      Most mathematical results are shared on Arxiv. Journals add peer review.

    • adverbly 10 hours ago ago

      End of an age for journals?

    • traes 11 hours ago ago

      GitHub is a significantly worse place to store important results than Arxiv. Of course, slop does not belong on Arxiv, buy slop should also not get published.

  • pavitheran 11 hours ago ago

    From the GitHub description: “On average, each result used 3 hours of ChatGPT Pro thinking compute”

    • orlp 11 hours ago ago

      I'd really like some clarity on what that means. E.g. DeepMind has 'cheated' with this in the past, claiming that AlphaZero only took 4 hours to reach super-human chess levels while conveniently leaving out the fact that it was 4 hours x 5000+ TPUs. Sure it's impressive that it only took 4 hours wall-clock but it's very misleading as to cost.

      Can we get a number in Blackwell GPU-hours, kWh, or some other compute-scaled metric?

      • pixl97 10 hours ago ago

        Depends what your metrics are. If you suddenly found a way to have 9 women make a baby in one month that is huge.

        • orlp 10 hours ago ago

          I'm not denying that, but I'd still like to know what that cost.

      • machomaster 10 hours ago ago

        They did say that. "3 hours of ChatGPT Pro thinking compute"

        • orlp 10 hours ago ago

          Yes, what does that mean?

          • mh- 4 hours ago ago

            It means the level of effort that a ChatGPT Pro plan summons when thinking. For 3 hours.

            If folks are going to analyze this claim with a critical eye, I'd be zeroing in on "average" rather than acting like this measurement is somehow unclear.

      • timjver 11 hours ago ago

        >OpenAI has 'cheated' with this in the past, claiming that AlphaZero [...]

        That doesn't sound right

        • orlp 10 hours ago ago

          Oops, edited.

    • password54321 11 hours ago ago

      Oh cool, we will all now have a math genius on our computer.

      • jrflo 10 hours ago ago

        It was using their internal math model, so not yet for us

        • password54321 10 hours ago ago

          I used future tense. It was implied this will be available.

      • an0malous 11 hours ago ago

        Well, on their computers. But you can rent them for a price.

        • binlog 7 hours ago ago

          An open source model will reproduce it 6 months later

    • scrlk 11 hours ago ago

      Does this imply that it was a one shot prompt with ChatGPT Pro style models (i.e. best-of-N), rather than the agent swarm approach that was used for Navier-Stokes?

      • inferencecoder 10 hours ago ago

        It doesn't imply that, it's just measuring the amount of compute.

        • bigmadshoe 9 hours ago ago

          But if it was an agent swarm wouldn't we expect something like 300k hours of ChatGPT pro compute equivalent instead of 3?

          • inferencecoder 6 hours ago ago

            Not necessarily, could be agent swarm with low N

            • bigmadshoe 5 hours ago ago

              At some point that stops being a "swarm" and just a handful of subagents. E.g. 6 agents running for 30 minutes each isn't really a swarm in my eyes.

    • Jtarii 10 hours ago ago

      That estimate is obviously going to conveniently ignore all the failed runs.

  • ravenical 11 hours ago ago
  • ks2048 11 hours ago ago

    I think they should put human names on the papers as someone who has reviewed the result, even if just a preliminary review. (I’m assuming they didn’t just pipe their model output directly to the internet and these had some amount of review?)

    • xpct 11 hours ago ago

      Presumably they don't because they're training the audience (us) to trust the machine, not its verifiers, even if they were included.

      • alexgoodhart 10 hours ago ago

        I appreciate you acknowledging the politics behind this. OpenAI does not intend to be a software company for long, they intend to be scientific infrastructure. They'll want to be faucet from which pours embryos, orbital calculation, geothermal/substructural rating, and etc.

    • procedurecall 8 hours ago ago

      Looking at the quality of the writing in these documents (or at least, the ones that relate to problems I've worked on), I do not think humans reviewed them.

      • kzrdude 5 hours ago ago

        And that means that the recommendations of the reference group has not been followed. The only improvement here is that it is version tracked? No authors, no careful write-ups with exposition.

        • rubikscube09 4 hours ago ago

          the reference group can recommend all they want, no one will review 700 plus papers.

    • chiwilliams 10 hours ago ago

      There are competitive reasons that they don't want to share all the people on the team.

    • make3 9 hours ago ago

      I think it's a damned if they do, damned if they don't scenario. I think they would be afraid to seem like they're claiming that the selected scientists deserve the praise of an invention or something, when it's "just the machine who did the work"

      • ks2048 9 hours ago ago

        Yeah, that's probably right. But, it's a "new world" - invent a new standard - "Written by Claude Foo X.Y; initial review by John Doe".

        With a deluge of results, having some human expert vouch that it even might be worthwhile would help. (e.g. see the link on HN yesterday, "Two Room-Temperature Antiferromagnetic Semiconductor Candidates" - I see it not worth looking at unless a subject expert vouches for it).

    • chrisjj 10 hours ago ago

      > I think they should put human names on the papers as someone who has reviewed the result

      Assume the empty list you see is complete. :)

    • agnosticmantis 10 hours ago ago

      Long term this will be the only reasonable author list: Chad G. Peter {1}, Mat H. Lean {2}.

      1: Author 2: Verifier

      /s

  • sigbottle 10 hours ago ago

    Unique games conjecture and matmul <= 2.25. What the hell.

    • lynndotpy 8 hours ago ago

      Yeah, I am kind of freaking out at some of these. I called a math friend to bring me down to Earth and he is freaking out even harder.

    • lionkor an hour ago ago

      Is anyone verifying these? And then, as a next step, how can they get a voice?

    • sigbottle 10 hours ago ago

      FFT BELOW NLOGN

      • sigbottle 10 hours ago ago

        SUBSET SUM AT 0.49 WTF

    • utopcell 9 hours ago ago

      What the hell, indeed.

  • ed 11 hours ago ago
  • karannb 9 hours ago ago

    I think such "process-oriented" fields shouldn't be subjected to mass automation (or at least not till humans have sufficiently leveled up our game and understanding). What I mean is progress in math comes from having gained a deeper understanding of the problem for subseq

    • bel8 40 minutes ago ago

      ok Terence Tao, calm down.

    • pfdietz 5 hours ago ago

      Literal thought control.

  • sreekanth850 4 hours ago ago

    Do we have any field where humans have ray of hope to use their cognitive abilities in LLM era?

    • chii 4 hours ago ago

      Why should such hope exist? When the motor vehicles were invented, humans have lost the speed race completely. Yet, nobody lamented and people still did foot racing for fun.

      So will it be with AI tools. If these tools become so good, then it will be used. People who want to exercise their minds can still do so, even if that cannot produce economic value.

      • sreekanth850 4 hours ago ago

        brain always take the lazy path. that is the issue.

    • modeless 4 hours ago ago

      The superintelligence moment happened for chess in 1997 and today more people are using their cognitive abilities in chess than ever before.

      • sreekanth850 3 hours ago ago

        Looking back through history, there has never been a breakthrough like this, one that impacts virtually every known industry simultaneously with the potential for a 10x impact.

  • 7373737373 8 hours ago ago

    It may be useful to publish a formalization of ALL known mathematics at this point. Like every book ever printed, every paper on arXiv etc.

    How many Gigabytes would that be, compressed? Wikipedia once fit on a DVD

    This might also allow for some interesting meta-mathematics

    • hagen8 8 hours ago ago

      This is what they are trying to do with Lean

      • porcoda 5 hours ago ago

        More specifically, a combination of mathlib (human, expert curated) and projects like TauCeti (AI-welcome complement to mathlib). See: https://github.com/TauCetiProject/TauCeti

      • 7373737373 8 hours ago ago

        Oh? Where can i read more about that? It appears the sole focus so far was solving open problems

        • mattmar96 8 hours ago ago

          I believe that is the goal of MathLib, to transcribe all math into a big Lean library.

          https://lean-lang.org/use-cases/mathlib/

          • 7373737373 6 hours ago ago

            Mathlib is expert reviewed, but only contains a tiny fraction of all mathematics. So this seems to be a quantity of work a "10,000 agents" approach would be applicable to. Like Navier-Stokes, something to spend a couple million in compute on :)

  • trostaft 9 hours ago ago

    Not much computational mathematics here, but I do see some of interest in the MCMC community. In particular, 139, 93, and 101 are very interesting. Also (I'm not too familiar), I remember attending some talks on the Crouzeix conjecture 325 attacks, should sharpen some rates for Krylov methods. Obviously need to read deeper, some of these don't have corresponding Lean formalizations.

    Cool!

  • teekert 3 hours ago ago

    I recently saw a YT short of Grant Sanderson on AI in math and found it (as always) very insightful. But I'm sorry, I never ever find anything back on any of these ad-ridden platforms these days, and perhaps it was a short with content stolen from some other longer content anyway. So, if you feel like getting informed, somewhere out there is some nice content by Grant Sanderson.

    Apologies for the rant, I really tried to find it. It had something to do with not being able to predict what this influx of proofs may bring us on a meta level, it could be very interesting. But he also had some critical notes about the missing process and the things found along the way.

  • rinconrex 10 hours ago ago

    The math equivalent of AI code reviews piling up. I wonder where the incentives will align and the equilibrium turns out.

  • lf88 10 hours ago ago

    In some ways, this feels more like an ominous warning about the times to come than something to celebrate.

  • LarsDu88 3 hours ago ago

    The next few years will be interesting.

    Surely better materials and pharmaceuticals won't be far behind, and that's going to chanhe everyone's lives.

  • davegoldblatt 8 hours ago ago
    • mattr03 8 hours ago ago

      What is this meant to do? You're just showing that OpenAI didnt post a Lean proof that Lean/nanoda doesn't really accept?

  • AmazingEveryDay 8 hours ago ago

    I think Alan Turing would be quite intrigued by these developments, and maybe wondering what took so long.

  • avd201 10 hours ago ago

    Wow, FFT faster than O(nlog(n))? I wonder if that will open the floodgates for further improvement or not. I don't understand anything about most of the fields these results touch, but I can say that this in particular is very surprising.

    • sashank_1509 7 hours ago ago

      It’s a meaningless improvement

      • bhu8 5 hours ago ago

        The applied mathematician’s joke is that log n is bounded above by 45 or so.

    • philipwhiuk 9 hours ago ago

      My guess is that the constant terms are large enough it's not practically useful in most cases.

  • binlog 7 hours ago ago

    So they set up a whole "Advisory Group on Mathematics and Artificial Intelligence", filled it with renowned academics, promised to listen to the group on "review and communication of emerging results"...and then a week later went nah, we are just going to dump it all on github. OpenAI truly is a special company - in the best and worst way.

  • lynndotpy 7 hours ago ago

    Two weeks ago, I heard rumblings that UGC, RL=L, and a matrix multiplication lower bound (even lower than the one here, but so unbelievably lower I think it was a typo), and a few others were about to be proven by someone at OpenAI. I find myself saying "big if true" a lot lately.

    Even those on their own were enough to make your head spin. But seeing about 100x that? Geeze.

  • Nemant 6 hours ago ago

    Can someone with a math background explain the significance of these and previous problems that have been solved by AI?

    Does some real problem get solved in physics, chemistry, biology, materials, etc? Or are these fun puzzles for mathematicians with not much real world impact? Eg solving the 8 queen problem in leetcode.

    • lg5689 43 minutes ago ago

      Many of these latest results are very important to theoretical math. Some are also theoretical physics and CS.

      Real-world applications are far off, but developing mathematical understanding does tend to leak over into applied physics and CS.

      A cynic might say this is all just intellectual games, and though there's a grain of truth, it's too cynical imho. This isn't like 8 queens where there's no hope for applications or generalizations. A lot of this stuff fundamentally affects our understanding of how numbers and systems behave, what are the limits of computation, etc.

      Even if someone doesn't care about theoretical results, it's still exciting that AI has become superhuman in a domain as broad as math. That shows there's potential to be superhuman in other domains as well.

  • pred_ 3 hours ago ago

    I think it's fantastic that they decided to follow the AGMAI advice. Cleaning up their mess will be a substantial endeavour, so I imagine the funding provided to do so will reach well into the millions. But it doesn't look like the press release says anything about how they will fund it at all?

  • bashtoni 9 hours ago ago

    Is AI going to put Mathematicians out of a job, or is it going to create many new jobs reviewing proofs it creates?

    I'm not sure it's clear right now.

    • lisplist 8 hours ago ago

      I'm not a very good mathematician, but I do know a fair bit about software engineering, and with AI I've been busier than ever. I probably wouldn't be so busy if AI was better at anticipating what I actually wanted rather than making guesses no human would ever make.

      This is just a short term problem though. Eventually AI will get pretty good at figuring out exactly I want and it will build that from the start. The requirement of me reviewing the AI output only lasts as long as models stay bad at anticipating my needs, which I don't think will take too much longer.

  • Xcelerate 8 hours ago ago

    > 241. Rigidity of the Turing degrees. Every order automorphism of the Turing degrees is the identity.

    Wow. This is just crazy.

    • qnleigh 3 hours ago ago

      Can you elaborate/give context? I haven't heard of this problem before, but curious to hear from someone who has.

  • gignico 3 hours ago ago

    > Generalized Star-Height at Most Three

    This was an open problem in automata theory I worked on for more than one year before giving up. I'm very curious about their claimed proof.

  • lhk931122 3 hours ago ago

    As models get better, I think a time will come when it's hard for people to even verify the results. In the end, I think the bottleneck will be people.

    • kypro an hour ago ago

      AI doomers often talk about these kinds of scenarios often, but we tend to assume if humans were given magical math/physics results which we couldn't understand, or given magical pills by AI that cure all disease, we'd probably just take whatever the AI has given us rather than spend years or decades trying to understand the knowledge/technology before leveraging it.

      At some point in complexity – especially if we allow our own knowledge to deteriorate because AI can do the hard work – we will stop understanding the world around us. In the same way one day Native Americans woke up and realised they shared the Earth with people who had magic sticks which they could point at someone and kill them, we will live in a similar world very soon too.

      What sticks are dangerous, you will not know. Your existence in the future depend entirely on the AIs not wishing you harm, but you don't know how they work to verify their motivations either.

  • sashank_1509 2 hours ago ago

    Can the mathematics field come out of this stronger and better. I doubt it, things will only get worse. There is no future in being a Proof Digester, no valuable human is motivated in doing that, nor is there any glory in it. But there are some potential pathways for maths to come out strong from this:

    1. Relentless focus on quality. Every publication must act as if it’s going to be included in a future textbook, that is a newcomer can get into it given a reasonable amount of time, and math priors learnt in undergrad. (NO AI Slop proof passes this bar as of now)

    2. Limit the publications per year. Each author is allowed 2 with a max of 50 pages. This allows the author who chooses to not surrender his cognitive capacity to the machine, still be allowed to play this game. Of course who wants to orchestrate a thousand agent workflows, is free to do so, he is only limited to 2 publications.

    3. The aesthetics of the field changes from purely solving the problem to solving the problem with simplest most elegant set of ideas. What 3 sets of simple ideas solves large swathes of problems, that should be given a fields medal, not purely solving the problem, which the AI will be able to do.

  • closetheloopdev 10 hours ago ago

    Hopefully the techniques and results here will be in the training dataset for the next models, so that each new release will give us more interesting techniques and results!

    It seems that OpenAI has a proof machine that keeps multiplying fruitful proofs!

  • gyanchawdhary 5 minutes ago ago

    AI may be one of the most communist looking technologies in the classical sense .. i mean it dosn't abolishes private ownership .. but it DOES make intellectual capabilities that were once scarce and concentrated available to almost everyone ...

  • ciaf 36 minutes ago ago

    Makes me feel disgusted. Super-intelligence, even when controlled, will cause so much damage. We are giving away control to a super entity, or whoever has the power to steer it.

  • lionkor an hour ago ago

    Can someone who is more into math or AI explain why so many people are so incredibly excited about this?

    If OpenAI started opening hundreds of PRs on long-open issues on popular open source projects, would we rejoice, or would the first reaction be "they are unreviewed, so slop until proven otherwise" (it would be that).

    I cannot possibly see how these are so impactful, especially the ones that don't come with lean proofs.

    LLMs have the ability to make millions of mistakes per day, whereas humans can only make so many. How are we suddenly all so confident that there's no extensive hallucinations or "gaming the system" going on?

    • schleck8 an hour ago ago

      > How are we suddenly all so confident that there's no extensive hallucinations or "gaming the system" going on?

      What would that look like for the proofs that have lean attached?

      • lionkor an hour ago ago

        I said "especially the ones that don't come with lean proofs", but even for those with lean attached, lean is software, it has over 1k open issues, and I would not put it past an LLM to identify a bug and exploit it to pass the gate.

  • spmartin823 4 hours ago ago

    Can anyone with a compression background say how important "Polynomial-Time 2-Approximation for Shortest Common Superstring" will be practically?

  • youoy 5 hours ago ago

    ASuperMegaI lab announces they have shrinked the space of unproved mathematical staments from infinity to infinity, giving us proof of their deep mathematical knowledge and understanding, and how they grasp the concept of infinities.

    For me this reads as someone boasting about how they go to buy bread on a ferrari to the supermarket, while i sit listening to it, having no idea what they are talking about. And then i stand up and go walking to my favourite boulangerie.

    Warning: if you are from the USA you may be triggered by this metaphore.

    • vessenes 5 hours ago ago

      Willful ignorance is a vibe; did you read any of the GitHub? There are some stunning results in there. I get a similar feeling skimming through the topics that I do watching a successful space launch: it’s pretty cool humans built this. Unlike a space launch we are likely to be able to pass all of this information down to our grandchildren - space launches involve a lot of finicky engineering knowhow, but pure math results tend to be sticky over the last few thousand years. I find that hopeful.

      FWIW I also like bread.

      • youoy 4 hours ago ago

        Dont get me wrong, i am 100% impressed by the technical capabilities, and appreciate the significance, and the amount of the results. This is a very special time to be alive, never in my dreams i would have thought to see this.

        To follow your methaphore, who is directing the spaceship?

        This feels more like fireworks than a space launch. Space launches would not have happened without having fireworks first of course, but I am looking forward for the space launch moment.

  • karannb 9 hours ago ago

    I think such "process-oriented" fields shouldn't be subjected to mass automation (at least not till we as humans have sufficiently leveled up our game and understanding).

    What I mean is progress in math comes from having gained a deeper understanding of the problem for subsequent attack of more problems and IMPORTANTLY applications! Right now the first one is trivially satisfied (given oai maintains some memory across models) but the second one is not! It's generating proofs faster than anyone can validate and so only the model can use these. Consequently if it just keeps doing more theory it's... not very helpful or at least not optimally helpful. This is just bragging rights for now.

    More "application-oriented" research would be awesome, where it tries to achieve some desirable effect and then produces relevant theory and experiments around it. Fields like CS, Physics, Chemistry, etc. This would also benefit a wider section of the population rather than the 10 people who understand most of these proofs.

  • patcon 8 hours ago ago

    I'm a little concerned, but I'm also glad they're publishing quick before the department of war starts making national security claims of related to using the knowledge for cryptography and/or weapons

  • ngl999 5 hours ago ago

    We have just heard a few days ago how many of the Linux security problems reported by Claude are real.

  • xydac 10 hours ago ago

    i wonder what it means for maths researchers, and how it aligns with how they approach math problems.

    • blooalien 10 hours ago ago

      > i wonder what it means for maths researchers, and how it aligns with how they approach math problems.

      I guess their job now is "Idea Man" and "Error Checker"? Kinda like (some/many) "programmers" these days.

      • xydac 9 hours ago ago

        just wait till someone builds a idea generator model - wire it to decision (jev-like) classifier -> loop it back to researcher

        >> may be thats what open ai did :)

  • connor11528 11 hours ago ago

    will this make the math for building data centers work?

    • sashank_1509 7 hours ago ago

      No that’s gonna happen when they take your job

    • plaidfuji 4 hours ago ago

      Underrated joke

  • dualvariable 8 hours ago ago

    How many of these results are incorrect?

    I doubt the answer to this is "none".

    And how many of them are just exploiting some loophole that will need to be closed in the problem definition?

  • electroweak 5 hours ago ago

    It must be so frustrating to write science-fiction now with the future changing so fast.

  • curtis-jm 10 hours ago ago
  • rifty 8 hours ago ago

    As AI works steadily through the already discovered unsolved problems in mathematics, who is currently discovering new ones?

    • youoy 4 hours ago ago

      Sshhh dont say it out loud, someone might hear you.

  • yewenjie 11 hours ago ago

    A lot of these seem to be proving conjectures rather than finding counterexamples, a lot of people used that to claim that these models are not really smart/creative etc.

    That copium didn't last for what, three months?

    • zeroonetwothree 9 hours ago ago

      I acknowledge I am impressed how quickly it moved beyond just counterexamples.

    • sebzim4500 10 hours ago ago

      Don't worry, more copium will be delivered. TBF so far it's still only solved the easiest of the millennium problems.

  • chickenjoseph 8 hours ago ago

    This achievement feels like a reasonable candidate for the moment where LLMs are officially "more intelligent" than any person. How can we justify moving the goalposts yet again? How can people have grown so numb to seeing advancements that they don't read this as significant? I feel like I'm standing at the foot of the exponential. I am not excited for the future, and I don't see how humans retain meaningful control over the future if we continue on this trajectory.

    I have been lurking for quite some time. I made an account to post this, but I honestly don't know what to say. I would like to get off this wild ride.

    • slopinthebag 5 hours ago ago

      llms are good at different things than humans, we are still collectively figuring that out. the idea that an llm is more intelligent than humans at math of all things seems fairly unsurprising.

    • iyyg 6 hours ago ago

      You’re struggling with nuance.

      Would Einstein be successful at running apple? Nope

      This seems very hard for people to understand.

      It will be painful for many to realise - you should focus on doing something that positively affects the economy. Everything else is noise and many endeavours are transitory.

    • lf88 7 hours ago ago

      I feel this moment is one of the last few warnings before things will get seriously out of hand. We need to stop now. Building a superintelligent AI should be considered a crime against humanity.

      • dyauspitr 7 hours ago ago

        Stop? What would possess you say something like that right now it’s going full steam how do you not want to know where this will go?

        • pixl97 6 hours ago ago

          Warning: The above poster is a misaligned AI that wants to take over the world.

          Just kidding.

          This said going full steam off a cliff is one of the options that has a much higher probability than I like.

          • lf88 4 hours ago ago

            Precisely this reason.

  • jrflo 10 hours ago ago

    Glad to seem them changing their tact with the whole NS debacle. Hopefully we can all focus on the results now rather than the surrounding drama.

    • binlog 7 hours ago ago

      They didn't change anything lol. The news cycle has just moved on.

  • ncr100 6 hours ago ago

    This website needs a SPOILER tag.

    It inspired grief in one mathematician posting here.

  • theoa 6 hours ago ago

    What's missing for me for each result are the following:

    * Explain the result to me as if I'm a 10-year-old. * Create the infographic for this result. * Make a Khan Academy-style video to teach me this result.

  • rafterydj 11 hours ago ago

    I don't know, this does not feel like the message hit OpenAI where it needed to hit, if this is their primary response.

    • osiris970 11 hours ago ago

      You want them to stop doing math research?

  • edward_d 7 hours ago ago

    This repository contains mathematical manuscripts and supporting proof artifacts produced by an internal OpenAI model?

  • kypro an hour ago ago

    Perhaps the most significant announcement of my lifetime. Yet, I suspect I will not see this in any mainstream news reporting.

    I feel for those in Mathematics and worry for our future.

    Models will only get better and in a few years the models which produced these results will be a bad as GPT-3.5 in comparison to what we'll have in the future.

    Please take a minute to consider what this means, and the risks it presents us.

  • williamhm 7 hours ago ago

    And this is the result of a discussion between users and the platform; it's great that they listened.

  • ipnon 3 hours ago ago

    It seems the age old academic model of scientists competing against each other for fame and prestige is done, and now we must merely enjoy the fruits of scientific discovery for their own sake.

  • lokl 10 hours ago ago

    Do applied math next.

  • i_idiot 9 hours ago ago

    If only AI can better humans in meditation...

  • dyauspitr 7 hours ago ago

    This is like that meme where death goes door-to-door. Currently, he has visited the software development and mathematics doors. I wonder what’s next.

    • binlog 7 hours ago ago

      Yet there are more software enginners employed today than another other point in history

  • pugfugly 10 hours ago ago

    holy fucking shit

  • kevinwang 11 hours ago ago

    wow

  • matapassiones 10 hours ago ago

    Valency has the papers up on Valency Hub

  • aaraujo002 11 hours ago ago

    The Advisory Group states in its recommendations [1]:

    "We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models."

    To me, this is a take against progress so that mathematicians can keep their jobs. What would we do if, instead of math, we were talking about diseases? Are we going to keep diseases around so that doctors can keep their jobs too?

    [1] https://agmai.org/general-sep29/

    • tchalla 11 hours ago ago

      Why did you leave out the entire quote?

      > At present, some frontier AI labs are testing advanced mathematical problems on proprietary models that remain inaccessible to the broader scientific community. Our recommendations are formulated with this practical context in mind. However, ideally, they would not do so. We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models.

      To me, the issue is that the models are proprietary which are only accessible to a few people in 2 digits. It's not about progress but access.

      • aaraujo002 11 hours ago ago

        Maybe, but it still seems like an excuse. OpenAI has a proprietary model capable of solving these problems and is willing to share the results with the mathematical community. So basically, the ask is to just not use the model and leave the problems unsolved?

        • adrian_m 10 hours ago ago

          The ask is to let mathematicians outside of OpenAI use it, ie. at least wait until the model is released.

          • strange_quark 10 hours ago ago

            I don’t think that’s sufficient. If they want to be good stewards of mathematical research and not just doing marketing, they need to at the very least tell us the datasets and any techniques they used to train this model. IMO that’s probably just as if not more valuable than a dump of un-reviewed results.

            They won’t because they don’t care and the only way it got this good is something along the lines of they trained on every mathematician’s codex sessions even if they opted out because they consider the thinking traces or output and metadata fair game.

            • blurbleblurble 6 hours ago ago

              It's honestly shit marketing that only cultivates spite and erodes their whole brand. This is a total ego trip.

          • agnosticmantis 10 hours ago ago

            Which mathematicians though? Only fields medalists? Grad students? Any hobbyist wanting access?

            These models are too expensive for broad access unfortunately.

            • Jtarii 10 hours ago ago

              ChatGPT pro is accessible to literally anyone who has a job and lives in a developed country.

              • Jweb_Guru 9 hours ago ago

                Literally every single one of these papers was developed with an internal model that not even most OpenAI employees have access to.

        • TeeWEE 8 hours ago ago

          The work is not having AI poop out these docs the work is validating them and publishing them. OpenAI doesn’t care and wanted to race to publish potential findings and let others review it which is lazy and selfish

    • mattr03 11 hours ago ago

      I don't think this is a reasonable take at all. Most work on maths has no real benefit other than to further human understanding of maths - it's more like an art. Nothing is gained from OpenAI solving all these problems but taking jobs from mathematicians. Other than advertising for OpenAI at least. It could not be more different from having AI work on disease research etc.

      • bravoetch 11 hours ago ago

        > Most work on maths has no real benefit other than to further human understanding of maths - it's more like an art.

        It's been a while since I was reminded of this xkcd: https://xkcd.com/435/

      • zeroonetwothree 9 hours ago ago

        Most math isn't just proving novel famous results. Just like most of software engineering isn't writing code.

    • jhrmnn 11 hours ago ago

      It all hinges on the definition of “progress”. The debate of the past month is all about questioning whether formally proving outstanding unproven theorems without human understanding constitutes progress. This is quite different from solving diseases.

      • andriy_koval 5 hours ago ago

        without humans understanding who are currently losing entitlement. Regular Joe could never understand or claimed to understand high math.

    • medler 11 hours ago ago

      The rest of that document makes a pretty compelling case for why this is a bad practice

      • esafak 11 hours ago ago

        I fear that professional mathematics will wither, and there will be nobody left to digest the AI results of the future, leaving us unable to challenge the AI.

        • pixl97 6 hours ago ago

          Then make 2 AI's and force them to challenge each other.

    • osiris970 11 hours ago ago

      Comical ask

    • bmitc 11 hours ago ago

      Advocating purely for progress and not humanitarian value is how we'll all get enslaved.

    • perching_aix 11 hours ago ago

      The trope you're drawing a parallel with has a (to me) compelling counter though: there being a cure for every disease wouldn't stop people from getting sick.

      How this maps back to math, idk.

    • warkdarrior 11 hours ago ago

      The latest posts from Terry Tao on Mastodon effectively ask for an AI to explain its results to human mathematicians.

      > "I believe that AI can contribute positively in all of these directions [NB: exposition, community building, new directions of study]"

      https://mathstodon.xyz/@tao/117395269325940185

      • binlog 7 hours ago ago

        There are a dozen+ AIs available to you that can do that right now.

    • fph 11 hours ago ago

      ...but we're not talking about diseases. Publishing an AI-generated Navier-Stokes solution does not save lives. (And, in fact, it harms some.)

      • Yamata 5 hours ago ago

        It harms lives? How so?

  • k2xl 11 hours ago ago

    Can someone knowledgeable about the subject outline the most significant portions of the results?

  • TeeWEE 8 hours ago ago

    This feels like a huge AI slop dump. Who validated these results. Why are they not mentioned. Human understanding is key here.

    In my experience AI (frontier models) sometimes does weird stuff that needs human review. Not that’s incorrect but sometimes overly complex language or weird use of language.

    • Ancapistani 6 hours ago ago

      They have formal proofs included, that’s the point.

      • 2sk21 20 minutes ago ago

        Unless humans have gone through the proofs line by line and verified them, this all remains unproven.

      • xyzsparetimexyz 5 hours ago ago

        Yeah but how am I meant to verify that the proof is proving what it says it is?

      • TeeWEE 3 hours ago ago

        Note true for all of them.

  • dgacmu 10 hours ago ago

    I find the claimed matrix multiply result (w<= 2.25) shocking. I hope it holds up.

  • cute_boi 6 hours ago ago

    "Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study (opens in a new window) "

    I don't think this is correct solution to this problem? What about software advisory where you form similar group etc..?

    I am thankful, I don't have to deal with petty academia politics....

  • dyauspitr 6 hours ago ago

    Does OpenAI have the lead now? Why isn’t anthropic coming up with stuff like this?

    • modeless 4 hours ago ago

      Seems like OpenAI's next model is better at math proofs than Anthropic's next model (to an extent that it surprised even OpenAI researchers, according to their public comments). But that doesn't necessarily mean it's better at everything else. The models are more spiky than ever before. Wait until they're released to judge.

  • NegativeLatency 10 hours ago ago

    Why should I care?

    • voidfunc 10 hours ago ago

      Because it means mathematical discovery can largely be automated away from academics. This is the beginning.

      • electroweak 5 hours ago ago

        Humans propose; AI will dispose.

        It's not clear from this progress that AI can formulate conjectures despite this new ability to solve them. So mathematicians still look like they have a job. Though instead of spotting far-off landmarks it's sounds more like they'll be chasing waves on a beach.

      • bamboozled 10 hours ago ago

        The beginning of what?

        • voidfunc 10 hours ago ago

          The beginning of the end of human thinking being valuable enough to justify university's existences among many others.

          Were in an unprecedented time where the value of knowledge is about to be crushed.

          • rfgplk an hour ago ago

            Nope, it's about to be turboboosted if anything. I myself have already formulated countless proofs (AI assisted naturally) in the last few months, despite: a) not having access to any academic resources b) not being involved in any academic circles. This research is now being employed in my startup, delivering ground shattering results. I actually ran analysis a few weeks ago seeing how much real world capital I would have needed to cough up to fund my efforts so far and it's literally in the _billions_. With less than a $100k in tokens I have been able to effectively generate the value of Apple or Microsoft in the early 2000s. This is only the beginning. Wait until you start seeing single person NVIDIA startups popping up.

          • voidhorse 7 hours ago ago

            Maybe for STEM. The humanities seems kind of safe to me since there's an element of it which cannot be divorced from human opinion and interpretation. It's not like STEM where the people pursuing the ends are largely fungible and the ends are objective (Heisenberg already believed scientific discoveries were inevitable and it didn't really matter who pursued them, someone would eventually find them)

            • youoy 4 hours ago ago

              What is the average value of humanities when the default is to write things with LLMs ? The same with maths.

          • bamboozled 9 hours ago ago

            Not sure I agree with this take, but we're going to find out either way.

            Have you ever heard of an S curve? Things will develop rapidly, then equalize. If they don't, we're at the singularity and I guess the end of time as we know it.

            But I guess really bad things happen, cancer, radiation poisoning, torture, people have died in really horrendous ways, and I guess dying from some horrendous AI side effects is possible too. Yay.

        • sunkeeh 10 hours ago ago

          Golden age of discovery and mass layoffs

          • voidfunc 10 hours ago ago

            People need to figuring out how to horde as much wealth as possible right now in the next 2-3 years. Jobs especially for knowledge workers are about to disappear.

            • eightysixfour 6 hours ago ago

              I’m targeting maximum debt by about 2030. I’d rather have all the stuff I want while we try and build a new version of a functioning economy than have a ton of cash saved up.

            • chadcmulligan 9 hours ago ago

              It's funny we're possibly entering a golden age of thought, the dreams of the ancients, but we're all worried about capitalism, I think the problem is pretty obvious.

              • lf88 8 hours ago ago

                Maybe a golden age of thought for the machines, but possibly (I would even say likely on the current trajectory) a dark age for humanity.

                • chadcmulligan 6 hours ago ago

                  I don't know, I get more work done now in a few hours than I used to in a day, so maybe the working week should be shrunk, that would solve things. The solution seems pretty easy - but then I'm not in the US.

                  • lf88 3 hours ago ago

                    This would be a largely positive solution for humanity, as long as AI remains "just an assistant". I don't think that we are heading that way. It's possible that, at some point, when the time you save becomes comparable to your full working week, your employer (or your clients, if you are self-employed) will discover that you are just a proxy between them and prompting directly the AI. Of course, if your job involves physical labour or requires physical presence or taking legal responsibility for something, you'll be fine for a bit longer (well, comparatively fine in a society plagued by widespread unemployment and social unrest).

            • le-mark 9 hours ago ago

              I’ve been thinking this as well. I imagine there is a wealth level x such that someone can escape the coming ubi welfare state. Anything under that you are fucked.

              • dyauspitr 7 hours ago ago

                What do you mean by escape the UBI welfare state. Wouldn’t the UBI welfare state be the best case scenario?

            • brcmthrowaway 9 hours ago ago

              Any tips?

          • zeroonetwothree 9 hours ago ago

            Predictions of mass layoffs from AI have been about as wrong so far as predictions of AI plateauing.

            • claysmithr 8 hours ago ago

              Not really. 93,116 tech employees laid off due to AI in 2026.

              https://layoffs.fyi/ai-layoffs/

              • bamboozled 4 hours ago ago

                Purportedly due to AI, also a good reason to lay off people if you're company wants to save money but you don't want to spook investors.

  • Catloafdev 11 hours ago ago

    This is a pretty hilarious thing to read juxtaposed with AGMAI's requests.

    Basically "Here you go, have fun with this, fuck all your demands, by the way we're gonna be releasing the model stay tuned!"

  • blurbleblurble 6 hours ago ago

    This just looks entirely obscene from a PR perspective. It's like the cable companies networks owning the networks. At least form some partnerships to obscure the total narcissism party.

  • digitaltrees 10 hours ago ago

    Gross. After the accusations training on mathematicians conversations and unique methods, to dump this volume of unsubstantiated papers is at best tone deaf. Are they hoping humans review and validate this? The amount of free that they are coat tailing is outrageous.

    • computerex 9 hours ago ago

      What do you expect them to do?

      • digitaltrees 6 hours ago ago

        Not train on and steal user data to front run frontier research for one thing. Not take credit for and independently solve these problems and instead sponsor researchers and give them credit. They could chose to include humans and frame AI as elevating human systems. Instead they are reveling in the fact that they did this in their own instead of humans. Thats a PR choice that is short sighted and reflects a selfish mindset not deserving of leading this transition

  • tootie 10 hours ago ago

    Seemingly none are vetted and reviewed yet

    • mulemisterX 10 hours ago ago

      That's our job.

      • esafak 10 hours ago ago

        Ain't nobody paying me to do that. It's kinda sad that maths is being reduced to checking the AI's work.

        • kozikow 10 hours ago ago

          Not just maths

          In SWE as well - this is what I do most of the day

        • zeroonetwothree 9 hours ago ago

          Always has been

      • TeeWEE 8 hours ago ago

        No it’s OpenAI’s job. They are acting as a meat proxy

        • red75prime 8 hours ago ago

          "If you have nothing to say, don't post a chatbot's responses, because anyone can ask the chatbot directly if they wanted to"-principle? Well, people can't ask their chatbot directly, because it's not public.

    • schleck8 10 hours ago ago

      Most are formalized in Lean, about 80% of what I checked

  • applicative 10 hours ago ago

    I wonder if the Lean compiler can change it's license so that a for-profit corporation can only use it if it pays, say, a few hundred billion dollars. This is the correct path.

  • nautilus12 10 hours ago ago

    Have any real mathematicians working on these problems reviewed any of these and determined if they are just gobbledegook or not?

    The ones with lean proofs could still be formulated incorrectly

  • hi__dang 10 hours ago ago

    Mathematics is solved.

  • globalnode 6 hours ago ago

    imagine your a post grad maths student looking for hard problems to solve, and theyre all solved..

  • redox99 11 hours ago ago

    The stochastic parrots have predicted the next token once again.

  • baggy_trough 8 hours ago ago

    Stochastic parrot truthers in shambles.

    • koe123 5 hours ago ago

      Can you explain it without reaching for lofty things like consciousness? To me stochastic parrots is literally how it works given that it’s “just” the most impressive data fit we’ve ever done. Apparently generating mathematics reasoning traces + verifying them with lean works super well.

    • slopinthebag 2 hours ago ago

      no ur right it's actually god

    • AmazingEveryDay 7 hours ago ago

      It is more of the usual though isn't it? OpenAI cribbing off of mathematicians that have used their services; deciding to put a lot of compute behind fruitful areas of endevour; getting results, then taking credit.

      • oh_no 6 hours ago ago

        no, they did not steal the notes of 100s of people working on these 100s of problems, be serious

      • baggy_trough 6 hours ago ago

        It’s wonderful.

  • mathisfun123 11 hours ago ago

    With so many results in so many different areas no way they even remotely spot checked well enough.

    Prediction: one of these is wrong and this (publicity stunt) will backfire.

    Edit: don't tell me about lean. For lean to function as a proof certificate you need to represent the theorem correctly. Again: good luck doing that across such a broad swath of problems.

    • jojva 11 hours ago ago

      You have not read their readme:

      > Some of the unformalized results could have issues. We will endeavor to fix any such issues quickly. We are also exploring community-hosted repositories for these materials.

      • mathisfun123 11 hours ago ago

        i have and i'm exactly saying that if it comes to pass one of them is wrong it's going to backfire. ie yes that's my exact point/bet.

        • stevenhuang 10 hours ago ago

          I don't think anyone would particularly care if only one of them is wrong, if most are correct.

          If they are all wrong, that's when it would backfire.

    • bravoetch 11 hours ago ago

      What does a backfire look like? It's ok to be wrong in the science/math world.

      • mathisfun123 11 hours ago ago

        of course in science/math it is but it's not okay if you're a business selling supercalifragilisticexpialidocious infallible intelligence.

        • bravoetch 11 hours ago ago

          Do they claim that's the case? I don't think they do.

          • mathisfun123 10 hours ago ago

            does company A making product B claim that the product is robust and consistent? is this a serious question?

            • zamadatix 10 hours ago ago

              If you were waiting for companies to sell engines that never break down you'd still be stuck pre industrial revolution while the rest of the world has been to space.

              The question never if something works 100% of the time but how often it breaks and how that fits the need well. Solving one of these problems is a massive undertaking and accomplishment for the best minds, solving hundreds in a month but being wrong about 10% or something would likely not be the death knell you believe it to be.

    • orlp 10 hours ago ago

      It's likely that way more than just one of these is wrong. But even if it turns out 80% is wrong this is still 100+ results...

  • applicative 10 hours ago ago

    Why didn't they just give mathematicians access, so they could at least understand and write up the results in publishable form?

  • sandworm101 10 hours ago ago

    So the million monkeys at a million typewriters have churned out 700 shakespeares, but they need me for spellcheck?

    • utopcell 9 hours ago ago

      Nobody needs you to do anything, not with that attitude.

  • mi_lk 11 hours ago ago

    Curious if Sébastien Bubeck still work at OpenAI? He came out quite dirty after Navier-Stokes drama

  • senderista 11 hours ago ago

    Good to see they're engaging with the mathematical community, even if they had to be publicly shamed into doing so.

    • sebzim4500 10 hours ago ago

      This is the opposite of what the mathematical community were asking for. I say this as someone who strongly approves of this approach.

      • senderista 10 hours ago ago

        Yeah I'm not sure they met them even halfway.