49 comments

  • oefrha a day ago ago

    Loosely related, you may want to check your own privacy settings at

    https://chatgpt.com/codex/cloud/settings/data#settings/DataC...

    https://claude.ai/new#settings/data-privacy-controls

    I just realized I've been happily "improving the model for everyone"...

    • tristanj a day ago ago

      That OpenAI setting helps, but there is a better way to do it. To completely opt-out of training, submit a request via the OpenAI privacy portal.

      Visit this website https://privacy.openai.com/policies/en/ , click "Make a Privacy Request", choose "Do not train on my content", and complete the form. That submits a formal objection to training on your data, as required by GDPR/your local legislation.

      • rich_sasha a day ago ago

        Haha. I am 100% these companies will ignore this if they choose to. Just as they played fast and loose with copyright rules.

        They would do it, the say “ah sorry chaps, impossible to extract it from the dataset by now, anyway we anonymized it so can’t tell what’s what, and we can’t risk losing to China. Oh look - did you see Superman fly outside?”.

        • nicce a day ago ago

          European users have right to be forgotten. Waiting for the court order to delete all models.

        • tristanj a day ago ago

          This form is legally binding and has more legal weight than just clicking a toggle. If they still train on my data, they can get sued, and I'll get a payout.

          • vuurmot a day ago ago

            Fighting legally against one of the richest companies in existence, with the entire American apparatus behind it, is a brave move

            • bot403 a day ago ago

              The world is pretty pissed at the u.s. and seeking sovereign models and ai. Richest or not if the companies are wise they will try a bit not to piss off the entire rest of the world. At this point many countries are willing to cut off the u.s. even if they take a short or medium term hit.

          • m4rtink a day ago ago

            As legaly binding as all the data the license and ToS of which they ignored when scrapping - before trying to sell it back with an eternal subscription to their plagiarism machine ?

          • 63stack a day ago ago

            Are there any precedents that companies got sued on this? Not that I doubt you, but it would be nice to see if it actually has teeth or not.

          • moljac024 a day ago ago

            How would you prove they trained on your data specifically?

            • Ancapistani a day ago ago

              Showing that a novel mathematical approach was in your prompts shortly before the model “proposes” that approach to relative amateurs would be ideal, if only we had a case like that…

              • itemize123 a day ago ago

                and still they denied it - it only shows that we have little recourse.

                • Ancapistani a day ago ago

                  It’s possible that they didn’t train on it, and the approach was derived by the LLM.

                  What’s a problem is that they haven’t outright denied it. That could be caution and them doing their due diligence first, it could be that it was intentional and they didn’t expect to get caught, or it could be because they have no way of knowing themselves.

        • undefined a day ago ago
          [deleted]
      • simianwords a day ago ago

        > but there is a better way to do it.

        This looks incorrect. OpenAI have flat out said that there's no need of doing this and both ways are equivalent

        > We respect our users' choice whether to use their data to “improve our models for everyone” regardless of where they express that choice. Users can opt out in the in-app settings or indeed also in our privacy portal. They do not need to opt out in both places, and we will make this clearer in our Help Center.

        https://x.com/thsottiaux/status/2097746417012166816

        • my-huge-pony a day ago ago

          Even if that's true right now, if one of the options is legally binding and the other is "we promise not to eat your data... for now", the first option is still better.

          • simianwords a day ago ago

            so what option do you leave them with? they are giving both alternatives and clearly stating both do the same thing underneath. there's literally nothing else they could have done here.

            • tristanj a day ago ago

              Option A can be accidentally undone with a few clicks.

              Option B is legally binding and permanent.

              • simianwords a day ago ago

                I get it, but what option does OpenAI have here other than to present the two options?

        • tristanj a day ago ago

          They're not equivalent.

          The one in ChatGPT settings only applies to ChatGPT. The one on the OpenAI privacy portal applies to all OpenAI products, present and future.

  • tetrisgm a day ago ago

    The problem is that eventually, there will be no humans who can follow the results AI will give. That’s the endgame for this tech: to produce knowledge at speeds and quality beyond what we can

    • rsfern a day ago ago

      That’s the prevailing narrative, but I think this controversy calls it into question to some extent. If the OpenAI result wouldn’t have been possible without experts seeding the training data with feedback on promising solution routes, there’s less reason to believe this, IMO. More information and transparency is needed

      • bot403 a day ago ago

        That's the pickle isn't it? AI solved it with human help. But it's standing on their shoulders. Neither AI nor the humans got it alone.

        Humanity is better having solved this issue. But which humans were credited and benefited is the issue.

        • rsfern a day ago ago

          I agree (and so does Buckmaster based on his written statement) that we are better having solved this.

          But I disagree that which humans were credited is the heart of the issue in this particular controversy. The question is what do you need to bring to the table for a result like this. A pre-release frontier model trained on the open literature and $15 million of inference? Or all that plus a year of the experts finding the path to the solution for the model to run with?

          I think it makes a huge difference in terms of what we think the future of mathematical research will be like, and whether we should still encourage students to go into this field, which was the original topic of this thread

        • bot404 a day ago ago

          [dead]

    • uargos a day ago ago

      Which asks the question of what knowledge is. Can knowledge be super human ? Or is knowledge a human matter ? If so (like i believe), then what those ai labs are doing is far from the end of the story. Because the goal of science is not to produce a certificate of something, but more to produce an explanation that can fit in a human brain, that can be reasoned on, and that can be retargeted. In this sense, producing a million lines proof is not really producing knowledge, even less so doing science.

      • michaeljx a day ago ago

        Knowing something is proven true is knowledge, even if the mechanics of the proof are not understood.

        • trolleski a day ago ago

          In maths, you go through proofs every step of the way. Same in physics, you start by performing even the most basic experiments, and build your way up.

        • rented_mule a day ago ago

          If the mechanics of the proof are not understood, how do we know it's been proven? Math proofs are not a "trust me, bro" kind of thing.

        • uargos 21 hours ago ago

          It's not knowledge, it's belief.

      • curt15 a day ago ago

        Suppose there is an oracle that can tell you whether a particular result in maths or physics is true. Would research still have any value?

    • menaerus a day ago ago

      In certain, more complex, software engineering domains this almost became true as of today.

    • smitty1e a day ago ago

      Yeah, but then you can use another round of AI to break these things down to crayon level, no?

      Or is it just all PFM? (Pure, Fanciful Magic)

  • tristanj a day ago ago

    This post is out of date. OpenAI quietly updated the references on their paper earlier today and added several authors.

    • kzrdude a day ago ago

      The main claim, that "they do not seem to have any mathematicians capable of understanding what they put out", was also corroborated by Sebastien (OpenAI) who explained they don't have any experts on Navier-Stokes.

  • japgolly a day ago ago

    > our hypodissipative result (which is not public), but as I understand it part of their training data

    How did it become part of their training data if it wasn't public? /confused

    • tristanj 4 hours ago ago

      It's baseless speculation and unfounded accusations of plagarism. OpenAI now categorically denies this: https://www.nytimes.com/2026/09/10/science/tristan-buckmaste...

      “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.”

      “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way.”

      Buckmaster reached the key result on August 15, far after the training cutoff.

    • vessenes a day ago ago

      This is speculation right now. The idea would be that if someone used the product and granted training rights, which is the default for many subscription levels, then some knowledge would have been imparted into the general weights of the new model.

      oAI has made clear they did not specifically pull in any user data to context for this run.

      • dpiers a day ago ago

        Tristan Buckmaster’s post cited extensive use of LLMs in the process of his collaboration with Levent:

        “We used several LLMs throughout: Anthropic’s Claude, OpenAI’s Codex, especially with GPT-5.6 Sol and, more recently, Astra.

        The latter was only used for writeups and auditing our arguments. For most of the past year progress was slow. We worked through the literature and upgraded various preliminary results, up to obtaining finite time blow up for the Incompressible Porous Media equation (with smooth forcing).

        This was until about a month ago, when we had real progress: on August 15th, we obtained the blow up results, with smooth forcing, for both Boussinesq and Euler. I can say the first LLM generated proof Levent sent me was the most horrendous I have ever read; we verified it on Lean on August 22nd. Since this point, we have been working around the clock to understand this proof and turn it into something readable.”

        The OpenAI research post states they began training GPT-6 internally on August 28th, and that user chats are used to train models.

        “We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .”

        For an incredibly niche topic like this, I believe it’s extremely likely that Buckmaster/Levant’s work would influence the direction of OpenAI’s agents’ work even as a de-identified drop in the overall bucket of training data.

        • vessenes 20 hours ago ago

          Yep, it's possible. But, we have literally no idea how much a set of prompts would impact training as far as general usefulness. I don't think we even know if Tristan's said he allowed training on his prompting or not. This is about money, ego, primacy, all the usual mathematician priority disputes.

      • rramadass a day ago ago

        Tristan Buckmaster, the mathematician at the center of it (https://cims.nyu.edu/~tristanb/) put out a public statement that everybody should read (pdf) - https://cims.nyu.edu/~tristanb/statement.pdf

        So what might have been the incentive for OpenAI to do all this shenanigans? It might have to do with getting its models certified for AGI and getting out of lockin with Microsoft - https://deadneurons.substack.com/p/the-quiet-unwinding-of-mi...

        • vessenes 15 hours ago ago

          Shenanigans is an inaccurate word; it implies underhanded behavior that's hidden / concealed. I think "to act so aggressively" is more balanced.

          Here's my answer: If you think we're getting to AGI in the next 9 months, then you believe, with all your heart, that these problems will fall soon. However, there's an ocean to boil in terms of what you could point your limited clusters at. In the meantime, the market is desperate for any sign your company might be first to AGI. Therefore, news of tractability with current models might focus an organization intensely - internally they have a huge leg up on the public, and therefore it's minimal compute to check - and if they are successful, they get approximately $50 million of free PR, likely adding 10-20% to their valuation.

          Likewise someone like Tristan is fighting for his (metaphorical) life right now, hoping to preserve his claims of primacy and have a shot at some of that prize money, despite being only partway to a full solution for N-S.

          I don't think we see any behavior at all that isn't simple to understand and well described by the setup here, but tell me what you see differently.

    • WoodenChair a day ago ago

      One of the researchers was using the product. They train on your private interactions unless you explicitly opt out in the settings.

      • rrobukef a day ago ago

        It wouldn't surprise me that even if you opt-out they still train on 'de-identified' chat data. After all, facts cannot be copyrighted and one cannot be sued for unethical behaviour.

        • ozgung a day ago ago

          I think Sam Altman already admitted that is a possibility.

  • simianwords a day ago ago

    Can anyone explain to me why Tristan simply didn't go to settings page and turn off the thing? Especially when Levent was collaborating with him and Levent is completely aware of how the training data is used and the implications thereby?

    If it is oversight, then that's ok and OpenAI can volunteer to make him the lead author which they did. But he's pissed that OpenAI is not letting Levent as well, who had access to internal Anthropic models. So this guy thinks

    1. oh my bad i forgot to turn off the consent thing in settings page

    2. also i'll collaborate with a literal Anthropic employee who has access to their internal models

    3. i'll also reject OpenAI's deal to be the lead author because i want an employee of the competitor to be a part of it

    I don't get the mindset.