The Navier–Stokes Millennium Prize Problem

(simonwillison.net)

132 points | by tosh 1 hour ago

33 comments

  • sdcfgy 53 minutes ago
    My take home from this entire drama is that one should not use LLM services for confidential or proprietary information as they all seem to be run by assholes. And you’re sending them everything you are doing. Would you send your lab notebook to an asshole? Hell no.

    I say that as a mathematician (on paper) who perhaps surprisingly doesn’t give a crap about the problem itself.

    • pansa2 4 minutes ago
      > one should not use LLM services for confidential or proprietary information

      That’s obvious, isn’t it? Just like you wouldn’t upload your confidential documents to an online spellchecker, or your proprietary code to an online compiler?

      • sdcfgy 3 minutes ago
        Everyone, including the security services in my country, uses OneDrive and O365.

        I don’t get it.

    • johanvts 14 minutes ago
      Or use Lumo from Proton. Are there any other privacy first companies offering LLMs?
    • walrus01 18 minutes ago
      > as they all seem to be run by assholes.

      Wait until you learn how many other SaaS and web 2.0 and cloud based things are also run by assholes.

      You know this metaphor? https://en.wikipedia.org/wiki/Turtles_all_the_way_down

      But instead of turtles, it's assholes.

      But more seriously, no, none of what I wrote above is an attempt to excuse or play down the specific role of assholes in large AI companies.

      • sdcfgy 4 minutes ago
        Of course. I work for assholes. The thing is they’re getting out assholed by several orders of magnitude here.
    • Blikkentrekker 12 minutes ago
      That they do it is just concerning to me in that it says that home-ran models just aren't good enough. Surely researchers like this have the processing power to run them at home, they just don't have the processing power to train models of comparable level.

      This is something I feared would happen and where open source would be left behind. Maybe they can do something with crowd-sourcing computational power from volunteirs. They were after all able to get Leela Chess Zero to be comparable to AlphaZero by training from volunteer processing power but it seems to me we live in a world now where the best models keep their stuff closed.

      In imagine generation too. I'm not sure how well Stable Diffusion can compete in following instructions with all those advanced models that are kept secret.

    • junofan 48 minutes ago
      Their privacy policy for normie subscribers says in plain English they use your Personal Data for research. I think it’s pretty unreasonable to use the service and expect otherwise.
      • hansvm 39 minutes ago
        My doctor's privacy policy is a bit more abusive than OpenAI's. It exists mostly because of a $%^&&* legal framework rather than malice, and I've grown accustomed to "if I don't want to die then I sign away these rights." Despite my having theoretically signed my soul away, my doctor isn't selling personal information to my exes or to life insurance companies (though they could in the US; that extremely personal information is no longer mine). OpenAI is engaging in the "technically legal maybe we'll see but obviously unintended" side of this transaction, and maybe that works out for them, but I wouldn't personally choose to be a shill for "it's unreasonble to expect somebody with 'legal' permission to do something other than the maximum 'legally' permitted" if I were in your shoes.
      • grey-area 23 minutes ago
        I think those are pretty unreasonable terms and it’s sad that you’re defending them.

        Does google docs own the content of docs you make with it? Does Apple claim ownership of discoveries made using their tools?

      • ZeWaka 43 minutes ago
        >implying most users read them
      • sdcfgy 42 minutes ago
        If that is the case, why on earth would you use it in any professional setting?
        • eru 32 minutes ago
          Depends on your profession? I sometimes work on open source code as part of my professional duties. Nothing that goes on there is necessary to keep private.
          • sdcfgy 21 minutes ago
            Vulnerabilities, responsible disclosure etc?
        • ragebol 28 minutes ago
          Because it's cheap and easy. Hold for many such questions...

          Why is all of the world dependent on tech an ever more hostile US? Same answer.

        • Simran-B 31 minutes ago
          Aren't business consulting firms even worse? They are explicitly for business and there are known cases where they shared confidential information of one of their customers with another one.
      • AlphaSite 15 minutes ago
        There’s an opt out so I’m blade for now.
      • shiandow 20 minutes ago
        I think it’s pretty unreasonable to use the service.
      • andersmurphy 43 minutes ago
        You'd think theft would still be illegal regardless of what a privacy policy says.
        • georgemcbay 30 minutes ago
          > You'd think theft would still be illegal regardless of what a privacy policy says.

          I wish that were true, but I live in the United States and it is 2026.

          The President of the United States rug-pulls memecoin crypto and regularly pardons people like Paul Walczak (who was convicted of massive payroll fraud) in exchange for large donations.

          I wouldn't make any assumptions about what is considered theft anymore, at least not when it is being committed by people who have enough money to be above the law.

  • keremk 25 minutes ago
    Occam's Razor says: "They heard this problem is solved or about to be solved amongst the rest of the other problems. They prioritized this and put substantial compute with their newest model and solved it." I know everyone loves juicy rumors, theories etc. but honestly that is the simplest and most plausible explanation given the state of AI improvement now. Obviously spending 15 million on a problem is not a slam dunk decision even for a company like OpenAI but if it has a significantly high chance of solving it and their competitor will be claiming they solved it, then it raises the stakes and they go after it. In fact this is the most rational and also curiosity-driven thing to do and totally what I would have expected from any frontier lab. Of course if one wants to prove their confirmation biases that they train on sessions or be able to identify individual users, the non-zero chance of that being also another explanation is attractive enough to wet their appetites.
    • PowerElectronix 1 minute ago
      Funny how they did not solve any of the other problems, just the one where there was already solutions to the NS with some restrictions in their chats, and their solution seems to derive from those.
    • l5870uoo9y 17 minutes ago
      Given that the solution took a somewhat “unusual” approach, I find it even more unlikely that an AI model would have come up with this on its own.
    • npiano 23 minutes ago
      Why is that simple or plausible? Why is simpler or more plausible than lifting an almost-finished solution from a researcher's account?
      • ghshephard 20 minutes ago
        Because they solved different problems, and where there was overlap, the solutions look different?
      • karmasimida 13 minutes ago
        Because it is not finished. Their follow up claim is that OpenAI’s approach looks like another proof they had been working on the side, but hasn’t published yet
        • 20k 7 minutes ago
          OpenAI have admitted their new model they used was trained on prompts at around the time that researcher was working on it, so it seems self evident that it was used as part of the millennium solution
    • alangibson 12 minutes ago
      This is not the simplest explanation.

      The most direct line from problem to proof is OpenAI building off of conversations the mathematicians had with their AI.

  • tosh 1 hour ago
    > My two favourite hypothetical questions regarding this used to be:

    > If I'm running Codex and one of my API keys accidentally gets consumed in the context, what are the chances that someone else might ask for an API key in the future and get mine back? (I asked someone at OpenAI once and they called this the "regurgitation" problem and assured me that they take great pains to prevent that... but wouldn't describe how.)

    > If I brainstorm with ChatGPT about potential new directions for my company, what's the chance that information might be exposed to a competitor in six months' time who asks "what might company X plan to do next"?

    > My new preferred hypothetical for this is:

    > If I use ChatGPT to help me partially solve a Millennium Prize problem, what are the chances that my work will influence training such that a later model helps someone else solve it first?

    • feverzsj 32 minutes ago
      LLM can't be trained that easily. More like actual human are checking your logs and stealing valuable things from you.
      • rzzzt 11 minutes ago
        They wouldn't appear in weights but could be added to the context. My conversations regularly go "regarding your Java problem"... which was a separate item in the history from earlier. As long as I only see these (and nobody else sees mine), it can be helpful.
      • grey-area 14 minutes ago
        Or searching anonymised logs for mentions of this problem and using that as part of the context or training.

        This would work just as well and have plausible deniability.

      • FeepingCreature 8 minutes ago
        LLMs can learn from one sample.
    • weinzierl 39 minutes ago
      Maybe just a rumor of a high value target having their API keys accidentally consumed in the context...
  • bob1029 46 minutes ago
    > ... we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors ...

    I've observed this exact effect last week. I made a discovery regarding a stepwise performance improvement in a codebase. I shared the benchmark results with a peer and within 12 hours they replicated the same. We had both been looking for this for years.

    I think giving someone hope that an answer exists might as well be the same thing as giving them the answer these days. Competition is a hell of a drug, and frontier LLMs aggressively compound that energy.

    • skylurk 33 minutes ago
      > I think giving someone hope that an answer exists might as well be the same thing as giving them the answer these days.

      If you read the history of major scientific discoveries, this has been the case for a long time. There are many things that were independently discovered by different people at nearly the same time. Once people know something is solved or solvable, it gets a relentless amount of focus.

      • teiferer 15 minutes ago
        Maybe that shows how scientific discoveries come to be. It's not a genius sitting alone in their chamber for a decade and then suddenly they emerge with this huge thing. That's Hollywood fiction. Scientific progress is the colaborative effort of countless researchers over long periods of time, communicating, exchanging ideas, many of them wrong, tweaking, trying, thinking, arguing. When a breakthrough happens then it's the tip of a mountain of work that came before it. If two individuals stand on that mountain and feel there is something somewhere then it's not too strange that they take the last step at roughly the same time because conditions were right. The preconditions were in place at that time, the results required for this were available and the focus was on this specific thing.

        I'd say it illustrates well that this last piece, the person celebrated for the achievement, is disproportionally overvalued and the rest of the work they are standing on is disproportionally ignored.

        • a_bonobo 11 minutes ago
          This ties back into AI, too, right?

          I always like to bring up how many decades of research, how many hundreds of years of entire PhD-theses, how many sleepless nights were used up to generate all the protein structure data that made up the corpus of Protein Data Bank - that was then hovered up by the AlphaFold team, and guess who got the Nobel Prize...

      • eru 30 minutes ago
        This seems to be the norm rather than the exception.

        On a tangent, the genius of people like eg Einstein is not so much that he came up with all these things: other people were close, but that he was a singular individual that did all of these discoveries, instead of five different guys all making some breakthrough here or there.

    • adg33 40 minutes ago
      It's something that happened before LLMs - multiple discovery. Calculus is a classic example.
      • jansport123 19 minutes ago
        Yes but in this case, the allegation is Leibniz literally looked into newtons notebooks
    • avs733 28 minutes ago
      If you believe the totality of the document, there was more shadiness in how OpenAI acted than just timing. Save other things they are accused of, the progression from rumors to replication would attract much less scrutiny. With those in mind timing begins to look suspicious at best.

      Would anyone be surprised if major model companies had tagged the accounts of competitor employees for extra tracking? Given the concerns about distillation and bench marking it hardly seems irrational, but how it is used matters quite a lot.

  • burrish 34 minutes ago
    >While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.

    What do you mean, as OpenAI employee, you cannot tell that his work has entered the training data ?

    But also correct me if I'm wrong, if the two mathematician were really close to finish this problem, and their conversation were used by OpenAI, shouldn't the Agent have succeeded way faster/efficiently instead of using "4.9 million messages and used about 300 billion output tokens."

    • jansport123 24 minutes ago
      I’m not a mathematician so take this with a grain of salt. Apparently terry tao commented that the approach used for the Euler paper can “probably” be used for solving NS but it’s still technically challenging and can probably be done with an LLM with a lot of compute. To me the crux of the issue is whether the insight were stolen so that the problem becomes something that is in the domain of LLMs. This is much different than LLMs coming up with the insight. OpenAI wants everyone to think the LLM came up with the insight and solved the thing by itself even though they have perhaps an army of researchers.
    • CSMastermind 19 minutes ago
      Having the chat logs enter the training data and having them have a meaningful influence on the ultimate result the model produces are very different things.

      The text for all the Goosebumps books are certainly in the training data and to some small amount influenced the solve. But their contribution was so vanishingly small it would seem absurd to say R L Stein should have recourse for contibuting to the solve.

    • feverzsj 29 minutes ago
      It's almost as if it was actually found by manually written brute-force algorithm running on OpenAI's massive computer cluster.
  • shellfishgene 1 hour ago
    What's missing from the story I think is the part about "...then had a breakthrough on August 15th. The mathematical rumour mill kicked into gear...". If only the two of them were working on the problem in secret, how did their breakthrough become a rumor?
    • tyre 57 minutes ago
      People talk. If you have a breakthrough solving one of the most famous problems outstanding, you're going to tell people.

      You'll say, "Don't tell anyone", which they will ignore because they get a rush and perceived status by sharing it. So then they tell someone, along with "Don't tell anyone", etc.

      It's a small enough world (both in academic math, one at Anthropic) that you get to OAI in very few hops.

      • 20k 3 minutes ago
        "The route to the Clay problem through a smooth force, options c and d in Fefferman’s statement of the problem, is the route Luis and Diego opened and the one Levent and I had quietly chosen to attack. Almost nobody else I know of was working on it. It is not the direction one arrives at in a few days by giving a model the problem statement. When I heard “forced,” it was a bright red flag."

        We know that OpenAI trained on their prompts, plagiarism is incredibly likely. The only thing we don't know is whether or not it was deliberate plagiarism yet

      • dguest 39 minutes ago
        The reality is also that most of your colleagues have no interest in stealing your work: they have their own work to do anyway, and having a colleague effervescing about whatever they are working on is kind of the norm in pure research. Just because they are making progress it doesn't mean they are about to do anything interesting.

        Also it's not like you're looking for a lost pair of car keys: just getting to the level where you can understand a problem well enough to "steal" it takes a huge amount of work. People are going to know if you're at the level where you could be a competitor.

        So in general it's pretty safe to talk generally about whatever you're doing.

    • dboreham 1 hour ago
      They weren't working in secret?
  • geraneum 27 minutes ago
    Having in mind the allegations by Apple against OpenAI, I don't find it unthinkable that there could've been some form of misconduct happening there.
  • _bobm 18 minutes ago
    Hah, what is the infrastructure which takes user sessions (chats with API keys, directions, navier-stokes math/progress) and regurgitates this into pre-training, RL, fine-tuning data? Or better, in-context data?

    People talk about the "compute" but what about the "storage"? Is storage exponentially greater, or soon to be, than the compute? Is the storage going to slow down growing to some constant rate, i.e. all people on earth using chatgpt, or no, on the contrary, it will keep growing?

    If there were any shady business, I do not condone it, but technologically we are not there yet for said shady business to happen.

    • awestroke 14 minutes ago
      AI companies use heuristics to filter sessions, then llms to further filter, then use various techniques too anonymize the session, then process it and add it to various datasets for further selection and refinement. they don't need huge storage for this.
  • feverzsj 41 minutes ago
    It could be much worse.

    OpenAI can easily identify these outstanding human behind their accounts. Human in OpenAI constantly check their logs for breakthrough. When they find something interesting, they brute force the result using their massive computing power.

    No LLM is even needed.

    • grey-area 18 minutes ago
      Or, using anonymised data, search for anyone seriously trying to tackle this problem - probably about 10 people in the entire world and use their ideas as a starting point. They wouldn’t even need to be watching specific accounts or using de-anonymised data if they know what they’re looking for.
  • maciejzj 3 minutes ago
    I know that this may be somewhat dramatised and even infantile, but my reflection is that in the world run by these reckless AI companies everyone looses. Navier-Stokes is solved but it feels like no one has won anything, controversy prevails, there is no glory in the math breakthrough. There is hardly anything to cherish, and even the guys at the top of it in OA who sit on the (supposedly) superhuman intelligence come across as massive losers and frauds.
  • civvv 51 minutes ago
    LLM’s seem very good at solving mathematical problems of which there is an enormous amount of exisiting work/attempts in their training data. This is an amazing capability, but does not convince me that these models are «thinking» or «reasoning» in the way a human does. A human mathematician could in theory categorize/discover an entirely new field of mathematics tomorrow, based purely on their «human intelligence», I wonder if we will see similar examples by LLM’s soon. It seems to me currently impossible that LLM’s can replace human mathematicians, because of their (assumption) likely dependence on human input in the sense of enormous amounts of pre-existing attempts/data.

    If an entirely new problem, within a new field of mathematics were to appear tomorrow, I highly doubt an LLM would be useful at all on their own. Is this the «ultimate ASI test»?

    • krona 33 minutes ago
      Extreme temperature levels (>2.0) can push a GPT of its manifold, essentially producing predictions barely distinguishable from random noise (it flattens the probability distribution of the next token). In theory this could predict anything including the next field of mathematics (infinite monkey theorem) but realistically that would never happen.

      However, how to we know the next field of mathematics isn't a novel combinations of several other sub-fields? That level of mathematics would be indistinguishable from magic to most people and so in their eyes the GPT did something truly inventive.

    • eru 24 minutes ago
      A lot of what humans do is combining old ideas.

      And an LLM could in theory also stumble upon entirely new ideas: there's randomness in how they generate their reasoning and answers after all.

      I suspect that we are seeing a lot of advances coming from the combination of existing but somewhat obscure knowledge coming from LLMs at the moment, because LLMs are really good at this. At least compared to humans.

      Even before our AI friends became good, they were already known for having read approximately every paper and every textbook published in any language. You only need to increase intelligence a fairly small amount from there to get to something like the 'convex hull' of human knowledge.

      Compare https://slatestarcodex.com/2016/11/17/the-alzheimer-photo/

      The gist is that basically whenever anyone comes up with a new method you get a big burst of activity of picking up all the now lower hanging fruit, that was previously out of reach.

  • karmasimida 11 minutes ago
    Use Bedrock or any kind of big tech hosted version of the frontier labs can be a solution. I think if secrecy is of utmost importance to you, then do not send data to first parties
  • jwr 28 minutes ago
    I always thought it was enough to switch off the "Improve the model for everyone" setting on chatgpt.com:

    "Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy."

    But apparently there is also an entire completely different route "Do not train on my data"?

    Does this mean that before I submitted the "Do not train on my data" request, my data was used for training in spite of "Improve the model for everyone" being turned off?

    We are getting to facebook/meta-levels of privacy settings obfuscation.

    • teiferer 9 minutes ago
      That's because what people enter into LLMs is the last gold there is out there. Everything else is already scraped or ensloppified.

      Maybe next step is to filter your input client side through an unknown number of obfuscators where you ask LLMs to rephrase your question (onion router idea) such that no single provider can be certain that this is human input and not some slop feedback loop.

  • tu26muwu 23 minutes ago
    > The discovery is somewhat overshadowed by accusations of skulduggery from Tristan Buckmaster [...]

    I think you should not say that. Buckmaster did only state his version of events and was very clear on that he did not make any accusations at all.

    To quote from his statement pdf:

    > I am not accusing anyone of anything.

  • aadyachinubhai 50 minutes ago
    LLMs can't contribute good code to some of the good OSS math libraries, How is it even solving these problems?
    • feverzsj 21 minutes ago
      No one knows if it's actually LLM doing the heavy weight. It could be just human written brute force algorithm running on their massive computer cluster.
    • emil-lp 40 minutes ago
      That's a good question.

      It is able to contribute code, but maybe not good code.

      It's the same in math: it's able to solve problems, but not necessarily in a good way with a human readable code.

      Math papers are a lot like software:

      - theorems are like API

      - lemmata like internal/private function API

      - definitions are like types

      - the proofs are the implementation

      The proofs of ChatGPT are not necessarily readable or maintainable.

    • linkgoron 27 minutes ago
      You don't need a GOOD proof, just A proof.
    • matrix2596 45 minutes ago
      search, verifiability and compute
  • rao-v 50 minutes ago
    It’s reasonable to wonder about what chat usage data gets into models (to be honest probably quite little - carefully curating training data and creating higher quality synth data seems to be the current approach) and the implied risk to privacy and creativity (every new patent filed this year probably touched a model before filing).

    What I cannot reconcile is the timeline and the concern in this specific case.

    I don’t think training pipelines are anything close to the level of continuous training needed to incorporate Aug 15th ideas into a model that generates a breakthrough early Sept. Either OpenAI nakedly had someone with mathematical understanding dig into a specific user’s chats (a massive red flag) or this really is poor handling of a more classic parallel discovery situation (with one party clearly having worked on it longer)

    • KeplerBoy 34 minutes ago
      I would assume a lot of codex data goes back into training. A well steered session is extremely valuable data.
  • caughtinthought 1 hour ago
    Basically no new info here, not really sure why this post needed to be written tbh.
    • kzrdude 49 minutes ago
      On the contrary, a level-headed summary that gathers information from all the different sources is necessary.
      • caughtinthought 30 minutes ago
        Sounds like a great use case for an LLM
        • teiferer 2 minutes ago
          They are busy generating pelicans on bicycles. The summary therefore needs to be written by a human.

          (This was a joke. I value Simon's role in the community.)

    • fimi 55 minutes ago
      Yes, you have ability to get information very quickly, but not everyone does.
    • iso1337 1 hour ago
      Agreed. This person just loves to shill their blog and their pelican benchmark.
      • jdlshore 54 minutes ago
        There’s no shilling here. The person who posted the link isn’t the person who wrote the blog.
        • sk4rekr0w 29 minutes ago
          He is in the HN ingroup of whenever he posts a blog, a prominent HN mod or poster with 1 billion karma will inevitably post it, and so it goes.

          It's sort of like the music industry.

          • neuroticnews25 26 minutes ago
            Does poster karma boost the post in any way?
            • sk4rekr0w 24 minutes ago
              I don't believe HN boosts their algo but I think people recognize posters like celebs. I recognize tosh, simonw, etc.

              Also it's an easy way to farm karma. The pelican guy posted a pelican, let's upvote it? It's at least a bit of psychosis / mass effect in this.

              For what it's worth, I like the pelicans but just making an observation

  • Simran-B 26 minutes ago
    I don't get the sales pitch, spend 15 million dollars to win a 1 million dollar price?

    Showing of the model's capabilities - okay, but it's not like it solved the problem on its own, and apparently not particularly efficient. Are there practical applications that justify the investment?

    • teiferer 4 minutes ago
      How is that any different from any other academic research? Every PhD candidate solves problems essentially nobody cares about. They don't even get $1M, they get nothing.

      There are two benefits though.

      One is recognition. Cred. The PhD candidate gets to put a ", Ph.D." behind their name, opening doors to future academic employment or other endeavors where people value titles. The AI lab gets to say their tech solved sth that humanity wanted bad for a long time. Both cases with substantial financial upside (higher income for Mr. PhD and higher company valuation for the AI lab).

      The other one is that this is how scientific progress works. $1M or not. That number was just a PR campaign by the math community to point to some goals. It's clear that it would cost more than $1M to get there.

    • grey-area 20 minutes ago
      The answer to this is obvious: the effect on the multi-billion dollar valuation in the imminent IPO.
  • protoman3000 31 minutes ago
    If they wanted to solve a Millenium prize problem so much, why did they not try to solve P-NP instead? It boggles the mind.
  • l5870uoo9y 43 minutes ago
    I run a small SaaS[1], like so many others, that uses AI to generate and optimize SQL. Getting this to perform optimally has been a lot of work and now I wonder if OpenAI is outright stealing this knowledge, which without a doubt is highly valuable to them.

    [1]: https://www.sqlai.ai

    • KeplerBoy 31 minutes ago
      SQL is so ubiquitous and the use case so obvious, there's no way they have not already been tracking performance and benchmaxxing on SQL queries for years.

      But I don't think openai will bother to release a competitor, the real threat is that anyone with a decent LLM and a harness to try a few queries will land at the same or a better query within minutes.

    • biorach 24 minutes ago
      It's extremely unlikely that there is anything interesting or novel in the optimisation of a small SaaS SQL
  • cs_throwaway 38 minutes ago
    Maybe someone on the NYU team forgot to opt out of “improve the model for everyone”.
  • throwaway63467 25 minutes ago
    Well OpenAI wants to maintain its edge and solving a Millennium problem is of course fantastic PR, and given that Anthropic seems to be working on that it’s not hard to imagine they wanted to be first. I don’t get how people buy into that whole “oh we heard models can solve Millennium prize problems now so we thought why not give it a go…” story - it’s a bit funny. This whole AI bubble is about hype and solving such problems is probably one of the best ways to keep this hype up so you can safely assume both OpenAI and Anthropic are using significant resources on these areas.
  • nicce 49 minutes ago
    Can’t wait to see the human verifying the results and then figure out that the AI model actually cheated and the results are not correct.
  • caidan 27 minutes ago
    I mean this is just the logical end game. The company/ies that control the uber mind will devour ALL useful or valuable work. Unlimited intelligence at unlimited scale means the value of humans for knowledge work goes to zero. We will eventually not have meaningful access to the uberminds because we will be pointless. At which point our Silicon Valley luminaries will really have no choice but to extinguish as many of us as possible for the greater good. After all if we are all pointless then that has to be weighed against are cost to Mother Earth. Clearly the only moral solution is to cull or allow to be culled some 98% or so down to a more sustainable, manageable population of curiosities. The end game of ai is the remnants of humanity in a zoo, and that’s if humans control the outcome… the machine minds might be more charitable as they would be less afraid…
  • josalhor 48 minutes ago
    I think this drama was blown up a bit out of proportion. The entire discourse I am seeing online seems to revolve around this:

    > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models

    I mean... yeah? What do you expect? What else can they say? How could you prove a negative in this case? I do not want to comment on specific OAI employee chat messages, but on the actual OAI discovery here.

    • 405error 23 minutes ago
      I can give you some context. 1. Terence Tao's mastodon explains the way this problem was solved does not in itself contribute much. LLMs (and in this case) produce massive, often unintelligible proofs that do not further understanding. It is often that in pursuit of solving these problems, many other discoveries are made. 2. There is a more serious question about scooping. If OAI is using chat data from researchers to make discoveries, essentially every researcher who chats with an LLM can get scooped. You could be 80% of your way to solving a problem, and LLM could solve the remaining 20%, and get all the credit. Years of your work could be scooped in an instant. If you're a PhD student, this is even worse. Here it's a world famous problem. But imagine you're a PhD student, working on your small but extremely career/progression critical problem, and you get scooped by an AI you talk to. No one is even going to care.
      • josalhor 16 minutes ago
        1. I know, but that is somewhat irrelevant to my question 2. I am indeed a CS PhD student (well, I am finishing now)

        > If OAI is using chat data from researchers to make discoveries, essentially every researcher who chats with an LLM can get scooped

        But they make very clear that they do train on this if you do not disable the setting. We can comment on the fact that this is opt out instead of opt in, but this discourse of OAI sniping the solution out of some researchers hands seems to be running on the best case speculation of the researchers having perfectly handled all their chats and discussions with other researchers and the worse case of OAI not having full pipeline control and I think that is an unfair assumption.

    • Fordec 30 minutes ago
      A very simple "these two pipelines don't connect up in our architecture, here's our internal high level network diagram combined with our data ingestion opt-out feature flag that we will stand by in court" as opposed to "yeah, we don't even entirely know how our own customer facing systems are connected to our training pipeline, but it probably didn't happen".
      • josalhor 21 minutes ago
        Have the other researchers opted out? On all their accounts? Through the entire time? And did they discuss this with anyone else? And did those people ask ChatGPT stuff? And did they disable it? If I was OpenAI, I would be very careful about my wording here when making claims of "we have never trained on any of their ideas directly or indirectly".
        • Fordec 15 minutes ago
          All very good questions that could end up in a court of law with a Millenium Prize on the line. In a competent world, these are very answerable from logs and considering the news cycle this is creating, should be able to be pulled up and made into a public postmortem in short order. You know, if it's all been above board that is. And if the researchers don't wish for that information to be public knowledge, it can be shared with the researchers promptly for a retraction of their statements lest some libel gets litigated.
    • emil-lp 38 minutes ago
      If they have zero-retention, then it is not possible.

      So what they are saying is that they don't have zero retention.

      • josalhor 33 minutes ago
        But why is this news? This is clearly described in their ToS and in the Settings to improve their models.
  • tyre 59 minutes ago
    Yeah, this was pretty shitty by OpenAI. Not surprising, sadly.

    Them being assholes, trying to exclude an author just because he worked at Anthropic, shows the kind of culture within (that part of) their organization. The focus wasn't on supporting academics or expanding research. It was on getting great marketing.

    If they had to burn millions of dollars solving a problem _that they thought was already being solved_ to do so, they'd do it.

  • sk4rekr0w 31 minutes ago
    Time to repeat the same angry mob style discussion again. Great job to the mods.
  • kdavis 50 minutes ago
    All your datum are belong to us!
  • dboreham 55 minutes ago
    The concept that "knowing something has been done" allows others to find the solution to an previously unsolvable problem is an old proven one. For example when Germany launched a rocket (V2 prototype) the British knew its rough trajectory and from spying it's rough size. Although they had previously believed that ballistic missiles weren't possible because no engine could provide the necessary thrust to weight ratio, given the obvious German launching of one, they went through all the known chemical compounds to arrive at the combination (Ethanol and LOX) used. [Story from RV Jones "Most Secret War"].
  • octocop 50 minutes ago
    Is there a good tldr on this topic?
    • emil-lp 38 minutes ago
      This article is the tldr.
  • avs733 53 minutes ago
    I’ll go back to the point about authorship. I’m Not a mathematician but I am in academia. if you are fucking around with authorship you are immediately suspect.

    That aspect alone would/should be unthinkable to any serious academic. Authorship reflects who did the work and changing it for business competition reasons should be a red flag for multiple different reasons. They include, the sheer tactlessness of treating a major theoretical advancement as a competitive posturing first, the norms of academia second, and all the misunderstandings of the culture of the disciplines culture that people will now suspect are hiding beneath the visible surface (insert topography joke).

    Math as a field is fairly unique even in how they list authorship. It was long the norm that authorship to be alphabetical because the idea of first, second, senior etc authorship is harder to define than many other fields.

    “The stated rationale for alphabetical order is that it treats co-authorship as intellectually joint work: every listed author’s name carries equal weight, and no one has to negotiate, or be seen to negotiate, over billing. That is a genuine advantage over position-coded conventions, where disputes over who is “first author” are one of the most common sources of authorship conflict in fields that use them” [0]

    That norm is changing, slowly, but one option people are pursuing is notable: randomized author order. Their is a perception that alphabetical is too biased…that’s the world OpenAI is stepping into when they make that offer of authorship to one scholar with a demand that he exclude his partner.

    I can’t speak to the facts of anything else in this, but if a grad student came to me and said someone made them that offer, I would tell them to run and if they were brave report it.

    [0] a to the point lay description of the history of math authorship can be found here: https://casrai.org/guides/mathematics-alphabetical-authorshi...

    • sobellian 34 minutes ago
      OAI's side of the story is that they discovered the approaches (and indeed solved problems - Euler equations vs NS equations) differed. They then offered Buckmaster lead authorship of OAI's proof, without Alpöge. But they never demanded that Alpöge be stripped of coauthorship on resolving the regularity of the Euler equations. At least that's the claim.

      https://xcancel.com/SebastienBubeck/status/20973794116915163...

    • emil-lp 37 minutes ago
      Yes, there is no doubt about scientific misconduct. I think the entire math community agree on that.
  • krapcys 27 minutes ago
    [dead]