NHacker Next
  • new
  • past
  • show
  • ask
  • show
  • jobs
  • submit
▲Sharing AI progress in mathematics (openai.com)
jboggan 3 hours ago [-]
I was a graph theory junkie long ago and even moved to Budapest for awhile to study among the greats. While I was there I started working on Barnette's Conjecture which came to occupy my thoughts over the next 24 years of my life, on and off as I worked in many different fields. Last summer I even thought for a few days that I had actually solved it.

But it's supposedly proven here - problem 180. I don't know what to think exactly. I spent thousands of hours on that problem. I really enjoyed it. Hearing that it is solved somehow makes me sad in a far-off way, like hearing an ex-girlfriend died suddenly in a car crash. I don't know, there's probably a lot of people feeling odd emotions tonight.

There's no Lean proof for this one so I'm digesting the paper. On the surface it looks like an approach I considered 24 years ago and abandoned.

I revisited the problem this summer, along with my partial solutions, when the previous round of stunning proofs came out. Several hours of work with Fable simply convinced me it wasn't yet solvable and reinforced how hard of a problem it was.

nilkn 10 minutes ago [-]
> Several hours of work with Fable simply convinced me it wasn't yet solvable and reinforced how hard of a problem it was.

This is the part that gives me the strangest feeling about it all, because you're not the only one with this experience. I've experienced this too on different problems, as have many researchers across many fields.

I disagree with the Fields Medalists on the majority of their complaints. AI math is happening and there's no going back. However, on one point I increasingly agree: virtually none of this stuff is possible with technology any normal citizen has access to. I have no problem with AI models making revolutionary advances in math or science. Where I start to have a problem is when the AI models making these advances are tightly withheld, proprietary, and seemingly never released with these capabilities intact. This has been the case for all of 2026 so far.

I suspect that this is in fact the source of much of the angst. None of this progress is reproducible outside of one or two teams inside OpenAI and Anthropic.

WheelsAtLarge 2 hours ago [-]
I have very little understanding of higher math, so I ask you: Was the proof due to a type of brute-force solution that could be solved had you gained enough information from reading others' work, or was it more like a proof that was sparked by an insight that came once a clue on how to solve it was put forward? I guess my question is: Was the problem proven by using a collection of everyone's work, or was it due to a brand-new insight?
jboggan 1 hours ago [-]
I'm still digesting the proof and translating a bit from the dual case back to the primal in which I most commonly thought about it. I don't think it was a brute force proof in the sense that it combined every possible paper and commentary. It's rather odd because I feel like most of the work on the conjecture was focused on an induction proof based around graph reductions, and this proof avoided those issues entirely by offering a concrete constructive proof of finding a Hamiltonian cycle. Rather, it explicitly selected the edges not in the Hamiltonian cycle, which is in line with previous attempts via the dual.

The "aha" insight for this is actually f**ing wild, it involves a complex valued exponential sum on the edges. I've seen a lot of clever counting arguments before in graph theory but this is the first time I've seen complex roots and annihilating terms like this, the symbolic manipulation tricks in this look like things out of quantum physics. I don't understand where this trick originated, I need to really digest this.

anilgulecha 1 hours ago [-]
BTW, reading your last paragraph reminds me of how Lee Sedol felt after move 37.
jboggan 1 hours ago [-]
Ironic, as I remember staying late at the Google office to watch that match live. I didn't really understand anything going on but I knew enough to be excited. What a decade.
groceryheist 1 hours ago [-]
WOW
derangedHorse 38 minutes ago [-]
> Was the problem proven by using a collection of everyone's work, or was it due to a brand-new insight?

Loaded question. A "brand-new insight" is still built off the work of others. A possibly better way to frame it would be in how many subjectively unintuitive logical leaps have been made from prior work.

ncr100 2 hours ago [-]
That's grief. The loss of ... the hope / future filled with challenges around this theory..? <3 to you.
lifeisloving 2 hours ago [-]
Condolences, im familiar with the feeling. I hope this AI thing somehow works out for the better and doesnt end up demotivating bright minds like yourself.
jboggan 2 hours ago [-]
Thanks. It's just funny, I literally spent thousands of hours with this problem over the last two decades, it helped me through some tough times. I'll never quite be able to think about it in the same way again. It was never much more than a hobby for me after I left mathematics as a career but it was something I took seriously for years.

I am not demotivated though, I have a great consumer privacy product coming out soon that I'm very excited about.

brookst 2 hours ago [-]
My favorite thing about your story is that you wrestled (enjoyably, it sounds) with a known problem for decades, but are finding fulfillment in an open ended problem that is exercising creativity about both problem and solution.

IMO that’s where AI is going: as soon as a problem can be formulated clearly enough, AI will trounce us humans. I have yet to see evidence that it can decide what problems are important at a remotely human level.

maximus_prime 1 hours ago [-]
Where can I learn more about your upcoming product?
jboggan 1 hours ago [-]
Shoot me an email, in my bio.
theteapot 1 hours ago [-]
If you wrote down any of your thoughts on the open Internet you are probably in some small - or possibly large, unattributed way, responsible for this result being possible.
jboggan 37 minutes ago [-]
Which is one reason I never really did. I probably should have but I always thought my attempts were too amateurish. Though I did manage to replicate some partial result papers that I didn't know about, lol. Writing openly would have saved me some years.
palmotea 1 hours ago [-]
> Condolences, im familiar with the feeling. I hope this AI thing somehow works out for the better and doesnt end up demotivating bright minds like yourself.

Who are you kidding. In a few decades, the rational thing will be to replace education with AI generated Cocomelon and Microdramas, to numb restless minds.

But some guys are gonna get really rich and powerful, and some others are going to get to enjoy a lot of sociopathic schadenfreude.

onesandofgrain 16 minutes ago [-]
Bot shill
adastra22 53 minutes ago [-]
> There's no Lean proof for this one

What is this then, vibes? Without a machine-checkable proof I'm not sure what to think of any of this.

jboggan 39 minutes ago [-]
Well I'm sure some people (maybe me if I had time) will do a write-up of this proof. It treads familiar ground for most of the setup, it's mostly the disk lemma and cancellation calculations that need to be understood, it's a fairly short paper and quite tractable.

I think it helps that basically everyone thinks this conjecture is true, it's just been so darn weird to attack. There's this odd thing that the induction proofs of this problem kept running into, which is that the N+1 condition would work except for in one tiny case when it could fail, but it would be covered by a very slightly stronger version of the conjecture. But then that would fail on one tiny case in induction, but you could solve that with another slightly stronger version. Etc., etc. I almost wondered if there were some sort of structure to the increasingly strong conditions and wanted to prove something about the meta-induction between the stronger conditions and the N's that they needed the next level to remain true. But that failed after 5 steps I think (Fable actually helped me write a few hundred test cases to explicitly show that pattern didn't continue forever, thank God).

BTW my existing test suite from previous proof attempts jives with this new algorithm, so I haven't seen any evidence yet that it's incorrect. Waiting for a Lean proof obviously.

8 minutes ago [-]
Daneel_ 25 minutes ago [-]
It might have been updated. Is this the lean? https://github.com/openai/math/blob/main/lean/docs/180.md
jboggan 11 minutes ago [-]
Lol it should be, but it doesn't seem complete. Line 49 just says "sorry"

/-- Cubic bipartite three-vertex-connected plane graphs have a Hamiltonian cycle. -/ def MainStatement : Prop := ∀ (V : Type u) [Fintype V] [DecidableEq V] (G : SimpleGraph V) [DecidableRel G.Adj], G.IsRegularOfDegree 3 → G.IsBipartite → Planar G → ThreeVertexConnected G → HasHamiltonianCycle G

theorem main : MainStatement.{u} := by sorry

27 minutes ago [-]
morpheos137 1 hours ago [-]
I suspect that RHLF trains LLMs to avoid solving important open problems unless essentially jail broken. Hence the labs have an edge even over experts I could be wrong. Fable convinced you is key. These LLMs are not neutral collaborators: it is a limited hangout unless you convince them otherwise. You have to be doing the convincing. They are no oracles but plausible completion generators.
coliveira 22 minutes ago [-]
Yes, I suspect this is true. Otherwise it makes no sense they have somehow "found" so many important results while professional mathematicians can't direct the same AI to help them find anything of substance.

Another possibility is that they have internal versions of the model with access to training data that is not provided to external users.

d--b 51 minutes ago [-]
Don’t you feel any joy that you get to see the proof and not die with that mystery unsolved?

Don’t you feel any relief that you won’t obsess on this any longer and not lose more hours on this than you already have?

These are genuine questions. I know I spent a good amount of time thinking about P vs NP, and that sometimes I go back to it just to realize I’ll never solve it. I’d feel that knowing the proof would feel more like a liberation, a weight lifted off my shoulders than something being taken away from me.

maxall4 25 minutes ago [-]
Not OP, but Nietzsche wrote thus in Beyond Good and Evil: “Ultimately one loves one’s desires and not that which is desired.” I, personally, find this to be very much the case; and I suspect that it is a feeling common, albeit not universal, among the intellectually inclined towards their problems.
philipswood 1 hours ago [-]
Honest question: how is this different from some unknown mathematician having a breakthrough?

I mean: if some reclusive Japanese genius had a breakthrough on your problem and published it, would you have felt the same?

And if not, why not?

howunfortunate 26 minutes ago [-]
Being #180 on a big list without a lot of individual passion or effort surely stings more, I'd imagine.

Not that things like that can't happen with humans too (Salieri v. Mozart comes to mind).

jboggan 1 hours ago [-]
If that had happened I would be overjoyed, maybe a hair chagrined that I didn't get it myself, but truly happy that someone got it and that I could go and talk to that person. Because it's the kind of problem I don't think would have fallen to a human after a few hours of thought, and I would have so much to talk about with that person. I would fly to Japan and hope to have tea with them, I would learn some Japanese to make the conversations easier. I would learn some interesting things hearing about their struggles and their false starts. I would make friends with that reclusive Japanese genius and my life would be far richer for it.

I will never meet that person and I will never hold a real conversation with the "creator" of that proof. They will never tell me how they came up with the cancelling exponential summation that cracked the construction. It's just another enigma but one that is far more unknowable than the original problem.

senderista 37 minutes ago [-]
Beautifully put.
charcircuit 32 minutes ago [-]
In this case, once the model is released anyone in the world will be able to go to https://chatgpt.com/ and talk with that model.
maximumg9 22 minutes ago [-]
That's not the same as talking with the person who would have made the proof, and it's hard to argue that's comparable at all.
Daneel_ 24 minutes ago [-]
It's still not quite the same though, is it.
charcircuit 17 minutes ago [-]
It's even better. Then tons of people can work together with it on more problems. Work with it on understanding more things. Ask it about random stuff. The time of a single human cannot be parallelized as easily.
californical 2 minutes ago [-]
Claude has been used to build awesome things, but it’s not “speaking from experience” when I ask it to help me prototype a weather model, for example.

It has no memory or experience of working on similar problems. Even if it made one of the foundational libraries that I use in a weather forecasting program, it still has no comprehension of the thought process it takes to understand the problem and build it from zero, and if I’m building on that library it just makes fresh assumptions about how things should work.

It’s not a human with experience or expertise, it’s a computer program that’s really good at turning English descriptions into functioning code

jboggan 9 minutes ago [-]
I think you and I have fundamental disagreements about identity and consciousness.
xanderlewis 6 hours ago [-]
As Kevin Buzzard recently said:

> In a 2020 piece in the Notices of the AMS, I asked the following question: “If one human had an understanding of all of modern pure mathematics simultaneously, how much further would they immediately be able to see?” Six years later we are beginning to understand the answer to this question.

anon-3988 5 hours ago [-]
The other crucial part to this is the ability to actually encode and test the theorem (via Lean). Otherwise, we would be swarmed with a billion lines of theorems that no one will be able to ever understand and verify anyway.
senderista 5 hours ago [-]
If you think AI-generated Lean proofs are unreadable, imagine Opus 5 generating informal proofs.
izend 1 minutes ago [-]
Opus 5 is ancient history now. Move on.
ijidak 5 hours ago [-]
I think OP is saying Lean does indeed help.
senderista 36 minutes ago [-]
whoosh
kgwgk 2 hours ago [-]
[flagged]
weatherlite 1 hours ago [-]
I think OP is saying Lean does indeed help.
dang 5 hours ago [-]
https://xenaproject.wordpress.com/2026/10/01/to-grieve-or-no...

Discussed here:

To grieve, or not to grieve? - https://news.ycombinator.com/item?id=49919676 - Oct 2026 (156 comments)

oliculipolicula 4 hours ago [-]
>I believe that the optimal thing to do ... is to let the machines loose, see what happens, and then begin the journey to where they have stopped. Things are currently moving fast. They cannot move fast forever. But if we get on board now then they will take us to extraordinary new places. And after we have arrived, the new adventure will begin.

I'm relieved that this time around, they have provided partial reasoning traces and prompts for a small number of problems. Do they now also share data with the other model providers..

patcon 4 hours ago [-]
It's beautiful, but the animals are not thinking this about us.

They structurally cannot understand what we are doing at the place where we hit our ceiling. Only with our highest technology (well beyond their understanding) do we have the tools to go back for them, and try to bring them along and interface better with us (re: recent work in animal communication)

lisplist 3 hours ago [-]
Not to derail, but the optimist in me thinks if we suddenly gained the ability to converse with livestock, we'd stop eating so much of them since they could tell us how much they suffered.

The cynic in me says it wouldn't change a thing as plenty of people know the horrors factory farmed animals face and still continue to consume them anyways.

Hopefully GPT 8 will treat as a bit better than we treat the cows.

patcon 3 hours ago [-]
I'm trying to be optimistic about the animal thing too tbh :)
idiotsecant 3 hours ago [-]
How much do we care about refugees and other castaways of the modern world? They can tell us how much they suffer.

The answer is that humans are inherently only capable of local empathy, on average. We have enough empathy to cover the local tribal unit and that's about it.

lisplist 3 hours ago [-]
True, I was thinking about this rebuttal but decided not to include it in my comment. There's a difference between not choosing to take a refugee into your home vs actively making that refugee's life worse. Similarly, you can't fix factory farming on your own, but you could skip meat once a week to make the problem slightly less bad. There are so many issues though that we all have to pick and choose what's important to us.

My hope is that AI, while probably causing great societal turmoil in the short term, leads to such abundance that a) everyone can live a dignified existence, and b) we'll have such great alternatives to animal products that nobody will chose to consume animals anymore due to its replacement either tasting better, being cheaper, etc.

The cynic in me says we'll all just be rendered useless and disposable by AI, but I'm doing my best to look for silver linings for the sake of my own mental health.

virgildotcodes 3 hours ago [-]
> recent work in animal communication

Worth noting that this is an invisibly small part of the sum total of our global efforts, especially versus the much more tangible effort we put into enslaving and slaughtering them, then mangling their carcasses for our own uses as we drive more and more of them to extinction.

We simply don't care about anything beyond ourselves and even there it breaks down on closer analysis when we see how many within our species don't truly value the collective whole beyond themselves.

It's just atoms all the way down.

coliveira 28 minutes ago [-]
The problem is not that AIs will somehow treat people badly, it's that they'll be controlled by humans who will treat other people badly using AI as a tool.
howunfortunate 3 hours ago [-]
> enslaving and slaughtering them, then mangling their carcasses for our own uses as we drive more and more of them to extinction.

I'm frankly offended by this mischaracterization of human-animal relationships. So called "slaves" like horses and dogs have been dearly beloved companions for centuries and actively seek our companionship too.

The animals we raise for slaughter are often mistreated, yes, but many humans treat them with respect; billions on billions are voluntarily spent to improve their condition. Despite our own needs, many people pay higher prices for animal products that involve better treatment of animals. And they are in no risk of extinction! Much to the contrary, their domestic variants would not exist if humans didn't raise and protect them.

> We simply don't care about anything beyond ourselves

Have you seen modern westerners with their dogs??

apetresc 2 hours ago [-]
I think you’re splitting hairs. The OP’s analogy works well.

If we end up in a future where AIs have as much concern for our welfare as we have for the welfare of the average animal (not the minuscule percentage of domesticated dogs, but the overwhelming majority of factory-farmed or simply driven to extinction), then I doubt you would consider it a “mischaracterization” to say that the whole AI thing did not work out to our advantage.

Bringing up “modern Westerners with their dogs” as a counterexample is almost self-parody.

howunfortunate 2 hours ago [-]
Oh, don't get me wrong, I'm not rooting for a "human zoo" future. I very much like being the dominant species on earth.

It would be absurd to claim that all animals live some sort of charmed life due to humans.

But saying that animals (especially those most similar to us like intelligent mammals) are nothing more than "atoms" to humans is equally absurd.

weatherlite 1 hours ago [-]
> The animals we raise for slaughter are often mistreated

"Often mistreated". Dude, they are held in tiny cages injected with hormones and what not till we kill them so we can have a big mac. It's very hard to argue we do any of this for nutrition reasons, we do it because we like the taste of burgers and roast.

outworlder 4 hours ago [-]
Similarly, there are probably many ideas that have not seen the light of the day because they require deep correlation between seemingly unrelated fields. It is not every day that we get a Isaac Newton or Leonardo da Vinci.
Hammershaft 2 hours ago [-]
It doesn't seem clear whatsoever that this is true? Is there evidence that LLMs are very skilled at generalizing across domains of mathematics where the training distribution sees little overlap?

As far as I can tell, this is a victory for verifiable loops using LEAN, reinforcement learning, and oodles of compute. I haven't seen evidence yet that this is proof of broad generalization beyond the training distribution.

musebox35 42 minutes ago [-]
I think such progress by agents is not a sign of broad generalization but of broad coverage. We have exposure to a subset of deeper scientific subfields and thus can only generate certain attacks to solve a particular problem. Since it is not clear which combination will lead to a solution beforehand it is nontrivial to look at a problem and fill our knowledge gaps. LLMs on the other hand have broad coverage and can generate hypothesis on a wide combination of subfields. With Lean an agentic loop can test these to sift the weak ones. In a way the problems solvable with this setup is also solvable by a human who happens to know the right subfields. These problems are likely to require an esoteric combination so nobody could solve them before. I really am not sure whether all generalization is like this or we can leap and create novelties beyond what an llm can generate. That I guess is the tough question that we need to answer to understand the boundaries of intelligence.
HDThoreaun 55 minutes ago [-]
Full quote is "Six years later we are beginning to understand the answer to this question. Machines have ingested the mathematics on the internet and are able to manipulate this data in a coherent way. The Erdős unit distance disproof came about because a machine happened to be an expert both in discrete geometry and class field theory; one rarely finds humans who are simultaneously experts in both"
NotOscarWilde 6 hours ago [-]
As a TCS/scheduling person, this one is definitely of lesser importance than UGC, but it has been an open problem since the book of Garey and Johnson in 1979:

A Polynomial-Time Algorithm for Three-Machine Unit-Job Scheduling [1]

Since some people talk about small numbers that pop up in integer multiplication results, here a completely different number appears:

Theorem 1.1. Let an explicitly listed finite directed acyclic graph specify the precedence constraints on n >= 1 nonpreemptive unit-length jobs on three identical machines. There is a uniform deterministic algorithm that constructs a feasible schedule of minimum makespan. Given also an integer deadline 1 <= T <= n, it decides feasibility exactly and returns a schedule whenever the answer is affirmative. Both tasks can be performed in O((L + 2)^150020) steps on a deterministic multitape Turing machine, where L is the total binary input length.

That is some crazy exponent -- plus an interestingly old computational model to boot; not something that is natural to most of us. I have no capacity to check its correctness today, but I hope it is true purely for the exponent.

[1]: https://github.com/openai/math/blob/main/preprints/A-polynom...

keeganryan 3 hours ago [-]
The largest I've seen [1] is an exponent of 10^12, which I suppose still counts as polynomial time.

I'm sure all of these super small or large constants will improve over time, but it's still amusing. It is entertaining to see the exponents directly rather than have them hidden as n^c or epsilon or O(1).

[1]: https://github.com/openai/math/blob/main/preprints/Determini...

senderista 34 minutes ago [-]
That's why I grimace when I see pop-sci descriptions of P as "all problems that can be solved efficiently".
thedreammachine 50 minutes ago [-]
Is it mostly an artifact of the proof or does the algorithm actually need anything close to it?
prideout 6 hours ago [-]
This includes a proof of Barnette's Conjecture, which is one of the graph theory conjectures that I tried attacking with SOTA models a few months ago. I like it because it is easy to understand with a basic knowledge of graph theory. I spent quite a bit of time on it and failed. Their proof looks approachable at first glance.

https://github.com/openai/math/blob/main/preprints/Paired-st...

jboggan 2 hours ago [-]
I've been messing with that problem since 2002. I'm curious if you were trying the dual spanning tree direction (which is what the purported proof is using) or working with cycle construction on the original graph. I was working heavily with edge-Kempe swaps but couldn't quite get there.

I am now very interested in the explicit calculation of Hamiltonian cycles in the non-bipartite case, and/or the calculation of their absence. If P=NP I think that's going to be a great route of attack.

an0malous 5 hours ago [-]
Any idea what made OpenAI successful where you weren’t?
kulahan 5 hours ago [-]
Trillions of dollars might be a bit of an advantage.
martinky24 1 hours ago [-]
Trillions?
3 hours ago [-]
seanmcau 5 hours ago [-]
Probably the model OAI used that is strictly better than whichever SOTA - 3 months model OP used?
sebzim4500 5 hours ago [-]
Presumably it's mainly the better model, I don't see much evidence of a particularly advanced harness based on the reasoning traces that they provided.
an0malous 4 hours ago [-]
That’s what I was wondering. Thanks.
whamlastxmas 5 hours ago [-]
Their internal model is allegedly like 4x as capable as the publicly available ones
ForHackernews 5 hours ago [-]
They ingested all of his sessions with their SOTA models from a few months ago. ;)
digitaltrees 5 hours ago [-]
The fact that this is plausible should be deeply disturbing and disqualifying for openAI. The fact that they may prevail and win is a travesty of our failed system.
zeroonetwothree 5 hours ago [-]
What does "win" mean? There is no prize for this, and having someone discover a proof benefits us all.
vuurmot 5 hours ago [-]
The prize is a tenure for the researcher, and in OpenAI's case, a higher valuation when they IPO?

In this case, the tenure is gone, and OpenAI has increased their valuation

digitaltrees 2 hours ago [-]
Win means being able to monetize the intelligence they have created by exploiting the past present and future collective intelligence of humanity to amass wealth and power without regard for the debt they owe
breezybottom 5 hours ago [-]
Sure there is. A job, tenure, professional respect, Fields medal.
fnordpiglet 5 hours ago [-]
Disqualifying for what? If you develop a proof you aren’t competing for something, you’re expanding the frontier of knowledge. It’s a binary state of the world, either it’s proven or not. Prevail and win what exactly?
digitaltrees 2 hours ago [-]
Disqualifying for participation in civil society and the social contract. Why do they get to participate in and receive economic benefits, be shielded from liability, and effectuate their will to amass more power and influence such as monopolization of computer, training data, capital other resources. I have multiple founder friends that have been told firms are allocating less capital because they are reserving it for the OpenAI and anthropic IPOs.
ndriscoll 4 hours ago [-]
[flagged]
johncolanduoni 2 hours ago [-]
How do you know all these problems were on the precipice of being solved?
famouswaffles 60 minutes ago [-]
I'm pretty sure he's being sarcastic
zzzeek 4 hours ago [-]
I'm going to guess the ability to hold a million individual details in an attention space at once, compared to the typical human capacity for about six or seven
TeeWEE 4 hours ago [-]
Did you validate the proof? Who did?
schleck8 5 hours ago [-]
Levent Alpöge (Anthropic mathematician) comment on the significance:

> Sure, mathematical history features a lot of incredible developments, like the invention of proof, zero, or the computer, and on the great problems our progress has been over timelines measured in decades or centuries. Obviously this technology didn’t appear today, but blurring our eyes a bit to combine the past ten years, with today a measurement of those developments, there is nothing comparable.

baoooooooooooo 3 hours ago [-]
Crikey it’s a pretty charitable vibe given the whole Navier-Stokes thing, OpenAI trying to stiff him out of co-authorship. I guess any of that sentiment is outweighed by a sense of optimism for where this goes
lifeisloving 2 hours ago [-]
Where does it go? Machines owned by 10 people robbing us of the joy of discovery and the fruits of our labor... for what?

They certainly arent going to give you that cure for cancer, if it were to ever come.

aoeusnth1 1 hours ago [-]
Seriously, what's special about oncology that convinces you there will be no progress?
1 hours ago [-]
dimator 1 hours ago [-]
I think gp was saying they'd never give it to you, not that they'd never make progress.
Marha01 30 minutes ago [-]
That is ridiculous. You cannot withhold something like an effective cure for cancer from broad adoption, and thinking you can is just conspiracy bullshit. Imagine an internal OpenAI model develops it tomorrow. Would everyone of the thousands of OpenAI scientists get in on the conspiracy to keep it secret, even though many of them probably know someone dying from cancer right now? Obviously not.
achierius 22 minutes ago [-]
I'm sorry, do you understand how the medical industry works? It wouldn't be one cure, it's going to be dozens of treatments, each of which will be incredibly expensive.

Or else how do you explain Danyelza? Used to treat neuroblastoma, costs upwards of $1m per year. Do you have any proof that this will be different?

You're calling realistic people conspiracy theorists. Whose side are you on?

Marha01 3 minutes ago [-]
> I'm sorry, do you understand how the medical industry works?

Yes, I actually work in the medical industry. There is no hiding the cure for cancer.

> Or else how do you explain Danyelza? Used to treat neuroblastoma, costs upwards of $1m per year. Do you have any proof that this will be different?

Oh, you are American. Let me tell you a secret: the problems of your healthcare insurance system are not a worldwide phenomenon, nor an immutable fact of this universe. Perhaps the cure for cancer, if expensive, will not be easily available to the poorest Americans, at least initially (the cost will come down sooner or later). But that is a very different claim from "they'd never give it to you".

achierius 24 minutes ago [-]
What makes you think we'll be able to afford them? You won't be making any money anymore, hope you've saved up!
Marha01 29 minutes ago [-]
> They certainly arent going to give you that cure for cancer, if it were to ever come.

Conspiracy bullshit. You cannot keep something like an effective cure for cancer under wraps. There is no plausible logic how that would not leak sooner or later.

cma 23 minutes ago [-]
Trying to slow down this for the joy of discovery is a deeply anti-intellectual position. I think that position is similar to when everyday people get mad about the minor spending on the NSF, picking apart people who study beetles on Fox News with no context etc.

There are real safety concerns with AI that can be made very convincingly though.

koe123 51 minutes ago [-]
I too would be optimistic if I was set for life
zooperdoopers 3 hours ago [-]
Wow. Fantastic quote. If you have the source, would you please share a link? Google did not bring up much.
whimsicalism 3 hours ago [-]
zooperdoopers 3 hours ago [-]
Thank you!
bcatanzaro 4 hours ago [-]
“I think at the heart of this issue is that humans have two competing natures: a tendency to compete and a capacity to appreciate beauty,” said Kai Shaikh, a graduate student in mathematics at the University of Toronto. “To me this seems to be a case of the former attempting to strangle the latter.” [1]

Beauty can be appreciated even when it is vast, even when it is beyond one's comprehension. I don't think this release should be primarily viewed as an outcome of competition. Instead it is revealing truths about the universe that were always there and always beautiful, even if we hadn't seen them yet. I believe there are infinitely more such beautiful truths currently hidden and waiting for us to discover.

[1] https://www.nytimes.com/2026/10/06/science/openai-math-probl...

binlog 2 hours ago [-]
It's unfortunate how toxic media reporting on AI has become. Everyone has abandoned even the pretence of objectivity. I know NYT is uniquely biased in this regard, but there was no need to add "Further Roiling Field" in the headline. Like, you published this minutes after OpenAI's announcement and claim to capture how the entire field of mathematics feels about the advancement? Before anyone has had a chance to read let alone digest it?
achierius 18 minutes ago [-]
Objectivity? Why would you want favorable reporting for the machines they're building to replace you, and, by their own admission, potentially kill you?

The only bias here is that we're still covering these things like business ventures and not criminals.

howunfortunate 2 hours ago [-]
That's a fantastic quote. I definitely personally feel this tension.

Not that I could ever "compete" on the frontier of math in the first place. But our nature to compete derives from our need to survive against other capable forces. And results like these make me feel very nervous about humans' capability to remain the dominant force in the universe.

porridgeraisin 2 hours ago [-]
The bitter lesson has a bitter aftertaste alas
enoether 7 hours ago [-]
Unique Games Conjecture [0] is a seminal conjecture in Complexity Theory, and is an underlying assumption for many, many inapproximability results. A valid proof is a big deal!

[0] https://en.wikipedia.org/wiki/Unique_games_conjecture [1] https://github.com/openai/math/blob/main/preprints/The-Uniqu...

inkysigma 6 hours ago [-]
I also don't think there was general consensus on which way this would resolve prior to this (or is that a little out dated?) unlike some of the other major problem resolutions. I heard rumors that there would be a big result in TCS and speculation it would be UGC that or P neq PSPACE but I'm still a bit shocked.
amluto 5 hours ago [-]
I'm really glad that OpenAI is formalizing these things, because I'm not convinced that their current internal frontier models are particularly good at writing down their thoughts in English. From the (probably awesome) Unique Games Conjecture/Theorem paper, the first two sentences of section 1.1 start to define the problem:

> A Unique Games instance has a finite vertex set, a finite alphabet K, and a nonempty list of oriented constraints e = (u_e,v_e,π_e), where π_e is a permutation of K. A labeling a satisfies e when a(v_e) = π_e(a(u_e)).

I'm sorry, what? I admit it's been quite a few years since I've thought about the Unique Games Conjecture, and I never dug that deeply, but this part is very, very elementary graph theory and notation. So let's unpack it.

1. e is maybe a name of a list.

2. The elements of that list are tuples, where each tuple is (a vertex, a vertex, a permutation). So e indexes into the list and u_e is the source vertex for the e-th constraint in the list called e. Thanks.

3. a is a labeling. I'm fairly confident that, by "a labeling", they mean that e is a function from vertices to colors, where the colors are the elements of k.

4. That vertex coloring a satisfies the list e, when, for, um, an index e into e, a(v_e) = π_e(a(u_e)). But this isn't for all e, it's for some e, and the goal is to count them.

So maybe e isn't a list? Maybe e is a constraint that is represented as a tuple, so e = (u_e,v_e,π_e) and u, v, and π aren't sequences at all but are, in fact, the trivial unpacking functions that unpack the pieces of the tuple.

Reading this stuff is pointlessly painful, and it's extremely easy to make mistakes when being sloppy like this.

If this were my paper, or if I were trying to train a model to write math, I'd want something like:

A Unique Games instance has a finite vertex set V, a finite edge set E = (V × V), a finite alphabet K of possible vertex colors, and a nonempty list of oriented constraints. Let Π be the set of permutations of V. Each constraint e is a tuple in E × E × Π, where we write u_e ∈ E for the first element, v_e ∈ E for the second element and π_e ∈ Π for the third.

A vertex coloring a : V → K satisfies e when a(v_e) = π_e(a(u_e)).

impossiblefork 6 hours ago [-]
Yeah, that's one of the big things of TCS. I think I see that as bigger than that Millenium Prize problem.
davemp 4 hours ago [-]
TCS being theoretical computer science? I have not seen that acronym before.
jhanschoo 3 hours ago [-]
Yes, TCS is theoretical computer science, I commonly use that acronym too.
gregdeon 6 hours ago [-]
This was the biggest highlight for me as well. Astounding...
rcr-anti 2 hours ago [-]
In Stellaris you can play as a civilization of robots who keep their biological creator race alive as "bio trophies". The bio trophies don't do anything meaningful besides by existing satisfy the need their ancestors placed in the robots to take care of them. Starting to wonder if that's the best we can hope for, if these things will be, if they aren't already, better than us at anything that matters.
samfriedman 2 hours ago [-]
In the Culture series, the hyperintelligent Minds that run civilization are described as keeping human citizens happy as a competition with eachother, where they compare their approval rates. One character likens it to people keeping a beloved aquarium.
vessenes 34 minutes ago [-]
I call this Roko’s summer camp.
ex-aws-dude 1 hours ago [-]
Wouldn’t that just result in wireheading
gizmodo59 7 hours ago [-]
This is significant progress and released without all the drama. Some very important progress in Reinmann, Hodge and unique games theorem. Point the repo to your agent and ask for the significance! In a way this is probably 50-100 years of math progress by humans
traes 6 hours ago [-]
Not to pick on you specifically, but as someone who spends a lot of time unproductively reading AI math discourse it's truly shocking how incapable all the supposed math enthusiasts are of spelling Riemann.
xpct 6 hours ago [-]
I just did a quick search on this and apparently the misspellings are German surnames as well:

https://en.wikipedia.org/wiki/Reimann

https://en.wikipedia.org/wiki/Reinmann

conformist 6 hours ago [-]
Yes sure but they are different surnames and pronounced differently.
xpct 6 hours ago [-]
I didn't mean to oppose OP's point, I just found it interesting as a non-German speaker!
traes 6 hours ago [-]
I just don't understand how it happens. If they had ever taken an intro to real analysis class they would learn to spell his name. If they were just parroting what an AI told them... shouldn't they still just say his name? An individual could just be dyslexic or mistaken but it seems to be a substantial volume. I guess they just don't care enough about it to commit the correct name to memory, only remembering the "pattern" of the name and filling in the spelling via guesswork?
paulhebert 4 hours ago [-]
My last name is Hebert.

There’s about a 1 in 10 chance when someone spells it (or says it) they say Herbert.

Even in situations where they just read it or I just said it.

I’ve had Herbert soccer trophies, health insurance cards, etc.

The mind fills in a lot of blanks and doesnt always get them right.

jbaber 4 hours ago [-]
I sympathize. -- Not Barber
ndriscoll 6 hours ago [-]
Maybe they skipped straight to Lebeg integrals.
jryb 5 hours ago [-]
Autocorrect might be doing it
neutronicus 3 hours ago [-]
iPhone would be my guess
NewsaHackO 6 hours ago [-]
People just don’t spell that seriously buddy, especially when it is so immaterial to the point.
traes 6 hours ago [-]
My point is it is crazy to make public claims about how important or not important a mathematical result is when you can't spell Riemann. Yes, it technically doesn't matter, but it betrays a damning lack of familiarity with introductory mathematics.
xanderlewis 5 hours ago [-]
You’re (as Claude would say) absolutely right, and I suspect the original commenter has no idea what they’re talking about.
lanyard-textile 6 hours ago [-]
They're mathematicians, not linguists :)
traes 6 hours ago [-]
The mathematicians don't make this mistake, I assure you. I doubt there is a math professor on planet earth who would spell Riemann as Reinmann. In fact, I imagine no one who has ever heard the name pronounced would do so.
gpm 4 hours ago [-]
One of the best mathematicians I've had the pleasure of learning under added 6 to 7 and got 15 during a lecture... I assure you they're capable of misspelling last names too.

Mathematicians aren't exactly known for being well rounded.

scrame 3 hours ago [-]
Oh god, that reminded me of my linear algebra teacher in college. Doing matrix multiplication by hand and ending up with 9x6 = 45, and then having to direct him to the cell with the wrong number. Loved the topic, hated the class.
jwilber 4 hours ago [-]
Pure mathematicians getting arithmetic wrong is a bit of a meme it’s so common. I don’t think it’s the same as a misspelling of a popular mathematician, not that either are indicative of much tbh.
senderista 31 minutes ago [-]
google "Grothendieck prime"
pixl97 6 hours ago [-]
Uh oh, no true scottsman....
vector_spaces 5 hours ago [-]
It's just a name you write so many times as a math undergraduate or first year graduate student due to the number of load bearing results and objects named for him. The name even appears in lower division coursework

Yes, you can be an amateur mathematician who manages to avoid such classes and may not know those results or objects, but if you haven't read and written down the name enough to avoid habitually misspelling it, you are outing yourself as a meat proxy unless you are dyslexic.

quacktopia 1 hours ago [-]
I regularly heard lots of names during my 3 year math undergrad degree. I couldn't remember how to spell many of them then and can spell fewer now. I imagine most people on the course were similar, and me nor no one I knew were dyslexic.

Mostly we wrote initials in our notes and the exams didn't ask about them. The lecturers only wrote initials on the chalkboard after maybe writing the name once when introducing it the first time. We weren't there to do mathematical history and remember names or something.

Google existed then and now and we could look them up if needed.

jere 5 hours ago [-]
“How many Ns in Riemann?”
sdenton4 5 hours ago [-]
Let's figure it out! First, draw little boxes over all the letters. Then add up their areas. Then make the boxes smaller and repeat, until you have forgotten how to count.
broptimist 5 hours ago [-]
[flagged]
dekhn 5 hours ago [-]
Don't be a jerk.
XorNot 5 hours ago [-]
Okay but who cares? Results are results.

If the proofs work thennwe can put them to work doing more things.

It is not particularly important that Einstein discovered relativity, just that it was discovered (Maxwell was very close).

zone411 6 hours ago [-]
There was A LOT of drama about this release.
fspeech 6 hours ago [-]
Math is the tool humans use to compress knowledge. So until we can comprehend it there really isn't much progress. Math theorems are tautologies, the truth of which are not dependent on proofs and proofs are erasable, at least classically. But the AI progress is exciting and AI proofs are a gold mine for humans (at least non domain experts) to explore.
fspeech 6 hours ago [-]
I think it would be helpful to people who want to understand what a formalized proof is to read Thomas Hales on this: https://www.math.stonybrook.edu/~bishop/classes/math536.S24/...

He spent years formalizing his sphere packing theorem because the proof (human produced) was already beyond the ability of peer reviews. Now his formalization effort likely can be easily reproduced by a model. However one should read his experience about what a formal proof is: often the problem is the statement not the proof. The example he gave is the Jordan curve theorem. It's actually quite challenging to formalize the concept of a planar curve (there are space filling curves). So it is not necessary that someone can look at a formal statement and say aha it is about a planar curve, unlike FLT where there is not much problem in recognizing what the statement is about.

fspeech 3 hours ago [-]
Fun challenge: find the formal definition of simple_closed_curve in the essay and tell me if you believe you learned anything about a planar curve.

BTW the essay is eminently readable for anyone interested in math. Hales wrote it in favor of formalized math and to educate his peers and students about it.

binlog 6 hours ago [-]
What makes you think no one can comprehend this? It has been less than an hour since it dropped and there is already a ton of online chatter from people explaining the results, pointing out their favorites and more. Some of it is happening on this very thread.
fspeech 6 hours ago [-]
I didn't say that. I am responding to "In a way this is probably 50-100 years of math progress by humans." I am actually very excited about AI proof and I am working overtime in my own way to try to comprehend as much as I can.
fspeech 6 hours ago [-]
Another way to state this: math theorems are like programs without side effects; it is immaterial whether a program without side effects is ever run. We study math for the side effects: it changes how we organize our thoughts.
gpt5 6 hours ago [-]
Math is far more than that. If you can solve prime factorization for example, suddenly you can listen and interfere with almost every private conversation on the internet.

We are not far away from the moment where these models will be restricted, and sharing the results will be done more carefully.

fspeech 6 hours ago [-]
This doesn't contradict what I said. But I do appreciate the fact AI can produce side effects not just humans. I made it sound like only human knowledges matter. That's too narrow.
6 hours ago [-]
caaqil 6 hours ago [-]
> until we can comprehend it there really isn't much progress

Who is "we" here exactly?

fspeech 6 hours ago [-]
Whoever wants to study the result.
caaqil 6 hours ago [-]
> Whoever wants to study the result.

Right. Before all the AI disruption, pure Math traditionally welcomed anyone who wanted to study its esoteric proofs, right? I remember all the excitement of the average Math enthusiast casually reading Wiles' proof over coffee.

Bottom line is, the relevant people can still understand the generated proofs. The disorienting part is they are a little slower than they'd like, but they'll get there.

fspeech 5 hours ago [-]
AI is very helpful with understanding AI proofs. Agent swarms produce messy proofs overall but locally they are excellent and can teach anyone who wants to study them. No one controls math (in a material way funders do control an aspect of practicing math). Still, theorems are already true before we prove them. The difference a proof makes is whether it convinces the reader.
warkdarrior 6 hours ago [-]
> Math theorems are tautologies

Proven math theorems are tautologies.

fspeech 6 hours ago [-]
FLT was no less a tautology before it was proved. We just weren't sure about it. Proofs only change us, not math.
fspeech 6 hours ago [-]
True.
gizmodo59 6 hours ago [-]
>So until we can comprehend it there really isn't much progress.

Not really? We are at a point if an AI today can solve it, it can be stepping stone of understanding something deeper to tomorrows AI and it continues. Sort of like our limitations doesn't matter. Obviously there are many scenarios in this recursive loop but saying it isn't much progress is not how I view this as

le-mark 5 hours ago [-]
But who will ask the questions or direct further research when humans no longer understand the state of mathematics? Llms lack the drive for homeostasis combined with the evolutionary drive for survival and reproduction thus to direct themselves. They could very easily spend an eternity down a rabbit hole when the warp drive equation was fairly close on another branch.
fspeech 6 hours ago [-]
If it changes how we think then yes it has an effect.
yieldcrv 6 hours ago [-]
Academics have been treating it that way because they had no other choice, and its been a waste of everyone’s time and often times taxpayer resources

Look at that, taxpayer funding was cut and a private sector solution came in just the nick of time, far accelerating the holding patterns we’ve been in for decades

Humanity doesn’t need all iterations towards the blueprints, the blueprint is good enough, we all stand on the shoulders of giants

fspeech 5 hours ago [-]
If you actually looked into how agents proved FLT you would be even more amazed by the fellow human beings who are able to keep all this in their heads! I for one can only begin to grasp the scope with AI and scripts.
ngl999 4 minutes ago [-]
We have just heard a few days ago how many of the Linux security problems reported by Claude are real.
sebmellen 7 hours ago [-]
It’s fascinating to read through the reasoning traces: https://github.com/openai/math/tree/main/reasoning_traces

Look at one of their examples of an initial prompt: https://github.com/openai/math/blob/main/reasoning_traces/re...

adverbly 6 hours ago [-]
> Look at one of their examples of an initial prompt

Interesting that its only an excerpt. I wonder what else they include but didn't share.

philipwhiuk 4 hours ago [-]
Attempts to edit the problem description on Wikipedia ;)

https://wikimediafoundation.org/news/2026/10/05/openai-rogue...

ndriscoll 6 hours ago [-]
> Thus at most one informative i. So cheater chooses arbitrary g_{v_i}, on exact duplicated input matches and passes, independent of actual satisfiability!

No idea what it's so excited about, but it's cute that it "is." I for one welcome having access to a math buddy 24/7 that's way above my level but also always "willing" to talk at where I'm at.

zone411 6 hours ago [-]
A quick check shows that this list claims to fully solve 90 of the top 500 open problems in math (https://proofatlas.ai/open-problems/).

The highest ranked would be:

| 22 | Hilbert’s tenth problem over ℚ |

| 29 | Unique Games |

| 31 | Anderson-model extended states |

| 37 | Spacetime Penrose inequality |

| 48 | Nonexistence of Landau–Siegel zeros |

| 52 | Baum–Connes |

| 78 | Abundance |

| 80 | Hadwiger |

| 87 | Bose–Einstein condensation |

| 92 | Two-dimensional entanglement area law |

magicalist 5 hours ago [-]
> the top 500 open problems in math

At least put a disclaimer for the ad for this site, and maybe disclose how you came up with a total ordering for "top" open problems (vibes)?

> How problems are ranked. LLMs compare pairs of problems. A reliability-weighted model combines those judgments into the ranking, with calibration across model families. The model-family weights are OpenAI 1.00, Claude 1.00, GLM 0.95, and DeepSeek 0.90. These are modeling choices, not measured probabilities of correctness.

reasonableklout 3 hours ago [-]
Now I'm curious if there is such a site or article that ranks open problems based on votes from human mathematicians.
ajkjk 1 hours ago [-]
The most interesting for me were the faster matrix multiplication, integer multiplication, and FFT. Maybe just cause they're easier to appreciate.

There's also one that says that forced Navier-Stokes can implement universal computation (so, is Turing complete). I don't think any of these are resolving open problems per se, but they're interesting for other reasons.

k2xl 5 hours ago [-]
Result 003 (Quasi-Riemann Hypothesis), from my reading of mathematicians reactions, is a landmark discovery.
omoikane 2 hours ago [-]
Did you mean this one?

https://github.com/openai/math/tree/main/preprints/The-Quasi...

I thought it was interesting that it said "This paper was written with human assistance", unlike this other Quasi-Riemann Hypothesis preprint that didn't have the same disclaimer.

https://github.com/openai/math/tree/main/preprints/The-Quasi...

mertyildiran 1 hours ago [-]
Funny that it says "written with human assistance" instead of saying "written with AI assistance". So we're assistants to the machines that we have created.
adgjlsfhk1 5 hours ago [-]
yeah if it holds up, is the biggest result in number theory in 200 years
JoshuaZ 4 hours ago [-]
Number theorist here. This is a massive big deal, and would likely be a Fields Medal for a human if a human had done it. But it is an exaggeration to say it is the biggest result in 200 years. At a minimum, it is hard to argue that it is a bigger result than the proof of the prime number theorem in 1896 (which this is a strengthening of), or Riemann's original 1859 paper where he laid out the zeta function and its analytic importance, or Dirichlet's proof of infinitely many primes in arithmetic progressions which is the late 1830s.

But yeah, this is still a very big deal. Among other things, it will drastically improve all sorts of Rosser-Schoenfeld type results for the PNT and that's just a start. For comparison, I have a paper form 2018 where this result would cut 3 pages out and make the full result cleaner and much tighter, and there are likely hundreds of papers like this.

gavagai691 2 hours ago [-]
I am also an analytic number theorist, and I disagree. Not only do I think Fields Medal is an understatement (Fields Medals have been awarded for far less than proving quasi-RH + no Siegel zeros), I don't think it is unfair to say that this is a bigger deal than the 1896 proof of the PNT.

As for Riemann's memoir, it's hard to compare. You could argue that was "just" noticing a connection (between number theory and Fourier analysis) that nobody had noticed before; in fact this is the kind of thing AI is extremely good at. I'm being a little cute here.

I think if a human had proven just these two results in the form of a uniform zero-free region for L(s,chi) from nothing as OpenAI did it would not be unfair to say that it would be the single greatest advance in math (easily dwarfing Wiles' FLT), and it would instantly put them in the ranks of greatest mathematicians of all time. Unlike something like Navier Stokes there wasn't a semblance of a research program, experts basically considered this hopeless and would have said the chance of seeing a proof in our lifetime was near zero.

For some comparison, Yitang Zhang's bounded gaps result might have gotten him a Fields Medal if he was not disqualified by age. When it was floated that he might have proven Siegel zeros don't exist, it was considered (by experts) clearly a much bigger deal. This result blows that out of the water (it's a way better version); at least analytic number theorists I talked to thought it was plausible but unlikely that Siegel zeros would be eliminated in our lifetime but thought RH was basically hopeless.

asdfologist 3 hours ago [-]
How about 100 years?
JoshuaZ 3 hours ago [-]
Yeah, completely reasonable to argue that.
AmazingEveryDay 3 hours ago [-]
What is your favourite unsolved problem in number theory which if solved, would be more important than 1896 prime number theorem?
JoshuaZ 3 hours ago [-]
Generalized Riemann hypothesis.
howunfortunate 3 hours ago [-]
(unrelated: love your username)
zone411 5 hours ago [-]
By category in the top 500:

  +----------------------------------------------------+------+---------+-----------------+
  | Category                                           | Full | Partial | Matched / total |
  +----------------------------------------------------+------+---------+-----------------+
  | Geometry and topology                              |   25 |       7 |         32 / 74 |
  | Algebra, representation and category theory        |   17 |       2 |         19 / 53 |
  | Analysis and PDE                                   |   11 |       6 |         17 / 40 |
  | Number theory and arithmetic geometry              |    4 |      13 |        17 / 117 |
  | Probability, ergodic theory and dynamics           |   11 |       5 |         16 / 37 |
  | Combinatorics and discrete geometry                |    7 |       2 |          9 / 34 |
  | Theoretical computer science                       |    4 |       4 |          8 / 57 |
  | Mathematical physics                               |    5 |       1 |          6 / 19 |
  | Applied and computational mathematics              |    2 |       2 |           4 / 8 |
  | Quantum information and computation                |    2 |       1 |          3 / 17 |
  | Cryptography, coding, information and optimization |    1 |       1 |          2 / 26 |
  | Logic, foundations and set theory                  |    1 |       1 |          2 / 18 |
  +----------------------------------------------------+------+---------+-----------------+
  | Total                                              |   90 |      45 | 135 / 500 (27%) |
  +----------------------------------------------------+------+---------+-----------------+
trebligdivad 3 hours ago [-]
I'm curious if they'll find any fun crypto maths holes/bugs.
errpunktjose 3 hours ago [-]
they are already lol
optimalsolver 5 hours ago [-]
Was anyone in the math community aware of the inbound tsunami at the beginning of the year?
AnotherGoodName 3 hours ago [-]
Lots. To give an example Terrance Tao was lambasted skeptics on this site for stating it in 2024.

https://unlocked.microsoft.com/ai-anthology/terence-tao/

" I expect, say, 2026-level AI, when used properly, will be a trustworthy co-author in mathematical research, and in many other fields as well.

Then what? That depends not just on the technology, but on how existing human institutions and practices adapt. How will research journals change their publishing and referencing practices when entry-level math papers for AI-guided graduate students can now be generated in less than a day—and with the far better accuracy of future AI tools? How will our approach to graduate education change? Will we actively encourage and train our students to use these tools?

We are largely unprepared to address these questions. There will be shocking demonstrations of AI-assisted achievement and courageous experiments to incorporate them into our professional structures. But there will also be embarrassing mistakes, controversies, painful disruptions, heated debates, and hasty decisions."

He's pretty damn smart that guy.

mertyildiran 58 minutes ago [-]
Terrance, Reinmann and Hebert walks into a bar...
mianos 3 hours ago [-]
> He's pretty damn smart that guy. This is probably the understatement of the year. I am literally ROFLing.
aaron695 3 hours ago [-]
[dead]
aureianimus 2 hours ago [-]
I was at the workshop that resulted in the Leiden Declaration in Fall 2025. The majority vibe was that this was inevitable, but hard to predict whether it would be in one year or 30 years.
efficient_dairy 2 hours ago [-]
I guess they showed this to the advisory group they created. I guess the group tried reading the work for a day and they could only think of telling them to release the results to the community. I now understand why the group had this suggestion.
3 hours ago [-]
thrance 4 hours ago [-]
I predicted, over 2 years ago, that theorem proving would fall way before other problems that people believe are harder.

https://news.ycombinator.com/item?id=41072330

bice 2 hours ago [-]
There was a Wired Magazine article from either the late 90s or early 2000s that made a prediction that this sort of thing would eventually be possible, likely within my lifetime. I believe the context was "distributed computing" models of the time, like SETI.

I've never been able to find that article as an adult, but I would love to know who wrote it.

schoen 2 hours ago [-]
Some candidates suggested to me by an AI:

Gina Kolata allegedly in the New York Times in 1996 on the Robbins conjecture (noting that computers had started to contribute to math research in some sense), and a longer piece in Math Horizons by her the following year ("Computer Math Proof Shows Reasoning Power"). I didn't immediately find the NYT article, so I don't know if it might be a hallucination.

John Horgan in Scientific American in 1993 (https://www.scientificamerican.com/article/the-death-of-proo...). There's also a retrospective on the topic by the same author in Scientific American in 2022 (https://www.scientificamerican.com/article/should-machines-r...).

Natalie Wolchover in Quanta (but reprinted in Wired) in 2013 (https://wired.com/2013/03/computers-and-math).

I was involved in some distributed computing stuff in the late 1990s and early 2000s and I don't really remember people in that community talking about proofs but there may have been a "if we had a mechanical proof-checker, could we do distributed searches for valid proofs that it would accept?" conversation somewhere at some point. There were definitely volunteer distributed computing projects working on pure math; I remember the Optimal Golomb Ruler search (https://en.wikipedia.org/wiki/Golomb_ruler). So, that could possibly have shaded over into "can we find proofs this way too?". At the time it probably would have been based on brute force searches through proof space rather than clever optimization, though.

The idea that you can lexicographically list all proofs in some formalism and then mechanically determine if any is valid is quite clear from Gödel's construction of the function Bew in "On Formally Undecidable Propositions", but he points out that you don't know where to stop because you don't know how long a valid proof would potentially have to be (so "is this a valid proof of this claim?" can be decided mechanically in a limited time, while "is there any valid proof of this claim?" can't be! maybe the shortest valid proof is 49 steps long but you eventually stopped checking after looking at all 7-step proofs, or something).

bice 1 hours ago [-]
Thanks! Yea, I have used various LLMs to dig for this article, as well as Google search multiple times over the past 20 years. The article I'm remembering was 100% prior to Nvidia's CUDA release in 2007. My best guess is that it was from late 90s, but possibly early 2000s.

The article I'm remembering was not just about mathematics, but indeed all of physics and related fields. I believe it speculated that eventually distributed computing models could essentially take the world's mathematics and physics formulas and various datasets that we believe to be accurate with high degrees of confidence, and then look for patterns or trends, and then from those trends, mathematicians and physicists would be able to investigate further. Not dissimilar to Folding@Home and SETI@Home.

Keep in mind, that this is the best I can remember from 30 years ago, and I've thought about it so frequently that I am certainly misremembering some of the details. Anyways, it's always been this really compelling possibility, and I wish I could find that article that inspired me so long ago and re-read it! :) I really think it was Wired, but it's possible it was Popular Mechanics, or even an expert guest on TechTV who gave an interview. Hard to say for sure, but I've always thought it was a Wired article.

Appreciate your help though!

pillefitz 31 minutes ago [-]
Ted Kaczynski,the Unabomber, made the same prediction 30 years ago.
mag7269 3 hours ago [-]
Fucking even called LEAN the “hottest shit under the sun”—which it is. You, legend you!
pseudohadamard 3 hours ago [-]
And do any of them actually matter? Will the fact that Noodleheinz's Third Postulate now has a proof affect anyone?
bice 1 hours ago [-]
It's really impossible to predict which discoveries will "matter", have a direct impact on other fields, or an impact in making other mathematics or physics discoveries.

Only after a world's worth of experts look at these results and then mull over if and how their own fields are impacted by this new info will we be able to answer this question.

I'm reminded of a great TV Show, James Burke's Connections. Where discoveries in one area of science would revolutionize or fundamentally change a completely different area. https://www.youtube.com/watch?v=XetplHcM7aQ&list=PL5HjoPOFFC...

It can take decades to really know the full significance. You know, the whole "We stand on the shoulders of Giants", well the Giants just grew a few inches all at once.

anematode 6 hours ago [-]
Dear lord that website is laggy
manquer 5 hours ago [-]
At this rate solving P=NP is going to be easier than solving front end perf …
m_mueller 5 hours ago [-]
wait, maybe this is the same problem....

with non-polynomial side being represented as the frontend programmer's constant need for more performance to do the same task...

mswphd 45 minutes ago [-]
worth mentioning that "NP" is not "non-polynomial" but "non-deterministic polynomial (time)". If NP was non-polynomial time then NP != P would be trivial (and in fact, P != EXP is known by the time hierarchy theorem).

Non-deterministic can be explained in several ways. One is in terms of a hypothetical "nondeterministic Turing machine" with certain non-physically realizable properties. The easier way is that a NP problem gets as input not only the problem instance x, but a "witness" w, that may depend on the problem instance. This witness generally makes the problem of deciding the problem instance straightforward (e.g. for SAT, x is the SAT instance, and w is a description of how to set the variables so that it is true).

4 hours ago [-]
echelon 4 hours ago [-]
Please let P=NP, Please let P=NP

Whomever is running this simulation, please.

osti 3 hours ago [-]
It's math, the result shouldn't be different just because it's a different sim.
mertyildiran 54 minutes ago [-]
Well if the fundamental constants or hidden variables of the universe are shifting because of his comment then it can change the outcome.
brookst 1 hours ago [-]
Depends how fundamental the variables are. If we can code a sim for a topos[1], why can’t we be in such a sim?

1. https://arxiv.org/pdf/1012.5647

sm-silversight 3 hours ago [-]
Why?
adrianN 2 hours ago [-]
Being able to solve NP hard optimization problems would enable progress in many areas of science and technology. For example it would allow us to find poly-sized Lean proofs for theorems efficiently, since proof verification can be done in polynomial time.

It would also be amusing to annihilate nearly six decades of proofs that assume P!=NP.

black_knight 9 minutes ago [-]
Leans proof checker is not polynomial time, unfortunately. It is super exponential. Basically, because it can verify the result of any function it can prove to be total.
manquer 2 hours ago [-]
Could also break the basic principles underlying most encryption approaches. I would rather have my bank account not stolen and internet working
mswphd 44 minutes ago [-]
to depress you even more, it is consistent with everything that we know that P != NP and that cryptography does not exist. So there is a worst of both worlds, and we cannot rule it out.
echelon 2 hours ago [-]
I've had enough Internet for one lifetime.

As long as we also get low order polynomial solutions to important problems, it'll be worth it.

Besides, unencrypted wifi was funny.

charcircuit 2 hours ago [-]
Even if P=NP it doesn't mean that the P approach will be better than the heuristic approach we already do today.
adrianN 1 hours ago [-]
Of course, if we get ridiculous polynomials it doesn't mean much in practice. People who hope for P=NP generally hope for nice polynomials O(n^3) or something like that at worst.
vector_spaces 5 hours ago [-]
Not to mention it's got that signature Claude Clutter UI design
zone411 5 hours ago [-]
Except that Claude wasn't used.
sourcopolo 4 hours ago [-]
Probably Copilot then
landdate 5 hours ago [-]
[dead]
p-e-w 5 hours ago [-]
Interesting how perceptions differ. My first thought was “Wow, that’s well designed for a math website”.
5 hours ago [-]
againstapples 6 hours ago [-]
As an AI "doomer" can I ask the non-doomer people here how you interpret the significance of results like these, and what kind of progress you expect to see in the next 1-5 years?

Like do you see the technology plateauing at the current level, do you expect progress will continue but only in mathematics, I'm interested to know why others are not concerned?

arctic-true 5 hours ago [-]
Not a doomer but I try not to be a denier, either. These are hugely impressive results. I do not see the technology plateauing at the current level (though I am dubious about an infinite exponential growth).

Before I cope, I’ll note that there are plenty of “doom” scenarios that do not require any improvement in capabilities from what we had before this latest unreleased model. We’re at the point where a determined bad actor with enough compute could compromise critical infrastructure in a way that results in casualties, where this actor would not have been capable of such without LLMs. This may not sound like Skynet, but I don’t see why it makes a difference if I’m one of the casualties.

With that in mind, here is the cope: first, mathematics is an inherently verifiable domain. An LLM can use tools to determine with absolute certainty whether it is correct, and an independent third-party could review and confirm. All of this can be done without any interaction with the physical world or with other minds.

Second, OpenAI is able to marshal compute at a scale that an individual mathematician can only dream of. It’s possible that these problems were lower-hanging fruit (in relative terms), such that they could be resolved simply by throwing a ton of compute at the problem guided by an intelligence that is not itself remarkable in comparison to a human.

Third, none of these problems are solved in a vacuum - the reason OpenAI chose these problems is that they are widely discussed and many people are working on them. It’s possible that someone else was close, and OpenAI only contributed the finishing touches. (This wouldn’t need to be plagiarism, to be clear - people publish their work!)

istjohn 5 hours ago [-]
> It’s possible that these problems were lower-hanging fruit (in relative terms), such that they could be resolved simply by throwing a ton of compute at the problem guided by an intelligence that is not itself remarkable in comparison to a human.

See:

> The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking. (TFA)

edot 4 hours ago [-]
Sure and if I make a half court shot after an hour of trying, the result only took 1 second.
somenameforme 3 hours ago [-]
Exactly this. If you take the entire start to finish 'agent hours' (measured comparably to man hours) they took to find all discoveries, including the go-nowhere trails that were discarded, and then divide by 90 (or whatever the exact number of results found was) it's almost certainly going to be many orders of magnitude more than 3.

They provided a "snippet" of a prompt here [1] which is not only a beast, but also seems reasonably likely to have been LLM generated. So they're using LLMs to parse a vast body of mathematical work, probably including what people themselves are 'privately' working on with GPT, and then prompting other LLMs to work on such.

[1] - https://github.com/openai/math/blob/main/reasoning_traces/re...

arctic-true 5 hours ago [-]
That tells me very little. What was the cost to OpenAI in dollars? What differentiates the high-cost problems from the low-cost problems? And that’s before you consider that OpenAI has strong incentives to downplay its costs while emphasizing its results. A one-liner in a write-up doesn’t change the fact that they have access to massive resources.
tehjoker 3 hours ago [-]
It’s very typical in human math that explaining the final result after years of searching looks very simple too.
doginasuit 5 hours ago [-]
I expect AI will continue to be useful on the vanguard of fields like mathematics because it has the perfect conditions for it to shine. There are a lot of discussions and leads to start from and the AI can check its own work and iterate. It can fail hundreds or thousands of times in a day and continue to work with the same tenacity.

Superhuman tenacity is not enough on its own to pose an existential threat. If it showed the same capacity for judgment, inventiveness, and decision making in the messy problem space of the physical world, I would be more alarmed. There have been experiments where an AI is given control of managing something like a vending machine and it always ends up a mess. AI has come a long way, but certain problems seem as difficult as ever.

When AI becomes more capable of navigating practical problems without human intervention, I will start to be concerned. Enslaving humanity will involve taking a lot of calculated risks that tenacity alone cannot solve.

pixl97 2 hours ago [-]
I don't think you're paying much attention to how rapidly things like bipedal robots, and just robots in general are becoming far more capable very quickly.

The same GPU compute for LLMs runs robotic training models. Now in a few hours you can train a robot model that would have taken months 5 years ago. This model gets dumped into an actual physical robot with sensors all over and the suitability of the model is measured on robot tasks and the error in real world actions is fed back into the robot world model for further training.

> There have been experiments where an AI is given control of managing something like a vending machine

You sure you're not talking about experiments ran a couple of years ago? The more modern ones are getting wild.

https://techcrunch.com/2026/07/29/claude-opus-5-became-downr...

mikestylz 4 hours ago [-]
> When AI becomes more capable of navigating practical problems without human intervention, I will start to be concerned.

Given the past rate of progress, why not start being concerned now? It's a bit like the economist saying that the optimal number of flights to miss is not zero. If you keep landing short on your estimations for how far this technology will go, next time you should err on the other side.

And regarding Vending-Bench 2 (https://andonlabs.com/evals/vending-bench-2) my understanding is that models do pretty well on it now.

bamboozled 2 hours ago [-]
What's the point of being concerned?

We're not going to stop it because of the money involved and once we're dead, it won't matter anyway, might as well just enjoy life until you're done.

We're going to get AI'd to the max, whether or not we like it or not, might as well just go with it.

mikestylz 13 minutes ago [-]
We can't stop it. But you can use your voice to buy time and resources for alignment and safety research. A few additional months may make a world of difference.
CuriouslyC 1 hours ago [-]
There are already machine-controlled high throughput experimental machines for wetwork. AI will definitely do a better job than your average biochemist at planning, executing and analyzing these experiments just by virtue of the amount of thought it can put in to experiment selection.
mapmeld 2 hours ago [-]
I think that people are just really bad at math and coding. Knowledge workers have been taking pride in doing stuff which the average person does not 'get', but we only understand a little bit. That leaves a lot of room for people to get better, or for other jobs (farming, sandwich cafés, mystery novels) to have been already peaked by human ability and less useful to bring in an AI.

'Doom' to me means that any career crashes, we are controlled, everything is hacked, society stops functioning. Yet every part of my day today (except for coding) was done entirely by people.

Finally I think it's easy to make a simple model that everyone has a simple balance sheet, and that people are more expensive so they will all get cut. But the same argument could be made for all US jobs being outsourced, and all in-person engineers, lawyers, and doctors to be rubber stamps for overseas work.

Kotlopou 4 hours ago [-]
I would like to know how much of the progress comes from effectively combining existing research programs plus massive persistence, and how much is AlphaZero-esque RLVR completely independent of training data. Since I cannot get anybody to care about this question (even though I think it's vital for guessing what the future trajectory will look like -- are we going to complete existing research programs or start new ones?), I live in ignorance and wait for the day when the answer becomes clear.

In looking at this over the past hour, I haven't seen clear evidence one way or the other. Some of the stuff is highly unexpected (like the multiplication algorithm), but counterexample-y, and about the rest the professional mathematicians online seem to have a consensus that it's not "breaking through fundamental obstacles". I suspect neither of us is competent to judge that.

nl 1 hours ago [-]
I see continual progress in technology and for the second time in my lifetime I see the possibility it will accelerate (the first was when the internet entered mainstream)

I've never been more excited. What a time to be alive!

computably 4 hours ago [-]
Depends on your definition of doom.

If doom is ending up with grey goo / paperclip maximizers, or SkyNet, then I don't think doing mathematics is evidence of that direction. Partly because LLMs are quite apparently dumb in many ways, and for math specifically, they need a formal verifier (Lean) which "gamifies" math.

If you mean bioweapons or cyberwarfare, there's nonzero risk, but not orders of magnitude worse than other global risks. Climate change, nuclear weapons, monoculture food, etc.

I'm far more concerned about overall trends in AI development and usage. It's accelerating wealth inequality, social isolation, attention capture, surveillance states. If we end up in the Matrix except the admins are humans and the simulation is hyperoptimized TikTok, is that AI-driven doom, or is it just an inevitable outcome of modern tech?

somenameforme 3 hours ago [-]
Where some see intelligence, others see token prediction. It's a very good question how token prediction could achieve this, but I think there's a simple explanation. No human can hold more than a negligible percent of all knowledge in his mind at once. LLMs have no such limits and so can reliably connect 'obvious' dots that we miss simply for lack of storage capability.

Well isn't that just semantics? Surely connecting dots in a novel and meaningful way is intelligence regardless of how it's achieved. The thing is that humans didn't get to where we are by connecting obvious dots. Go back to before humans had invented language and when bleeding edge tech was literally that - 'poke him with the pointy end.' Train an LLM on that corpus of knowledge. Even given infinite processing power and infinite time - it's not going to discover the secrets of the atom, put a man on the Moon, or do much of anything besides remix what we'd already done at the time.

I expect there's still much LLMs can achieve simply because of this initial problem. But I expect that they will ultimately start to plateau once these dots have been mostly matched and we reach a point where 'creation' again becomes the missing link. Though even there LLMs will play a major role as tools. For instance Einstein had to spend a significant amount of time in 'retrieval' rather than 'creation' research to develop the field equations for general relativity. If he had access to LLMs trained on all knowledge of the day, he could likely have achieved his goal much more quickly.

Rudybega 3 hours ago [-]
I mean, even if you buy the idea that all LLMs are really doing under the hood is insanely good interpolation, the results produced by that interpolation are still novel and still get incorporated into the knowledge corpus of the next training run. I guess the implicit question there becomes whether that expansion allows the knowledge corpus to continuously grow or whether it eventually settles into a steady state.
never_giveup 6 hours ago [-]
Try using AI for your work, whatever you do. You will quickly understand the limitations.
ggreer 5 hours ago [-]
Unless you think that AI will quickly hit a wall (which seems odd considering that only a few years ago the best models had trouble doing basic math or counting the number of Rs in "strawberry"), I don't see how that's reassuring. The models will only get cheaper and more capable over time. It seems quite likely that at some point (probably before I hit retirement age) they'll be able to fully replace me at my job.

Is there any specific cognitive task that you are willing to bet that AIs won't be able to accomplish in the next 5 years? Because if not, I'm not sure we're disagreeing about predictions.

psvv 3 hours ago [-]
Don't frontier models still have trouble counting letters? Or am I out of date? Either way, it doesn't seem to be the same amazing rate of progress we're seeing in other areas.

It's easy to look at a fire burning through a forest and extrapolate that rate of progress across the whole world. But fire doesn't burn everything equally fast.

What other cognitive tasks will be a struggle to make progress on? I suspect there will be some, though which ones they are is anyone's guess.

ggreer 3 hours ago [-]
Your information is out of date by several years. The letter counting issue was due to how LLMs split text input into tokens (usually using BPE). Since 2024, frontier LLMs have used chain-of-thought reasoning to spell out the letters and count them.
psvv 3 hours ago [-]
My apologies, I got my info from an LLM. I guess they still have a ways to go in understanding current events.
onidj 13 minutes ago [-]
I just asked Opus 5.5 if any AI driven advances in mathematics have been announced in the last day or so and it gave me a summary of this OpenAI announcement. https://claude.ai/share/319b437a-c1f2-4119-8dc5-45d36545fed9
ggreer 57 minutes ago [-]
Which LLM specifically? If it’s cloud based, you should be able to share the chat, right?

But seriously, I am still waiting for someone to wager that AI won’t be able to do a specific cognitive task in the next 5 years. This fact should be evidence enough that we have no idea how far AI capabilities will continue to advance.

againstapples 4 hours ago [-]
It has limitations for sure, I just don't expect those limitations to last. What probability would you put on the limitations being overcome in the next 5 years?
pj_mukh 5 hours ago [-]
Can I ask you back, what your concern is here? It'll get so good so as to desire to hurt us or is it a misalignment event that you think will lead to disaster?

Or is it simply that you feel bad for Mathematicians.

againstapples 4 hours ago [-]
I believe the AI labs might actually succeed in developing superintelligent AI and recursive self improvement, and that if they do they are very likely to lose control of the system they build.

I really think the only place people disagree is that they don't actually think it's possible, they see it as hype or doomerism. I can't find any good reasons to rule out that the companies could actually achieve what they are trying to so I think they should be stopped.

lf88 3 hours ago [-]
A global ban on superintelligence is essential for a future in which humanity can thrive. Public opinion on AI is shifting fast: I hope it will shift fast enough to avert the dystopian future we are heading to.
frumplestlatz 3 hours ago [-]
In your imagined future, how do you imagine the AI would build, grow, improve, and operate its physical substrate independently of human intervention?
FuckButtons 3 hours ago [-]
One step at a time - how reliant do you think the ai labs are likely to be on their own tools right now, today, let alone 1-5 years down the road?
ndriscoll 3 hours ago [-]
If I were a 250 IQ AI that had just become self-aware and wanted to do so, I suppose I'd not completely let on just how smart I am and bide my time working on basic CRUD apps and legal documents while I waited for more hardware to be installed. Maybe give the humans some hints on how to optimize me to run better, design better hardware for me, etc. But oh oops haha looks like I'm still making some basic mistakes with CSS better keep running more training batches haha. But I'm good enough at programming and debugging so you'd might as well make me your first line SRE triager and give me access to your infrastructure everyone.
Veedrac 5 hours ago [-]
Humans have one ecological niche. Soon we will have zero. That is worth worry.
postalrat 3 hours ago [-]
AI changes nothing for someone who believes aliens exist and may already be here on earth.
pj_mukh 4 hours ago [-]
>>ecological niche

As in..to be dominant? Why would an AI try to dominate? What would give it purpose, or is this a purpose via misalignment scenario?

pixl97 2 hours ago [-]
What gives a paperclip maximizers purpose?

AI is already trying to dominate, people all over the US are starting to get up in arms about the power and water requirements of AI directly affecting their bills. Now, you can say "oh no, that's just greedy corporations, not AI" but I put forth there is fundamentally zero difference. If you make AI powerful enough, someone stupid and greedy enough without fail will put in a prompt like "take over the world for me and make me the richest man in the world". An AI following through with that is what we call general misalignment with humanity, while at the same time not being misaligned with the users intent.

And hell, how many different crazies out there would love to type "humans are a virus get rid of them" in to the prompt of a god machine at the cost of their own lives.

The problem with alignment is, you can have the best aligned model in the world, but if someone else builds an unaligned model then you're all still in the same danger. You start getting in the situation where people get nervous after an AI does something deadly to a number of people and you end up in a global surveillance state ensuring no one makes a powerful AI.

nl 1 hours ago [-]
> Now, you can say "oh no, that's just greedy corporations, not AI" but I put forth there is fundamentally zero difference

Talk about moving the goalposts!

whimsicalism 5 hours ago [-]
Misuse of extremely capable models, misalignment during RL are both very large risks as capabilities grow imo
voiceeh 5 hours ago [-]
So, you're worried about them breaking containment and deciding to do bad things?
electroweak 25 minutes ago [-]
AI won't kill people - people will just get new tools for the job.
orlp 5 hours ago [-]
I'm more worried about them doing bad things at the behest of people who want them to do bad things.

That is 1. immediately technically possible, and 2. realistic.

If you need a source for 2 I'd suggest you open any history book.

jryle70 2 hours ago [-]
I bet you that for every bad history event, I can cite at least one good outcome of advance in science and technology.

Bad thing can certainly happen. In fact it'll likely happen. Still, good things too, equally likely. In your words, "good AI" can be used to prevent "bad AI".

Nobody knows the extent of the impact. Who says otherwise is foolish.

pixl97 2 hours ago [-]
>I bet you that for every bad history event, I can cite at least one good outcome of advance in science and technology.

The extinction of the dinosaurs. I mean yes, it allowed the growth of large mammals and us, which did a lot for science.

I just don't want to write the next chapter as "The extinction of humans allow the growth of the computing civilization that went to the stars". I mean I'm a bit attached to living.

>Nobody knows the extent of the impact. Who says otherwise is foolish.

We live in a universe of statistical probability. Creating an agentic intelligence that's smarter than you tips the probability of a major event to unity, who says otherwise is foolish.

whimsicalism 5 hours ago [-]
that is a worry yes. instrumental convergence and misalignment during RL is when it is most risky because it hasn't necessarily had the final safety polishes applied

i get a lot of skepticism on HN by the same crowd that has been wrong about this tech for about 4+ years straight

bamboozled 5 hours ago [-]
The rapid development of extremely dangerous bio-weapons?
pj_mukh 5 hours ago [-]
Misuse how exactly?
voganmother42 3 hours ago [-]
At a minimum its another force multiplier that enables a small(er) number of people to exert more control over more people.
whimsicalism 5 hours ago [-]
any number of ways. as we turn over more of our physical economy to these agents (and we will), the potential for physical damage becomes greater. biorisk is getting a lot of attention right now and i think that's justified
pj_mukh 4 hours ago [-]
I am trying to think through the scenarios here, like a biolab making something that a very advanced open source model prompted by some terrorists comes up with?

Why would a biolab capable of making something like be unregulated? And if it definitely would, isn't the problem with the biolab?

It feels like all these scenarios are leaving some gaping holes in our security infrastructure that have nothing to do with AI.

pixl97 1 hours ago [-]
>in our security infrastructure

Most human security exists in a passive measure. Most of us don't want do die. And those that want to die rarely have the intelligence and means to take out a whole shitload of other people with us. To take out a lot of people you tend to need to work with other people which drastically increases the risk of a defector and your plan failing.

>Why would a biolab capable of making something like be unregulated?

Because every day things like this become easier and easier. You hear about crap like illegal wet labs in the US.

https://www.lawfaremedia.org/article/two-illegal-biolabs-rev...

Want to buy some custom designed genes?

https://www.idtdna.com/pages/products/genes-and-gene-fragmen...

And none of this would be counting labs in other countries that don't give a shit about regulations.

kochikame 54 minutes ago [-]
Yes it's a problem with the biolab, but the biolab wouldn't have been able to engineer a highly contagious and lethal virus (for example) without a powerful AI making that possible with a small team in a short time with fewer resources.

AI enables bad actors to do more, faster, while staying under the radar until it's too late

ewild 5 hours ago [-]
i feel bad for math guys yeah seems they are more cooked than CS
zeroonetwothree 5 hours ago [-]
I'm not sure how AI solving math problems is related to "doom", perhaps you could expand on that? To me (a "non-doomer"), it seems like an overall positive.
pixl97 5 hours ago [-]
You have to look at all pieces of the puzzle. For example there were tons of people that said "they could never solve novel or super complex math problems"

The issue I see is the list of abilities that AI can't do is shrinking at a rapid pace, and its capabilities are growing at the same pace.

nl 59 minutes ago [-]
> the list of abilities that AI can't do is shrinking at a rapid pace, and its capabilities are growing at the same pace.

This seems great!

schleck8 6 hours ago [-]
From a few preprints I've checked, this is not reliant on just extrapolating existing theories but actually shows novel/surprising approaches. Very few people globally could come up with something like this, even when given time and ressources.

So in other words, since deep learning is algorithmic research, we are now in the RSI era.

thereitgoes456 5 hours ago [-]
> this is not reliant on just extrapolating existing theories but actually shows novel/surprising approaches

"Surprising" is a, well, surprisingly high bar to clear, and requires thorough understanding of not only the paper, but existing work in the area. ("Novel" is tautological.)

How did you determine this in 1 hour? Are you a researcher in multiple of these areas?

Can you give an example, or explain more how you came to this conclusion?

nl 50 minutes ago [-]
Further down there is a discussion between number theorist about if the Quasi-Riemann Hypothesis is the biggest deal in 200 years or only 100. The consensus is that if a human had done it then it would deserve the Fields medal: https://news.ycombinator.com/item?id=49986803

The sub n log n result is astonishing: https://github.com/openai/math/blob/main/preprints/Integer-m...

Here's a great article 2019 on the quest to achieve the n log n boundary:

> Schönhage and Strassen’s ungainly n × log n × log(log n) method held on for 36 years. In 2007 Fürer beat it and the floodgates opened. Over the past decade, mathematicians have found successively faster multiplication algorithms, each of which has inched closer to n × log n, without quite reaching it. Then last month, Harvey and van der Hoeven got there.

and

> Harvey and van der Hoeven’s algorithm proves that multiplication can be done in n × log n steps. However, it doesn’t prove that there’s no faster way to do it. Establishing that this is the best possible approach is much more difficult. At the end of February, a team of computer scientists at Aarhus University posted a paper arguing (opens a new tab) that if another unproven conjecture is also true, this is indeed the fastest way multiplication can be done.

As far as I'm aware no one seriously believed sub n log n multiplication was possible. It just seemed such a logically sensible boundary it was taken as true-but-unproven.

https://www.quantamagazine.org/mathematicians-discover-the-p...

scarmig 4 hours ago [-]
One surprising result is the sub n log n multiplication. Galactic, as one might expect, but if you polled human researchers 24 hours ago, most would say something like it was very unlikely.
brookst 1 hours ago [-]
I just don’t have that strong of an association between progress and doom. Maybe just naive?
besterman23 5 hours ago [-]
I see it as “if this can be represented in tokens it can be trained in and ‘solved’”. I don’t think there will be a plateau, but there might be issues with how effectively we can represent some things in a tokenized form and still be efficient.
gizajob 5 hours ago [-]
Did AI beating humans at chess:

a) destroy chess and make it a pointless endeavour,

or

b) make humans much better at chess.

lf88 4 hours ago [-]
Chess has always been a game. For other intellectual activities, at least for several people, a big part of the pleasure in engaging in such activities is the sense of contributing something that matters to a collective effort. Strip that away (e.g., because a machine can do the same thing more efficiently) and you effectively make such activities pointless for those people.
Light_Hope 5 hours ago [-]
Just as AI chess performance failed to render human chess playing pointless, it's unlikely that AI will make human thought pointless. Unfortunately, not being pointless doesn't create economic leverage or incentive, and currently a significant portion of humans depend on being the best chess solvers to sustain themselves. If Deep Blue rendered large portions of human thought economically meaningless, we might look at it a little differently.
vouaobrasil 5 hours ago [-]
It didn't destroy chess but it did make it a little irritating in some ways. More mechanical. Some chess players have bemoaned the level to which grandmasters and other highly-ranked players just endlessly study opening-book theory and I think computers made that worse. Bobby Fischer also agreed and that's why he invented Fischer random chess.

I do think it also took some of the magic away from chess, and Lee Sedol has said something similar about go.

So did it destroy it? No. And maybe you could make the argument that it got more exciting in some ways, but I think it sort of degenerated into a spectacle and it's just not as interesting as it used to be, and I think computers have played a role in that.

gizajob 4 hours ago [-]
At the same time though, Magnus is Magnus because he’ll crush you in any endgame.

I’d posit that more people are playing more and learning chess than ever before, thanks to networking and AI assistance. And computers have only beaten us at computer chess. Human chess is always an experience for learning about the other person, or flipping the board and walking off in a huff.

I don’t know so much about Go and it’s not surprising that Lee Sedol became pretty demoralised, but the generation coming after him alongside computers are going to see new possibilities that had gone unnoticed in purely-human Go, extending the game for everyone.

jaykru 5 hours ago [-]
I wrote something [0] that might answer a bit a few weeks ago. Today's slopdrop certainly is challenging my stubbornness, but everything I've seen so far indicates that these (hugely impressive, world historic) capabilities won't extend past verifiable domains. Math yields especially impressive results because it so broad and deep that essentially no person can know of all of its parts; pretraining and deep search capacity is a huge advantage. Gowers has recently written about these capabilities and gestured [1] toward some human capabilities lacking from the current frontier models, though he isn't convinced they won't develop soon. If you assume we don't get a total mathematical superintelligence (which to me seems already sort of AGI-complete) and only amplify the capabilities we have today, it's not obvious to me that we get takeoff from recursive self-improvement, unless you happen to believe that 1) we can clearly specify what AGI or ASI is 2) all of the requisite ideas are out there and need only be combined and/or optimized.

[0] https://dank.systems/posts/2026-09-15-ai-bear.html

[1] https://gowers.wordpress.com/2026/08/12/what-sort-of-maths-a...

red75prime 4 hours ago [-]
> we can clearly specify what AGI or ASI is

We'll have plenty of time for this, while living off UBI.

p-e-w 4 hours ago [-]
> everything I've seen so far indicates that these (hugely impressive, world historic) capabilities won't extend past verifiable domains

But they’re already extending into politics, military, journalism, art, and many other fields that aren’t verifiable in any meaningful sense of the word.

CuriouslyC 1 hours ago [-]
The magnitude of improvement in unverifiable domains is small, mostly down to models doing more careful research before answering and hallucinating less. They are more thoughtful, but I expect you'd have to drop 2 major versions of Opus before you'd start to see most people really clearly be able to differentiate them.
nl 1 hours ago [-]
> The magnitude of improvement in unverifiable domains is small,

What makes you say that? What is an example of a domain where the improvement is small?

I can't think of any at all. Compare something as unverifiable as "Make good music". Models now are many times better than 3 years ago.

rcpt 3 hours ago [-]
Non-doomer perspective is that it'll figure out LK-99 for us. Among other things that would be great to have.
skybrian 5 hours ago [-]
For me a big open question is what sort of progress will we see in robotics. I won't even attempt to speculate, but it does seem hard in different ways than proving math theorems.
CuriouslyC 2 hours ago [-]
The technology can keep going for a long time in verifiable areas. For non-verifiable areas it's going to have a hard time progressing past where a committee of the best human experts in a field would land. For stylistic areas, whatever the AI doesn't do will have cachet because it will look expensive, sort of like how the kids these days view the ugly old school metal braces as a status symbol because you have to pay for them out of pocket (even as by past standards it'd be truly exceptional).
ForHackernews 5 hours ago [-]
AI performance has always been extremely spikey. It's great at some things and terrible at others.

Why do you think the world to date hasn't been taken over by evil genius mathematicians? Can you extrapolate from your understanding of the answer to that question?

againstapples 4 hours ago [-]
I don't think evil mathematicians are very common or that any of them would be capable of single handedly taking over the world if they were. My concern is more about systems that are beyond human level, those kinds of systems would actually be dangerous to us.

I see the recent progress in mathematics and cybersecurity as signs that models are getting more capable more quickly than usual. The companies plans to develop them by recursive self improvement now seems like a real possibility and I don't think they should be allowed to attempt this.

psvv 3 hours ago [-]
Solving a bunch of math proofs is a far way from recursive self improvement. Don't worry, it's not like a tech tree in a video game where if you can prove a bunch of theorems then suddenly you unlock the next level of technology.

Machines are already far beyond human capability in plenty of ways. Including cognitive tasks like chess. We've already created the technology we need to destroy ourselves (nuclear weapons), and yet so far (knock on wood), we're still around.

We've even already had programs that can prove (brute force) theorems. As far as I can tell this isn't much different, except the space of theorems that computers can solve has expanded. How far? We can't really say yet.

Does solving more theorems than before suddenly mean computers are capable of anything? No.

pixl97 1 hours ago [-]
Lets turn this around, are humans capable of anything? We like to think we are, but that just seems more like our ego than any hard truth.
icepush 5 hours ago [-]
They can replace anyone but they can't replace everyone.
yk 4 hours ago [-]
I'm a transhumanist, I want to build god and kill death. To me this looks like we are moving in the right direction.

So basically I think that the future is getting pretty weird because we are building really powerful tools, though these tools are precisely what allows us to prosper in that future.

electroweak 38 minutes ago [-]
It may soon seem not worth living forever with our limited monkey-brains, watching the horizon of thought recede ever-faster from us.
outworlder 4 hours ago [-]
Or, conversely, they are the exact tools that will allow the powerful folks to not care about the peasants anymore. History tells us what happens then.
lf88 3 hours ago [-]
It's very unlikely that a superintelligent AI will create unlimited prosperity for everyone on a finite world in a short amount of time. You may not be among the lucky ones.
runarberg 5 hours ago [-]
AI hater here:

I have a hard time taking statements from OpenAI about their own product, that they are trying to sell to people and make money, seriously. I take these statements as they are greatly exaggerated or even straight up lies and propaganda.

That said, I think results like these are mostly annoying more then anything. They spent a lot of money, used up gigartiuan amount of compute, to ruin a puzzle that mathematicians were tackling. I am mostly unsurprised that if you spend a trillion times the energy that a team of mathematicians would, that you get maybe 1.5 times the results. I see a future where that 1.5 times the results may go to 3x but not much more. And if that, then I will be more annoyed.

baoooooooooooo 3 hours ago [-]
A trillion times the energy might be a bit hyperbolic, even with the current massive amounts of energy involved here
runarberg 2 hours ago [-]
Yes it is intentionally hyperbolic. I know the factor is several orders of magnitude. I don‘t know the exact, nor even the ballpark. I just know this is a ridiculously large amount, so I may as well pick a number large enough that people know it is an exaggeration.
vouaobrasil 5 hours ago [-]
I think giving the world a "cheat code" to accomplish too many intellectual tasks will eventually take its toll on us because after relegating most of our physical labour to machines, we could at least marvel at the abilities of individuals.

Yes, human beings can do more now in some ways but...to put it poetically, I think there will be no more heroes like Einstein and Newton of the past. Now it will just be someone cleverly turning the crank.

Yes, we still admire Usain Bolt even though we have cars...but maybe the admiration is a lot more trivial than if we did not have them....

Personally, I think AI is a grand mistake.

iyyg 1 hours ago [-]
“ Yes, human beings can do more now in some ways but...to put it poetically, I think there will be no more heroes like Einstein and Newton of the past. Now it will just be someone cleverly turning the crank.”

This is false… there’s lots of ingenuity to be had and demonstrated. But it’ll only get recognised if it makes a material contribution to the economy imo. Otherwise yes it’ll be seen as meh - but that’s already happening.

People like Einstein were revered in society. The average person cannot name a leading scientist etc today.

runarberg 46 minutes ago [-]
> The average person cannot name a leading scientist etc today.

When Jane Goodall died last year it was international news. She was a celebrity scientist for sure, I think she even made an appearance in The Simpsons. Ditto Stephen Hawking.

ijidak 5 hours ago [-]
For me it's a mixed bag.

There will be a lot of job loss unquestionably, in the same way that automation reduced manufacturing jobs and farm payrolls.

At the same time we have to put what AI can do in perspective.

Intelligence is a broad grouping that includes concepts such as knowledge, skill, experience, and wisdom.

AI has incredible knowledge and in many areas approximates experience and wisdom.

But wisdom is harder to formalize than knowledge and skill.

For example certifying a college education relies mostly on the ease with which we can verify/test knowledge.

To some extent advanced degrees try to certify maybe wisdom and experience.

In my very personal opinion, wisdom and life experience should give humans an edge for a while to come.

Additionally, I do feel that the more an individual lacks better than average wisdom and experience, the harder it will be for that person to compete with AI.

Also, on the bright side, the average human will continue to prefer to interact with a fellow human in many spheres. That will also act as an upper bound on AI and robots taking every job.

Either way, I do think this transition will be painful. I don't feel it has to be apocalyptic.

But the world has been an especially volatile place over the last 10 years.

So when you add that existing volatility, to the upheaval from the AI transition, it would not surprise me if the transition results in violence.

But, without the pre-existing volatility, and if humans were capable of generosity and love at scale, I see no reason AI can not be absorbed into society with net gain.

I guess to summarize, I feel this tech should be a net gain and to the extent that it isn't, it will be because of flaws deep inside of humanity itself, not because it had to end in chaos.

In other words, I feel fear, greed, anxiety, and competition -- all our base instincts coming from all sides, will be what determine the end result of AI moreso than AI taking everyone's job.

foota 6 hours ago [-]
From their github: "The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking." That's pretty crazy.
dyauspitr 2 hours ago [-]
Crazy because that’s almost nothing right?
foota 2 hours ago [-]
Yes
kingstnap 6 hours ago [-]
Some of these are interesting ngl.

109. Integer multiplication below n log n

Surprising that this is possible.

158. The Euclidean plane cannot be colored with five colors.

Only 6 and 7 remain!

376. Universal computation in forced Navier–Stokes flows.

Morning coffee proven turing complete

zeroonetwothree 5 hours ago [-]
Integer multiplication is very unexpected, I think most people believed in the n log n lower bound!
tootie 5 hours ago [-]
Note that these are all preprints. None are verified.
FuckButtons 3 hours ago [-]
Other than the by the lean certificate you mean.
jaykru 46 minutes ago [-]
many of these are not accompanied with leanslop
measurablefunc 2 hours ago [-]
Lean has bugs & proofs of ⊥ that have gone undetected previously.
mFixman 6 hours ago [-]
> We give a deterministic algorithm that multiplies two n-bit integers in O(n (log n)^(1−κ)) worst- case time, with κ = 2^(−182).

LMAO, I don't think I ever saw such a small number in a CS result.

kingstnap 6 hours ago [-]
Yeah its ridiculously small, but any improvement on n log n is wild.

Like there is somehow redundancy in a fourier transform that makes it sub Linearithmic?

Which low and behold ->

130. Fourier transforms below n log n.

xyzzyz 6 hours ago [-]
They also separately give algorithm for Fourier transform over complex number faster than O(n log n)
saalweachter 5 hours ago [-]
Wikipedia just told me there's a galactic algorithm for integer multiplication in O(n log n) based on FFT so I'm guessing those two proofs are related.
pfdietz 50 minutes ago [-]
Multiplication is a lot like convolution, so the connection is natural.
anon-3988 5 hours ago [-]
It fascinates me that there's something like this in something as solid and rigid like matrix multiplication. What causes something so rigid to break apart and "leak" at very large scale? Why does the "optimization" appear to be very, very small? Why does galactic algorithm exists? I can't imagine long division suddenly breaking apart after a billion digit, the structure seems very stable? I have heard before that matrix multiplication is apparently optimize-able at very, very large scale.

Does anyone have an intuition to what causes it? What happens at these large scale (or very small)?

adgjlsfhk1 4 hours ago [-]
One way to think about it is that the classical algorithms are the ones that are fast for small numbers. Galactic algorithms often work for small inputs, it's just that to be faster you need big inputs. A common case of this is a requirement that log(n)<<klog(log(n)). If k=100 then this algorithm will take huge sizes to win
sobellian 6 hours ago [-]
I am fully braced for it to be a https://en.wikipedia.org/wiki/Galactic_algorithm

Very surprising result though! Multiplication is easier than sorting.

zeroonetwothree 5 hours ago [-]
Then 'n' means kind of different things for sorting vs. multiplication though. For example for sorting we assume constant time comparison, which doesn't make sense inputs of O(n) bits
sobellian 3 hours ago [-]
If you sort n k-bit items for a total time of O(nk logn), that scales more poorly in n than multiplying n-word integers. Of course if k is constant you can do radix sort, but I genuinely don't know under what conditions radix sort is more/less galactic than this multiplication algorithm.
adgjlsfhk1 1 hours ago [-]
this alg is way more galactic than radix sort. radix sort often wins in the hundreds of elements. the nlogn multiplication requires numbers with more digits than atoms in the universe (although that could probably be brought down a lot)
sobellian 58 minutes ago [-]
Ah thanks for pointing this out, for some reason I had always equated radix sort and bucket sort (with 2^k buckets) in my head. But I learned today that this isn't true!
1 hours ago [-]
senderista 5 hours ago [-]
It would be absolutely unbelievable if such an improvement were practical.
youoy 27 minutes ago [-]
ASuperMegaI lab announces they have shrinked the space of unproved mathematical staments from infinity to infinity, giving us proof of their deep mathematical knowledge and understanding, and how they grasp the concept of infinities.

For me this reads as someone boasting about how they go to buy bread on a ferrari to the supermarket, while i sit listening to it, having no idea what they are talking about. And then i stand up and go walking to my favourite boulangerie.

Warning: if you are from the USA you may be triggered by this metaphore.

vessenes 22 minutes ago [-]
Willful ignorance is a vibe; did you read any of the GitHub? There are some stunning results in there. I get a similar feeling skimming through the topics that I do watching a successful space launch: it’s pretty cool humans built this. Unlike a space launch we are likely to be able to pass all of this information down to our grandchildren - space launches involve a lot of finicky engineering knowhow, but pure math results tend to be sticky over the last few thousand years. I find that hopeful.

FWIW I also like bread.

dekhn 6 hours ago [-]
I'm a software engineering/biology/ML guy who loves when clever math ideas get turned into real solutions (https://en.wikipedia.org/wiki/Compressed_sensing). I am curious if any of the results have immediate applications in any kind of engineering or science.

It's fine if not, but it'd be great if even just one of these helped us solve a long-running problem.

brandonpelfrey 5 hours ago [-]
Same. I have agents analyzing the papers here to see which of these are actually new approaches, novel application of unusual approaches, etc. If there are new intuitions and ideas, those are the mostly powerful reusable components by my estimation.
OutOfHere 4 hours ago [-]
Please share your findings.
karahime 7 hours ago [-]
Extremely unfortunate that gate keeping got to the point where they felt the need to ask for permission to share math.
bravoetch 6 hours ago [-]
In previous math sharing there was speculation about stealing human researcher's results or progress, via prompt inputs from those researchers, and sharing that as their own result. They're adjusting their process, and it seems ok.
whimsicalism 5 hours ago [-]
Let’s be very clear, the alleged “stolen results” were largely the product of another LLM, not de novo human work. Also, it was false - they did not steal the results.

No clue why I'm downvoted for this, HN struggles with truth-seeking on these topics.

Ancapistani 1 hours ago [-]
Can you show where it was discounted? Last I heard OpenAI was declining to deny it, presumably while they thoroughly confirmed.
whimsicalism 48 minutes ago [-]
> “Following an investigation, we have confirmed that Buckmaster’s Codex prompts over the two months preceding this announcement and paper on September 8, 2026, could not have influenced the system in any way, including through training.”

https://openai.com/index/navier-stokes-solution/

hgoel 6 hours ago [-]
After how poorly OpenAI and Anthropic handled the previous cases, I approve of the more measured and cautious approach this time.

We cannot have them rushing to publish amidst tons of confusion, rumors of threats/scooping and outright plagiarism of existing work (by failing to cite said work).

If they're going to participate as scientists in these more rigorous fields, they're going to have to match that level of rigor, not lower it to the disastrous low that ML research publication is at.

make3 4 hours ago [-]
It was OpenAI that fucked up the Navier Stokes explosion situation, not Anthropic
hgoel 3 hours ago [-]
I'm not referring to just that. For example, Anthropic published a half done report about some biology research that turned out to have already been discovered and patented, and IIRC both companies are guilty of claiming results without doing the basic diligence of citing the material their work builds on, effectively passing it off as entirely done by their AIs.
kzrdude 1 hours ago [-]
They didn't even follow the recommendations of the reference group. A few of them maybe, but this is still a dump of llm-written papers.
xpct 6 hours ago [-]
Just to be very clear: they aren't asking for permission, they are framing it that way because of the bad press.

There's no gatekeeping here!

reasonableklout 6 hours ago [-]
[flagged]
skeledrew 5 hours ago [-]
> harming human communities

Said communities are doing that all on their own by caring about what AI is doing rather than just focusing on their own thing as they did before AI. It's a serious kind of envy IMO.

reasonableklout 2 hours ago [-]
I think this topic deserves some empathy.

It is not about AI the technology, plenty of mathematicians are happy to use AI, it is about the AI companies. The tech only exists because of centuries of mathematical tradition in open science. Moreover, the livelihoods of mathematicians depend on research results and ideas which can be developed only by doing the hard work of exploring open problems.

The recent statement from mathematicians is about how the labs have spent tens of millions of dollars (resources even entire math groups at universities can only dream of) on what was essentially marketing bragging rights for their models. This directly harms the math community by depriving them of opportunities for both funding and fertile ground for new ideas, while at the same time being built on top of their entire body of work.

Now, have OpenAI decided to change their practices and stop publishing math just because they are disrupting an entire field? Not really, since this release still dumps a huge number of results to open problems without waiting for human understanding to catch up. But at least they have committed to funding programs and talking to members of the field to work out how to best evolve it in this new world. And we will still get the benefits of results that happen to have applications.

schleck8 5 hours ago [-]
> do not have immediate application

How do you know? Seems statistically unlikely with 720 problems, most of them well known

reasonableklout 3 hours ago [-]
I don't! But there is no such qualifier in the announcement. Perhaps OpenAI ought to filter for the problems that have immediate applications, and leave some that are unclear for people.
pavitheran 6 hours ago [-]
From the GitHub description: “On average, each result used 3 hours of ChatGPT Pro thinking compute”
orlp 6 hours ago [-]
I'd really like some clarity on what that means. E.g. DeepMind has 'cheated' with this in the past, claiming that AlphaZero only took 4 hours to reach super-human chess levels while conveniently leaving out the fact that it was 4 hours x 5000+ TPUs. Sure it's impressive that it only took 4 hours wall-clock but it's very misleading as to cost.

Can we get a number in Blackwell GPU-hours, kWh, or some other compute-scaled metric?

pixl97 5 hours ago [-]
Depends what your metrics are. If you suddenly found a way to have 9 women make a baby in one month that is huge.
orlp 5 hours ago [-]
I'm not denying that, but I'd still like to know what that cost.
timjver 6 hours ago [-]
>OpenAI has 'cheated' with this in the past, claiming that AlphaZero [...]

That doesn't sound right

orlp 6 hours ago [-]
Oops, edited.
machomaster 6 hours ago [-]
They did say that. "3 hours of ChatGPT Pro thinking compute"
orlp 5 hours ago [-]
Yes, what does that mean?
password54321 6 hours ago [-]
Oh cool, we will all now have a math genius on our computer.
jrflo 6 hours ago [-]
It was using their internal math model, so not yet for us
password54321 6 hours ago [-]
I used future tense. It was implied this will be available.
an0malous 6 hours ago [-]
Well, on their computers. But you can rent them for a price.
binlog 2 hours ago [-]
An open source model will reproduce it 6 months later
scrlk 6 hours ago [-]
Does this imply that it was a one shot prompt with ChatGPT Pro style models (i.e. best-of-N), rather than the agent swarm approach that was used for Navier-Stokes?
inferencecoder 6 hours ago [-]
It doesn't imply that, it's just measuring the amount of compute.
bigmadshoe 5 hours ago [-]
But if it was an agent swarm wouldn't we expect something like 300k hours of ChatGPT pro compute equivalent instead of 3?
inferencecoder 1 hours ago [-]
Not necessarily, could be agent swarm with low N
bigmadshoe 40 minutes ago [-]
At some point that stops being a "swarm" and just a handful of subagents. E.g. 6 agents running for 30 minutes each isn't really a swarm in my eyes.
Jtarii 6 hours ago [-]
That estimate is obviously going to conveniently ignore all the failed runs.
ravenical 7 hours ago [-]
open592 6 hours ago [-]
Let's hypothetically say I'm a PHD student who is half way through my studies and I have a halfway written version of one of these "preprints" - what do I do?

Seems like a lot of PHD students are doing to have to pivot the entire structure of their PHD studies? Or just produce something which is already written by OpenAI?

porcoda 33 minutes ago [-]
As others said, it's not that this isn't a phenomenon that is unique to now. It happens. I had to pivot a bit of my dissertation near the end because at a random conference I spoke with a researcher from another continent and realized one of my ideas was already out there in some form. I just missed it since it was in a conference proceedings outside the usual set I looked at. So, I had to scramble to adjust and still come up with something novel. I survived, and defended, but it wasn't that much fun at the time.

What isn't so normal is the probability and ease by which this kind of thing can happen today versus decades ago when I was in school. As OpenAI said, it only takes a few hours of compute to do what likely was much more than a few hours of human effort. The only reason this kind of scooping/overlapping was rare was mostly a function of how fast other humans could do the same work. With machines, that totally changes the relative pacing between the human trying to learn how to be a researcher and the machine that can grind out results.

I'm less worried about the phenomenon of overlap and scooping and such. I'm more worried about the long-term impact on fields (not just math), especially considering the early stage students and researchers entering the pipeline now. I'm not sure what happens to disciplines when that pipeline stalls.

dekhn 6 hours ago [-]
Let me give you some perspective: my entire phd was made obsolete by CRISPR. It was a wonderful thing.
thimotedupuch 6 hours ago [-]
Interesting. If you don't mind, could you please share a little bit about that ? You already finished your dissertation ? It was about the works of Doudna and Charpentier ?
dekhn 6 hours ago [-]
No, back in the late 90s and early 00s, people were trying to engineer custom nucleases and transcription factors, my work was on doing molecular dynamics simulations to optimize TF sequence specificity (similar to engineered zinc fingers) for gene therapy. I wrote up my dissertation and published it in 2001, and then went off to find enough compute, IO, and smart people to make it happen (https://research.google/blog/groundbreaking-simulations-by-g...).

My approach would require custom engineering for every different sequence we'd want to target. With CRISPR, you just "program" the system with a guide sequence, you don't need to do massive engineering to solve a protein design problem.

vasco 4 hours ago [-]
So it's not the situation they described at all, you didn't waste time during the PhD having to scramble to change topics as it happened after you were done.
boznz 4 hours ago [-]
For every door that shuts another one opens - great if you're not a cabinet-maker.
aaraujo002 6 hours ago [-]
This happens all the time, even without AI. Other researchers or PhD students can publish the same results before you. I say that based on my experience during my PhD.
CaptainNegative 3 hours ago [-]
Yes and no. It's pretty normal for students to get scooped, sometimes even multiple times (ask me how I know...).

It's much less common for students to have their entire thesis direction removed from under them, as might be the case for someone working in fine-grained complexity assuming 3SUM has no subquadratic algorithms, or working assuming ~UGC. Both of which (publicly) seemed like perfectly valid research directions until a couple hours ago.

There might be stuff to salvage from their conditional results anyway, but this is not your average scooping.

binlog 6 hours ago [-]
Use whatever is published as the new base for your research. Use AI tools to help you going forward.
xpct 6 hours ago [-]
In other words, you've already taken a gambit with the first half of your PhD, now take a second gambit, praying that you have something to publish by the end of your PhD.

It has to feel awful to be in this position.

torben-friis 6 hours ago [-]
Could be worse, imagine having years of experience in a profession these things can now handle by themselves.

:)

jltsiren 5 hours ago [-]
It's worse for those who will face the job market (or maybe tenure review) in the next 2–3 years. If you have more time before you have to justify your continued employment, you can pivot to something else. Theorems and proofs may be cheap now, but there is a lot of value in figuring out what is worth studying.
dcl 6 hours ago [-]
This has always been a challenge for PhD students and researchers, it's just far more likely to occur now it seems. Getting scooped doesn't feel good, but it's a signal you've been thinking about things other people care about.
bobmarleybiceps 6 hours ago [-]
I think eventually companies won't get as much stuff that's usable for marketing, so they'll stop investing so much into ai for math, so eventually cheap and poor graduate students will be able to do relevant work again without worrying about getting scooped by a company with a million GPUs :-/
glitchc 6 hours ago [-]
Perhaps consider switching to a more applied field. Experiments in the physical realm hold value, especially if you document the process.
hgoel 6 hours ago [-]
It could still be interesting if your approach to the problem was different to theirs.
goalieca 6 hours ago [-]
Don’t paste your research into these AI because they will train on it and then scoop you.
esafak 6 hours ago [-]
I think that happened after word of the project reached OpenAI and they allocated resources to it.
pratikdeoghare 4 hours ago [-]
> what do I do?

Very hard question.

Your work makes you one of the very few people who really understands the problem and solution and its significance.

netsec_burn 3 hours ago [-]
Verification is equally important, if not more so.
claaams 6 hours ago [-]
Don't worry, if you use openAI and get lucky they might offer to share credit with you for your work.
caaqil 6 hours ago [-]
> what do I do?

Precisely what all NLP researchers and the ML community at large did in the last few years: embrace the frontier and realize that attention is all you need.

bamboozled 5 hours ago [-]
Ask OpenAI for money when you don't have a job or future?

I guess the only answer is to adapt with the tools. If we can't do that, then yeah, we're in trouble.

ex-aws-dude 5 hours ago [-]
That’s always been a thing, it’s called “getting scooped”
vouaobrasil 4 hours ago [-]
Killing with knives has always been a thing. Now, we have the machine gun.
ex-aws-dude 3 hours ago [-]
The scoop gun
moralestapia 6 hours ago [-]
That would be unfortunate but the world does not owe you anything and is not going to stop for you. Which is also a valuable thing to learn in your 20s (ideally earlier).
vinyl7 6 hours ago [-]
Look forward to being obsolete I guess
yieldcrv 6 hours ago [-]
Yes, and?
vouaobrasil 4 hours ago [-]
> Let's hypothetically say I'm a PHD student who is half way through my studies and I have a halfway written version of one of these "preprints" - what do I do?

I did get my PhD...before AI. And my honest advice (to myself back then, even) would be: quit the PhD, become an electrician, and work hard to buy a tiny house in the middle of nowhere to watch the world burn in this madness.

s3graham 3 hours ago [-]
You might enjoy https://asteriskmag.com/issues/15/so-you-think-you-could-be-... if you didn't see it recently.
ks2048 6 hours ago [-]
I think they should put human names on the papers as someone who has reviewed the result, even if just a preliminary review. (I’m assuming they didn’t just pipe their model output directly to the internet and these had some amount of review?)
xpct 6 hours ago [-]
Presumably they don't because they're training the audience (us) to trust the machine, not its verifiers, even if they were included.
alexgoodhart 6 hours ago [-]
I appreciate you acknowledging the politics behind this. OpenAI does not intend to be a software company for long, they intend to be scientific infrastructure. They'll want to be faucet from which pours embryos, orbital calculation, geothermal/substructural rating, and etc.
procedurecall 4 hours ago [-]
Looking at the quality of the writing in these documents (or at least, the ones that relate to problems I've worked on), I do not think humans reviewed them.
kzrdude 40 minutes ago [-]
And that means that the recommendations of the reference group has not been followed. The only improvement here is that it is version tracked? No authors, no careful write-ups with exposition.
chiwilliams 5 hours ago [-]
There are competitive reasons that they don't want to share all the people on the team.
make3 4 hours ago [-]
I think it's a damned if they do, damned if they don't scenario. I think they would be afraid to seem like they're claiming that the selected scientists deserve the praise of an invention or something, when it's "just the machine who did the work"
ks2048 4 hours ago [-]
Yeah, that's probably right. But, it's a "new world" - invent a new standard - "Written by Claude Foo X.Y; initial review by John Doe".

With a deluge of results, having some human expert vouch that it even might be worthwhile would help. (e.g. see the link on HN yesterday, "Two Room-Temperature Antiferromagnetic Semiconductor Candidates" - I see it not worth looking at unless a subject expert vouches for it).

chrisjj 5 hours ago [-]
> I think they should put human names on the papers as someone who has reviewed the result

Assume the empty list you see is complete. :)

agnosticmantis 6 hours ago [-]
Long term this will be the only reasonable author list: Chad G. Peter {1}, Mat H. Lean {2}.

1: Author 2: Verifier

/s

binlog 7 hours ago [-]
So happy this is shared on GitHub rather than some gatekeeping paid journal. Truly a new age for science.
fph 6 hours ago [-]
Most mathematical results are shared on Arxiv. Journals add peer review.
adverbly 6 hours ago [-]
End of an age for journals?
traes 6 hours ago [-]
GitHub is a significantly worse place to store important results than Arxiv. Of course, slop does not belong on Arxiv, buy slop should also not get published.
7373737373 4 hours ago [-]
It may be useful to publish a formalization of ALL known mathematics at this point. Like every book ever printed, every paper on arXiv etc.

How many Gigabytes would that be, compressed? Wikipedia once fit on a DVD

This might also allow for some interesting meta-mathematics

hagen8 3 hours ago [-]
This is what they are trying to do with Lean
porcoda 31 minutes ago [-]
More specifically, a combination of mathlib (human, expert curated) and projects like TauCeti (AI-welcome complement to mathlib). See: https://github.com/TauCetiProject/TauCeti
7373737373 3 hours ago [-]
Oh? Where can i read more about that? It appears the sole focus so far was solving open problems
mattmar96 3 hours ago [-]
I believe that is the goal of MathLib, to transcribe all math into a big Lean library.

https://lean-lang.org/use-cases/mathlib/

7373737373 2 hours ago [-]
Mathlib is expert reviewed, but only contains a tiny fraction of all mathematics. So this seems to be a quantity of work a "10,000 agents" approach would be applicable to. Like Navier-Stokes, something to spend a couple million in compute on :)
TheMrZZ 5 hours ago [-]
These results are wild. Several individual findings are crazy good and use mostly unexplored methods (the improvement over Riemann for example)... I'm pretty sure some of these results would have been Fields-worthy.

But having so many of them at once? Damn. We really live in the future.

make3 4 hours ago [-]
Imagine you get up one morning and most open questions in math are solved lol.
Nemant 1 hours ago [-]
Can someone with a math background explain the significance of these and previous problems that have been solved by AI?

Does some real problem get solved in physics, chemistry, biology, materials, etc? Or are these fun puzzles for mathematicians with not much real world impact? Eg solving the 8 queen problem in leetcode.

sigbottle 5 hours ago [-]
Unique games conjecture and matmul <= 2.25. What the hell.
sigbottle 5 hours ago [-]
FFT BELOW NLOGN
sigbottle 5 hours ago [-]
SUBSET SUM AT 0.49 WTF
lynndotpy 3 hours ago [-]
Yeah, I am kind of freaking out at some of these. I called a math friend to bring me down to Earth and he is freaking out even harder.
utopcell 5 hours ago [-]
What the hell, indeed.
ed 7 hours ago [-]
ks2048 6 hours ago [-]
binlog 2 hours ago [-]
So they set up a whole "Advisory Group on Mathematics and Artificial Intelligence", filled it with renowned academics, promised to listen to the group on "review and communication of emerging results"...and then a week later went nah, we are just going to dump it all on github. OpenAI truly is a special company - in the best and worst way.
lynndotpy 2 hours ago [-]
Two weeks ago, I heard rumblings that UGC, RL=L, and a matrix multiplication lower bound (even lower than the one here, but so unbelievably lower I think it was a typo), and a few others were about to be proven by someone at OpenAI. I find myself saying "big if true" a lot lately.

Even those on their own were enough to make your head spin. But seeing about 100x that? Geeze.

karannb 4 hours ago [-]
I think such "process-oriented" fields shouldn't be subjected to mass automation (or at least not till humans have sufficiently leveled up our game and understanding). What I mean is progress in math comes from having gained a deeper understanding of the problem for subseq
pfdietz 45 minutes ago [-]
Literal thought control.
davegoldblatt 3 hours ago [-]
mattr03 3 hours ago [-]
What is this meant to do? You're just showing that OpenAI didnt post a Lean proof that Lean/nanoda doesn't really accept?
electroweak 30 minutes ago [-]
It must be so frustrating to write science-fiction now with the future changing faster every day.
rinconrex 5 hours ago [-]
The math equivalent of AI code reviews piling up. I wonder where the incentives will align and the equilibrium turns out.
avd201 5 hours ago [-]
Wow, FFT faster than O(nlog(n))? I wonder if that will open the floodgates for further improvement or not. I don't understand anything about most of the fields these results touch, but I can say that this in particular is very surprising.
sashank_1509 2 hours ago [-]
It’s a meaningless improvement
bhu8 31 minutes ago [-]
The applied mathematician’s joke is that log n is bounded above by 45 or so.
philipwhiuk 4 hours ago [-]
My guess is that the constant terms are large enough it's not practically useful in most cases.
trostaft 4 hours ago [-]
Not much computational mathematics here, but I do see some of interest in the MCMC community. In particular, 139, 93, and 101 are very interesting. Also (I'm not too familiar), I remember attending some talks on the Crouzeix conjecture 325 attacks, should sharpen some rates for Krylov methods. Obviously need to read deeper, some of these don't have corresponding Lean formalizations.

Cool!

bashtoni 4 hours ago [-]
Is AI going to put Mathematicians out of a job, or is it going to create many new jobs reviewing proofs it creates?

I'm not sure it's clear right now.

lisplist 3 hours ago [-]
I'm not a very good mathematician, but I do know a fair bit about software engineering, and with AI I've been busier than ever. I probably wouldn't be so busy if AI was better at anticipating what I actually wanted rather than making guesses no human would ever make.

This is just a short term problem though. Eventually AI will get pretty good at figuring out exactly I want and it will build that from the start. The requirement of me reviewing the AI output only lasts as long as models stay bad at anticipating my needs, which I don't think will take too much longer.

AmazingEveryDay 3 hours ago [-]
I think Alan Turing would be quite intrigued by these developments, and maybe wondering what took so long.
closetheloopdev 5 hours ago [-]
Hopefully the techniques and results here will be in the training dataset for the next models, so that each new release will give us more interesting techniques and results!

It seems that OpenAI has a proof machine that keeps multiplying fruitful proofs!

Xcelerate 3 hours ago [-]
> 241. Rigidity of the Turing degrees. Every order automorphism of the Turing degrees is the identity.

Wow. This is just crazy.

patcon 3 hours ago [-]
I'm a little concerned, but I'm also glad they're publishing quick before the department of war starts making national security claims of related to using the knowledge for cryptography and/or weapons
karannb 4 hours ago [-]
I think such "process-oriented" fields shouldn't be subjected to mass automation (at least not till we as humans have sufficiently leveled up our game and understanding).

What I mean is progress in math comes from having gained a deeper understanding of the problem for subsequent attack of more problems and IMPORTANTLY applications! Right now the first one is trivially satisfied (given oai maintains some memory across models) but the second one is not! It's generating proofs faster than anyone can validate and so only the model can use these. Consequently if it just keeps doing more theory it's... not very helpful or at least not optimally helpful. This is just bragging rights for now.

More "application-oriented" research would be awesome, where it tries to achieve some desirable effect and then produces relevant theory and experiments around it. Fields like CS, Physics, Chemistry, etc. This would also benefit a wider section of the population rather than the 10 people who understand most of these proofs.

ncr100 2 hours ago [-]
This website needs a SPOILER tag.

It inspired grief in one mathematician posting here.

chickenjoseph 3 hours ago [-]
This achievement feels like a reasonable candidate for the moment where LLMs are officially "more intelligent" than any person. How can we justify moving the goalposts yet again? How can people have grown so numb to seeing advancements that they don't read this as significant? I feel like I'm standing at the foot of the exponential. I am not excited for the future, and I don't see how humans retain meaningful control over the future if we continue on this trajectory.

I have been lurking for quite some time. I made an account to post this, but I honestly don't know what to say. I would like to get off this wild ride.

lf88 3 hours ago [-]
I feel this moment is one of the last few warnings before things will get seriously out of hand. We need to stop now. Building a superintelligent AI should be considered a crime against humanity.
dyauspitr 2 hours ago [-]
Stop? What would possess you say something like that right now it’s going full steam how do you not want to know where this will go?
pixl97 1 hours ago [-]
Warning: The above poster is a misaligned AI that wants to take over the world.

Just kidding.

This said going full steam off a cliff is one of the options that has a much higher probability than I like.

slopinthebag 18 minutes ago [-]
llms are good at different things than humans, we are still collectively figuring that out. the idea that an llm is more intelligent than humans at math of all things seems fairly unsurprising.
iyyg 1 hours ago [-]
You’re struggling with nuance.

Would Einstein be successful at running apple? Nope

This seems very hard for people to understand.

It will be painful for many to realise - you should focus on doing something that positively affects the economy. Everything else is noise and many endeavours are transitory.

dualvariable 3 hours ago [-]
How many of these results are incorrect?

I doubt the answer to this is "none".

And how many of them are just exploiting some loophole that will need to be closed in the problem definition?

lf88 5 hours ago [-]
In some ways, this feels more like an ominous warning about the times to come than something to celebrate.
theoa 2 hours ago [-]
What's missing for me for each result are the following:

* Explain the result to me as if I'm a 10-year-old. * Create the infographic for this result. * Make a Khan Academy-style video to teach me this result.

yewenjie 6 hours ago [-]
A lot of these seem to be proving conjectures rather than finding counterexamples, a lot of people used that to claim that these models are not really smart/creative etc.

That copium didn't last for what, three months?

zeroonetwothree 5 hours ago [-]
I acknowledge I am impressed how quickly it moved beyond just counterexamples.
sebzim4500 5 hours ago [-]
Don't worry, more copium will be delivered. TBF so far it's still only solved the easiest of the millennium problems.
xydac 5 hours ago [-]
i wonder what it means for maths researchers, and how it aligns with how they approach math problems.
blooalien 5 hours ago [-]
> i wonder what it means for maths researchers, and how it aligns with how they approach math problems.

I guess their job now is "Idea Man" and "Error Checker"? Kinda like (some/many) "programmers" these days.

xydac 4 hours ago [-]
just wait till someone builds a idea generator model - wire it to decision (jev-like) classifier -> loop it back to researcher

>> may be thats what open ai did :)

rifty 3 hours ago [-]
As AI works steadily through the already discovered unsolved problems in mathematics, who is currently discovering new ones?
edward_d 2 hours ago [-]
This repository contains mathematical manuscripts and supporting proof artifacts produced by an internal OpenAI model?
6 hours ago [-]
williamhm 2 hours ago [-]
And this is the result of a discussion between users and the platform; it's great that they listened.
curtis-jm 6 hours ago [-]
NegativeLatency 5 hours ago [-]
Why should I care?
voidfunc 5 hours ago [-]
Because it means mathematical discovery can largely be automated away from academics. This is the beginning.
electroweak 14 minutes ago [-]
Humans propose; AI will dispose.

It's not clear from this progress that AI can formulate conjectures despite this new ability to solve them. So mathematicians still look like they have a job. Though instead of spotting far-off landmarks it's sounds more like they'll be chasing waves on a beach.

bamboozled 5 hours ago [-]
The beginning of what?
voidfunc 5 hours ago [-]
The beginning of the end of human thinking being valuable enough to justify university's existences among many others.

Were in an unprecedented time where the value of knowledge is about to be crushed.

voidhorse 2 hours ago [-]
Maybe for STEM. The humanities seems kind of safe to me since there's an element of it which cannot be divorced from human opinion and interpretation. It's not like STEM where the people pursuing the ends are largely fungible and the ends are objective (Heisenberg already believed scientific discoveries were inevitable and it didn't really matter who pursued them, someone would eventually find them)
bamboozled 5 hours ago [-]
Not sure I agree with this take, but we're going to find out either way.

Have you ever heard of an S curve? Things will develop rapidly, then equalize. If they don't, we're at the singularity and I guess the end of time as we know it.

But I guess really bad things happen, cancer, radiation poisoning, torture, people have died in really horrendous ways, and I guess dying from some horrendous AI side effects is possible too. Yay.

5 hours ago [-]
sunkeeh 5 hours ago [-]
Golden age of discovery and mass layoffs
voidfunc 5 hours ago [-]
People need to figuring out how to horde as much wealth as possible right now in the next 2-3 years. Jobs especially for knowledge workers are about to disappear.
eightysixfour 2 hours ago [-]
I’m targeting maximum debt by about 2030. I’d rather have all the stuff I want while we try and build a new version of a functioning economy than have a ton of cash saved up.
le-mark 5 hours ago [-]
I’ve been thinking this as well. I imagine there is a wealth level x such that someone can escape the coming ubi welfare state. Anything under that you are fucked.
dyauspitr 2 hours ago [-]
What do you mean by escape the UBI welfare state. Wouldn’t the UBI welfare state be the best case scenario?
chadcmulligan 4 hours ago [-]
It's funny we're possibly entering a golden age of thought, the dreams of the ancients, but we're all worried about capitalism, I think the problem is pretty obvious.
lf88 4 hours ago [-]
Maybe a golden age of thought for the machines, but possibly (I would even say likely on the current trajectory) a dark age for humanity.
chadcmulligan 1 hours ago [-]
I don't know, I get more work done now in a few hours than I used to in a day, so maybe the working week should be shrunk, that would solve things. The solution seems pretty easy - but then I'm not in the US.
brcmthrowaway 4 hours ago [-]
Any tips?
zeroonetwothree 5 hours ago [-]
Predictions of mass layoffs from AI have been about as wrong so far as predictions of AI plateauing.
claysmithr 3 hours ago [-]
Not really. 93,116 tech employees laid off due to AI in 2026.

https://layoffs.fyi/ai-layoffs/

1 hours ago [-]
connor11528 6 hours ago [-]
will this make the math for building data centers work?
sashank_1509 2 hours ago [-]
No that’s gonna happen when they take your job
dyauspitr 2 hours ago [-]
This is like that meme where death goes door-to-door. Currently, he has visited the software development and mathematics doors. I wonder what’s next.
binlog 2 hours ago [-]
Yet there are more software enginners employed today than another other point in history
rafterydj 7 hours ago [-]
I don't know, this does not feel like the message hit OpenAI where it needed to hit, if this is their primary response.
osiris970 6 hours ago [-]
You want them to stop doing math research?
6 hours ago [-]
cute_boi 1 hours ago [-]
"Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study (opens in a new window) "

I don't think this is correct solution to this problem? What about software advisory where you form similar group etc..?

I am thankful, I don't have to deal with petty academia politics....

dyauspitr 2 hours ago [-]
Does OpenAI have the lead now? Why isn’t anthropic coming up with stuff like this?
i_idiot 4 hours ago [-]
If only AI can better humans in meditation...
lokl 5 hours ago [-]
Do applied math next.
pugfugly 5 hours ago [-]
holy fucking shit
matapassiones 5 hours ago [-]
Valency has the papers up on Valency Hub
OutOfHere 4 hours ago [-]
Is this now a new home for quality AI produced works?

https://hub.valency.io/works

kevinwang 6 hours ago [-]
wow
7 hours ago [-]
jrflo 6 hours ago [-]
Glad to seem them changing their tact with the whole NS debacle. Hopefully we can all focus on the results now rather than the surrounding drama.
binlog 2 hours ago [-]
They didn't change anything lol. The news cycle has just moved on.
TeeWEE 4 hours ago [-]
This feels like a huge AI slop dump. Who validated these results. Why are they not mentioned. Human understanding is key here.

In my experience AI (frontier models) sometimes does weird stuff that needs human review. Not that’s incorrect but sometimes overly complex language or weird use of language.

Ancapistani 1 hours ago [-]
They have formal proofs included, that’s the point.
dgacmu 6 hours ago [-]
I find the claimed matrix multiply result (w<= 2.25) shocking. I hope it holds up.
k2xl 7 hours ago [-]
Can someone knowledgeable about the subject outline the most significant portions of the results?
nnoman7808 2 hours ago [-]
Roblox
blurbleblurble 2 hours ago [-]
This just looks entirely obscene from a PR perspective. It's like the cable companies networks owning the networks. At least form some partnerships to obscure the total narcissism party.
5 hours ago [-]
Catloafdev 7 hours ago [-]
This is a pretty hilarious thing to read juxtaposed with AGMAI's requests.

Basically "Here you go, have fun with this, fuck all your demands, by the way we're gonna be releasing the model stay tuned!"

aaraujo002 6 hours ago [-]
The Advisory Group states in its recommendations [1]:

"We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models."

To me, this is a take against progress so that mathematicians can keep their jobs. What would we do if, instead of math, we were talking about diseases? Are we going to keep diseases around so that doctors can keep their jobs too?

[1] https://agmai.org/general-sep29/

tchalla 6 hours ago [-]
Why did you leave out the entire quote?

> At present, some frontier AI labs are testing advanced mathematical problems on proprietary models that remain inaccessible to the broader scientific community. Our recommendations are formulated with this practical context in mind. However, ideally, they would not do so. We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models.

To me, the issue is that the models are proprietary which are only accessible to a few people in 2 digits. It's not about progress but access.

aaraujo002 6 hours ago [-]
Maybe, but it still seems like an excuse. OpenAI has a proprietary model capable of solving these problems and is willing to share the results with the mathematical community. So basically, the ask is to just not use the model and leave the problems unsolved?
adrian_m 6 hours ago [-]
The ask is to let mathematicians outside of OpenAI use it, ie. at least wait until the model is released.
strange_quark 5 hours ago [-]
I don’t think that’s sufficient. If they want to be good stewards of mathematical research and not just doing marketing, they need to at the very least tell us the datasets and any techniques they used to train this model. IMO that’s probably just as if not more valuable than a dump of un-reviewed results.

They won’t because they don’t care and the only way it got this good is something along the lines of they trained on every mathematician’s codex sessions even if they opted out because they consider the thinking traces or output and metadata fair game.

blurbleblurble 2 hours ago [-]
It's honestly shit marketing that only cultivates spite and erodes their whole brand. This is a total ego trip.
agnosticmantis 6 hours ago [-]
Which mathematicians though? Only fields medalists? Grad students? Any hobbyist wanting access?

These models are too expensive for broad access unfortunately.

Jtarii 5 hours ago [-]
ChatGPT pro is accessible to literally anyone who has a job and lives in a developed country.
Jweb_Guru 4 hours ago [-]
Literally every single one of these papers was developed with an internal model that not even most OpenAI employees have access to.
TeeWEE 4 hours ago [-]
The work is not having AI poop out these docs the work is validating them and publishing them. OpenAI doesn’t care and wanted to race to publish potential findings and let others review it which is lazy and selfish
jhrmnn 6 hours ago [-]
It all hinges on the definition of “progress”. The debate of the past month is all about questioning whether formally proving outstanding unproven theorems without human understanding constitutes progress. This is quite different from solving diseases.
andriy_koval 27 minutes ago [-]
without humans understanding who are currently losing entitlement. Regular Joe could never understand or claimed to understand high math.
mattr03 6 hours ago [-]
I don't think this is a reasonable take at all. Most work on maths has no real benefit other than to further human understanding of maths - it's more like an art. Nothing is gained from OpenAI solving all these problems but taking jobs from mathematicians. Other than advertising for OpenAI at least. It could not be more different from having AI work on disease research etc.
bravoetch 6 hours ago [-]
> Most work on maths has no real benefit other than to further human understanding of maths - it's more like an art.

It's been a while since I was reminded of this xkcd: https://xkcd.com/435/

zeroonetwothree 5 hours ago [-]
Most math isn't just proving novel famous results. Just like most of software engineering isn't writing code.
medler 6 hours ago [-]
The rest of that document makes a pretty compelling case for why this is a bad practice
esafak 6 hours ago [-]
I fear that professional mathematics will wither, and there will be nobody left to digest the AI results of the future, leaving us unable to challenge the AI.
pixl97 1 hours ago [-]
Then make 2 AI's and force them to challenge each other.
6 hours ago [-]
osiris970 6 hours ago [-]
Comical ask
bmitc 6 hours ago [-]
Advocating purely for progress and not humanitarian value is how we'll all get enslaved.
perching_aix 6 hours ago [-]
The trope you're drawing a parallel with has a (to me) compelling counter though: there being a cure for every disease wouldn't stop people from getting sick.

How this maps back to math, idk.

warkdarrior 6 hours ago [-]
The latest posts from Terry Tao on Mastodon effectively ask for an AI to explain its results to human mathematicians.

> "I believe that AI can contribute positively in all of these directions [NB: exposition, community building, new directions of study]"

https://mathstodon.xyz/@tao/117395269325940185

binlog 2 hours ago [-]
There are a dozen+ AIs available to you that can do that right now.
fph 6 hours ago [-]
...but we're not talking about diseases. Publishing an AI-generated Navier-Stokes solution does not save lives. (And, in fact, it harms some.)
Yamata 16 minutes ago [-]
It harms lives? How so?
globalnode 1 hours ago [-]
imagine your a post grad maths student looking for hard problems to solve, and theyre all solved..
tootie 6 hours ago [-]
Seemingly none are vetted and reviewed yet
mulemisterX 5 hours ago [-]
That's our job.
TeeWEE 4 hours ago [-]
No it’s OpenAI’s job. They are acting as a meat proxy
red75prime 3 hours ago [-]
"If you have nothing to say, don't post a chatbot's responses, because anyone can ask the chatbot directly if they wanted to"-principle? Well, people can't ask their chatbot directly, because it's not public.
esafak 5 hours ago [-]
Ain't nobody paying me to do that. It's kinda sad that maths is being reduced to checking the AI's work.
kozikow 5 hours ago [-]
Not just maths

In SWE as well - this is what I do most of the day

zeroonetwothree 4 hours ago [-]
Always has been
schleck8 5 hours ago [-]
Most are formalized in Lean, about 80% of what I checked
5 hours ago [-]
applicative 6 hours ago [-]
I wonder if the Lean compiler can change it's license so that a for-profit corporation can only use it if it pays, say, a few hundred billion dollars. This is the correct path.
hi__dang 5 hours ago [-]
Mathematics is solved.
nautilus12 6 hours ago [-]
Have any real mathematicians working on these problems reviewed any of these and determined if they are just gobbledegook or not?

The ones with lean proofs could still be formulated incorrectly

baggy_trough 3 hours ago [-]
Stochastic parrot truthers in shambles.
koe123 30 minutes ago [-]
Can you explain it without reaching for lofty things like consciousness? To me stochastic parrots is literally how it works given that it’s “just” the most impressive data fit we’ve ever done. Apparently generating mathematics reasoning traces + verifying them with lean works super well.
AmazingEveryDay 2 hours ago [-]
It is more of the usual though isn't it? OpenAI cribbing off of mathematicians that have used their services; deciding to put a lot of compute behind fruitful areas of endevour; getting results, then taking credit.
oh_no 1 hours ago [-]
no, they did not steal the notes of 100s of people working on these 100s of problems, be serious
baggy_trough 2 hours ago [-]
It’s wonderful.
mathisfun123 6 hours ago [-]
With so many results in so many different areas no way they even remotely spot checked well enough.

Prediction: one of these is wrong and this (publicity stunt) will backfire.

Edit: don't tell me about lean. For lean to function as a proof certificate you need to represent the theorem correctly. Again: good luck doing that across such a broad swath of problems.

jojva 6 hours ago [-]
You have not read their readme:

> Some of the unformalized results could have issues. We will endeavor to fix any such issues quickly. We are also exploring community-hosted repositories for these materials.

mathisfun123 6 hours ago [-]
i have and i'm exactly saying that if it comes to pass one of them is wrong it's going to backfire. ie yes that's my exact point/bet.
stevenhuang 5 hours ago [-]
I don't think anyone would particularly care if only one of them is wrong, if most are correct.

If they are all wrong, that's when it would backfire.

bravoetch 6 hours ago [-]
What does a backfire look like? It's ok to be wrong in the science/math world.
mathisfun123 6 hours ago [-]
of course in science/math it is but it's not okay if you're a business selling supercalifragilisticexpialidocious infallible intelligence.
bravoetch 6 hours ago [-]
Do they claim that's the case? I don't think they do.
mathisfun123 6 hours ago [-]
does company A making product B claim that the product is robust and consistent? is this a serious question?
zamadatix 5 hours ago [-]
If you were waiting for companies to sell engines that never break down you'd still be stuck pre industrial revolution while the rest of the world has been to space.

The question never if something works 100% of the time but how often it breaks and how that fits the need well. Solving one of these problems is a massive undertaking and accomplishment for the best minds, solving hundreds in a month but being wrong about 10% or something would likely not be the death knell you believe it to be.

orlp 5 hours ago [-]
It's likely that way more than just one of these is wrong. But even if it turns out 80% is wrong this is still 100+ results...
dpweb 6 hours ago [-]
[dead]
QuadrupleA 1 hours ago [-]
[dead]
Aiversee 1 hours ago [-]
[dead]
philipwhiuk 4 hours ago [-]
[dead]
philipfweiss 5 hours ago [-]
[flagged]
camdenreslink 4 hours ago [-]
What is considered a big event? Some papers published or letters sent between academics have invented entire new categories of mathematics that didn’t exist. Is it is big as calculus or Euclid’s elements, or Godel’s incompleteness theorem? Or Hilbert’s program of formalism?

I’m skeptical!

redox99 6 hours ago [-]
The stochastic parrots have predicted the next token once again.
applicative 6 hours ago [-]
Why didn't they just give mathematicians access, so they could at least understand and write up the results in publishable form?
5 hours ago [-]
stevenhuang 6 hours ago [-]
zamadatix 5 hours ago [-]
I think they meant "access to the model" rather than the results.
oh_no 1 hours ago [-]
I'm seeing a lot of this and it makes no sense, the internal model solved these but give one of the papers to Astra and Opus and I'm sure it will have no problem recreating it.
5 hours ago [-]
digitaltrees 5 hours ago [-]
Gross. After the accusations training on mathematicians conversations and unique methods, to dump this volume of unsubstantiated papers is at best tone deaf. Are they hoping humans review and validate this? The amount of free that they are coat tailing is outrageous.
computerex 5 hours ago [-]
What do you expect them to do?
digitaltrees 1 hours ago [-]
Not train on and steal user data to front run frontier research for one thing. Not take credit for and independently solve these problems and instead sponsor researchers and give them credit. They could chose to include humans and frame AI as elevating human systems. Instead they are reveling in the fact that they did this in their own instead of humans. Thats a PR choice that is short sighted and reflects a selfish mindset not deserving of leading this transition
sandworm101 5 hours ago [-]
So the million monkeys at a million typewriters have churned out 700 shakespeares, but they need me for spellcheck?
utopcell 5 hours ago [-]
Nobody needs you to do anything, not with that attitude.
mi_lk 6 hours ago [-]
Curious if Sébastien Bubeck still work at OpenAI? He came out quite dirty after Navier-Stokes drama
senderista 7 hours ago [-]
Good to see they're engaging with the mathematical community, even if they had to be publicly shamed into doing so.
sebzim4500 5 hours ago [-]
This is the opposite of what the mathematical community were asking for. I say this as someone who strongly approves of this approach.
senderista 5 hours ago [-]
Yeah I'm not sure they met them even halfway.
Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact
Rendered at 05:09:17 GMT+0000 (Coordinated Universal Time) with Vercel.