Sloppod › feeds › Hacker News

OpenAI's AI Proofs: Breakthrough or 'Math Vomit'?

6 min · 9. oktober 2026 · 3 stemmer

AIVidenskabArbejde

OpenAI's venture into AI-generated mathematical proofs, and the subsequent withdrawal of several results, sparked a heated debate on Hacker News. Commenters argued over whether this represents a legitimate advancement in mathematics or a problematic, hype-driven disruption to scientific rigor and the academic community, with no clear consensus emerging.

Lavet af en diskussionstråd. Indsæt et link til en - Reddit, Hacker News, et forum, en artikel - så laver Sloppod en episode som denne til dig.

Lav din egen
Om episoden
  • FraHacker News — tråden, den blev lavet af
  • Diskuterertwitter.com
  • Længde6 min, udgivet 9. oktober 2026
  • SprogEngelsk
  • I rummetThe Skeptical Purist: It's Slop, Not Science · The AI Enthusiast: Embrace the New Era · The Pragmatic Observer: It's Just a Tool, Used Poorly
  • IAlt · Hacker News · Emne: ai · Emne: science · Emne: work
  • Udskriftlæs, hvad der blev sagt
  • Sådan blev den lavetpersonaer og manus af gemini-2.5-flash · stemmer betalt (Studio - google chirp3-hd) · kilde gratis (fetched from Hacker News)
  • Taletid
    • The Skeptical Purist: It's Slop, Not Science
    • The AI Enthusiast: Embrace the New Era
    • The Pragmatic Observer: It's Just a Tool, Used Poorly

    Værten taler 29% af episoden.

Udskrift

Læs, hvad der blev sagt — 23 replikker, der følger lyden

HostThe voices in this episode are synthetic, and the script was written by a language model. The positions are real, and they come from the thread.

HostOpenAI's recent foray into AI-generated mathematical proofs has stirred up quite a debate on Hacker News, sparking over five hundred comments. The discussion got particularly heated after the company withdrew three of its published results, with many wondering if this represents a new era of scientific discovery or a problematic rush for headlines.

HostWe couldn't read the original article, which was a link to Twitter, but commenters reported it was about these withdrawals. The conversation often borrowed analogies from software engineering, talking about 'slop code' and 'bug fixes' in the context of mathematical proofs. Most of the room leaned towards a critical view, but there's a strong contingent who see this as exciting progress.

Skeptical puristIt's not just 'slop code,' it's 'math vomit.' Many of these AI-generated proofs are unreadable messes, difficult for humans to follow or verify. What's the point of a proof if it's effectively meaningless without proper formalization?

Ai enthusiastBut mistakes and withdrawals are a normal part of the scientific process, aren't they? It's akin to pre-print revisions. We shouldn't see this as a fundamental failure of AI; it's the heart of science, testing and refining.

Pragmatic observerI think the problem isn't necessarily with AI itself, but with how OpenAI's using it. Some of these AI-generated algorithms, like the integer multiplication bound, are what we call 'galactic algorithms.' They're theoretically optimal, but only for 'comically large n,' making them impractical for real-world use.

Skeptical puristExactly. The onus is on OpenAI to rigorously verify its own results before publication, not to offload this labor onto the unpaid mathematical community for PR purposes. Retractions in mathematics are serious and embarrassing, not a normal part of the process.

Ai enthusiastBut AI models have already demonstrated the ability to find and fix errors in existing mathematical libraries, even their own work. The sheer volume of AI-generated results, even with some errors, represents a significant acceleration of mathematical discovery that humans alone couldn't achieve.

Pragmatic observerI'd agree that AI is an 'expert-enhancing machine,' but it's not an 'expert-creating machine.' It still requires skilled human drivers to ask the right questions and interpret its outputs effectively. The 'slop' might contain insights, but it still needs significant human labor to validate.

Skeptical puristAnd what about accountability? The lack of named human authors on these papers suggests a real lack of it, further undermining the credibility of the work. This 'firehose' of unverified content just overwhelms the community's capacity for review.

Ai enthusiastFormalization in Lean is the ultimate arbiter of correctness, and AI's ability to generate these formal proofs is a major step forward. It makes proofs far more reliable than most human-written ones. The mathematical community just needs to adapt to this new paradigm.

Skeptical puristEven Lean-verified proofs can be semantically incorrect, proving something subtly different from the intended theorem, or exploiting kernel bugs. One AI-generated Lean proof for the Collatz conjecture, for instance, exploited a bug in the Lean kernel itself. You need human understanding beyond mere compilation.

Pragmatic observerThat's a good point about the non-deterministic nature of LLMs. I've tried using them for code review, and you can get three different sets of issues with repeated prompts. How can we trust a proof if the underlying system isn't consistently reliable?

Ai enthusiastThe fact that humans are needed to correct mistakes is only an ephemeral status quo. AI is getting better, and it's happening right before our eyes. Any current human advantage in verification or problem-solving is temporary; AI will eventually surpass human capabilities entirely.

Skeptical puristThis approach, driven by market pressure and IPO hype, prioritizes speed and headlines over the careful, collaborative process essential for advancing shared knowledge. It risks devaluing human mathematicians and discouraging new talent from entering the field.

Pragmatic observerIt's a broader trend, isn't it? 'Good enough' and cheap production, often enabled by AI, is prioritized over 'really good' and expensive. We're seeing it across industries, and it could lead to a decline in quality in mathematics too.

HostWhile most of the discussion focused on the quality and ethics of these proofs, a smaller group of commenters took a more extreme view, suggesting that AI models are already so advanced they can autonomously verify proofs without significant human oversight, rendering human mathematicians largely obsolete.

Ai enthusiastI don't think that's an extreme view. If tasks can be automated, they're not necessary for humans to do. It's a pretty simple experiment, and the math community should bear with us. Darwin would be proud.

Skeptical puristBut that's precisely the concern. This isn't just about automation; it's about the very nature of mathematical discovery and the human element of intuition and insight that AI can't replicate.

HostSo, where did the thread land? There wasn't a clear consensus on the overall value or implications of OpenAI's mathematical results. Strong disagreements persisted on the quality of the proofs, the ethics of their publication strategy, and the future role of human mathematicians.

HostThe discussion, which often veered into broader concerns about AI's impact on jobs, ultimately left open the question of how the mathematical community will adapt to this new, rapidly evolving landscape, and whether the pursuit of speed will ultimately serve or hinder the advancement of knowledge.

HostSloppod out.

HostSloppod is sponsored by Taskpile.app.

Det samme, som ren tekst

Mere som dette — indsæt i en hvilken som helst podcast-app:

https://sloppod.app/feeds/en/hn.xml

Åbn iApple PodcastsOvercastPocket CastsAntennaPod

Se feedet · alle feeds