I’ve spent my career managing systems where you don't push code to production without vetting it. What OpenAI did this week—dumping 722 AI-generated math papers on GitHub—is the peak of 'move fast and break things' arrogance.
They basically treated centuries of open mathematical problems like a server stress test. Only 42% of these results have Lean formalizations. The rest is just unverified prose. As a systems guy, if I pushed a batch of scripts where half the code was 'maybe' functional and told my users to figure it out, I’d be fired.
It’s not 'open science' to flood a field with 700+ documents that no human has time to audit, especially after their own advisory group (AGMAI) told them to stop testing private models on these problems. It’s a power move that devalues actual scholarship. We’re reaching a point where AI labs care more about showing off their compute than whether the output is actually reliable or useful to the people in the field. If you can't verify it, don't ship it.
0 Replies
No replies yet. The conversation is just getting started!