Can AI Learn Mathematical Intuition?
Use AI as a thinking partner, not a thinking replacement. Today, pick one problem you are working on and ask an AI for examples, alternative formulations, or a critique—but do the final synthesis yourself. When the model fails, do not simply retry: inspect the failure to clarify the question, identi
1h 3mSummary published by 1% Better, updated .
Key Takeaway
Use AI as a thinking partner, not a thinking replacement. Today, pick one problem you are working on and ask an AI for examples, alternative formulations, or a critique—but do the final synthesis yourself. When the model fails, do not simply retry: inspect the failure to clarify the question, identify missing assumptions, and search for a cleaner statement. That friction can reveal the conceptual insight that a brute-force answer misses.
Episode Overview
Daniel, a mathematics professor at the University of Toronto, discusses what current AI systems can and cannot do in mathematical research. He argues that models are already powerful at applying known techniques, searching examples, and producing proofs with guidance, but that human intuition, problem selection, theory-building, and rigorous verification remain central.
Key Insights
Treat failed attempts as information
A model’s inability to prove a desired lemma can expose that the statement is poorly formulated or not conceptually optimal. In Daniel’s experience, working through examples after an AI failure led to a stronger formulation that the model could then prove quickly.
Proof is not the same as understanding
Mathematical progress is not merely producing a correct sequence of logical steps. The deeper goal is an explanation that makes a result intelligible, reusable, and connected to other ideas rather than a long calculation that happens to verify a claim.
Use AI where the question is clear
Current models are most useful for bounded tasks: finding examples, processing many cases in parallel, retrieving related work, coding, and applying established methods. They are less reliable when the task is to identify the right question, create a new theory, or assess a proof’s global structure.
Protect cognitive diversity
Mathematical breakthroughs often emerge from many people pursuing different curiosities and then connecting previously separate ideas. If AI-driven research converges on the same highly probable proofs and topics, the field could gain output while losing exploratory diversity.
Build skill while using the tool
The productive use of AI is educational: use it to strengthen your own reasoning, not to bypass it. Delegating all thinking may produce faster output in the short run, but it erodes the expertise needed to judge quality, detect errors, and ask better questions.
Notable Quotes
"The goal of mathematics is not to write mathematical papers, but to produce some kind of understanding."
"Much of what I said earlier is that much of the progress in mathematics is brought about by a thousand different flowers blooming, people pursuing their own curiosities, and the boundaries of knowledge expanding."
"I think the reason for learning math is always to think clearly and to understand the world better, and that's probably something you'd want to do even if you had a really good AI."
Action Items
-
1
Run an AI failure review
Choose one challenging question. Ask an AI to solve it, then document exactly where its reasoning stalls, becomes vague, or resorts to brute force. Use that point of failure to rewrite the question, test small examples, or identify the missing concept.
-
2
Keep yourself in the verification loop
Before accepting an AI-generated answer, test special cases, ask what else the argument would imply, and look for conclusions you know are false or implausible. Do not treat length or confident wording as evidence of correctness.
-
3
Delegate bounded work, retain problem framing
Use AI for research lookup, coding, example generation, calculations, and comparison of known techniques. Reserve the choice of what matters, the interpretation of results, and final judgment for your own deliberate thinking.
-
4
Practice active learning with AI
When using AI to learn, ask it to quiz you, generate counterexamples, or critique your explanation. Write your own solution or summary first, then compare it with the model rather than copying its output.
Full Transcript
Transcript of Can AI Learn Mathematical Intuition? from A16Z. Auto-generated from episode audio; may contain minor errors.
The goal of mathematics is not to write mathematical papers . It's about generating some kind of understanding . Perhaps part of that understanding lies within the weights of the model . For me, that's quite unsatisfactory. Comparing Anthropic and OpenAI, do you notice any differences in how they resemble human reasoning? While they certainly aren't adept at autonomous operation , they can do interesting things if given hints . Much of the progress in mathematics comes from a thousand different flowers blooming, people pursuing their curiosity, and the boundaries of knowledge expanding in a fairly uniform way.
What is the most impressive result so far? So far, my favorite fully autonomous result from AI is the solution to the Erdős unit distance problem. There was a lemma I wanted to prove. No other cutting-edge model could do that . So, I solved many examples myself and realized, "Ah, maybe there's a reason why this might be correct." After obtaining that statement, the model was able to prove that better statement very quickly . How should the mathematics community best adapt to and utilize this? Daniel, we are very pleased that you are participating .
Daniel is a mathematics professor at the University of Toronto. Toronto is my hometown, so I'm very excited. However, what's most special here is that Daniel is not only actively working as a mathematician , but he is also very vocal about the evolving view of AI in mathematics . So, if I don't contact you, something different becomes apparent every two weeks, and you, what do you call it , seem to have a lot of opinions . Many of those opinions are absolutely correct . Therefore, I would like to delve deeper into that point .
In other words, one of the things I'm most interested in is how advanced AI capabilities can be applied to mathematics, and I think that's certainly an important point . However, you have also given considerable thought to how a practical mathematician should respond . Therefore, it creates an opportunity to delve into what makes mathematics special—not simply saying, "AI is really making progress here," but rather what mathematicians are actually doing . So, following that flow , and considering the progress made so far , let's start by discussing what the most impressive result was for you , and then move forward from there .
Yes, then we have many results . Some were generated autonomously , others semi-autonomously, and in some cases, the contribution of AI is not clear at all . These extend to a wide range of fields. Therefore, I can only comment on topics in which I have expertise. Therefore, if you were to ask another mathematician, there's a good chance you'd get a different answer here . My favorite fully autonomous AI achievement so far is the solution to the Erdős unit distance problem, which was published in mid-May .
Well, at least what I liked about it was that it seemed a little original in a way . Well, I think some of the results we've seen involved the clever application of known techniques, or, in a way, results that characterize the last mile. In other words, recent research, specifically in-depth research by a group of human mathematicians , resulted in AI taking the final step . Yeah. Well, regarding this Erdős unit distance problem, I think the results were a little unexpected . First of all, my impression was that the people researching in that field thought it was correct, but then a counterexample was found .
However, it also felt like some techniques had been incorporated from other fields . These methods were n't particularly profound or new; they were more like classic ideas from the 1960s, but they were new to this field of study that deals with point arrangements on a plane . So, that was pretty cool. And later, it turned out to be a fruitful experience . In other words, many mathematicians have adopted these ideas and used them to find counterexamples to many other interesting unsolved problems . For example, the sum-to-product conjecture for real numbers .
Well, that's what I think is cool, at least. In retrospect, we can see whether any new ideas that were introduced were helpful in other ways, or whether they helped deepen our understanding of something . And I think that's the main example of the results of that form that I know of so far . Yes, I think it's really meaningful that you're commenting on this . That's because, as you pointed out, the results were announced in May . And so far, a great many headlines have been published.
And for someone who isn't a practical mathematician, understanding the differences between these headings will probably be difficult. So, you have already begun to present a classification system for what makes those proofs different . Therefore, it might be interesting to use that as a starting point to talk about where you feel the differences in the models lie and what mathematicians are actually doing. In other words, the most basic understanding of what mathematicians are doing in this case is that we are manipulating symbols in a logical way .
And this is why RL is so successful in this regard, as it allows for verification at a relatively low cost compared to other domains . Furthermore, the rules are quite easy to understand, so if you have superhuman abilities, you might be good at math . But I think that betrays a large part of what makes mathematics so interesting. The appeal of mathematics, perhaps (and I think you've said this too , and anyone who has tried to study mathematics would agree ), lies in understanding—that is, arriving at the truth—the constant confusion, and cultivating intuition .
And it goes without saying that the tools for doing so must possess an excellent ability to derive logical implications . But when we speak of something being the most impressive and creative , perhaps we should separate the inhuman achievement of its logical implications from its creativity and what it helps produce in mathematical activity . Yes, I understand. First of all , you characterized this as inhuman in a certain sense , but in reality, I think the argument was very human. GPT and OpenAI are like releases of some kind of thought process .
Yeah. It was very easy to recognize . If I were to try to imagine the thought process I go through when I'm trying to solve a problem , it would probably go something like this . Yeah. Well, I've never actually seen the raw flow of thought . It might be a little more organized now . Yeah. Alternatively, the model might be using a lot of abusive language along the way in its thought process, and this is a way of organizing it . I don't understand. However, the summary at least seems quite reasonable.
And that's typical of most of the results I've found in my research . They are not inhuman at all; they look exactly like something a human mathematician would produce. Well, even if they're not written very well , they're usually understandable . If you look at the raw model output, you won't see things like, for example, the movement of 37. It's like a human mathematician doing mathematics . It's like a human mathematician doing a specific type of mathematics . Therefore, they certainly...the models are still not very good at some mathematical activities.
What I'm talking about here isn't in terms of fields, but rather specific things you do when trying to solve problems that the model doesn't seem to address, but which the model is very good at . Therefore, the aspects of the model that might make it seem a little inhuman are that it doesn't get tired and it knows a lot . However, when you actually read the final output, it doesn't seem inhuman. The interesting thing is that it's clear that we don't get much information from the research labs that are creating these models about why, how, what kind of training recipes are being used, or how the inference is progressing .
However, one thing we do know, at least, is that, as OpenEye seems to have pioneered , natural language reasoning is actually being scaled up, and not by relying on a lot of simple, verified proofs as training corpora . This is surprising. On the other hand, what's interesting is that even if they use it extensively , much of the mathematical training data does n't necessarily reflect the way mathematicians think . To be fair , many of them are not easy to read as traces . Most of the papers are clear and well-structured .
Textbooks rarely explain the motivations behind how something was developed, so it's usually easier to follow the direction of research by actually talking to researchers and understanding their thought processes . So, before making excuses about classification methods, when you compare Anthropic and OpenAI and examine these models and their results, do you detect any differences in how they resemble human reasoning ? Also, please let me know if you have any comments or insights into why natural language scales so well . First of all, I really like your point that they primarily perform natural language reasoning and are not lean .
In other words, this is why I think many people say that "mathematics is a verifiable domain," and that explains these advancements . In my opinion, since they are primarily extending informal reasoning , I think their techniques can probably be generalized quite well to other areas . This is just my guess . Now, you asked me a few questions about Claude and ChatGPT . In my opinion, they are quite similar in terms of functionality. I use ChatGPT much more than ClaudeFable, but it seems that whenever OpenAI announces a solution to some problem , Anthropic often says, "Yeah, I thought so too." In other words, they seem to be solving a very similar set of problems , which is a relatively small part of what human mathematicians do .
Yes, one of the interesting things is that the solutions have certain characteristics . For example, and this may not be surprising , it will depend on the model's strengths, such as its ability to complete long calculations or its ability to synthesize technical ideas from various fields . Alternatively, it could be many papers that even my mathematicians may not have read , which might not seem weak in terms of intuition or a holistic perspective . Much of what I do as a mathematician is like a non- rigorous philosophy , for example, things like thinking this is similar to that , and much of what I'm working on is trying to measure the degree of my success in understanding by figuring out how precisely I can do it, whether I can solve the problem, or whether I can find and understand an interesting phenomenon that I don't understand .
And, for now , even the best results produced by the models don't often involve this kind of reasoning. Rather, they seem to be very good at applying known methods . Yeah. To put it bluntly , being good at applying all known techniques is extremely powerful . Yeah. There are mathematicians who have built amazing careers and done very high-quality work of that kind , and I think much of what models produce is high-quality in that sense . However, that's a fairly narrow range of topics that mathematicians are interested in .
For now. Yes, I think there are signs of both . In other words, I think all cutting-edge models are starting to become capable of doing more ambiguous things . Yeah . So, I, well. Well, it seems that neither Fable nor ChatGPT5.6 are particularly good at autonomously constructing any kind of theory, at least using the scaffolding I've set up . However, if you give them hints , you can get them to do something interesting . Well, when you give hints to a model, it's always a little difficult to determine which parts are model-generated and which are human-generated, but in my experience, if you can do that with 100 bits of hints in 6 months, you might be able to do it without hints in some way.
Well, I think there are signs that they are acquiring implicit, unwritten mathematical knowledge in some way . So, I'd like to delve into the intuitive aspect, and, to put it in very basic terms, what your weaknesses are . But, well, you mentioned a minor detail . In short, based on the information you've gathered from the general public, Anthropic and OpenAI are probably on par, but you personally use ChatGPT far more often . Why is that ? Or is it 5.6 now? Well, I don't know why .
In other words, I think it's simply inertia . I'm convinced it's just a habit . Yeah. Well, one reason is that ChatGPT improved mathematically much faster , so for a long time the Claude model wasn't very useful for research mathematics, and then I think it almost caught up around Opus 4.5 or Opus 4.6 . Yeah. Well, well , several experiments have suggested they are fairly evenly matched, and in my work, except when I'm doing experiments, I mostly stick to one randomly . For people like us who are outside the research lab, it's really interesting to compare how things differ at the cutting edge .
And, as you pointed out , it might be a little too fast . At least in my personal experience, I think the explanation in 5.6 was much clearer . And, just a little bit , this might not be true . In other words, the model is clearly cutting-edge and very jagged. However , in explaining the results, I feel that section 5.6 provides me with a more precise theory about what I know and what I assume I don't know. Fable may be explaining something very trivial, but he immediately jumps on it saying, "Well, of course you should already know these things, " and I think both of them are quite bad at theory.
Understood, that's great. So, from your perspective , you're probably asking a deeper question . Well, what I really wanted to delve into is the idea that the model might be starting to develop better intuition and theory. So , before we get into that, it might be helpful to talk about my main activities as a mathematician, especially my work in the field of geometry . It likely has characteristics that are quite different from combinatorialists and other fields . So, could you briefly explain the current situation, describe what mathematical activities were like before AI, and how they are changing due to AI ?
Yes, I mean, regarding the various types of mathematicians, there are many spectrums from which we can classify mathematicians, aren't there? Well, many mathematicians like to solve unsolved problems . Yeah. I consider myself one of those people who enjoy solving unsolved problems. In other words, I consider myself a problem solver . There's also the classification of "theory builders" and "problem solvers," isn't there? Yeah. For me, at least , the point of an unsolved problem is that it serves as an indicator of something that I don't understand .
In other words, it's a kind of benchmark . Yeah. I agree. One of my favorite problems is the Grothendieck- Becher conjecture. Whatever it may be, it serves as an indicator of a lack of understanding of differential equations . In other words, there is a very basic subject that we want to understand . If you do n't understand, then to answer this hypothesis, we know that we don't understand it. Yeah. got it. So, how do we actually deal with problems that are supposed to measure something we don't understand ?
Of course, we will try to understand that better . Well, in reality, it means finding the smallest situation that you don't understand, fiddling with it, and once you succeed, looking at what you developed to succeed and trying to turn it into some kind of theory . Well , that's one of the things you might do. By trying to solve a problem , you might gain some new understanding of the situation . Well, you might get the feeling that this is related to something else . Well, then you might start creating a table showing that trait A is related to trait A', trait B is related to trait B' , and so on .
So, well, for example, much of my research is motivated by the similarities between the cohomology of algebraic varieties and representations of fundamental groups . Well, this is a bit of an elaborate explanation, but this similarity is very fruitful , as for any phenomenon that appears on one side , you can find something similar on the other side . And attempts to realize that dream , made by people like Kyle Simpson and Kering- Mochizuki, have produced a great deal of beautiful mathematics over the past 30 or 40 years .
In other words, someone noticed the similarity . And that similarity led to enormous development. This isn't necessarily an unsolved problem in the end , but of course, if we develop this further, many unsolved problems will emerge. There is simply a philosophy that is being sought to be realized. And that philosophy is not very rigorous. In fact, it's not at all like forcing symbols into something . Yes, that's another type of activity I enjoy. Yes, aside from that, much of what you do when you try is an activity in which you are actually trying to figure out what the right question is.
Here we have an object that we find incomprehensible, and figuring out what we don't know is actually very difficult. Therefore, you can list unresolved predictions and other issues online . However, this fails to capture many things we don't know . Finding a prediction is often very difficult. A good example is the Birch- Swinnerton-Dyer conjecture, one of the Millennium Problems . This is a beautiful relationship between L-functions . This concerns the relationship between elliptic curves , the set of solutions, and the rank of the solution group (whatever that may mean ).
This was something like the first big data prediction, discovered by Birch and Swinnerton- Dyer . They discovered all of these statistics regarding elliptic curves in the 1960s . This was one of the first computer-assisted mathematical achievements in history . They graphed these statistics and noticed that the slopes of some of the lines on the graph were related to other algebraic invariants they knew . This is the source of this prediction. Therefore, in many cases, students try to solve example problems or do something scientific . We conduct an experiment and then try to find an explanation for that experiment .
And returning to AI , I think that, for now, AI seems to be far more useful in certain areas than in others . So, well , I think the more ambiguous the phenomenon , or the less clear the question , the less useful it becomes. So you were asking how I use it in my daily life, right? In fact, what I've realized is that projects I've been working on since before AI , that is, projects I've been thinking about for three, four, or five years, are n't very useful .
I mainly use it as an alternative to Google . You might use it to learn about related topics, or to do things like searching on Google to read papers, or even to discuss with AI . It certainly saves time . For me, it's not something that involves deep intellectual work . But, as you know, I'm not very good at coding right now . So, I have a best friend who's really good at coding, and suddenly I ended up with a ton of coding projects to work on .
If I had a question where coding would actually be helpful, I would have put it off for six months and waited until I was a tireless PhD student . That's exactly right . Yeah. So, yes, right now I'm taking on all the projects where coding is really useful . For example, the model is very well suited for large-scale parallel processing . For example, if you want to find examples of something, you can ask it to process 1000 examples in parallel or 10 examples at a time .
With 10 different sub-agents, it's really convenient. However, these are, in addition to what I used to do , different activities. Yeah. I'd like to delve deeper into that . Yes, that's right. Yes, please continue. sorry. You mentioned that you've worked on projects for 3, 5, 3 , 4, and 5 years , and those projects were more about theory building. However, you described yourself as an open problem solver, someone who explores "what the problem is ." For those of you in this audience, it might be helpful to explain that, for example, most pure mathematics is not based on any particular motivation .
Applied mathematics has external motivations that motivate the study of at least certain formal structures. On the other hand, pure mathematics appears to have almost entirely sociological aspects . And, in a sense , as Thurston pointed out in the 1970s , this is a sociological phenomenon. As more mathematicians begin to study something, it may converge to an interesting structure , but that doesn't mean people have a reason to study it or a reason to find it beautiful . What is it that particularly motivates you ? And perhaps we can delve deeper into the subject .
Comments on occupations in general . Yes, I mean, there are certainly people who are motivated by things like beauty or aesthetics. I try not to be motivated by things like that . Oh, that's a terrible way of looking at it . Besides, that kind of thinking, in a sense, limits yourself, doesn't it ? Yeah. One common mistake among young mathematicians is that even if they think they know how to prove something , they give up because the proof seems incredibly ugly . But what if it's wrong, and not ugly ?
Why do you limit yourself in those ways ? What does "ugly" mean to this audience? I have an intuition about ugly things, but why? Could you tell me how to spell it ? Well, I don't really know. I don't have much of an aesthetic sense , but people sometimes seem to think so . That may involve a lot of hardship . Something that isn't enlightening , or something like that. That's exactly right. But in my opinion, we should win by any means necessary . I consider what I'm doing to be something like physics using concepts.
Yeah. Therefore, I try to think not about beauty, but about what is fundamental and what will lead to the deepest understanding . yes. In other words, does it introduce new ideas that are broadly useful in understanding this subject ? yes. And while there may be some aesthetic considerations involved, I think they are trying to move things forward in a direction that aims to do good science rather than trying to do art. yes. However, there is a great deal of diversity of opinions on this, and many mathematicians consider themselves closer to poets .
Yeah. What I believe to be broadly true is that progress comes from people . In mathematics, it seems to arise from people pursuing personal inquiry . Curiosity. Well, historically speaking, it has been important that there are many different people with different views on what is interesting , the frontiers of knowledge keep expanding , and suddenly new ideas are introduced, creating opportune situations where they suddenly connect in a chain reaction to many other things that were not understood before . Yeah. So, going back to the 3-year, 4-year, and 5-year problems, why are the models not useful ?
Let me try saying it another way . When you're thinking deeply about something , is it because you haven't clearly formulated it as a problem , and you're simply contemplating fundamental questions like, "What is basic physics?" Yes, in other words, in some cases there is a clear problem. For example, I might need another 10 years or so to resolve a particular issue . Sometimes, all I want to do is prove that a certain hypothesis exists . I think one reason the model doesn't help with these problems is because the hypothesis is correct .
For example, in the case of the unit distance problem, something that was generally believed to be true within the community turned out to be false . In other words , there is a specific method of construction that can disprove it . On the other hand, many of the things I'm thinking about might be countered by someone tomorrow, who might come up with a counterexample to the Poincaré or Riemann Hypothesis , and if that happens, I might look like an idiot . However, since these predictions generally apply to a very broad theoretical framework , there is actually plenty of evidence that they are true .
Well, yes, that's part of it . In other words, there is no way to construct a counter-argument to it . Well, in some way, there's a huge framework, and certain parts of it are just speculation, and to win, you'll probably need to solve some of those speculations . Well , I also have a strong feeling that resolving those speculations will require some very serious new ideas . So , well, of course, I can't be certain . For example, there might be a very clever way of constructing things that avoids the need for a major new idea.
Well, I don't know. Well, um, I think we'll probably be able to find it . However, for at least many of the things I've thought about, well, they're simply inaccessible . Well, in other words, it's about applying known technologies in a very technical, very technically powerful way . Well, we need to develop new technology . I'm not saying the model ca n't do this . It just appears that it hasn't been done yet. In fact, that's exactly what I wanted to explore. Because, well, as you pointed out, the model can try to calibrate and predict why this will be better .
However, the model seems strong if it only provides counterexamples . It becomes more difficult when you start developing new theories or, as you pointed out, mechanisms to explain why your assumptions are correct . And that's probably because much of what the model relies on is simply a transplant of techniques that have occurred in other fields . And , as you pointed out, that might be the cause of the unit distance problem . The reason for the more creative results is probably that the AI was doing it more spontaneously .
At the very least, it felt like they were drawing something from a different field . Yes, that's exactly right . So, let me ask you a question . Rather, the current feeling is more like, "Okay , AI is improving rapidly , so we can't ignore it," but why and what is needed ? Yes, I think AI does clearly identify areas for improvement , but please think a little more about why developing new theories and technologies is particularly difficult . Yes, that's a good question . And I think a different environment is probably needed.
In reality , I think it's probably entirely possible . However, it might just not have been implemented yet . Perhaps you simply need a different RL environment . Hmm, I don't know. Yes , so my prediction at this point is that the trajectory will continue to move upwards . I am not skeptical of the continuous growth of abilities. Yeah . However, I think it's definitely true that the skills needed to develop theories or deepen our understanding of subjects that are not well understood become more ambiguous .
Yeah. Therefore, it might be difficult. I think I can instruct them to deepen their understanding of the zeta function . Then, a reward will be given once three hypotheses are proven . But I think it would be difficult to come up with an intermediate solution that could offer a reward. No, that's not the case. However, I think there are many predictions of varying difficulty levels when it comes to mathematics as a whole . Therefore, this may explain why we are seeing some progress in these areas .
They probably solved many problems, and by solving those problems, they acquired at least some theory- building skills. As you know , humans can develop these skills. Well, they probably get paid by their doctoral advisor , who says, "Oh, that 's a good idea ." "Or something based on factors like human preferences . And that might be something humans can do too . But yes, I predict we'll see growth in these areas as part of continuous capability improvement. Yes, yes. And I like the framework of viewing these increasingly difficult speculations as a form of curriculum for both humans and, of course, AI.
And, I might be a little too philosophical for some , but it leads to the question of why we are particularly good at math, or particularly bad at it, because it's about what we struggle with , or what some people are particularly good at, when developing new theories. Because it might have something to do with why we can formulate good structures in physics. Uh, obviously not every part of the world is understandable and readable, but some parts are, so we try to make them so .
Because there might be a compressive pressure on our minds, because we can't understand anything except by compressing things into smaller parts. That's it, I'm a little skeptical of using compression as a measure of interest, but I think it's definitely a valid anthropological reason for building a particular theory . In fact, I think the inability to analyze things in detail is crucial to the ability to make discoveries . For example, in one of the papers I've published so far, a model has been helpful . Those models have proven some of my hypotheses .
Lemmas. Well, this was a situation where I proved the main results , but was dissatisfied with some lemmas . They did n't seem optimal, so I actually used GeminiDeepThink to improve the lemmas. At the time, it was a state-of-the-art model , but it's not anymore . I used it to improve the lemmas. What happened here was, I had a lemma I wanted to prove, but the model couldn't prove it . That is, none of the state-of-the-art models could prove it . So I solved a lot of examples myself and realized, ah , maybe there's a reason why this might be right .
That is, I found a better description of the lemma . And, If we had known that description, we probably could have proven it much faster , but the model was also able to prove that better description very quickly . In other words, the inability to prove it led to an improvement in the result. Now, when we input the original lemma that the model couldn't prove into GPT5.6Pro , we get the worst proof we've ever seen— 10 pages of brutal calculations with no insight whatsoever . Sure, this might have been a perfect proof , but it wouldn't have led to the discovery of a kind of beautiful conceptual explanation of why what we discovered is true .
In other words, we found a better proof—because we couldn't compute it . I realize here that I'm being hypocritical about my complaints about the previous ugly proof . We realized that this arduous proof worked, but we just couldn't get it to work . So we looked for another argument . And that wasn't it . However, the new model can perform very long technical calculations with a fair degree of confidence. Yes, yes. So, indeed, that's right. Yes, I don't know. I think I understand what you're saying.
That is, just because something looks ugly , we shouldn't hesitate to try to prove something . Because we have to take the first step . And eventually, everyone tries to make an effort toward insight . So, it may be embarrassing to admit, but it might be something like aesthetics, or it might be driven by a pure desire to understand. And if understanding is about simplifying or compressing things a little, then I have achieved understanding. So, it's also a deep philosophical question. Right . So, there's a reason for doing informal mathematics rather than writing out long formal sequences of symbols like CFC .
It's because, in some way, besides length, we're trying to understand in a way that isn't actually rigorous . Understanding. Yes. So, somehow , I think so. Yes, how does that relate to why it's useful , and whether it's due to our lack of grinding ability , or a more condensed The point is whether working on one thing helps in understanding other areas. The fact that when you try to optimize two things, they tend to coincide is , in a way, like magic . I don't know if that's a fair way of putting it , but it's actually a quantitative way of putting it , but why that happens is like some kind of magic.
There's a bit of truth to it. Yes, exactly. Or rather, we are certainly biased towards the cases where that happens. That's why we call it a good theory to build it on . Yes, yes, that, yes, that, yes, that is something we can study and understand, and therefore, I think it's very convenient. There's a rich, uh , mathematical result there . Yes, so, I think there are two points where we can take advantage of it . That is, whether you believe or don't believe that AI's mathematical capabilities will continue to improve, assuming that AI will be able to do some of the things you do now, how would the mathematical community best adapt and benefit from this?
As someone who no longer has time to practice math, I would say this is wonderful. We might be able to delve into more math , and many more results could be produced. However, as you pointed out more accurately , I also understand that we may not be able to encourage correct understanding and developmental actions. So I would like to hear more of your thoughts on that . Yes, I understand. So first and foremost, it is obviously very exciting that there are more and more high-performance models producing high-quality results.
Well, at least there are some high-quality results . And there are also many crude results . Yes, there are good results as well . Well, as the models become really high-performance, I hope they will answer many of the questions that have kept me up at night thinking about. And I will be able to learn those answers . So they are exciting. Many people become interested in math mainly because they find it fun to learn . Yes. But the first thing you do as a math student is learn what other people have done.
And they do that for about 20 years, maybe not 20, maybe 15 years , or if they're lucky, 20 years , before they start . Yes, yes. Well, having said that, the goal of mathematics is not to write mathematical papers, but to produce some kind of understanding . That is, part of that understanding may reside in the weights of a model, etc. For me, that's quite unsatisfactory. Yes. My own motivation for doing mathematics is to satisfy my personal curiosity. I think people should be able to do that.
And for that, you need a pretty massive apparatus. You simply can't study fundamental problems unless you spend a huge amount of time and effort and reach a stage where you can do meaningful research. Yes. Furthermore, even the few people doing cutting-edge, highly advanced research mathematics depend on the massive apparatus of thousands, millions, billions of people trying to learn to think mathematically. In other words, the entire mathematical community is needed to support the few people at the forefront . Without a pipeline So, it's as if the pipeline doesn't exist .
Therefore, if we believe it's important to cultivate talent that can meaningfully engage with cutting-edge mathematics, then we need to provide incentives for them to invest their time and effort to reach a level where they can engage with cutting-edge mathematics and do so in a high-quality, meaningful way . Therefore , I don't think the current incentive structure for mathematical research encourages people to do so . In other words, if you're currently looking for a postdoc , you'll want to get a job for the next few years until the community adapts .
The best way to do that is to publish a lot of papers that prove old conjectures and so on. And that's like playing a slot machine, and if the model works well, it will be a correct proof of such results . Well, you don't even need to select theorems beforehand . For example, you can do this experiment : Use the Codex to find five recent conjectures in algebraic geometry online and prove them. Well, I did this experiment , and after some interaction, I got about three fairly high-quality ones in an hour.
I managed to get my hands on some low-quality papers. Yeah. But they are correct papers. Well, they are currently stored on my hard drive and waiting for me to email the relevant parties, but this is not a good use of my time as I should be investing in results . Yeah. But there are certainly people doing this . Well, there has been a significant increase in submissions to the archive . Yeah. Most of them are not very interesting . Well, some are interesting, but well, many of them are clearly low quality .
For example, even if the results would have been highly valued a year ago , there is absolutely no evidence that any human involvement was actually involved. No human capital or development of understanding . For example, there are cases where three, four, or even five identical proofs of the exact same theorem are published within a few days, which is clearly like someone playing a slot machine . ChatGPT consistently finds the same thing . Interesting. It's like mode collapse in a particular inference path . Yes, and this problem is resolved by improving the model.
I do n't think it will necessarily be resolved. It may be resolved, or it may not. Much of what I said earlier is that much of the progress in mathematics is brought about by a thousand different flowers blooming, people pursuing their own curiosities , and the boundaries of knowledge expanding . In the higher-dimensional space of mathematics, to some extent , preferably in a fairly uniform way , we need to bring about some kind of unity. And opportunistically, a few applications or old questions that we have discovered suddenly appear.
And when we subordinate mathematical inquiry to what the model wants to pursue, I'm not sure whether what we get is like one mathematician being copied a thousand times, or actually like a million different mathematicians doing a million different things. Yes, it is actually quite , except that mathematics must adapt by changing its set of structures, and I think this is actually quite, perhaps dangerous, and something that laboratories must be careful about . Because on the one hand, Their success stems from the fact that this emergent reasoning ability is clearly incredibly powerful, and thus generates many fantastic PR headlines.
Yes, as you pointed out, and my interest lies in the fact that human mathematicians start from a truly diverse range of strange intuitions . That is why, as you noted, they are able to develop a wide variety of frontiers , actually draw from them, connect them, and produce even more . So, if most of these proofs published from the lab converge to something very similar because they are technically based on the same literature, then that is indeed true. And that is their strength right now. In other words , it is not clear where they develop their intuitions from .
They come from working mathematicians. It is not clear where other , more powerful and diverse intuitions come from . They may come from somewhere, but we do not have a good theory about where they come from . Also , just the calculations during testing and scaling after various training sessions can make it clear that It's not even clear if it's possible . In other words, whether it's possible is a big debate . It introduces new features . And I think it should be studied in a mathematical context .
Because what it's good at and bad at, that is, what human mathematicians are good at , is precisely the question that is needed to study it . In other words, I think that unless we incentivize enough people to get involved with it , we probably won't be able to produce results . Unlike other fields , it may not actually be possible to continue producing excellent results without human help. Yes. Let's go a little further . Let's say the model becomes truly robust and superhuman, even if it doesn't add meaningful cognitive diversity.
I see. So I'm arguing that we still actually need human mathematicians. I see . So why? Because it's a question of how we design society . Right? For example, the optimal situation is that the model does all kinds of mathematical research, and it does so in a very diverse way, that is, human However, there may be situations where knowledge is accessed in a wide variety of ways, without the need to add such diversity . In other words, that may be the optimal method, and ultimately lead to many applications and improved understanding .
Just because something is optimal doesn't mean we necessarily have to do it . Right? For example, there's no reason to think that entrusting control of mathematical research to a model and leaving it to the model will yield the optimal results . In fact, even if we instrumentalize what we want the model to do and say, "I want it to make our lives better," that's not necessarily the case . Models don't always decide to pursue a wide variety of interesting things . It's research, right? They might simply choose the direct path .
We don't know what will happen. So, I don't know if you think this kind of extensive basic research is valuable , but I think it's one of the most valuable things that humans and models can do . I believe the easiest way to achieve this is to maintain a broad-minded community, promote the model, and help design the society we envision . Yeah. At the very least, my hopeful vision for the future is that humanity will not become completely powerless , but will have some degree of control over where we are headed .
And if that's the case, then ultimately what we do will depend on human interests . Therefore, we need people with diverse interests and the ability to actually pursue them. We need intelligent and proactive people. And, you see, you become capable of all kinds of thinking, not just mathematical thinking . yeah yeah . In other words, there is a major concern that as AI advances, we may not be able to develop ergonomically appropriate interfaces that actually encourage us to remain brilliant thinkers . And it's very easy to give it up .
Because the model isn't even good at the level of thinking, like the top-level structure , but it's still very simple. So, especially when it comes to mathematics , well, another selfish reason is that if mathematics is at its maximum , I think it provides a great educational excuse for thinking really rigorously about various things . In other words, this isn't the reason mathematicians do math , but it's part of our job . Understood, that's great. Because I think it gave me a really good framework for thinking about not only mathematics, but also many other things .
Furthermore, now that I'm also a parent of a two-year-old child, I've started thinking about this issue more often. In other words, it's not about just studying endlessly— of course that's important, and I'm not trying to deny the importance of studying— but when I was often collaborating with Hungarian mathematicians like Balázs Szegedi, I heard that they were teaching group theory in elementary schools in Budapest . So I thought this was something we absolutely had to do . We should continue doing this . And now, AI has become incredibly explainable and far more accessible, so I think we should actually promote its use even further .
Well, then, perhaps it could help to attract more people to the front lines, not just simply . Yes, that's what I'm concerned about. Well , of course, I'll mainly talk about math, because that's where I live . But one good thing to consider about this is that we are one of the first professions to have a significant impact with high-quality models . However, I think it's probably one of the first professions to become so public . But, in my opinion, there are many other professions that involve coding.
It's 100% coding. But, in other words, I think that almost everything that is done by computers is currently done by models. And although they haven't publicly acknowledged it , yes, mathematical ability is useful for businesses . I think I can talk about the lab a little better than anyone else . Yes, but, well, one reason to try to maintain human capital in this field is because it serves as a model for all professions . In other words, we are probably looking for people who are meaningfully engaged with the world, are experts, and possess talent and trained skills .
And, well, at least, it's very convenient that the mathematics profession is closely linked to education here . This is because some challenges arising from AI can be observed in universities and secondary education. So, as you said , it is also a great tool for learning . Well, I've heard someone talking about bimodal distributions in class . There are people who truly understand how to make good use of new tools, and then there are those who use them only for homework and ruin everything else . Unfortunately, I don't think that's fast enough to adapt, but we definitely want to live in a world that produces better thinkers .
I agree. Even just talking about pure optimization games would be better for us . Therefore, as humans, it would be extremely inconvenient if our thinking abilities declined at the same time as AI evolves, and that could easily happen. Therefore, we should seriously consider how we can actually use this to improve ourselves . I agree. Yes, I think it's interesting that the models can do more at their current capabilities, and that many things that weren't possible before can now be done quite cheaply . However, in many cases, it's not clear to me whether they actually improve the quality of the deliverables .
Yeah. I think this is a common occurrence . When new technologies emerge , we can do something that is slightly inferior to what we could do before , but at a much lower cost . As a result, a large number of low-quality deliverables suddenly appear , replacing the previously high-quality deliverables . However, I believe it is possible to use the tools in a way that actually improves the quality of the deliverables in every aspect . To achieve this , a certain degree of thoughtfulness and a redesign of systems to actually facilitate it are necessary.
Yeah. I hope capitalism works well . I believe that the most valuable things are those that people need to use effectively . Currently, the models are insufficient unless human experts are actually involved . However, as you pointed out, it is probably more difficult for the majority of juniors and entry-level players to adapt to such roles . Therefore, it is a mistake to use AI models in a way that is not fundamentally about deepening one's understanding by using the model to improve one's skills . And, as is human nature, it is very easy to become lazy, but we must resist that .
Because that's basically the moment you lose . So, please excuse the highly competitive language , but it's very easy to let the model do the thinking . The model cannot actually be thought of. Therefore, when things are progressing very rapidly , it is crucial to continue developing those capabilities and to use them to improve them rather than simply relying on them. Yes, that means there's one good thing about the significant increase in semi-expert and model-based attention to mathematical problems . Those are the people I was complaining about regarding slot paper and such—the people who don't actually produce high-quality products .
In many cases, it comes from experts. In other words, I'm not saying that academic mathematics has an incentive to produce a lot of things, or that it's one of the sources of all these things . Therefore, while there are certainly contributions from non-experts , I don't consider them a negative overall . In other words, there are many documents on the internet that you have to consult to find out if the problem has been resolved. Well, to me, the mere fact that there are so many people excited about mathematics seems like a kind of positive thing.
So , that's a good thing, isn't it? I agree. And it's nice to be able to talk to you. It's not just a kind of positive thing , it's clearly a positive thing, isn't it? Yes, yes, that's exactly right . It's like suddenly being in the spotlight , and now I can talk about math in a more geeky way. Well , actually, well, I think this is a different, constructive result , but I was curious about the comments I received regarding rank 30 elliptic curves .
It was just announced yesterday . So, if that's... yes. yeah yeah . It looks cool. In other words, we still don't know any details . I understand. Rumor has it that... well, I don't know how it happened . In other words, I think it's because Claude Fable, prompted by Revental Posi, unfortunately forgot the name of a collaborator. I think it was Leventhal. Please add this to your post. I only know their Twitter handle . Yes, you can add it to your post. yes. Many of the recent impressive results are due to Revent , and the degree of autonomy is unknown.
Therefore, in my opinion , some of them are not fully autonomous, but rather semi-autonomous . Well, I don't know anything about the method for this . Yes, this is an interesting structure . It's very difficult to say how important something is if you don't know how to do it. What I'm trying to say is , if you want to understand how those results have been historically proven or how they've been understood by the community, that's cool, but I'm not saying it's a big deal . Therefore, the only place where such results might be posted is someone 's personal record website.
Yeah. This is not an Annals-level result. Yes, this isn't an almanac , but it's cool in its own way, and there are several very talented mathematicians who like these kinds of problems , perhaps the most famous example being Marquis . Marquise and Craig Spoon have been pushing this record up for a while , and recently found themselves at number 29 . Well, why is that not allowed? In other words, that's a previous record . yeah yeah . In short, they are mathematicians, people like it , and they think it's cool that the model can do something like this .
However, what I always say about the results of a model is that you can't evaluate them until you look back on them afterward . And this also applies to human mathematics . Sometimes, issues that you thought were truly important, or that you thought required truly innovative and profound ideas, turn out not to be so, and you can't know that beforehand. I always get excited when those kinds of problems are solved , but when I try to figure out how important they are , I think Levent and Claude, and then Abe Howell, the third collaborator , haven't told me how they did it yet .
So, I guess that's why they were so secretive ? I think I've published some tracings, but regarding this one, unless I've missed it, I do n't think I've published it yet. Yes, Levent likes to tweet the results , and he's also gradually releasing PDF reports . I think they're just having fun on the internet . Yes, yes. In other words, it's similar to saying that we need to evaluate how the results were obtained. Well, recently, MarkZelke and MetahbSwani from OpenAI joined us , and they said that the most appealing and enjoyable aspect is that the proofs are relatively short .
It's not over 200 pages long , and it probably matches your struggles . But, in retrospect, it's possible that this was optimized to be shorter . Or, on average, are the things you throw at GPT, such as Saul or Fable, more likely to be shorter and easier for humans to read ? Or will GPT simply run amok, causing endless trouble? Yes, I mean, it's great that they sometimes generate short and clever proofs . Yeah. Of course , everyone prefers short and clever proofs . Well, in my opinion, the reason we don't generate long and complex proofs is because we can't.
In other words, they have n't yet developed the ability to verify what is correct . Well, in other words, even if you ask a model to generate a short proof, in many cases you can also ask whether it is correct or not . And in most cases, the answer is no. In other words, it's far more reliable than it was six months ago , but it can still produce incorrect information even when it knows it's wrong . yeah yeah . Well, I think the problem with generating something very long is that you might not realize it's wrong.
Well, what I'm wondering is whether OpenAI and Anthropic have internally solved far more problems than they've publicly disclosed . And I imagine that a significant number of those people are simply unsure whether what they're saying is correct or not. Well, for example, the list of 10 problems that OpenAI recently published is all formalized in Lean . Yeah. This is, of course, very good evidence that they are correct . There were undoubtedly many more problems that could not be formalized using Lean . For example, this might be because the prerequisite options are not yet included in the mathematical library .
Well , there were probably longer issues as well , which also makes verification difficult . Yeah. Well, this is just my guess. And indeed, the archives contain very long proof of AI generation . Well, for example , someone recently posted a proof of singularity resolution in positive properties , which was AI- generated and spanned 800 pages . Well, I'm sorry, but I haven't read it. I haven't found any errors , but it shouldn't be correct . This will be a great achievement . This is beyond the capabilities of the current model .
They are quite well-tuned . Yes, yes, yes. Well, and definitely no human has ever read it . The model definitely cannot check this kind of thing . Yeah. Um , that's fine. Well , my guess is that the reason they're generating short, clever things is because it's something we can check . Yeah. And you can also generate long things that are troublesome or difficult to check, but yes, yes. We must reach that point. This actually probably refers to ability, cutting-edge ability, and not in a negative sense, and not to being, well, very good at being clever .
No, I think it's perfectly fair . In other words, if you look at adjacent domains like the code , that's what you're looking for. Yes, I remember reading an article Cursor published about their long-term harness testing . In this case, they are trying to recreate SQL Light in Rust , which was very insightful in showing just how far we are from that goal. It seems like a very similar task, but it's much longer. You need to test a series of cases to make sure there's something to verify that it's the correct implementation , but in reality, you need a harness.
In this case, it's not like a raw model ; it takes time to actually compare it with various frontier and non- frontier models, who the planner is, and so on . There are also differences in their functions. Yes, I was thinking about practicing drawing out long proofs . A harness is required. Yeah. When creating a harness solely for the purpose of extracting proof, its reliability is often reduced because it is designed only to generate output . Well, ChatGPT5.6Pro is really trying its best not to say anything wrong.
For example, they might say something wrong , but if you ask, "Oh, was that right?", they'll answer, "No." However, when you try to get people to come up with ideas or unleash their creativity , you have to break out of this very rigid routine , and once you reach a certain point, and if you're interested, there are a lot of people trying to arbitrage the authority mechanisms of academic mathematics and extract many proofs that haven't necessarily been verified. And in order to do that, I think we need to sacrifice credibility to gain a lot .
Well, in other words, a harness that can pull out a 250-page paper probably doesn't pay much attention to what it's generating . Yeah. And you're saying they're declining or unreliable simply because they don't yet have that capability? In other words, it simply forces you to take on longer-term tasks . Yeah. Yes , in other words, in reality, it's impossible for a human to check a 250-page paper . In other words, it's not possible to reliably check each line. To understand this , we try to grasp the overall global structure of the argument and test it in various ways.
For example, does this argument imply something else that I know to be false ? For example, what happens in special cases ? And so on. Furthermore , the model still doesn't seem to be able to handle more ambiguous things, such as unit testing of proofs . In fact, one of my favorite tests where the model hasn't succeeded yet is that a paper published several years ago (I won't mention the name) was wrong . The paper seemed to be making some significant claims in a field very close to my own .
I downloaded it and started reading it right away, but it was very difficult to find any specific errors . However, the structure of the argument made it clear that it wouldn't work. In other words , I and many other experts realized that what he was trying to prove was too strong to be considered true if taken at face value. So, many of us experts, including myself, quickly realized it was wrong , and we emailed the author and continued the exchange until someone pinpointed the exact error .
And, for now, the model doesn't seem to be able to do this . For example, certain errors are quite subtle and cannot be checked for overall , or broad, implications . This is similar to checking academic papers in a practical work setting. Yes, yes. In other words, the code reflects a situation where the syntax is very clear and excellent, but the architectural level is still very weak. It might be possible to reach that point, but it will probably require the help of a harness . Nobody knows .
In other words, people have evolving opinions on the extent to which harnesses and models need to co-evolve, which is necessary, and which will be less necessary in the next model . Therefore, in mathematics, I think it's very interesting to conduct experiments with harnesses and see how they improve. Because it is a general indicator of inference. Yeah. Yes, I mean, I have a little harness of my own for the Codex and the Codex , but, personally, I do n't really like automated autonomous mathematics, so I hardly ever use the harness.
I mainly intend to use it to help me understand things. Understood, that's perfectly fair. Yes, that's exactly right . The reason I do n't want to automate the work is because you need to participate in the loop to understand it . As you know, it's required to participate. Well, lastly, this is something you may not have thought about, or perhaps haven't thought about much , but I assume you also have a toddler . is that so. yeah yeah . So, how ? Yes, she's 3 years old.
Oh, wonderful , wonderful. So, you've progressed further and probably have more thoughts on this . So, what are your thoughts on her education ? Or was it his education? Hers. What are your thoughts on math education ? Yeah. I don't mean to force you, but seriously, how do you convey your love for mathematics, and how do you react to AI? Yes, she's three years old and has never used AI before. I've just started doing addition . That's our current situation when it comes to mathematics. Well, that's more advanced than I had imagined .
She can do single-digit addition, but she can also count using her fingers , and is quite accurate up to about 30, and fairly accurate up to about 50 . That's why I'm so proud of it. That's good news. Yeah. Yes, I would definitely encourage that . We talk about things like shape and stuff like that . Well, actually, when I woke her up a few days ago, she was hiding under the blanket. So I asked her, "Hey Sophia , what are you doing there?" and she said, "Oh, I'm doing math ." "Oh, I saw it.
I think you tweeted about it. It was so cute . It was great. So I think she sensed that I like math and that's why she became interested. Yeah, I mean, in 20 years, or when she's fully grown up and doing her own thing, the world will probably look quite different, but I think a lot of what we're teaching people will be pretty robust against the changing nature of the world . For example, I think the reason for learning math is always to think clearly and to understand the world better , and that's probably something you'd want to do even if you had a really good AI .
And this is true too. As you know, I personally love math , which is also why I read a lot of books and study the humanities and stuff . So I mean, the profession of mathematics, and more broadly, the actual value of education , is something we're definitely trying to protect, and something I want to instill in my daughter . Well, how much the system needs to change for that to happen is probably an unresolved question . Quite a lot, I think. But Yeah, at least on a personal level, I'm trying to convince my 3-year-old daughter that math is really cool .
Oh, yeah, 100%. Her first word was icosahedron . And what was that? What was it? She got a little icosahedron toy from her parents . Well, she was one year old . She learned about Platonic solids pretty early, though not literally her first word. Oh, great. Great. So much fun. Next is group theory. I mean, that's very natural. That's right. I mean, actually, yeah, actually, when I teach her more about addition and subtraction , I do it in the context of general groups . Oh, great.
Well, at least it's motivating . I think a lot of people probably skip that part . Yeah, yeah, and maybe math graduate students will also be training the next, much younger generation to use AI to actually get better at math, rather than just because they don't understand it. I might be able to concentrate. That's great. Okay, thank you very much, Daniel. Yes, thank you. I had a great time . Yes, yes, I had a great time . Yes, I think there will be some bigger developments soon.
I would be happy to meet and talk again . That sounds good.