-
Notifications
You must be signed in to change notification settings - Fork 106
Getting started with using AI for research mathematics
The wiki is no longer updated. The latest data is as of Jun 30, 2026.
These are some common pitfalls that we have seen. We hope that some of this advice may help with getting started.
Compare also with recommendations from the Leiden Declaration.
AI tends to be advertised as "magic"; this has led to understandable widespread skepticism.
In reality, it is a valuable and powerful tool. But like all such tools, there are associated learning curves.
If AI does not produce good results the first few times, do not immediately quit or blame the AI.
Persistence pays off.
Some practices are already outdated. For example, the role-playing frame e.g. "You are an expert mathematician", or whether saying "please" improves performance (it usually doesn't).
In general, many older practices are based on considerations that LLMs "predict" words. Newer practices that work tend to emphasize cognitive complementarities with humans, orchestrations of long-range processes, and techniques for improving correctness and accuracy.
But it is likely that no one has the best recipe, and so you should also perform your own experimentation.
It is somewhat common for beginners to treat all LLM as a single class.
In the first place, for research mathematics, non-reasoning models such as the free ChatGPT model are no-go.
Among frontier reasoning models, performance still varies quite a bit between models. We will not go into specifics here because there is no objective consensus and the situation changes frequently; do your own research.
We're not talking about profits or ulterior motives here. (And for the record, please do not assume such things beforehand; see also Hanlon's razor.)
For example, some features may have undesired (unintended) side effects that interfere with the AI's work. You are encouraged to look into what factors may have hindered the AI's performance and work to eliminate or reduce them. On the other hand, you are also encouraged to look into factors that enhance performance or otherwise have positive effects on the AI.
Some products may not work as advertised. For example, for a long while ChatGPT Deep research was powered by an outdated model, so much that a regular chat might do better for that specific use case.
At least one person consistently prefers the regular ChatGPT Thinking model over the expensive Pro model, despite this being the opposite of what is recommended.
In general, you know your use cases best.
When getting started, you should aim for using AI to produce results that you can confidently check.
This ensures that you're getting positive rather than negative value out of it.
Subsequently, you can become more ambitious, but note that there are always limitations, bound by what you can confidently verify with your existing processes.
When in doubt, stay more grounded rather than less grounded.
No matter the process used to produce the work, you are the one responsible. (Unless AI gains personhood in the future, then we can question this premise.)
This means that, for correctness, you have the responsibility to maintain it to a reasonable extent. To clarify, even published papers have mistakes, so the intention is not to paralyze you or restrict your options. Be pragmatic, and think of benefits to the mathematical community as a whole in making decisions.
No less important are qualities other than correctness. Concretely: 1) literature review should be performed to a reasonable extent; 2) writing and presentation should be up to the usual standards; and 3) results should be significant in the usual sense.
Most importantly, you should be proud of the work, and you should care.
Disruption to the old way invariably invites negative sentiments. See here for some analogy.
Nevertheless, stay positive.
Decide for yourself what is best.
Cheers!