AI and the future of science
Preamble on the meaning of life
Let me start by admitting something a little bit embarrassing. My job is a key source of meaning in my life. Not in an I-think-I-am-playing-an-important-role-in-many-lives kind of way, but something more existential.
When I was done with high school and had to figure out what I wanted to do with my life, I decided to study physics and philosophy. I picked those two because I yearned to find some kind of meaning of life … and felt that those two were my best bet.
I’ve long since given up on finding the meaning of Life, the Universe, and Everything. But studying physics gave me something related. It’s a less ambitious mission, it’s different, but it still gives meaning to my life: contributing to science (even if in my own small way) is a way of contributing to humankind’s grand project of understanding the universe and everything in it. And that’s meaningful.
We are not even at why we’re here yet; we’re still working on how the world works. But I think how is the first step, and it’s enough to give my life meaning. Even if I know that we won’t get anywhere near answers in my lifetime, and even if I know that my contribution is just a tiny blip in the grand scheme of the big project.
I think this understanding of why I find my job as a scientist meaningful is important to keep in mind when you read the rest of this post.
Getting started
The thing that set this whole thing in motion is the following post by Joe Bak-Coleman:
When I saw it, my initial reaction was something along the lines of “I get those scientists. I think it’s actually in the spirit of science to investigate AI Science 1.” It’d be amazing if we could speed up the rate of exciting discoveries or perhaps (at some point in the future) come up with new insights that humans couldn’t even have imagined.
Not only do I think it’s right for academics to investigate whether AI Science could become a reality, I also think it’s brave. (To be clear, I’m talking about academics here; the giga-corps have their own sinister reasons.) But for a working scientist, it is brave to work on something that might mean that the whole point of your life’s work could be eradicated, and that the thing that gives your life meaning (doing science) could be outsourced to computers.
So that’s what I’ll be doing here. I’ll be unpacking the thoughts above and investigating the idea of AI Science in detail.
This is a controversial topic, so one important assumption before I start: I am discussing the future scenario where AIs are actually better than human beings at doing research.
I get that there are arguments that we’ll never get to that point, and I think it’s possible that those arguments are right. But I won’t discuss the pros and cons here; it’s beyond the scope of this post 2. I’ll even admit that if AIs aren’t better than humans, then we shouldn’t let them anywhere near science, or at least only use them with the utmost skepticism. (There are a few more caveats, but I’ve put them in the fine print at the end, so you can get on with the reading.)
One last point on my assumption. Better doesn’t need to mean infallible. Human researchers aren’t infallible either, which is exactly why we have a scientific method to check and validate results. The question is what the checking looks like when the producer is a machine; more on that below.
The split
Before I get to science, a word on art (as in fiction writing, paintings, movies), because I think the connection is important. I’m an extremist when it comes to art. I personally only want human art and have more or less stopped reading books written after 2023 (see this post for why). To me, art is about sharing the human experience, so art made by a machine is not interesting by definition.
Science is different. Here it’s all about that project I talked about in the preamble. Let me explain.
I think that the activity of science can broadly be divided into two parts: a) a technology component: building gadgets and machines, curing disease, those kinds of things; and b) an epistemic component: expanding human knowledge: think quantum mechanics, the theory of relativity … figuring out the clockwork of the universe.
If we take the technology component first, I think it is somewhat uncontroversial that we would want the best technology ASAP and it wouldn’t really matter if a human or an AI did the research, as long as it works. If a loved one is dying of cancer, most people just want the scientific strategy that cures them 3. I know I would.
The epistemic component is where I want to focus my attention. I am of the opinion (related to my view of science explained in the preamble) that science should create insight for humans 4.
Insight that exists only as traces within or from an AI is pointless. I want the knowledge to be inside human brains. To make the universe (more) comprehensible to human consciousnesses. To humankind.
And if machines claiming knowledge first can help humans get to a new understanding of the universe, then I want that.
I love the thrill of discovery. I love pretty much every aspect of science and being a scientist. But as I see it, the scientific project is not about me. It’s about humankind understanding its role in the universe. And if it works better without me in the role of researcher, then it is my duty to get out of the way 5.
Sigh.
The new situation: moving from discovery to understanding
The brilliant mathematician Terence Tao has written a lot about AI. And his perspective is particularly interesting because mathematics is currently the place where the situation we’re discussing is closest to already being the case. So I will draw heavily on his thoughts as my starting point below.
Below is how Tao described the situation in the Atlantic in Feb 2026.
These problems are like distant locations that you would hike to. And in the past, you would have to go on a journey. You can lay down trail markers that other people could follow, and you could make maps.
This is a beautiful image. The scientist as an explorer of a new continent of knowledge. If we extend this to all of science, we get the image in Panel A below. In Figure 1A, I show our existing knowledge as a landscape (in the bottom left) and the frontier of what we don’t yet know winding down approximately along the diagonal.
Crucially, discovery and human understanding move together in the landscape. The way human knowledge grows is by discovery and understanding moving in lockstep, together 6.
The way I think about the new situation with AI Science (and what seems to be the current situation in math 7) is that the knowledge map gets a new region! I show this situation in Figure 1B.
In this region, we have a completely new situation where the discovery frontier and understanding frontier can separate. This opens up a deeply interesting new area where AIs have discovered new insights about the universe, but where no human being has yet understood it. It’s a complicated place and we’ll get to that later. But I’ll be talking about it a lot below, so let’s name it. Staying in the map metaphor, I’ll call it the “unclaimed territory” 8.
When we connect this back to the epistemic component of science, the presence of large not-yet-understood aspects of the universe could be construed as something horrible. In that same interview Tao intuits the presence of the unclaimed territory without making it explicit. He says:
AI tools are like taking a helicopter to drop you off at the site. You miss all the benefits of the journey itself. You just get right to the destination, which actually was only just a part of the value of solving these problems.
Here’s a similar point (from a post on Mathstodon from September of this year)
But modern AI tools, when directed without such expert supervision, can now be aimed all too easily at the ostensible goals of the field, without any incentive to take the slow, whimsical path. The flag is captured, the goal scored, and the problem is solved; but at the cost of lessons learned, insights gained, collaborations formed, and new targets located.
Tao also wrote A Severe Misalignment of AI in Mathematics 9, declaring that “the goals of the AI companies and the goals of the mathematical community are severely misaligned” and that the issues “must be addressed urgently”.
In the declaration, he says something really interesting that aligns with my own view:
In recent months, the success of AI in solving major mathematical problems has made headlines even outside mathematical circles. But solving problems is only a tool and proxy for achieving the primary goal of conceptual understanding and insight. (My emphasis)
That is exactly the epistemic component. Tao and I share the goal of human understanding and insight.
The unclaimed territory could be great for the Scientific Project!
But where we differ is where my argument perhaps gets interesting. I think that AI progressing beyond human understanding could be a great thing.
Tao worries about “proof indigestion” – proofs being generated and even verified without being digested (= understood).
But what’s the rush? Imagine a big open territory that we can take our sweet time digging into. We don’t have to keep up with discovery; we can proceed at a human pace, nibbling away until knowledge is fully incorporated.
Consider this idea: What if exploring the unclaimed territory turns out to be fun? Maybe it’s not so bad to be helicoptered in, landing on a peak. I explore this idea in Figure 2. In panel A, we have Tao’s view of undigested insight, the researcher helicoptered in.
First off, the ability to land on a peak doesn’t rule out the old ways. I think we could and should hang on to “insights gained, collaborations formed, and new targets located”. Moving forward, awareness that the aim is squarely to understand implies that we should strengthen those aspects everywhere, as science evolves in this new phase. (How we keep the old skills alive once nobody has to hike is a real question. I return to it in Appendix 2.)
Secondly, and this is what I think we need to think more about as we envision the science of the future. The presence of the unclaimed territory opens up many new interesting ways of supporting the epistemic component of science. There could be many new ways of supporting human understanding. To stay in Tao’s metaphor, the human being is still hiking; it’s just from a different trailhead.
So arriving by helicopter doesn’t necessarily skip the journey; it provides new places to start your journey. It might also tell you which journey is worth taking and give you a point to hike towards 10.
Knowing that something is true, and roughly how it fits into the landscape of knowledge, opens up a new doorway to understand that problem. Every student who has gotten stuck on a problem and worked backwards from the answer at the back of the book knows this.
I realize that working backwards from the answer has a negative ring to it, but in the case of the Scientific Project it shouldn’t have. We’re not vying for good grades in a class here – the aim is to wrestle secrets from the universe. Any method counts!
In fact, science has worked backwards from the results before, and not just through lucky accidents like penicillin, as I mentioned above. Experiments have handed us answers that took theorists decades to explain. The Balmer formula for hydrogen’s spectral lines (1885) waited 28 years for Bohr (1913) 11. Superconductivity (1911) waited 46 years for BCS theory (1957) 12. High-temperature superconductors (1986) are still not fully explained 13.
And the idea of results without proofs leading to exciting times in math is also not unheard of historically. There is Ramanujan, who died in 1920 at the age of 32 and left behind notebooks full of results with almost no proofs. Mathematicians have spent a century happily hiking back to his peaks; Bruce Berndt alone filled five volumes (1985–1998) proving what was in the notebooks. Or good old Fermat’s theorem. Around 1637, Pierre de Fermat scribbled in the margin of a book that he had a marvelous proof that the margin was too small to contain. That note was a result that fueled mathematicians for 358 years, until Andrew Wiles’s proof was published in 1995. (More Fermat below.)
The unclaimed territory is a new place for human minds to play
Now let’s get even crazier. I actually think that we should think of the unclaimed territory as an unclaimed playground. I argued above that it gives us completely new ways of adding to human understanding – and that having results to shoot for could invigorate human minds.
But most of all, it could also be good old-fashioned fun! There are lots of examples showing that we don’t need discovery or some kind of “cognitive primacy” to have fun.
- Puzzles. Human beings love puzzles. Sudoku, crosswords, wordle, the list goes on. They all have an answer key and nobody thinks that ruins them. Knowing that an answer exists can even be part of what makes the hunt fun.
- Chess. A machine beat the world champion in 1997, when IBM’s Deep Blue defeated Garry Kasparov. Today engines are far stronger than any human, but chess has never been more popular: Chess.com passed 100 million members in December 2022, and a whole generation of players learned the game by studying engines.
- Go. Cognitive scientists analyzed 5.8 million moves made by professional Go players between 1950 and 2021 (Shin et al., PNAS, 2023). After superhuman Go programs arrived in 2016, human players started making significantly better decisions. And the improvement went hand in hand with novelty: players began making moves that had never been played before, and those new moves were increasingly good ones. The machines didn’t make the game boring. They pushed the humans off the old trails and into new parts of the map.
The epistemic surveyor
So I argue that we’ll all be fine in the future. It might even be a golden age for the growth of human understanding, for the Scientific Project.
But one last question remains. Once the discovery part is gone, and the focus is on understanding, are we still researchers? My best answer: probably not.
As this page goes live, I am walking on stage for a debate where I have to argue that the job of researcher will disappear in the future. This position is dictated by the debate format.
My strategy is to lean into the arguments below: the job of the future “understander of unclaimed territory results” is so different from what scientists are doing today that it deserves a new name 14. And in honor of the main metaphor, I came up with “epistemic surveyor”, or just the humble “surveyor” for short 15.
It’s a rhetorical trick, but is it overkill to rename the whole profession? I actually don’t think so.
If we look at the history of work, plenty of jobs have vanished because technology made their essential task obsolete. Lamplighter. Switchboard operator. Elevator operator. The list is long.
But notice that most jobs don’t vanish when the tools change. A chef with an induction stove is still a chef. A surgeon with a robot is still a surgeon. An architect who has swapped the drawing board for software is still an architect.
So what is the difference between the jobs that disappeared and the ones that didn’t? In the jobs that stayed, the tools changed, but the central task of the job stayed the same: deciding what people will eat, cutting into a living body, imagining a building. When technology only changes the tools, the name of the job survives.
Jobs disappear when technology changes the central task itself. The example I want to feature is computer.
Before electronic computers existed, “computer” was a job description. A computer was a person who computed: ballistics tables, the positions of stars, the long columns of arithmetic behind early numerical science. When the electronic computers began to arrive in the 1940s, the task that defined the job was gone within a few years.
But here’s the thing. The (human) computers didn’t go home. They became programmers.
The first six programmers of the ENIAC computer were recruited straight from the pool of human computers working on ballistics tables: Kay McNulty, Betty Jennings, Betty Snyder, Marlyn Wescoff, Fran Bilas and Ruth Lichterman. At NASA’s Langley laboratory, Dorothy Vaughan saw the IBM machines coming, taught herself FORTRAN, and then taught her entire group of human computers to program them 16.
The job “computer” disappeared. The people moved into a new job that grew up around the new bottleneck: no longer doing the calculation, but deciding what the machine should calculate and making sense of what came out.
So the point is that the job changed name because the substance of what they did changed.
And what is a researcher? According to Merriam-Webster, research is
investigation or experimentation aimed at the discovery and interpretation of facts, revision of accepted theories or laws in the light of new facts, or practical application of such new or revised theories or laws
Notice the order and the focus. Discovery is the essential part. Interpretation is in there, but tucked in behind it, consistent with the preconceived notion of the two moving in lockstep.
In the new world, discovery rests mostly with the AIs. It is all about understanding and interpretation, a whole new job. That is why we won’t be researchers 14.
What does the surveyor do?
So what will the surveyors be doing? What does the work-life of a surveyor look like? I’ll mention some examples below, but I’m sure there are many yet to be discovered.
Let’s start with the post below, for example. I see this as Andrej Karpathy beginning to develop ways of working in the unclaimed territory.
A friend of mine called the process of learning how to work with and learn from frontier models dancing with the AI. I think our current experiences are already revealing the creativity of how human beings can tease understanding out of the unclaimed territory.
We can also attack the job changes more schematically. Today, research runs more or less like this:
question → search → discovery → understanding
In the future, I think it is going to be more like:
machine discovery → selection → interrogation → simplification → human understanding
Selection is about choosing the problem in the first place. There will be effectively infinitely many new theorems to choose from, each of them machine-checked and true, and most of them boring. A key task for the surveyor will be deciding which ones are worth spending time on 17. (More on the challenges this raises below.)
Interrogation is the dancing with the AI aspect. The helicoptering. The working backwards from the result. The part that I argue will be fun. What will it look like? We don’t know; this part is just beginning. All of our old tricks will be in there, but I reckon that it’ll be mostly weird and new. (One more example of what it might look like is in Appendix 1.)
Finally, simplification. Basically, I imagine this as the part where findings are connected more formally to the edifice of science. It’s about finding the frame that makes things simple (if they can be simple) and coming up with ways of communicating so other humans can more easily approach the result. It’s the sense-making aspect we remember (Balmer had a formula that fit four spectral lines; Bohr found the atom that explained why).
Somewhere along the path outlined above, understanding arrives. First as intuition, then something more firm, and then finally ownership. Think Bloom’s taxonomy 18.
Novel challenges for the surveyors?
The unclaimed territory isn’t necessarily all idyllic, flowers and open meadows. As many smart people have already worried, there are many new challenges interspersed with the opportunities. We talked about Terence Tao above. The anthropologist Lisa Messeri and the psychologist Molly Crockett wrote a paper in Nature on AI and illusions of understanding in science back in 2024, long before any of the results above existed.
Facing these challenges is part of the surveyor’s job description.
The “dark side” of the selection aspect from the previous section is overwhelm. And it’s not just frontier labs publishing new results. Human scientists with AI agents are also generating more (and blander 19) papers on an unprecedented scale. arXiv received about 105,000 new papers in 2015 and about 284,000 in 2025, and this September alone it took in more than 40,000, a new record 20. On 1 October, arXiv capped authors at two submissions per month, saying that AI tools make it easy to “flood arXiv” with thin papers. If the literature already feels like more than anyone can read, imagine it with the machines submitting.
Above, I mentioned Bloom’s taxonomy, a framework that shows the complexity of the term “understand”. And I agree that AI use can lead to (and has led to) illusions of understanding. Messeri and Crockett warn that AI tools make scientists believe they understand more than they do. In the unclaimed territory this is the occupational hazard. A result arrives with a proof, an explanation, maybe a tidy summary, and it feels understood.
Feynman’s blackboard had the test written on it: “What I cannot create, I do not understand.” Part of the surveyor’s job will be to truly “own” the results they claim. The ideas of interrogation and simplification are the first steps, and maybe we’ll develop many other tools for this part. Understanding is slow, and that is fine. What’s the rush?
Wrapping up
One last aspect of the whole thing: the new stage of science could be really exciting to witness.
That is because science as an exploration of the unclaimed territory could mean a dramatic speed-up of the Scientific Project described in the Preamble. Remember how I said that “[e]ven if I know that we won’t get anywhere near answers in my lifetime” … well, now maybe we will get somewhere wild!
Postscript (8 October 2026)
I wrote everything above before Tuesday evening (October 6th). That’s when OpenAI released 719 manuscripts in 372 families, each claiming to settle or substantially advance an open problem in mathematics or theoretical computer science, all produced by an internal model at about three hours of compute per result. Among them is a claimed proof of the Unique Games Conjecture. Scott Aaronson, whose wife Dana Moshkovitz has worked toward that conjecture for her entire career, called the day the Mathocalypse.
Consider this piece of news with the arguments above in mind. Maybe we just opened up our first big chunk of the unclaimed territory. Wild times indeed.
Post postscript (9 October 2026)
The checker is only as good as the translation. OpenAI’s Navier–Stokes blow-up result from September came with a Lean verification, the kind of guarantee I lean on above. A new paper by Alexander Bastounis, Fabian Circelli and Anders Hansen argues that Lean checks the formal proof, not the written proof it was translated from, that faithful translation is in general uncomputable, and that in the Navier–Stokes case the two proofs don’t match. I have no idea what’s what, but wanted to underscore the importance of human understanding and the sometimes-treacherousness of the unclaimed territory.
The fine print
A few more caveats, as promised.
- Before we can ever let AIs do science and implement technology, there are many complex legal and ethical questions to solve. I won’t get into any of those here. Let’s imagine for the moment that we manage to figure all that out too.
- I won’t get into anything related to energy consumption. There are valid reasons why we should worry about data centers, pollution, etc. And those reasons could mean that we shouldn’t use AIs to do science because it’s just not worth it. But I won’t worry about that here. Let’s imagine that energy problems have also been solved. Maybe models are better, maybe there’s fusion. Whatever. The point is that I want to focus solely on the topic defined above.
- Here I’m mostly discussing the hard sciences like math, physics, chemistry, computer science, etc. Things get muddier when we get to the social sciences and humanities that are about the human world. Within those fields, it’s not as clear-cut that AI Science will create a clean unclaimed territory. As someone doing quantitative work within social science where systems react to being understood 21, I spend a lot of time thinking about this question, but a fully fledged theory for these parts of the great scientific project is a job for another post 😅
Notes
-
By AI Science I basically mean building a computer that goes brrr and spits out science. ↩
-
And let’s be explicit that I mean we’re at the point where we have automated wet labs, etc., so the AI can also do autonomous research in biology. In that imaginary future, we’ve also figured out how to do human trials with AI-developed drugs. ↩
-
I get that many people would say “AIs will never cure cancer,” but here I refer back to the assumption above. We are currently discussing the situation where they actually could. ↩
-
And this is the connection to my view on art. In the end we’re doing all this for the benefit of humans. ↩
-
Thanks to Petter Holme for discussions on this topic! ↩
-
Sometimes understanding takes a while – think penicillin or many other examples. ↩
-
I know that some people will still argue that the AIs just had the proofs in their training data, or that a model only produces proof-shaped text. Maybe. But that doesn’t seem to be Tao’s view (see the next note), and in any case: whether a proof was remembered, guessed or reasoned out, the proof checker doesn’t care. Neither should we. ↩
-
One rule for the map. A result only counts as unclaimed territory once it has been checked. By a proof checker, an experiment, by new strategies that we haven’t even developed yet. Unchecked claims aren’t territory quite yet. Fog maybe? In mathematics the checker already exists, which is one reason math is where all this is happening first. But it’s tricky, cf. the post postscript. ↩
-
It was co-signed by 24 other Fields Medallists, so Tao is not alone. And for what it’s worth, I think the fact that so many other mathematicians signed this document suggests that what AIs are doing in math is real. It’s not just models regurgitating their training data … something fundamental is changing. ↩
-
Backcountry skiers who climb mountains on their own two skis have a slogan for this: “earn your turns.” Heli-skiers pay a small fortune to skip the climb and go straight to the part that’s actually fun – the way down. As far as I can tell, both groups have a wonderful time, and nobody has accused heli-skiers of not really understanding snow. ↩
-
Johann Balmer found his formula in 1885 by staring at the wavelengths of the four visible lines of hydrogen. Pure pattern-finding, no theory. It took Niels Bohr’s 1913 model of the atom to explain why it worked, which is one of the founding moments of quantum mechanics. Bohr got the Nobel Prize in 1922. ↩
-
In 1911, Heike Kamerlingh Onnes cooled mercury to about 4 kelvin and watched its electrical resistance vanish. The explanation came from Bardeen, Cooper and Schrieffer in 1957 and earned them the 1972 Nobel Prize. ↩
-
In 1986, Bednorz and Müller found superconductivity in a copper-oxide ceramic at around 35 kelvin, far higher than most theorists thought possible. They won the Nobel Prize the very next year. Now, four decades later, there is still no consensus on why these materials superconduct. ↩
-
But to be honest, I don’t think it matters much what we call the new role. Researcher, surveyor, who cares. The future will be very different from now – and for those of us who share the goal of science as defined in the preamble – it’ll be wild and hopefully fun! ↩ ↩2
-
And let’s be honest: without the discovery part, the job is going to be more humble, so a more humble name is a bonus. Surveyors have never been lauded the way discoverers are. Columbus gets a holiday; Juan de la Cosa, who sailed with him and in 1500 drew the first map that shows the new continent, gets a footnote (this one). But remember, turning discoveries into territory that other people can actually use was necessary, valuable work, and it still is. I’m fine with the modesty. ↩
-
If you’ve seen Hidden Figures you know this story; if you haven’t, Margot Lee Shetterly’s 2016 book of the same name is even better than the film. Vaughan is the one who walks into the room with the IBM 7090 that nobody at Langley could get to work … and gets it to work. So great. ↩
-
From a Simons Foundation piece on Lean’s impact on mathematics, June 2026. The same piece has the line that, with Lean, mathematicians can “trust AI-generated proofs of lemmas without worrying about AI hallucinations or ‘slop’.” That’s the whole trick: an unreliable generator plus a reliable checker gives reliable mathematics. ↩
-
Bloom’s taxonomy is the classic ladder of learning objectives from educational psychology (Benjamin Bloom and colleagues, 1956; revised by Anderson and Krathwohl in 2001). The revised rungs are: remember, understand, apply, analyze, evaluate, create. Checking that a proof is correct lives on the bottom rungs. Owning a result – being able to judge it and build with it – is the top. The surveyor’s job is to climb that ladder on behalf of humankind, one machine result at a time. ↩
-
Blander is a strong word, but there is data behind it. Qianyue Hao, Fengli Xu, Yong Li and James Evans analysed 41 million papers and found that scientists who use AI tools publish about three times as many papers and collect almost five times as many citations, while the collective set of topics that science studies shrinks, and scientists engage less with each other’s work. AI-augmented research drifts toward wherever the data is richest. More papers, fewer places. Nature 649, 1237–1243 (2026). ↩
-
From arXiv’s own monthly submission statistics, downloaded 7 October 2026: 105,280 new submissions in 2015, 284,486 in 2025, and 40,363 in September 2026. The 2026 total had passed the whole of 2025 before the year was three quarters done. How much of the recent jump is machines writing papers, I can’t tell you. Nobody can. The rate limit is in arXiv’s blog post of 1 October 2026, which also gives the September count and notes that submissions in the AI category grew more than sixfold between 2024 and 2026. ↩
-
Social scientists have a whole vocabulary for this. Robert Merton called it the self-fulfilling prophecy (a rumor that a bank is in trouble can be enough to sink it), Karl Popper called it the Oedipus effect, and economists know it as Goodhart’s law: when a measure becomes a target, it ceases to be a good measure. Planets don’t change their orbits because we’ve predicted them. People change their behavior all the time because of what we’ve predicted about them. In the social world, the map changes the territory. ↩
Appendix 1: More inspiration for what interrogation might look like
The post was getting long, but I wanted to include one more example of what I think the “interrogation” step might look like.
I put it here because it was run on known terrain, but I still think it counts: Claude agents’ complete machine-checked proof of Fermat’s Last Theorem. The proof is about thirteen million lines of Lean and some thirty thousand theorems. Nothing new about Fermat was discovered.
But look at the figure below. The dependency graph of the proof is a map. For the first time you can see the entire proof at once, every path from the axioms to the summit, with the theorems grouped by family. I think that is what surveying could look like.

Appendix 2: What other people are saying
I’m not the first person to think about AI Science, and some very smart people see it very differently. Rather than going through them one by one, here are the strongest objections to my argument, grouped by what they’re actually worried about. (Objections that are about some of the surveyor’s working conditions – overwhelm and illusions of understanding – I discuss above.)
”It won’t happen”
The most common objection is that the premise is fantasy. The always awesome Carl Bergstrom argues that point in this thread. He may well be right. But that’s the assumption I made up front: I’ve chosen to argue about the world where the machines actually are better. If they never get there, this whole post is moot, and I’m happy to say so.
”It’s about capital, not curiosity”
A point from Bergstrom that hits me harder comes later in the same thread. The actual choice, he argues, is about concentrated capital taking one more step along the path of the industrial revolution, mechanizing the knowledge that lives in skilled labor, like the Jacquard loom did for weaving. I think he’s right about this, and it worries me too. But it’s a worry about technology and power, not about insight.
”Understanding has to be earned”
This is a deep objection, and it’s what this post’s two figures are about. Tao’s helicopter is the cleanest version: land on the summit and you miss the benefits of the journey. Bergstrom and Joe Bak-Coleman (yes, the author of the post that started this whole thing) make a related point in a Nature column on AI and the human activity of science. Its best line, I think, is that “papers are not the products of scientific research in the same way that castings are the products of a foundry.” Papers are vehicles for “the human understanding that is forged through scientific activity”.
I couldn’t agree more on the goal, that’s the entire point of this post. And, as we saw above, neither could Tao. We agree on the destination; we disagree about the route. My answer is Figure 2B and the arguments about seeing the fun in the new situation. The way from a machine-found peak back to the mapped world is a journey, and it still forges understanding.
”Where will the surveyors come from?”
Bergstrom and Bak-Coleman also worry about de-skilling. Automation transfers knowledge from humans to machines, and “as that transfer occurs, human skills are lost, perhaps irrevocably.” This is the strongest objection to the surveyor, and I take it seriously: if nobody does the hard work of discovery, where do the people who can understand the results come from? I don’t have a full answer yet.
But I have the beginning of an answer, and it predates this post. In June, YY Ahn and I spent a good part of a podcast episode on exactly this. YY said it so well: we replaced eight to ten hours a day of hard physical labor with a couple of hours a week at the gym. Perhaps we need the same thing for the mind – deliberately doing the hard mental work, regularly, under suitable constraints.
It’s also important to remember that we always deskill. Plato warned that writing would destroy memory, and it did, for the kind of memory that held oral epics. Nobody computes logarithms by hand anymore. Losing the old skill is the point of a new tool. The real danger is what the medical literature has started calling never-skilling: not training the new skill at all. That’s an argument for designing the surveyor’s training now, not for keeping the old one going.
”It just makes things up”
A whole genre of criticism holds that language models don’t produce answers, only text that looks like answers. The philosophers Michael Townsen Hicks, James Humphries and Joe Slater made the case in a paper with the unimprovable title ChatGPT is Bullshit (2024), and Léon Bottou and Bernhard Schölkopf call the models a fiction machine (2025). On today’s models, for plenty of tasks, they have a point.
But notice that this is an objection about the generator, and science has never trusted its generators. It trusts its checks. In mathematics the check is now a program: a Lean proof is accepted because a small, open proof checker accepts it, not because anyone believes the thing that wrote it. An unreliable generator plus a reliable checker gives reliable mathematics. Everywhere else, for now, the check is the surveyor, which is what science has always done with unreliable producers. Humans included.
”It takes away the meaning”
Finally, the objection that I understand, but also must reject.
Julian Togelius wrote a post with the excellent title Please don’t automate science. He separates “weak” science automation (tools that make researchers more productive), which he’s happy with, from “strong” automation, where humans become redundant.
His core argument is about meaning: researchers “are here because they love research and want to contribute to advancing human knowledge,” and “if humans could not usefully contribute to science anymore, this would be a disaster.”
If you’ve read the preamble, you’ll know that I share the premise completely. We just land on opposite conclusions. Togelius thinks we should protect the meaning; I think the meaning was always in the shared project, not in my personal share of it.
Appendix 3: A post I found after finishing this one
The closest thing to this post that I’ve found is Jordana Cepelewicz’s Quanta essay Is AI the End of Math As We Know It?, published on 5 October, which I only found on 8 October. Cepelewicz arrives at a future of “interpreters” that looks a lot like the surveyor’s, and she is less sure than I am that anyone will want the job. If this post left you wanting the view from inside mathematics, read hers