5 More Mistakes About AI
From Baseball to Big Science
Last week I outlined five misconceptions about AI. Here are five more problems with how we think about it.
If everyone uses it, it stops being an advantage
In Moneyball: The Art of Winning an Unfair Game (2003), writer Michael Lewis describes how baseball teams like the Oakland A’s gained a competitive edge by valuing on-base percentage instead of traditional stats like batting average.1 A blockbuster movie starring Brad Pitt followed in 2011. As Lewis put it, the A’s had found an ace-in-the-hole, a new way of scoring players and games.
What does this have to do with AI today? Actually, a lot.
Quick detour here, the Moneyball story traces back to an amateur statistician named Bill James, who had started analyzing baseball based on relatively quotidian metrics like On-Base Percentage (OBP) and Slugging Percentage (SLG). James’s method worked because almost no one else was doing it at the time. His method was later coined “sabermetrics” and found its way to the major leagues with, famously, Billy Beane, the Oakland A’s general manager whose low-payroll club became the emblem of the strategy depicted in Lewis’s book.
The larger point eclipsed baseball: if a system systematically undervalues something, then recognizing its value before everyone else can produce outsized returns. Bill James’s statistical work mattered because it identified a blind spot in a competitive environment. Billy Beane’s genius was not that he liked numbers. It was that he operationalized them before the rest of the league caught on.
But the crucial fact about advantages of this sort is that they are often temporary. An edge derived from better information, metrics, or inference lasts only as long as it remains unevenly distributed. Once every front office uses the same data, hires the same analysts, and prices players through the same lens, the market adjusts. What was once an inefficiency becomes standard practice; what was once a strategic advantage becomes table stakes. The method does not stop working in an absolute sense. But it stops conferring asymmetrical advantage. It “saturates” in the market, we might say.
This brings us to AI.
Early evidence from software engineering shows a similar pattern. A controlled study of developers using GitHub Copilot found that participants completed coding tasks about 55% faster than a control group.2 The gains are real. But as adoption spreads, the advantage disappears. More talented programmers, relative to junior coders, remain valuable to organizations, as before. Human talent remains the centerpiece.
Right now, many people speak about AI as though merely using it grants an advantage. In the short run, in some contexts, it does. If you use a language model to summarize documents faster, draft competent boilerplate, accelerate coding tasks, or generate plausible first passes at routine intellectual work, you may indeed outperform someone who refuses to use it at all. But this is the easy part of the story. The harder question is what happens when everyone does the same thing.
The answer is: the advantage disappears.
If every student uses AI to draft essays, then AI-assisted drafting no longer distinguishes one student from another. If every consultant uses it to generate slides and memos, then faster slide production ceases to be differentiating. If every marketer uses it to produce copy variations, then the supply of acceptable copy rises and the marginal value of any one instance falls.
In each case, the technology may increase throughput. But increasing throughput is not the same as creating durable strategic advantage. More often, it simply resets the baseline.
This is what many people miss when they talk about AI as though adoption itself were a moat. It is not a moat if the capability is generic. A true moat requires scarcity, defensibility, or some difficult-to-replicate integration with talent, judgment, proprietary data, institutional process, or domain-specific expertise. Otherwise the tool behaves less like a secret weapon and more like a spreadsheet. Useful, yes. Transformative in certain workflows, yes. But once generalized, it becomes infrastructure.
A generalized technology does not make everyone exceptional. It simply adjusts the level where competition takes place.
That is why the most extravagant claims about AI-driven competitive advantage should be treated with caution. In the early phase of adoption, when some people use the tool well and others not at all, gains can look dramatic. But those gains often reflect what we might call a “diffusion lag,” rather than a deep transformation. They are advantages purchased by being early to a method, not by possessing a fundamentally new kind of intelligence.
“Moneyball” worked because the insight was not yet common knowledge. Once it became common knowledge, baseball did not stop being data-driven. It became more data-driven than ever. But the original edge disappeared into the structure of the game.
The same thing is happening with AI. As the tools spread, their benefits do not vanish, but their distinctiveness does. The organizations and individuals who continue to matter will not be those who merely use AI, but those who can do something with it that others cannot: frame better questions, exercise better judgment, integrate outputs into real expertise, or build systems around it that are not easily copied.
In other words, human talent will still discriminate on the new playing field—as always. This is one reason why talk of massive unemployment from AI adoption is wrongheaded. It is also why the usual patter about a coming AGI is—as always—off the mark.
Takeaway point: when everyone has the same statistical assistant, advantage returns to the human being. As always. How will companies train employees for the new forms of competitive advantage? That’s a human question, and boomers and doomers should listen.
“AI” is no longer a scientific concept. It’s a capital-intensive industry
AI is no longer primarily about ideas, but about resources. Training state-of-the-art AI models requires massive compute, specialized infrastructure, and large-scale data pipelines, which concentrate capability in a very small number of organizations.
The Large Hadron Collider
In this sense, modern AI extends a trend that stretches back decades. It’s been termed “Big Science,” and in fairness this approach has delivered scientific results and on occasion breakthroughs. For instance, the Large Hadron Collider, stretching 27 kilometers on the Swiss-French border, enabled particle physics experiments that confirmed a new particle, the Higgs boson in 2012, something no small lab could do. Large, coordinated efforts like the Human Genome Project mapped the entire human genome in the early 2000s.
But big science also has limits: it’s expensive, centralized, and agenda-driven. The funding requirements alone shift the focus from small innovative groups, like the famed “Rad Lab” that proved so effective in producing bleeding-edge technology in the fog of World War II, to major organizations that were once part of the military-industrial complex and now stem largely from the corporate world dominating Silicon Valley.
The innovative “rad lab”—MIT’s Radiation Laboratory
Big funding and Big Science also inevitably centralize research under hierarchies of management, even as generations of studies from business schools around the globe have shown that smaller groups with more freedom tend to innovate more quickly and effectively.3 And the very nature of invention and innovation means that a top-down agenda can find itself out of step with realities that change on the ground, as research and results change the understanding of the phenomenon investigated. AI? Big Science on steroids; problems included.
Frontier AI increasingly resembles this model, and adds yet another problem: it is not merely large-scale science. AI today is large-scale science fused to venture capital, cloud infrastructure, platform economics, and geopolitical competition. It’s a technology that has been quickly and perhaps carelessly woven into the fabric of society at pressure points both civilian, commercial, academic, bureaucratic, and military. The result looks less like an open scientific inquiry into intelligence, than a race to build ever more expensive systems whose value must be justified in commercial and strategic terms. We can, in theory, “defund” Big Science. We cannot defund AI.
This changes the character of the enterprise, evidenced by contemporary worries and discussions about AI. Media sources ostensibly agnostic about the value of AI, like The Economist, now routinely run pieces focusing not on the science but the logistical and financial aspects of the field.
Depressingly, the questions are “Big Money,” rather than idea-driven too: Who has the compute? Who has the data centers? Who can afford the chips? Who can pay the engineering teams, absorb the training costs, and sustain the burn long enough to remain at the frontier? AI has become a social and cultural bandwagon that we moderns cannot help but join. Yet the levers of power remain in the hands of the few.
One consequence here is that certain research directions become self-reinforcing, not necessarily because they are theoretically deepest, but because they are the ones that can absorb capital and produce visible benchmarks, demos, and products.
Scaling, for instance, is especially attractive in this environment, even as top researchers have begun abandoning it. It is legible to investors, journalists, and internal management: more parameters, more tokens, more compute, better benchmark scores. These are measurable, reportable, and fundable. They fit the logic of the industry.
What becomes harder to support are alternative approaches that do not scale straightforwardly, or that require more conceptual risk than financial magnitude. A small team with a new theory of learning or reasoning may have interesting ideas, but if the field’s center of gravity is moving toward trillion-parameter systems and industrial training runs, then those ideas struggle to compete for attention. Not because they are false, but because they are structurally outmatched by the incentives of the moment.
In an earlier period, one could plausibly speak of artificial intelligence as a scientific aspiration, even if that aspiration was often confused or overstated. The question was: what would it mean to build an intelligent system? Today, at the frontier, the operative question is often more practical and more industrial: how do we train larger and more capable foundation models, deploy them across markets, defend the moat, and capture value before competitors do the same?
Those are business questions before they are scientific ones.
If that is where AI now lives, then we should stop pretending the field is still best understood as a neutral, open-ended search for intelligence in the abstract. At the frontier, it is increasingly a competition among giant institutions to industrialize one particular vision of machine cognition.
Which means the central question is no longer just whether the systems work. It is also who defines what counts as working, what counts as intelligence, and which alternatives never receive the resources to be tried.
AI has always been, and will remain, a military technology
For all the talk about productivity and creativity, one of the primary drivers of AI has always been defense. Or war.
“Computers” were once human accountants—historically women—who calculated ballistics and artillery trajectories for the military by hand using differential equations and lookup tables. They transferred data onto punch cards, recorded results on paper, and later transferred them onto punch cards for tabulation by machines such as those developed by IBM and its predecessors. By the 1940s, a young mathematician named Alan Turing was working on codebreaking at Bletchley Park, where machines like the Colossus were used to help decipher German encrypted communications in the Second World War.4 Norbert Wiener, of cybernetics fame, and other mid 20th century luminaries like Vannevar Bush, Claude Shannon, and Julian Bigelow were seeking automated methods for flying airplanes (autopilot), and shooting them and missiles out of the sky (antiaircraft guns).
By the war’s end, John von Neumann would spur the development of programmable computers with storage (the so-called von Neumann architecture) to help compute the blast radii of nuclear weapons. A decade later, a commercial version of the early ENIAC computer, known as the UNIVAC, would help the Census Bureau process census data. IBM would develop the Univac into business machines, like the IBM 701.
But it all started with the needs of the military and the military-industrial complex. And, as artificial intelligence took root as a field of study in the 1950s and 60s, the military would continue to pour millions into AI to develop Cold War-era systems for fully automated machine translation and early command-and-control and surveillance systems.
I was funded by DARPA—I should know.
Fast forward to today, and the now four-year-long war in Ukraine makes the same point. Cheap, widely available drones, combined with real-time data processing, computer vision, and constantly evolving targeting systems are reshaping the battlefield (Ukrainian-designed drones are also increasingly used in the current Iran conflict). Systems costing thousands of dollars are now disabling or destroying artillery and other bread-and-butter military equipment that costs millions or billions. But “AI” confers an advantage on the battlefield, and the range of its uses will no doubt grow.
What makes this possible is not “intelligence” in any deep sense, but rather data-driven pattern recognition:
identifying targets from visual feeds
detecting and tracking movement
adjusting trajectories in real time
This is artificial intelligence redefined as statistical inference applied at scale. And unlike ongoing discussions about, say, the future of commercial self-driving cars, the evidence emerging from the battlefield confirms again and again that AI and war are a perfect fit.
But the problem, again, is that while AI has found a home on the battlefield, the conditions for success have little to do with intelligence—ostensibly the goal of artificial intelligence—and much to do with reliably mapping sensor input to outputs under uncertainty. In plain terms, AI seems good at killing enemies of the state.
Any workable “AI” will likely find its way into the military. Today’s Big Data/Big Iron AI is form-fit for the battlefield, where it will not only persist but become more central to future conflicts.
Prediction is not explanation
The history of AI is a history of succumbing to technical and conceptual challenges and narrowing the field to what “works.” By the 2010s, what worked was, in essence, black-box prediction from massive data and compute.
This vision of AI would have made little sense to the pioneers of the field, who dreamed of discovering the Rosetta stone of intelligence and programming it on a computer. Far from any such Rosetta stone, we’ve now successfully redefined AI not even as machine learning in general, but as a particular type of machine learning known as neural networks (technically: Artificial Neural Networks, or ANNs). With huge increases in data and compute, neural networks were rebranded this century as “deep neural networks.”
Yet researchers have known for decades that ANNs are poor candidates for true intelligence, since they have a notorious optimization problem in learning—the “local minima” problem associated with gradient descent. In layman’s terms, we can never know if the nets will converge on the correct answer, or one that merely “looks” correct given the flawed convergence of the gradient-descent training. Researchers liken this to stopping at a mountain lake, rather than reaching the summit.
Curiously, other forms of machine learning, like so-called wide margin classifiers, called Support Vector Machines (to take one example) don’t suffer from local minima problems and are mathematically guaranteed, if they converge at all, to converge on the globally optimal solution. No matter; other forms of machine learning, whatever their mathematical properties, were quickly abandoned when big data and big compute proved that the neural networks performed better anyway.5
Let’s hear it—not for theory, but for gobs of data from Flickr, or where have you.
NNs, in other words, undercut one of the principled reasons to use machine learning methods based on mathematics in the first place: guarantees of an optimal solution given some solution space. Though clever tricks like dropout have helped mitigate the bugbear of local minima, there is no general theoretical solution to the problem of non-optimal learning in neural networks.
So why does today’s generative AI work so well on so many problems? The answer is, perhaps ironically, Big Data and Big Compute—or the scaling hypothesis. At sufficient scale, a deep neural network’s loss landscape contains many acceptable solutions, and gradient descent reliably finds one of them. Global optimality is no longer required at scale.
But generative AI is sequential, and it cares only about the next token given a sequence. In some domains, the question of truth versus probability may be less germane. But with language models, those “next tokens” are answering questions, holding conversations, and writing your boss an email. Truth is never guaranteed by next-token probability. In other words, scaling may mitigate optimization failures, but even well-optimized language models ignore truth by design. Witness the now notorious “hallucinations” or “confabulations.”
Add to this that ANNs are perfect black boxes, opaque to human inspection. This has gone so far in recent years, with the advent of LLMs, that even experts who design and train these systems admit they don’t really know why they work.
Epistemological foundation for a new age? Hardly.
Black-box AI is a pale substitute for what we once aimed for: machines grounded in our best, most elegant theories of intelligence and cognition. Instead, the pat answer to questions about future performance is simply to add more data, or “scale.” The so-called “scaling hypothesis” has proven inadequate for squeezing more out of these new, energy-hungry black boxes. Prediction without explanation has been and will continue to be inadequate for a serious computational science. Or for AGI, for that matter.
And in the meantime, we’re no closer to understanding anything substantive about intelligence at all.
We don’t know what intelligence is, but we’re acting as if we do
Would Albert Einstein score meaningfully higher on an IQ test than members of MENSA? If not, it’s not clear what the test measures. Was Pablo Picasso “less intelligent” than J. Robert Oppenheimer? Would an IQ test decide? Unlikely.
There is no settled theory of intelligence. Over decades, cognate fields—cognitive science, neuroscience, machine learning—have uncovered partial accounts, but no unified theory of what intelligence is, how it works, or how it should be measured has emerged.
Psychometrics is the study of intelligence through tests—IQ scores, SATs, GREs, and the like. The measures are useful, but contested. After a century of work, the consensus is modest: the tests capture something about intelligence, but not everything.
The ballyhoo over proving how “smart” AI is has intensified the problem. AI researchers rely on benchmarks: curated collections of problems presented as questions, code prompts, or reasoning tasks, paired with answer keys or clear grading rules. The model is run against this fixed suite, and its outputs are graded according to predefined criteria.
Putative benchmark tests have proliferated of late. Consider MMLU (Massive Multitask Language Understanding), GSM8K (grade-school math word problems), and HumanEval (code generation via unit tests). More recent efforts like HELM attempt to aggregate performance across tasks into a single profile. Generic benchmark tests like François Chollet’s ARC (Abstraction and Reasoning Corpus), intended to measure something akin to what psychometrics calls “fluid” g—the ability to infer rules and solve novel problems from minimal examples—are also popular.
These are not trivial tests. But they remain closed-world evaluations, where the test problems are specified in advance, the scoring rules are fixed, and the space of acceptable answers is known.
This matters because success on a benchmark does not establish a general capacity. It establishes competence under particular conditions. A model that scores highly on MMLU has learned the statistical structure of question-answer pairs drawn from its training distribution. A model that performs well on GSM8K can reproduce solutions for a class of problems it has effectively seen before.
A model that passes HumanEval can generate code that satisfies software unit tests, which themselves define what counts as correctness. Even tests for commonsense or “general” intelligence show the telltale pattern of teaching to the test: scores on Chollet’s ARC have risen sharply as researchers optimize against the test, yet the systems are not acquiring commonsense.
The problem is not that benchmarks are useless. The problem is the inference we draw from them. We move from:
“the system performs well on this task”
to
“the system is intelligent in the general sense”
That step is not justified by the evidence.
In psychometrics, we at least proceed under the assumption—contested but operational—that different tests are imperfect measures of a latent general intelligence, often denoted g. In AI, we lack even that. There is no agreed-upon underlying construct that the benchmarks are measuring. Instead, researchers build systems to score highly on accepted benchmarks, and report those scores as indicative of underlying intelligence.
This is why benchmark gains so often fail to translate into robust real-world competence. When the problem changes—when the task is underspecified, when the data distribution shifts, when the criteria for success are not fixed in advance—performance degrades in familiar ways. The system produces answers that are locally plausible but globally meaningless.
Benchmarks are instruments. They measure what they are designed to measure. But they are not theories, and they do not justify claims about general intelligence. And treating them as such is not progress.
This confusion has been with the field from the beginning. If a machine can play chess, recognize speech, translate text, classify images, or solve a set of reasoning tasks, then those capabilities are taken as local stand-ins for intelligence. The move is understandable, but the danger is that the proxy hardens into the concept itself. We stop saying, “This system performs well on a narrow task we associate with intelligence,” and start saying, “This system is intelligent.”
That slippage is not merely linguistic, but unfortunately has shaped the entire public understanding of the field. Researchers at organizations like Google, Meta, OpenAI, and Anthropic are no doubt aware of this tension. But neither they nor their employers have much incentive to clarify it for a public eager for the next breakthrough.
The problem becomes even clearer once we step outside formal testing environments. Human intelligence, or “fluid g,” involves transfer across domains, learning from sparse and ambiguous evidence, and—crucially—the ability to determine what matters in the first place. It integrates perception, memory, action, and social understanding in ways that do not reduce to fixed tasks or scoring rules. None of these fit neatly into a benchmark suite.
Nor is intelligence exhausted by abstract problem-solving. Human beings are embodied creatures acting in a world. That is why the Einstein-Picasso comparison is so revealing. It is not merely that intelligence comes in different forms, though it plainly does. It is that the word itself sits over a heterogeneous field of capacities that resist reduction to a single numerical scale. Scientific genius, artistic genius, strategic genius, social genius, mechanical genius all overlap, but they are not identical. The attempt to compress them into one latent variable may be useful for certain purposes, but it is already an abstraction from the richness of the phenomenon.
AI’s incessant focus on “intelligence”—it’s in the name—ignores a basic fact: we lack a settled theory of intelligence in the first place. What are we testing? We end up defining machine intelligence by whatever machines currently do well. We talk as if building systems that perform intelligently on selected tasks gives us an account of intelligence itself. That is like mistaking an increasingly accurate map for a theory of geography. The map may be useful, but it does not explain the terrain.
A more sober view would begin from the opposite premise: we do not yet know enough. We do not know whether the current dominant methods in AI are converging on the relevant capacities, simulating some of them, or merely bypassing them with powerful statistical shortcuts. Those are very different possibilities, and a science of AI would have slowed the roll long ago. The obvious inference here is that we are not yet capable of such a science, and are settling for a dangerous simulacrum. One consequence is that governments, educators, politicians, the media, and the general public are getting snowed.
So the deepest irony may be this: at precisely the moment when public discourse is most confident that we are building intelligence, the underlying concept remains unsettled. We do not know what intelligence is, and yet we increasingly organize research agendas, investment flows, institutional priorities, and even civilizational rhetoric as though the matter were resolved.
The honest position is more demanding. Intelligence remains, in crucial respects, an open question. Any field that forgets this risks mistaking progress on proxies for understanding of the thing itself.
Erik J. Larson
The Oakland A’s did not go undefeated, but their 2002 team won 20 consecutive games, then an American League record, under Billy Beane’s data-driven approach. What followed is more telling: the methods spread across Major League Baseball, and the original advantage largely disappeared as other teams adopted the same statistical framework.
Peng, S., Kalliamvakou, E., Cihon, P., & Demirer, M. (2023). The Impact of AI on Developer Productivity: Evidence from GitHub Copilot. arXiv:2302.06590.
See, for instance, Wuchty, S., Jones, B. F., & Uzzi, B. (2007). The Increasing Dominance of Teams in Production of Knowledge. Science, 316(5827), 1036–1039; and Wu, L., Wang, D., & Evans, J. A. (2019). Large teams develop and small teams disrupt science and technology. Nature, 566, 378–382.
Colossus, developed in 1943–44, was among the world’s first programmable digital computers. Its existence was kept secret under British law until the 1970s, which helps explain why the United States’ ENIAC is often credited with this distinction. It is worth noting that Colossus did not use the von Neumann architecture with stored programs; instructions were supplied by switches and plugs.
As Chris Wiggins and Matthew L. Jones explain in How Data Happened: A History from the Age of Reason to the Age of Algorithms (2023), so-called ensemble methods, combining many different machine learning algorithms, dominated the field just prior to the “deep learning” revolution in 2012. Ensemble methods typically required large data and compute. Perhaps ironically, neural networks were still ignored, at least partly on grounds that they were too data and resource intensive. As it turned out, by the 2010s, Moore’s Law had largely erased such shibboleths of prior eras in AI.










This is a terrific piece. You state the problems and issues clearly, and address them with much relevant data. And your conclusions are very solid. This essay should be widely read.
"Mistaking an increasingly accurate map for a theory of geography" that line is going to stick. It names something I've been struggling to articulate for a while: the difference between a system that gets better at predicting outputs and a system that actually understands what it's predicting about. The map analogy works particularly well because a map can be almost perfectly accurate and still tell you nothing about why the terrain is shaped the way it is. In my work, there's a version of this that comes up constantly. A life pattern can be mapped with considerable precision. But the moment someone asks "why does this keep happening to me," you've left the map entirely. That's not a limitation of the map. That's a question the map was never designed to answer.