-
7 Alarming Warning Signs the AI Bubble Could Be Ready to Burst in 2026
Bringing the Full Picture Together
This series has traced nearly 800 billion dollars in annual hyperscaler AI infrastructure spending in Article 1, a 95 percent enterprise pilot failure rate in Article 2, a 750 billion dollar web of circular financing between Nvidia, OpenAI, and Microsoft in Article 3, and 1.2 trillion dollars in hidden lease obligations flagged by Moody’s in Article 4. Each of those articles examined one piece of the puzzle in isolation. This final article asks the question the entire series has been building toward. Taken together, do these four pieces of evidence describe a genuine AI bubble, and if so, what would its bursting actually look like.
The honest answer requires resisting both extremes that dominate public discussion. Dismissing all AI bubble concerns as reflexive skepticism from people who missed the boat ignores genuinely alarming, well-documented financial signals from serious institutions. Treating a bubble collapse as an inevitable, imminent certainty ignores substantial, equally well-documented evidence of real revenue growth and genuine underlying demand. What follows are seven specific, evidence-based warning signs, each drawn from credible financial reporting, followed by an honest look at the counter-arguments and what a genuine unwind would actually mean.
Warning Sign One: The Paper Wealth Problem
Bridgewater Associates founder Ray Dalio has issued what he describes as his most severe market warning yet, stating plainly that current conditions have pushed markets into AI bubble territory comparable to 1929 and 2000. His specific evidence is precise and easy to verify. Recent earnings from Amazon and Alphabet have been significantly inflated by unrealized investment gains from their stakes in unlisted AI companies including Anthropic, as private market valuations soared. Strip out these unrealized paper gains, and the S&P 500’s actual earnings growth rate drops sharply. Dalio’s core warning is simple and worth repeating exactly as he framed it. Stock market wealth is not cash.
Warning Sign Two: The IPO Wave Itself
Dalio identifies a second specific mechanism that has historically preceded bubble collapses. A surge in equity issuance combined with rising interest rates are the two forces that pop bubbles, and the current wave of IPOs from SpaceX, OpenAI, and Anthropic is, in his assessment, a classic warning sign. Anthropic closed a 65 billion dollar Series H round with a post-money valuation of 965 billion dollars, surpassing OpenAI, and is expected to formally launch its IPO process this fall.
Combined, the three pending mega IPOs could raise more than 200 billion dollars. Bank of America has characterized this specific pattern directly, stating that this epic IPO cycle is essentially a large-scale transfer of accumulated risk from early private investors to the public market, precisely the mechanism through which prior AI bubble style collapses have historically transmitted losses to a much broader set of investors.
Warning Sign Three: Burn Rates That Do Not Add Up
The clearest financial red flag underlying AI bubble concerns is the specific, quantifiable relationship between spending and revenue at the industry’s most prominent company. OpenAI is losing 12 billion dollars per quarter and expects 44 billion dollars in additional losses through 2029. Financial analyst Bittner summarized the arithmetic starkly, describing OpenAI as spending 2.25 dollars to make 1 dollar of revenue, and noting pointedly that no dot-com era company survived with that kind of burn rate. The comparison to the 2000 collapse is not incidental commentary. It is the specific historical benchmark analysts keep returning to.
Warning Sign Four: Extreme Revenue Concentration
A particularly concerning AI bubble signal involves how narrowly concentrated actual paying demand for AI infrastructure remains. OpenAI and Anthropic together consume roughly 70 to 80 percent of all AI compute revenue, yet both lose tens of billions of dollars annually. This concentration compounds the circular financing risk documented in Article 3 of this series. CoreWeave illustrates the downstream effect precisely. Its largest client is effectively Microsoft, purchasing capacity specifically to serve OpenAI, meaning CoreWeave’s revenue is highly concentrated in a chain that ultimately traces back to two companies, neither of which has demonstrated a clear path to profitability.
Warning Sign Five: The Debt Burden Documented in Article 4
The 1.2 trillion dollars in off-balance-sheet lease commitments and 460 billion dollars in direct debt detailed in the previous article of this series constitutes, on its own, one of the seven clearest AI bubble warning signs. Economists at the World Economic Forum have specifically flagged AI-related debt pressures as a worrying macroeconomic trend for 2026, and tech companies issued a striking 108.7 billion dollars in corporate bonds during a single recent quarter, a pace that has continued through the first half of 2026 without meaningful slowdown.
Warning Sign Six: Concentration at the Index Level
The AI bubble concern extends well beyond individual companies into the structure of the broader stock market itself. The so called Magnificent Seven technology stocks, Alphabet, Amazon, Apple, Nvidia, Meta, Microsoft, and Tesla, currently make up 33 percent of the entire S&P 500 index. AI-related investment accounted for over 90 percent of United States GDP growth in the first two quarters of the prior year, an extraordinary concentration of economic growth in a single sector. When any single theme drives this large a share of both an equity index and national economic growth simultaneously, the potential downside if that theme falters is proportionally amplified across the entire economy, not contained within the technology sector alone.
Warning Sign Seven: The National Security Bailout Framing
Perhaps the most novel AI bubble warning sign, one without a clean historical precedent from the dot-com era, is the increasing embedding of major AI companies directly into national defense contracts. Analysts have noted this could potentially lead to a future bailout request should financial conditions deteriorate sharply, since companies positioned as critical to national security infrastructure carry an implicit expectation of government backstop that purely commercial dot-com era companies never possessed. Scott Galloway has raised the same concern explicitly, noting that talk of a potential taxpayer bailout itself constitutes evidence that OpenAI lacks a sustainable financing strategy.
The Case Against a Bubble
Responsible analysis requires taking the counter-arguments equally seriously, and they are not trivial. Unlike many dot-com era companies that generated minimal revenue chasing speculative business models, today’s major AI companies show genuine, rapidly compounding revenue growth. The value of OpenAI subscriptions increased 18 percent in a recent year, while Anthropic’s Claude revenue grew nearly sevenfold over the same period.
J.P. Morgan projects 5 trillion dollars in additional AI infrastructure spending over the next four years, a figure that reflects institutional conviction in sustained demand rather than speculative excess alone. CoreWeave, despite its concentration risk, posted a substantial contracted revenue backlog, real signed commitments rather than merely aspirational projections. Chief research officer Sharyn Leaver captured the more measured institutional view precisely, noting that 2026 marks the point where the AI hype period ends as pressure to deliver real, measurable results intensifies, a description of a maturing market correcting its excesses rather than a market collapsing entirely.
Slow Deflation Versus Sharp Correction
Even among analysts convinced some form of AI bubble correction is coming, meaningful disagreement exists about its shape. Capital Economics has already observed that one narrower AI stock bubble, concentrated in smaller, less established companies, has already burst, while a larger, more consequential bubble specifically in mega-cap AI infrastructure stocks continues to grow. The firm’s own modeling anticipates a blow-off rally followed by a 21 percent S&P 500 decline once the larger AI bubble fully unwinds, a sharp correction scenario rather than a gradual one.
Benchmark’s Bill Gurley offered a similarly direct warning in March 2026, stating flatly that AI spending is about to reset. The specific trigger analysts are watching most closely is precise and observable, the moment any major hyperscaler, Microsoft, Google, Amazon, or Meta, publicly announces a cut to AI capital expenditure, an event that has not yet occurred but that multiple analysts identify as the single clearest signal an AI bubble correction has genuinely begun.
What OpenAI and Anthropic’s IPOs Could Actually Trigger
Financial analyst Joachim Klement has offered perhaps the bluntest characterization of what the pending OpenAI and Anthropic IPOs actually represent within the broader AI bubble debate, describing them as probably nothing more than a major transfer of investment risk from current private owners to retail investors, pension funds, and others willing to buy into the hype at a much later and more expensive stage of the cycle.
This framing matters considerably for anyone assessing what a genuine AI bubble collapse would mean practically. Unlike a purely private market correction, which primarily affects venture capital firms and wealthy early investors who can absorb losses, a public market collapse following these IPOs would transmit losses directly to pension funds, retail brokerage accounts, and index funds that millions of ordinary investors hold, a meaningfully different and more broadly damaging outcome than a private valuation reset alone.
What a Genuine Burst Would Mean
If the AI bubble concerns documented across all seven warning signs in this article ultimately prove correct, the consequences would extend considerably beyond the technology sector itself. Given that AI-related investment has driven over 90 percent of recent GDP growth, a sharp AI bubble correction would represent a genuine macroeconomic event, not merely a sector rotation.
Given the 33 percent index concentration in AI-adjacent Magnificent Seven stocks, the impact on retirement accounts and index funds held by ordinary investors would be immediate and significant. Given the 1.2 trillion dollars in debt and lease obligations documented in Article 4, a sharp revenue shortfall relative to expectations could trigger genuine credit stress at specific companies, particularly Oracle and CoreWeave, the two firms Moody’s already identified as facing the sharpest ratings pressure.
And given the deeply circular financing relationships documented in Article 3, distress at any single major node in that web, OpenAI, Anthropic, Oracle, or CoreWeave specifically, carries genuine potential to propagate rapidly through the tightly interconnected companies that depend on one another’s continued participation.
Conclusion
Across this five-part series, the evidence assembled points toward a genuinely mixed but increasingly concerning picture rather than a simple verdict in either direction. Real revenue growth and real infrastructure genuinely coexist alongside speculative excess, unsustainable burn rates, and dangerously concentrated financial exposure. Whether 2026 marks the beginning of the AI bubble’s gradual, manageable deflation, the kind that ultimately leaves behind genuinely valuable infrastructure the way the fiber optic bust eventually did, or a sharper, more disruptive correction triggered by the pending OpenAI and Anthropic IPOs, remains genuinely unresolved as of this writing.
What is no longer credible, based on the evidence traced across all five articles in this series, is the claim that no bubble exists at all. The specific question worth watching most closely, as multiple analysts have identified precisely, is straightforward and observable, whether and when a major hyperscaler is the first to publicly announce it is cutting AI spending. When that happens, this series suggests, the far larger unwind will already be underway.
-
Moody’s Sounds a Critical Alarm: Is Hidden AI Debt risk a Ticking Financial Time Bomb
A Warning From the Institution That Rates Trust Itself
When Moody’s Ratings, one of the three institutions the entire global financial system relies on to judge whether a company can be trusted to repay what it owes, issues a formal warning about a specific industry, markets tend to pay close attention. In July 2026, Moody’s did exactly that, stating plainly that unprecedented AI spending threatens the credit quality of six of the largest technology companies in the world, Microsoft, Amazon, Alphabet, Meta, Oracle, and CoreWeave.
The core of the AI debt risk Moody’s identified is not that these companies are spending enormous sums, a fact already well documented in Article 1 of this series. It is how that spending is being financed, and how much of it remains deliberately structured to stay off the balance sheets investors actually scrutinize.
Understanding the specific mechanics of this AI debt risk, and why Moody’s chose this particular moment to sound the alarm, requires examining three distinct categories of exposure: direct corporate debt, off-balance-sheet lease commitments, and the bond market’s own increasingly nervous response to absorbing all of it at once.
The 460 Billion Dollar Direct Debt Figure
The most straightforward component of AI debt risk is direct corporate debt, borrowed money that already appears plainly on company balance sheets. According to Moody’s own analysis, direct debt across the six hyperscalers tracked in its report has reached approximately 460 billion dollars. This figure alone represents a meaningful shift for companies whose historical financial identity was built specifically on the opposite characteristic.
As Moody’s own report observes, the current moves break a decades-long Silicon Valley formula that created the world’s most valuable companies. Software cost little to replicate, yielding fat profit margins and fortress balance sheets. Generative AI, by contrast, demands a vast physical footprint, warehouses crammed with expensive and energy-hungry servers and chips, and that physical footprint is now being financed increasingly through borrowed capital rather than the internally generated cash flow that once defined these companies’ financial character.
The 1.2 Trillion Dollar Shadow
The far larger and more structurally significant component of AI debt risk sits entirely off the balance sheet, and this is where Moody’s analysis becomes genuinely alarming. Lease commitments across the six hyperscalers Moody’s tracks have ballooned to 1.2 trillion dollars, of which more than 820 billion dollars is tied to data centers that have not even finished construction yet. Moody’s accounting analysts David Gonzales and Alastair Drake calculated that an earlier snapshot of this hidden obligation, 662 billion dollars specifically tied to leases that had not yet begun among just five hyperscalers, was equivalent to 113 percent of those companies’ most recent adjusted debt, larger than everything already sitting openly on their balance sheets combined.
The accounting mechanism behind this AI debt risk is legal and well understood, but its scale is what has changed dramatically. Rather than owning every new AI data center outright, hyperscalers increasingly sign long-term leases with specialized infrastructure developers. Under generally accepted accounting principles, these lease commitments are not required to appear as current liabilities until the underlying data center actually begins operating.
Moody’s is explicit that this does not constitute deception, a Moody’s spokesperson clarified directly that this is not a case of companies avoiding a liability through structuring, simply that the obligation has not yet reached the balance sheet under standard accounting timing rules. Nonetheless, Moody’s treats these lease commitments as debt-equivalent liabilities, obligations that will bind these companies to substantial rent payments for years regardless of how AI revenue actually develops, and a separate investigation by Nikkei Asia Review found that off-balance-sheet obligations across five hyperscalers have surged eightfold in just four years to 1.65 trillion dollars, a figure that now exceeds their combined on-balance-sheet debt entirely.
The Bond Market Is Already Showing Fatigue
Perhaps the clearest real-time signal of genuine AI debt risk comes not from Moody’s report itself but from how the corporate bond market has responded to absorbing this wave of new borrowing. S&P Global calculated that hyperscalers and closely related entities including Nvidia issued 225 billion dollars in bonds during just the first half of 2026, a 973.7 percent increase compared to the same period the prior year, and they remain on pace to issue roughly 400 billion dollars for the full year.
Corporate bond issuance from technology firms specifically exceeded 108.7 billion dollars in a single quarter of 2026, a volume that would have been considered extraordinary for the entire sector across a full year just three years earlier.
The market’s appetite for absorbing this AI debt risk is showing visible strain. S&P Global’s own analysis notes that hyperscalers are now paying a meaningfully higher premium compared with yields on risk-free government bonds than they were previously required to pay, direct evidence that bond investors are demanding greater compensation for what they perceive as rising risk. As S&P put it directly in its own report, market participants are growing leery of quickly rising leverage from issuers previously characterized by strong and reliable cash flow, a notably blunt assessment from an institution not generally given to dramatic language.
Alphabet’s Negative Free Cash Flow Quarter
The clearest individual illustration of how this AI debt risk translates into immediate market consequences arrived when Alphabet reported its first negative free cash flow quarter since its initial public offering, an event that stunned even seasoned analysts given the company’s historical reputation for financial conservatism, despite Google Cloud revenue simultaneously surging 82 percent. Alphabet’s stock dropped 7 percent on the news.
The company subsequently raised its 2026 capital expenditure guidance to 205 billion dollars and announced an 85 billion dollar stock offering, one of the largest equity raises ever undertaken by a technology company, a clear signal that even one of the cash-richest companies in corporate history is now seeking additional financial flexibility specifically to sustain its AI infrastructure buildout.
Average free cash flow margins across the hyperscaler group have compressed from roughly 28 percent down to 11 percent, a genuinely dramatic deterioration in the underlying financial health metric that has historically distinguished these companies from more conventional, capital-intensive industrial businesses.
Who Faces the Sharpest Credit Rating Pressure
Not every company carrying AI debt risk faces equal exposure, and Moody’s analysis draws a meaningful distinction worth understanding precisely. Oracle and CoreWeave face the most immediate ratings pressure among the six companies tracked, reflecting their comparatively weaker underlying balance sheets and heavier relative reliance on debt financing to fund their AI infrastructure commitments.
By contrast, the four largest players, Microsoft, Amazon, Alphabet, and Meta, retain what Moody’s characterizes as fundamentally strong balance sheets even accounting for this new leverage, a distinction that matters considerably for anyone assessing which parts of this AI debt risk landscape represent genuine near-term vulnerability versus which represent a more manageable, if still historically unusual, shift in capital structure among companies with substantial existing financial cushion.
CoreWeave in particular illustrates the sharper end of this risk spectrum concretely. As documented in Article 3 of this series, CoreWeave’s own credit default swaps have briefly implied pricing consistent with something close to a coin flip probability of default, a striking market signal for a company whose infrastructure underpins a meaningful share of current AI compute capacity.
The Circular Revenue Complication
Moody’s analysis explicitly connects this AI debt risk to the circular financing dynamics examined in Article 3 of this series, flagging what it calls a circular AI ecosystem in which tech giants invest directly in AI labs including OpenAI and Anthropic, which then route significant portions of that same capital back into purchasing cloud services from the very companies that funded them.
Moody’s view is that this concentration creates a specific and identifiable systemic vulnerability, most of the AI infrastructure spending documented across this entire series is ultimately serving a remarkably small number of end customers, with OpenAI and Anthropic alone representing an outsized share of total demand, and if AI adoption falls meaningfully short of current market expectations, the concentration of debt among this small number of interconnected companies could trigger broader financial pressure that spreads well beyond any single firm.
Moody’s also flagged a specific structural timing risk embedded directly in this AI debt risk picture, a two to three year lag between when capital is actually spent on data center construction and when corresponding AI-related revenue is realized. That lag means the true test of whether this debt was prudently deployed will not arrive immediately, and current financial statements cannot yet definitively confirm or refute whether the underlying investment thesis is sound.
A Reassurance Worth Taking Seriously
It would be inaccurate to characterize Moody’s report as predicting imminent financial collapse, and the agency’s own careful language deserves to be represented faithfully. Moody’s explicitly states that hyperscalers still maintain some of the most robust balance sheets in the entire corporate world, and their investment grade ratings, while under increased scrutiny, remain intact for the four largest players specifically.
The off-balance-sheet lease commitments driving much of the headline AI debt risk figure are legitimate, disclosed practices under standard accounting rules, not hidden liabilities in any deceptive sense, and much of what currently sits off balance sheets will simply migrate onto them naturally as data centers begin operations over the coming years, a normal and expected accounting transition rather than a hidden financial trap.
Conclusion
The AI debt risk Moody’s has documented in careful, methodical detail is neither a prediction of imminent catastrophe nor a dismissible non-issue. It is a precise, quantified description of a genuine structural shift, 460 billion dollars in direct debt, 1.2 trillion dollars in off-balance-sheet lease commitments, and a bond market already showing visible signs of fatigue after absorbing an unprecedented volume of new issuance in an extraordinarily compressed timeframe.
Whether this AI debt risk resolves smoothly as the anticipated two to three year revenue lag closes, or whether it becomes the mechanism through which the broader AI investment cycle experiences genuine financial stress, is precisely the question Article 5 of this series turns to directly, examining whether the full picture assembled across this series, staggering infrastructure spending, an unresolved ROI crisis, a circular financing web, and now a mounting debt burden, adds up to a genuine AI bubble approaching its limits.
-
Inside the Alarming AI Circular Financing Web: How Nvidia, OpenAI, and Microsoft Fund Each Other’s Growth
The Web Bloomberg Mapped
In January 2026, Bloomberg published a detailed graphics investigation that gave a name and a visual shape to a pattern industry watchers had been describing in increasingly alarmed terms for months. The map traces roughly 46 billion dollars in direct equity stakes and 879 billion dollars in multi-year purchase commitments moving between Microsoft, Oracle, Amazon, Google, Meta, OpenAI, Anthropic, xAI, CoreWeave, Nvidia, and AMD. At the center of this AI circular financing web sits Nvidia, whose market valuation reached 5.4 trillion dollars in mid-2026, a position investor Michael Burry has publicly described as sitting dead center of the entire structure.
The nearly 800 billion dollars in annual hyperscaler infrastructure spending documented in Article 1 of this series does not appear from nowhere; a significant share of it flows directly through the AI circular financing relationships examined here. Understanding whether this AI circular financing arrangement represents rational supply chain coordination in a genuinely constrained market, or the same structural warning sign that has preceded prior financial bubbles, requires tracing the actual mechanics of the deals, examining both sides of a genuinely contested debate among serious analysts, and being honest about what nobody yet knows.
How the Loop Actually Works
The mechanics of AI circular financing are, once traced carefully, straightforward enough to describe in a single sentence, even if the dollar figures involved are difficult to comprehend. Microsoft invested more than 13 billion dollars in OpenAI over several years. OpenAI committed to spending 250 billion dollars on Microsoft’s Azure cloud services. Oracle is constructing 300 billion dollars in Stargate data center infrastructure specifically for OpenAI under long-term contracts. Nvidia invested 30 billion dollars in OpenAI’s most recent 122 billion dollar funding round, while OpenAI simultaneously remains one of Nvidia’s largest chip customers. Nvidia has separately taken equity stakes in CoreWeave and other so-called neocloud providers, companies that are themselves major customers for Nvidia’s chips.
The pattern repeats with variations across nearly every major relationship in the industry. OpenAI’s total named compute commitments now sum to more than a trillion dollars across Azure, Oracle, AWS, CoreWeave, Nvidia, Broadcom, and a six gigawatt AMD deal running through 2035, a deal structured so that OpenAI is poised to become one of AMD’s largest shareholders. Amazon’s 50 billion dollar investment in OpenAI’s March 2026 funding round was structured partly as compute credits, with OpenAI simultaneously committing to spend 100 billion dollars on AWS over eight years.
Google agreed to backstop lease payments at five separate data center locations for Anthropic, effectively helping Anthropic obtain what amounts to a 35 billion dollar loan. Money moves from investor to startup and back to the investor’s own products and services through a loop that is, by construction, self-reinforcing.
The 750 Billion Dollar Escalation
Rather than slowing amid growing scrutiny, this AI circular financing pattern accelerated sharply through mid-2026. Nvidia is now working on a fresh round of infrastructure deals potentially worth more than 750 billion dollars. A partnership unveiled with South Korean conglomerate SK Group in late July 2026 alone represents more than 500 billion dollars in mutual business, tied to building more than two gigawatts of AI data centers on the Korean peninsula, enough electricity to power roughly 1.5 million homes.
More striking still, the Wall Street Journal reported on July 27, 2026, that Nvidia is in talks to provide a 250 billion dollar financing guarantee tied to OpenAI leasing a portion of a planned 500 billion dollar data center project in southern Ohio, led by SoftBank’s energy arm. The same reporting indicates the guarantee would help SoftBank raise debt on more favorable terms than it could obtain otherwise, precisely because OpenAI itself does not currently hold an investment grade credit rating.
Nvidia may also separately help finance roughly 350 billion dollars in chip purchases from OpenAI under a related arrangement. The market reacted immediately and visibly. Nvidia shares fell 4.5 percent on the news, and the price of credit default swaps on Nvidia’s own bonds, effectively a form of default insurance for bondholders, recorded their highest single day increase since active trading in the instrument began.
Jensen Huang’s Direct Rebuttal
Nvidia CEO Jensen Huang has responded to AI circular financing criticism with characteristic directness rather than deflection. Asked specifically about the vendor financing charge as Bloomberg documented the growing 750 billion dollar deal total, Huang stated flatly, the idea that it is circular is ridiculous. His underlying argument, echoed by supporters of the current deal structure across the industry, is that building frontier AI infrastructure is extraordinarily expensive and that the most advanced chips remain genuinely difficult to obtain even now.
In that kind of constrained market, Huang and his allies argue, companies do not simply place purchase orders and wait. They lock in scarce supply by pairing long-term buying commitments with financing, a practice with long precedent in genuinely capital-intensive industries from telecommunications to energy infrastructure.
The Virtuous Circle Counter-Argument
This defense of AI circular financing has a specific and influential institutional champion. Asset manager Janus Henderson has characterized the current wave of AI dealmaking as more accurately described as a virtuous circle, one that helps line up suppliers, builders, and customers to meet what the firm characterizes as genuinely exploding demand for computing power. Under this framing, what critics label circular financing is simply the efficient coordination mechanism a young, capital-intensive, rapidly scaling industry requires to align capacity investment with demand that outstrips what any single company could finance independently through conventional means.
There is a genuine kernel of truth in this defense that deserves acknowledgment. CoreWeave, one of the clearest examples of a company deeply embedded in this AI circular financing web, is at least a public company whose filings provide real numbers rather than speculation. Its first quarter 2026 results showed 2.08 billion dollars in revenue against a 740 million dollar net loss, alongside nearly 100 billion dollars in contracted revenue backlog. That backlog, if it converts to actual delivered revenue over time, represents real economic activity, not merely accounting fiction circulating between related parties.
Why Serious Analysts Remain Alarmed Regardless
Set against these defenses, a growing chorus of serious market analysts continues to treat AI circular financing as a genuine structural risk, and their concern rests on a specific, carefully stated argument rather than blanket skepticism of AI itself. As one detailed industry analysis put it precisely, none of this has to be fake to be dangerous. The revenue can be entirely real, the chips can actually ship, and the data centers can genuinely get built, all while resting on a financing structure in which the same small handful of companies are effectively supporting one another’s demand, obscuring how much of the total activity reflects genuine, independent end-user demand versus intra-industry financial engineering.
Michael Burry, the investor who famously anticipated the 2008 mortgage crisis, has invoked the AI circular financing pattern repeatedly and pointedly in 2026, sharing Bloomberg’s own diagram of the deal web as evidence of a structure he considers genuinely precarious. This concentration of financial risk sits uncomfortably alongside the AI ROI concerns detailed in Article 2 of this series, since much of the revenue circulating through this web has yet to translate into the kind of measurable enterprise value that would justify its scale.
Harvard Kennedy School senior fellow Paulo Carvao has drawn an explicit historical parallel to the late 1990s technology bubble, noting that circular deals during that era were often centered on advertising and cross-selling arrangements between startups, where companies bought each other’s services specifically to inflate the appearance of genuine growth. The concern is not that the parallel is exact in every detail, but that the underlying structural vulnerability, revenue and valuation that depend heavily on continued participation by a small, tightly interconnected group of counterparties, rhymes closely enough with prior bubble dynamics to warrant serious caution.
The Credit Risk Dimension
The AI circular financing debate connects directly to a parallel and increasingly urgent concern that will be examined in full in Article 4 of this series. Moody’s has explicitly warned that the scale of AI related spending threatens the credit quality of Microsoft, Amazon, Alphabet, Meta, Oracle, and CoreWeave specifically. CoreWeave’s own credit default swaps have briefly implied pricing consistent with something close to a coin flip probability of default.
Anthropic’s most recent major compute contract required underwriting through a bank letter of credit rather than being supported directly by its own balance sheet, a structural detail that suggests even sophisticated market participants are not fully confident in the standalone creditworthiness of companies deeply embedded in this AI circular financing web.
What Happened During the Late July Selloff
The genuine market sensitivity to AI circular financing concerns was demonstrated directly during the final week of July 2026. A significant amount of market value was wiped from global chip and AI hardware stocks between July 24 and July 29, 2026, coinciding with the intensified scrutiny of Nvidia’s expanding deal book. Notably, most of that lost value was recovered within the following week, a pattern that itself illustrates the deeply contested nature of this debate.
Investors sold first on the circularity concern, then substantially reversed course, suggesting the market itself remains genuinely undecided about whether AI circular financing represents a serious systemic vulnerability or simply the necessary financial architecture of a capital-intensive industry scaling at unprecedented speed.
Conclusion
The AI circular financing web documented by Bloomberg, and expanding rapidly through 2026 with Nvidia’s 750 billion dollar deal book at its center, is neither obviously fraudulent nor obviously benign. It is a genuinely novel financial structure, real revenue and real infrastructure resting on a foundation of relationships concentrated among a remarkably small number of counterparties, each simultaneously acting as the others’ customer, supplier, and investor.
Jensen Huang calls the circularity framing ridiculous. Michael Burry calls it a warning sign serious enough to invoke repeatedly and publicly. Both cannot be straightforwardly right, and the honest answer, at least for now, is that the structure has not yet been tested by the kind of demand slowdown or credit event that would definitively reveal which characterization is closer to the truth. Article 4 of this series turns directly to that credit risk dimension, examining Moody’s specific warnings and what a genuine stress event within this AI circular financing web would actually look like for the broader economy.
-
The Startling Truth About AI ROI: Why 95 Percent of Enterprise Projects Are Failing
A Number That Refuses to Go Away
Since its publication in mid-2025, one statistic has become the single most repeated, most contested, and most consequential figure in the entire enterprise AI conversation. MIT’s Project NANDA, in a report titled The GenAI Divide: State of AI in Business 2025, found that 95 percent of generative AI pilots deliver no measurable profit and loss impact. Only 5 percent of integrated AI systems create significant, measurable value.
Given the nearly 800 billion dollars in AI infrastructure spending documented in the first article of this series, the AI ROI question this statistic raises is not academic. It is the question on which the entire economic justification for the current investment cycle ultimately rests.
Understanding whether this AI ROI crisis is real, overstated, or something more nuanced requires examining the methodology behind the headline number, the deeper productivity paradox it sits inside, and, most usefully, exactly what separates the small minority of companies that are succeeding from the large majority that are not.
Inside the MIT Report
The GenAI Divide report, based on 52 executive interviews, a survey of roughly 150 business leaders, and analysis of 300 public AI deployments, draws a sharp distinction the authors call the GenAI Divide, a split between widespread adoption and genuine business transformation. Over 80 percent of organizations have piloted tools such as ChatGPT or Copilot, and nearly 40 percent report some form of deployment. Yet these systems overwhelmingly boost individual productivity rather than delivering measurable enterprise level AI ROI.
The report identifies four structural factors behind this divide. Disruption remains limited to just two of nine major sectors, technology and media, that show genuine business transformation from generative AI use. Large enterprises paradoxically lead in pilot volume but lag significantly in successful deployment, while mid-market companies move from pilot to full implementation in roughly 90 days compared to nine months or longer at large enterprises.
AI budgets are allocated in a way that actively works against AI ROI, with over 50 percent of spending in 2025 directed toward sales and marketing pilots, precisely the category the report finds delivers the weakest returns, while the strongest AI ROI consistently comes from back office automation in finance, compliance, and document processing, categories that receive comparatively little budget attention. Finally, tools built by external vendors succeed roughly twice as often as internally built systems, a genuinely important finding for any enterprise weighing a build versus buy decision.
Perhaps the most striking finding is the emergence of what the report calls a shadow AI economy. While only 40 percent of companies maintain official LLM subscriptions, roughly 90 percent of workers surveyed report daily use of personal AI tools such as ChatGPT or Claude for actual job tasks, tools that frequently deliver better performance and faster adoption than the sanctioned systems built specifically to replace them.
The Methodology Question Worth Taking Seriously
Before accepting the 95 percent AI ROI failure figure uncritically, it is worth noting that the report itself has faced genuine methodological scrutiny. The finding of zero measurable return was based on just 52 interviews that the report’s own authors describe as directionally accurate based on individual interviews rather than official company reporting. Marketing AI Institute founder Paul Roetzer has argued publicly that a closer reading of the study’s methodology reveals a considerably more nuanced picture than the viral headline suggests, noting the sample size and self-reported nature of much of the underlying data.
This caveat does not invalidate the broader AI ROI concern, particularly because the MIT figure has since been corroborated, directionally if not precisely, by entirely independent research using different methodologies. Gartner separately predicts that over 40 percent of agentic AI projects will be cancelled by the end of 2027, citing escalating costs and unclear business value.
RAND Corporation’s independent research puts the broader AI project failure rate above 80 percent, roughly double the failure rate of conventional enterprise IT projects, itself a meaningful baseline given how notoriously difficult large-scale enterprise software rollouts already are. When multiple independent research organizations using different methods converge on directionally similar conclusions, the underlying AI ROI concern deserves to be taken seriously even if the precise 95 percent figure carries some uncertainty.
S&P Global and the Abandonment Crisis
A separate and independently sourced data point adds further weight to the AI ROI concern. S&P Global Market Intelligence, surveying over 1,000 IT and business leaders across North America and Europe for its 2025 Voice of the Enterprise report, found that 42 percent of companies abandoned most of their AI initiatives in 2025, a dramatic jump from just 17 percent the prior year. The average organization scrapped 46 percent of its proof of concept projects before they ever reached production.
The mechanism behind this abandonment pattern is instructive for understanding the AI ROI problem more precisely. Organizations that budget six months for an AI project typically allocate roughly five months to building the AI capability itself and only one month to what practitioners call productionization, the unglamorous but essential work of hardening a system for real operational use.
Production infrastructure, if built properly, takes roughly as long as the AI capability itself. By the time this reality becomes apparent, usually around month five, the project is over budget, behind schedule, and executive confidence has eroded. The project gets abandoned, not because the underlying AI capability failed, but because the operational foundation required to sustain it in production was never adequately budgeted for in the first place.
The Productivity Paradox: Real Gains That Vanish at Scale
Perhaps the most conceptually important dimension of the AI ROI debate is what researchers now call the AI productivity paradox, the widening gap between clearly documented task level gains and the near invisible effect of those same gains on company wide and national productivity statistics. The paradox is genuinely puzzling because both halves of it are independently well supported by evidence.
Customer service agents using AI resolve 14 percent more issues per hour. GitHub Copilot users complete coding tasks 55 percent faster. BCG consultants using AI finish work 25 percent quicker with 40 percent higher quality scores. These task level AI ROI gains, ranging from roughly 14 to 55 percent depending on the specific study and task, are real, controlled, and repeatedly replicated.
And yet, at the organizational level, this AI ROI evaporates almost entirely. NBER researchers tracking AI adoption from 61 to 71 percent of surveyed firms between early 2025 and early 2026 found that 89 percent of managers reported no change whatsoever in sales volume per employee over that same period. Only 39 percent of enterprises can trace any measurable EBIT impact to their AI investments at all. Nobel laureate economist Daron Acemoglu has projected a strikingly modest 0.5 to 0.7 percent total productivity gain from AI over the entire next decade, a figure he describes candidly as disappointing relative to the promises the industry has made.
The explanation researchers increasingly converge on is that task level speed is simply not the same thing as firm level throughput. An individual worker completing a task 55 percent faster does not automatically translate into an organization producing 55 percent more output, because the surrounding workflow, approval processes, quality checks, and organizational structure were never redesigned to actually capture that individual speed gain at scale.
What the Successful 5 Percent Actually Do Differently
The most practically useful finding across this entire body of AI ROI research is not the failure statistic itself but the consistent pattern separating the minority that succeed from the majority that do not. McKinsey’s 2025 AI survey found that organizations reporting significant financial returns were twice as likely to have redesigned their end to end workflows before selecting any AI tool, confirming that organizational change, not the underlying technology, is the actual differentiator.
MIT’s own data on the successful minority is similarly specific. Pilots that blended internal AI specialists with external vendor expertise achieved a 67 percent success rate, compared to just 22 percent for projects built entirely in-house. The winning 5 percent consistently shared three traits: tightly scoped initiatives focused on a single, well-defined pain point rather than broad transformation ambitions, domain specific focus rather than generic tooling, and smart partnerships with vendors who understood both the technology and the specific operational context it was being deployed into.
As one MIT report author put it directly, describing successful startups, they pick one pain point, execute well, and partner smartly, a strikingly simple formula against the backdrop of billions of dollars in more diffuse enterprise spending that has failed to replicate it.
Conclusion
The honest verdict on AI ROI in 2026 is neither the total failure the viral 95 percent statistic suggests in isolation, nor the seamless transformation the marketing around generative AI has promised since 2023. It is a genuine and well documented paradox: real, measurable, repeatedly replicated task level productivity gains that are, for the overwhelming majority of enterprises, failing to survive the jump from individual workflow to organizational output.
The 5 percent of companies that are succeeding are not doing so because they have access to better models. They are succeeding because they redesigned the underlying work itself before deploying AI into it, a lesson that costs considerably less to implement than the infrastructure billions documented in Article 1 of this series, and one that most of the market has still not learned.
-
The Staggering $775 Billion AI Infrastructure Spending Race: Where All the Money Is Actually Going
A Number Larger Than Most National Economies
In 2026, the five largest hyperscalers, Amazon, Microsoft, Alphabet, Meta, and Oracle, are on track to spend between 775 and 800 billion dollars on infrastructure, according to CFA analysis published in August 2026. To put that figure in perspective, AI infrastructure spending in the United States now represents roughly 5 percent of national GDP, a level of infrastructure commitment that analysts describe as the largest in modern economic history, 2.5 times the scale of the fiber optic overbuild of the late 1990s and three times the peak of national electrification a century earlier. This is not a niche technology investment cycle. It is a capital deployment event on a scale usually reserved for wars, railroads, and national power grids.
Understanding where this staggering sum of AI infrastructure spending is actually going, and whether the historical parallel to prior infrastructure booms is reassuring or alarming, requires looking closely at the individual commitments, the financing mechanisms behind them, and the physical constraints that are already beginning to bite.
Breaking Down the Big Five
The scale of individual hyperscaler AI infrastructure spending commitments in 2026 is difficult to grasp in isolation. J.P. Morgan estimates aggregate hyperscaler capital expenditure will reach 697 billion dollars this year, while separate analysis from Goldman Sachs projects total hyperscaler capex from 2025 through 2027 will reach 1.15 trillion dollars, more than double the 477 billion dollars spent across the entire 2022 to 2024 period. Roughly 75 percent of this spending, approximately 450 billion dollars, is directed specifically at AI infrastructure, servers, GPUs, data centers, and specialized equipment, rather than traditional cloud computing capacity.
Individual company figures illustrate the intensity of this AI infrastructure spending race. Amazon has guided to approximately 125 billion dollars in 2026 capital expenditure, a 61 percent increase over the prior year, with 64 percent of that spending allocated to AWS and AI initiatives specifically. Alphabet has guided toward 75 to 85 billion dollars. Each of the four largest hyperscalers now individually exceeds 100 billion dollars in annual infrastructure spending, a threshold that would have seemed implausible even eighteen months earlier. Capital intensity, capex as a share of company revenue, has reached 45 to 57 percent at several of these companies, a ratio historically associated with capital intensive industrial and utility companies rather than software businesses.
The Stargate Project and Government-Backed Ambition
Layered on top of individual company AI infrastructure spending is Project Stargate, a joint venture between OpenAI, SoftBank, Oracle, and MGX announced in January 2025 and publicly backed by the Trump administration, with an ambition to invest up to 500 billion dollars in United States data centers and energy infrastructure over four years. J.P. Morgan’s John Servidea, global co-head of Investment Grade Finance, described the moment plainly: AI financing is the biggest secular theme in our professional lifetimes.
The Stargate project illustrates a broader pattern within AI infrastructure spending in 2026: the blurring of lines between corporate capital expenditure, sovereign investment, and government policy. Sovereign programs beyond Stargate itself, including a 40 billion dollar commitment from Saudi Arabia’s Public Investment Fund and roughly 200 billion euros in European Union AI infrastructure ambitions, push the true global figure for AI infrastructure spending considerably higher than hyperscaler capex alone would suggest.
Financing a Buildout That Exceeds Cash Flow
Perhaps the most consequential shift within this AI infrastructure spending cycle is how it is being financed. For most of the past decade, hyperscalers funded capital expenditure primarily from internal operating cash flow, a position of financial strength that distinguished them from more leveraged industries. That era has ended. Hyperscalers issued a record 428 billion dollars in corporate bonds during 2025 alone, with projections suggesting up to 1.5 trillion dollars in additional debt issuance over the coming years as AI infrastructure spending continues to outpace what internal cash generation can support.
This transition from cash funded to debt funded infrastructure spending represents a fundamental change in the financial character of companies that were, until recently, among the most conservatively financed in the entire economy. Analysts at IEEE ComSoc noted the shift directly, observing that hyperscalers are increasingly leaning on debt markets to bridge the gap between rapidly rising AI capex budgets and internal free cash flow, transforming historically cash funded business models into ones utilizing meaningful leverage, even while balance sheets remain nominally strong for now.
The Physical Constraints Nobody Can Spend Their Way Around
A critical dimension of AI infrastructure spending in 2026 that pure dollar figures obscure is the extent to which physical, rather than financial, constraints are now the binding limitation on deployment speed. Critical supply chain bottlenecks, including high bandwidth memory, advanced chip packaging capacity known as CoWoS, and transformer lead times for electrical equipment, threaten to constrain how quickly this enormous volume of AI infrastructure spending can actually translate into operational data center capacity.
Power availability has emerged as perhaps the single most significant constraint. The scale of the AI infrastructure spending buildout has pushed hyperscalers toward power sources that would have seemed exotic for a technology company just a few years ago. Meta’s nuclear power purchase agreement, Amazon’s expanding nuclear power offtake commitments, and Microsoft’s agreement to restart the Three Mile Island nuclear facility all confirm that nuclear power has become an operational requirement for AI infrastructure at this scale, not merely an environmental preference. This same theme, examined in detail in our earlier coverage of AI data centers and their environmental impact, is intensifying rather than resolving as spending accelerates.
The Historical Parallel: Reassuring or Alarming
The comparison between current AI infrastructure spending and prior infrastructure overbuild cycles cuts in two directions simultaneously, and reasonable analysts disagree sharply about which direction should dominate the interpretation. On one hand, every prior infrastructure overbuild cycle identified by historical analysis, the railroad network of the 1880s, the national electrical grid built around 1929, and the global internet backbone constructed during the fiber optic boom of the late 1990s, despite producing bankruptcies, market crashes, and significant excess capacity in the near term, ultimately produced infrastructure that became genuinely foundational to the next era of economic productivity.
Under this framing, current AI infrastructure spending, however excessive it may appear relative to near-term AI revenue, may simply be the necessary and historically consistent overbuilding phase that precedes durable long-term value creation.
On the other hand, the fiber optic comparison specifically carries an uncomfortable warning that industry commentators invoke repeatedly. As one industry analysis put it directly, referencing the stupendous increase in fiber optic spending from 1998 to 2001 until that particular bubble burst, the parallel is not merely rhetorical. Fiber optic capacity built during that boom did eventually prove valuable, but only after a wrenching financial crash wiped out the equity value of the companies that built it, transferred the physical assets to new owners at steep discounts, and left an entire generation of telecom bondholders with significant losses.
Whether the AI infrastructure spending cycle of 2026 follows the same trajectory, useful infrastructure ultimately, but only after a genuinely painful financial reckoning for the companies and investors who financed the initial buildout, is precisely the question this five-part series is built to examine.
What This Means Going Forward
The scale of AI infrastructure spending documented here sets the stage for the four articles that follow in this series. Article 2 will examine whether this extraordinary capital deployment is actually generating measurable returns for the enterprises purchasing AI capability, a question where the evidence, drawn from MIT, Gartner, and RAND research, is considerably more sobering than the raw spending figures might suggest.
Article 3 will trace the increasingly circular financing relationships between Nvidia, OpenAI, Microsoft, and Oracle that are helping fund this buildout, relationships that several analysts argue obscure the true underlying demand signal for AI infrastructure spending itself. Article 4 will examine the credit and debt risk this financing structure is creating, drawing on Moody’s own recent warnings. And Article 5 will bring the full picture together to assess whether the AI infrastructure spending boom documented in this article represents durable economic transformation or a bubble approaching its limits.
Conclusion
What is beyond dispute is the sheer scale of what is being built. Nearly 800 billion dollars in hyperscaler spending in a single year, financed increasingly through debt rather than cash, chasing physical constraints in power and semiconductor supply that money alone cannot immediately solve, and layered with sovereign and government backed commitments that add hundreds of billions more to the global total. Whether this AI infrastructure spending ultimately proves as foundational as the railroads and the electrical grid, or as painful in its near-term unwinding as the fiber optic bust, is a question that will be answered not by this article, but by the years of actual demand, revenue, and repayment that follow it.
-
The Critical Future of AI Mathematics: Literature Mining, Crisis Debates, and What Comes Next (Part 2)
This is Part 2 of a series examining how AI is transforming mathematical research. Part 1 covered automated theorem proving and formal proof assistants. Part 2 examines literature mining, the peer review crisis, existential debates within the mathematical community, and the future of AI mathematics as a discipline.
From Proving Theorems to Reading Everything Ever Written
Part 1 traced how AI systems learned to construct genuinely new mathematical proofs. But an equally consequential and less discussed capability underlies much of that progress: the ability to read, search, and synthesize the entire published mathematical literature at a scale no human researcher could ever match. This capability, often described as literature mining, has become central to understanding the future of AI mathematics, and it has already produced one of the field’s more embarrassing public controversies.
In October 2025, OpenAI claimed that GPT 5 had solved ten previously open Erdős problems. The claim was publicly refuted within hours. The model had not actually solved the problems from first principles. It had performed what researchers now call a super literature search, locating previously published but obscure papers that had already resolved the problems, papers that had simply escaped the attention of the mathematicians maintaining the Erdős problem database.
A similar pattern recurred when DeepMind deployed an agent called Aletheia at the end of 2025, which attempted 700 unsolved problems from the Erdős database and correctly resolved thirteen, but only four represented genuinely new mathematical work. The other nine were, once again, successful literature searches rather than novel proofs.
This distinction matters enormously for understanding the future of AI mathematics honestly. Locating a forgotten proof buried in decades of published papers is a genuinely valuable service to the mathematical community, since human researchers cannot possibly track every result published across thousands of journals. But it is a fundamentally different capability from generating new mathematics, and conflating the two, as several early press releases did, has become a significant source of friction between AI labs and the mathematicians whose trust they need.
The Peer Review System Under Genuine Strain
The most immediate and practically consequential challenge shaping the future of AI mathematics is not a technical limitation at all. It is institutional. AI systems can now generate a large number of proofs that appear correct on inspection, often within hours, while carefully verifying a single dense mathematical argument by hand can take a human expert weeks or longer. The number of mathematicians qualified to review highly specialized proofs in any given subfield is extremely limited, and this mismatch is creating what several researchers now openly describe as a peer review crisis specific to AI generated mathematics.
The concern is not hypothetical. Multiple instances have already occurred in 2026 in which AI systems or their developers announced significant mathematical results through press releases or preprints before the claims had received adequate scrutiny, only for errors or overstatements to surface afterward under closer examination. If a substantial volume of AI generated mathematical reasoning enters circulation as preprints or public announcements faster than the community can verify it, the practical effect is not merely wasted reviewer time.
It risks burying genuinely valuable human and AI assisted discoveries under a volume of unverified claims that erodes trust in the published mathematical record itself, a concern mathematician Jeremy Avigad has documented carefully in his own 2026 survey of the field, noting that automated reasoning tools including SAT solvers have already resolved open problems in combinatorics, algebra, and discrete geometry, alongside machine learning techniques that have identified new combinatorial objects and counterexamples to standing conjectures, all while formal verification systems like Lean’s Mathlib library are increasingly used to verify results even before or entirely outside the traditional peer review process.
A Genuine Crisis Essay and the Question of Authorship
The tension within the mathematical community reached a notably sharp point in early August 2026, when a widely circulated essay titled “The Crisis of AI-Generated Mathematics” argued for what its author called total opposition to the use of artificial intelligence in mathematics.
The essay’s specific example is illustrative of the deeper concern driving the future of AI mathematics debate: a mathematician working in matroid theory, before publishing a completed solo paper, offered the project as a test case for an AI system’s ability to prove theorems and autonomously write up results, raising a question the field has not yet resolved, namely what authorship and intellectual authority even mean once AI can generate publishable mathematical content without a human necessarily understanding every step.
The essay proposes genuinely radical institutional responses, including replacing traditional individual authorship with a model of co-ownership, in which any mathematician who can demonstrate authoritative understanding of a result, the kind of deep comprehension expected of a human author today, would be recognized as a legitimate steward of that work regardless of who or what originally generated it.
Whether or not this specific proposal gains traction, its existence signals something important about where the future of AI mathematics debate has moved: from a purely technical question about capability toward a genuinely institutional question about what journals, credentialing bodies, and the mathematical community itself will need to become in response.
The Existential Framing Emerging From Within the Field
Perhaps the most striking development shaping discussion of the future of AI mathematics is the emergence, from within the mathematical community itself rather than from outside AI safety circles, of essays explicitly framing rapid mathematical AI progress as a signal of broader existential risk. One widely discussed 2026 essay observes that career defining theorems are now being proven on a weekly basis by AI systems given only minimal guidance, and notes that internal frontier models at major AI labs are reportedly producing mathematical breakthroughs in batches, with the rate of serious AI proven theorems appearing to grow exponentially through the year.
The essay’s central argument is not really about mathematics as a profession at all. It uses the visible, measurable acceleration in mathematical capability as a legible proxy for a much larger and harder to observe acceleration in general AI reasoning ability, arguing that mathematicians are uniquely well positioned to notice this signal early precisely because mathematical correctness is so much easier to verify than progress in messier real world domains.
This framing has proven genuinely divisive. Some in the mathematical community view it as an overreaction that conflates competition style problem solving with the far broader, messier work most research mathematicians actually do. Others, including voices circulating informally on social platforms suggesting that a given year’s Fields Medal might be the last one awarded primarily for human insight, treat it as a serious and urgent signal.
Terence Tao’s own more measured position, discussed in Part 1, sits deliberately between these poles, acknowledging the genuine disruption while insisting that the deeper question mathematicians must answer is what mathematical research is actually meant to accomplish, a question that predates AI entirely and that AI has simply made newly urgent rather than newly created.
What the Career Landscape Actually Looks Like
For students and early career mathematicians, the future of AI mathematics carries direct practical stakes beyond the philosophical debate. Current labour market analysis suggests the discipline is bifurcating rather than simply shrinking. Roles centred on routine computation and mechanical proof verification are being genuinely automated, while demand is rising sharply for hybrid roles, AI research scientists who blend theoretical mathematical training with practical machine learning experimentation, computational mathematicians who apply numerical and AI methods to open scientific problems, and quantitative analysts who integrate AI driven techniques into financial and risk modelling.
Compensation data suggests these hybrid roles, which explicitly combine deep mathematical fluency with programming and AI systems knowledge, currently command a meaningful premium over more narrowly traditional theoretical positions, a trend that career analysts expect to strengthen rather than reverse as the decade continues.
The clear implication for mathematics education, a question raised explicitly in university seminars examining the future of AI mathematics through 2025 and 2026, is that foundational mathematical fluency, understanding what a proof actually establishes and why, rather than merely executing computational procedures, is becoming more valuable precisely because AI has made the procedural layer nearly free.
Whether mathematics curricula adapt quickly enough to reflect that shift, moving away from testing procedures AI now performs flawlessly and toward cultivating the judgement needed to specify problems correctly and evaluate AI generated arguments critically, remains genuinely unresolved and varies enormously between institutions.
Toward a Genuinely Balanced Outlook
Bringing the full picture from both parts of this series together, the future of AI mathematics is neither the triumphant, fully automated transformation suggested by the most breathless press releases, nor the wholesale crisis threatening the discipline’s survival that the most alarmed essays describe. The verified achievements are genuinely remarkable: medal level Olympiad performance, formally verified proofs of major theorems, and at least a handful of authentically novel contributions to open research problems accepted by leading mathematicians.
The genuine problems are equally real: a peer review infrastructure straining under a volume of claims it cannot verify fast enough, unresolved questions about authorship and intellectual credit, and a small but vocal contingent within the field itself treating the pace of progress as a warning sign for something considerably larger than mathematics.
What seems most likely, based on the trajectory traced across both parts of this series, is a discipline that reorganizes around a division of labour broadly consistent with what Terence Tao has already described, humans specifying problems and exercising judgement over what mathematics is worth pursuing and why, formal systems and AI handling an increasing share of the mechanical construction and verification of proofs, and an institutional structure, journals, credentialing bodies, and peer review itself, that will need genuine reinvention rather than incremental adjustment to remain trustworthy.
Whether that reinvention happens deliberately, through the kind of proposals now circulating in essays and conference discussions, or reactively, in response to a genuine crisis of confidence in the published mathematical record, is likely to be decided over the next several years, not decades, given the pace this series has documented throughout 2025 and 2026.
Conclusion
The future of AI mathematics is being written in real time, and unusually for a technological transformation, it is being written with genuine, careful participation from the very experts most qualified to evaluate it, rather than imposed on a discipline caught unaware. That is, on balance, a reason for cautious optimism rather than alarm.
Mathematics has weathered a genuine crisis of foundations once before, a century ago, and emerged with a more rigorous, more explicit, and ultimately more resilient understanding of its own methods. Whether the current moment produces a comparable resolution, or whether the strains identified across both parts of this series prove harder to reconcile than the logical paradoxes of the early twentieth century, is a question only the coming years of actual practice, not further speculation, will be able to answer.
-
The Powerful Rise of AI in Mathematics: Automated Reasoning, Proof Assistants, and What Comes Next (Part 1)
This is Part 1 of a series examining how AI is transforming mathematical research. Part 1 covers the core contributions in automated theorem proving, proof assistants, and pattern mining, along with the limitations and open debates currently dividing the mathematical community.
A Discipline That Prided Itself on Being Unautomatable
For most of computing history, mathematics was assumed to be the last stronghold that AI would conquer, if it ever could at all. Mathematical proof requires airtight, step by step logical rigor of a kind that resists the probabilistic pattern matching underlying most machine learning systems. Yet in the space of roughly two years, AI in mathematics has moved from a curiosity discussed at specialist workshops to a subject serious enough to warrant a dedicated public lecture at the 2026 International Congress of Mathematicians, delivered by Terence Tao, widely regarded as the most accomplished living mathematician.
Tao’s framing was direct: mathematics, he argued, is entering a second crisis in its foundations, comparable in scale to the crisis triggered by Russell’s paradox and Gödel’s incompleteness theorems a century earlier, except this time the disruption comes from artificial intelligence rather than internal logical contradiction.
Understanding what AI in mathematics has actually achieved, where it genuinely struggles, and what mathematicians themselves are saying about it requires examining three distinct but interconnected fronts, automated theorem proving, formal proof assistants, and pattern mining across the mathematical literature, each of which has developed at a strikingly different pace.
Automated Reasoning: From Olympiad Silver to Erdős Problems
The most publicly visible achievement of AI in mathematics has come from competition mathematics, precisely because Olympiad problems provide a clean, verifiable benchmark. In 2024, Google DeepMind’s AlphaProof, an AlphaZero inspired reinforcement learning system, combined with AlphaGeometry 2, solved four of six problems at the International Mathematical Olympiad, achieving a score equivalent to a silver medal, the first time any AI system had reached medal level performance at the competition.
AlphaProof trains by learning to find formal proofs through reinforcement learning on millions of auto-formalized problems, and for the hardest cases uses what DeepMind calls test time reinforcement learning, generating and learning from millions of related problem variants at the moment of inference itself, rather than relying purely on pretrained knowledge.
Progress since then has accelerated further. By the 2025 IMO, an advanced Gemini Deep Think framework achieved gold medal level performance, and OpenAI reported a comparable gold medal result from one of its own models. These results moved AI in mathematics from an interesting research direction to a genuine competitive presence in a domain long considered the exclusive preserve of the most gifted young mathematicians in the world.
The frontier has moved beyond Olympiad problems entirely into genuinely unsolved research mathematics. In January 2026, reports emerged that GPT 5.2 Pro, paired with the formalization system Aristotle, generated proofs for two specific Erdős Problems, open questions in number theory that had remained unresolved for years, and crucially, these proofs secured acceptance from Terence Tao himself after careful review. Separately, DeepMind’s AlphaEvolve system collaborated directly with Tao to find new approaches to previously unsolved mathematical problems, demonstrating that AI in mathematics is no longer confined to reproducing known results faster but is beginning to genuinely contribute novel mathematical insight.
Proof Assistants: The Infrastructure That Makes Trust Possible
Running parallel to automated theorem proving is a distinct and arguably more foundational thread of AI in mathematics: formal proof assistants, software systems such as Lean, Coq, and Isabelle that allow mathematical proofs to be written in a machine checkable formal language, verified line by line with the same rigor a computer applies to checking whether a program compiles. Tudor Achim, CEO of Math Inc, captured the significance of this approach starkly: when a formal system outputs a proof, nobody has to look at it, because you know it is correct by construction, addressing what he calls the verification problem, the bottleneck created when AI generates mathematical content faster than humans can check it.
The pace of progress specifically within Lean 4 based formalization has been extraordinary through 2025 and into 2026. HunyuanProver, a model fine tuned specifically for interactive theorem proving, achieved state of the art results on the standard MiniF2F benchmark and successfully proved several genuine IMO level statements. Using a system called Gauss, Math Inc completed a challenge originally set by Terence Tao and mathematician Alex Kontorovich to fully formalize the strong Prime Number Theorem in Lean, a genuinely significant undertaking given the theorem’s depth and historical importance.
Most recently, a system called AxiomProver, working with mathematician Ken Ono, reportedly solved all twelve problems from the 2025 Putnam Competition, widely regarded as the most difficult undergraduate mathematics competition in the United States, and went further, resolving four previously open conjectures that had stumped human mathematicians, including uncovering a connection to nineteenth century Jacobi symbols that had been entirely missed by the human researchers working on the problem.
Tao himself has tracked this progress with characteristic precision, introducing the concept of the de Bruijn factor, a measure of how much additional effort formalizing a proof in Lean requires compared to writing it informally. He estimated this factor at roughly twenty in 2023 and 2024, and noted by late 2025 and into 2026 that rapid advances in autoformalization, AI systems that translate informal mathematical writing directly into formal Lean code, had essentially emptied the queue of unclaimed formalization tasks on at least one major mathematical library project, a striking practical demonstration of how quickly this specific application of AI in mathematics has matured.
Pattern Mining and Mathematical Discovery Beyond Proof
A third and less publicly discussed application of AI in mathematics involves pattern mining and conjecture generation, using machine learning not to prove statements but to discover which statements might be true in the first place, a task that has historically depended on mathematical intuition built over decades of experience. DeepMind’s FunSearch system, which combines large language models with evolutionary program search, discovered new solutions to the cap set problem, a longstanding open question in combinatorics, and produced more effective bin packing algorithms than previously known, genuinely novel mathematical objects rather than reproductions of existing results.
A related system called PatternBoost used pattern recognition across large mathematical datasets to disprove a conjecture that had stood unresolved for thirty years, demonstrating that AI in mathematics can contribute not only proofs of true statements but also counterexamples that overturn long held mathematical beliefs. This lineage traces back to earlier systems such as Graffiti, which pioneered automated conjecture generation decades before the current wave of deep learning made such systems dramatically more capable. Comprehensive surveys of this emerging field now describe mathematical exploration and discovery at scale as a distinct research area in its own right, separate from both automated theorem proving and formal verification, focused specifically on using AI to identify which mathematical questions are worth asking.
Where AI in Mathematics Genuinely Struggles
Despite this rapid progress, mathematicians closest to the technology are notably careful about its current limitations, and Tao’s own analysis is instructive precisely because it avoids both dismissiveness and hype. He draws a sharp and important distinction between Lean as a formal proof assistant versus an automatic theorem prover, noting that Lean formalizes a proof a human has already found, and that on its own it is not all that useful in discovering genuinely new proofs.
The emerging best practice he describes divides labour deliberately: humans author or carefully review the statement of a theorem, since verification only certifies that a formal proof matches a formal statement, not that the formal statement actually captures the mathematician’s real intent, while automation increasingly handles the mechanical work of constructing the proof itself once the statement is correctly specified.
This human review bottleneck remains genuinely unresolved. As Tao and his co-author Tanya Klowden note in their 2026 preprint on mathematical methods in the age of AI, there are serious concerns that entire areas of academic mathematical discourse could be drowned out by a flood of low quality AI generated content, echoing a concern raised independently by mathematician Vladimir Voevodsky years earlier, that a technically dense argument by a trusted author, difficult to check and superficially similar to arguments already known to be correct, is hardly ever checked in careful detail, a human trust shortcut that becomes considerably more dangerous once AI can generate such arguments at essentially unlimited scale.
There is also a genuine philosophical unease circulating within the mathematical community that goes beyond technical limitation. Discussions at academic seminars, including a Fall 2025 mathematics and AI course at the University of Washington, have raised pointed questions that Tao’s own lecture explicitly grapples with: does mathematics lose value when computers become better at it than humans, is there an enfeeblement risk in incorporating AI into mathematical training and education, and should foundational skills such as long division or manual integration still be taught if AI in mathematics can perform them instantly and flawlessly.
Tao’s own answer, delivered at the ICM lecture, was that mathematicians need to articulate far more clearly what goals mathematical research is actually meant to serve, arguing that theorem proving and problem solving alone were never the complete picture of why mathematicians do mathematics in the first place, and that this question has become newly urgent precisely because AI has begun to threaten the sufficiency of the old, implicit answer.
Conclusion
AI in mathematics has progressed, in the space of roughly two years, from solving Olympiad geometry problems to contributing genuine proofs accepted by Terence Tao for previously open questions in number theory, while formal proof assistants have simultaneously matured into infrastructure capable of verifying mathematical claims with a rigor no individual human reviewer can match at scale. Pattern mining systems have begun generating and disproving conjectures independently, adding a third distinct capability to the toolkit.
Yet the mathematicians working closest to these systems remain measured rather than triumphant, emphasizing that formal verification certifies correctness against a stated formal claim, not that the claim itself captures genuine mathematical intent, and that the deeper question of what mathematical research is actually for has become considerably more pressing than the narrower question of what AI can currently compute.
-
Is AI Making Us Smarter After All? Part 2: The Balanced Verdict on Human Thinking
This is Part 2 of a two-part series examining whether outsourcing creative and cognitive work to AI is degrading human thinking. Part 1 reviewed the substantial evidence for cognitive offloading and skill decay. Part 2 examines the counter-evidence, the conditions under which AI making us smarter is genuinely possible, and what a fair verdict actually requires.
The Evidence Deserves a Second Look
Part 1 of this series presented a genuinely troubling body of evidence: EEG studies showing weaker neural connectivity, clinicians losing diagnostic skill after AI support was introduced, and a documented illusion of competence among AI users. None of that evidence was overstated, and none of it should be dismissed. But responsible engagement with any body of research requires looking at the full picture, including the studies, researchers, and institutions actively exploring whether AI making us smarter is not just possible but already happening under the right conditions.
The truth that emerges from a complete review of the literature is neither the alarmist story nor a naive optimism. It is something more specific and more useful: the outcome depends heavily on how AI is used, not merely on whether it is used at all.
The Same MIT Study, Read More Carefully
It is worth returning to the widely cited MIT Media Lab EEG study from Part 1, because subsequent, more careful engagement with its actual design reveals an important nuance often lost in headline coverage. The study compared three conditions: writing entirely from memory, writing with a search engine, and writing with an LLM that performed the bulk of the composition itself. The condition that showed weakened neural engagement was specifically the one in which the AI did most of the intellectual work for the participant, essentially replacing their thinking rather than supporting it.
This distinction matters enormously for the AI making us smarter question, because it points toward a specific, testable hypothesis: the harm observed in cognitive offloading research may depend less on AI use per se and more on whether the human remains an active, effortful participant in the cognitive task or becomes a passive recipient of a finished output. A 2026 paper in Computers in Human Behavior captured this distinction precisely in its title: AI makes you smarter but none the wiser, describing a genuine disconnect between measurable performance gains and accurate self-assessment of understanding, a finding that complicates rather than confirms a simple decline narrative.
The Cover Letter Study: A Case for Genuine Learning
One of the more carefully designed recent studies bearing on AI making us smarter comes from behavioural scientists at Wharton, led by Benjamin Lira Luttges. Researchers taught participants to edit poorly written cover letters using either AI-generated feedback or feedback from human professionals. After this training phase, participants were then asked to edit a new, poorly written cover letter entirely without any assistance, human or AI.
The results were genuinely encouraging for anyone hoping AI making us smarter is achievable rather than wishful thinking. Letters produced by the AI-trained group were just as likely to secure a job interview, according to blind human evaluators, as letters from the group trained by human professionals. Crucially, the AI in this study did not simply hand participants a rewritten letter to copy. It walked them through structured feedback, and the learning transferred to genuinely independent performance afterward. This is precisely the kind of evidence that distinguishes AI used as a teacher from AI used as a replacement for thinking, and the distinction turns out to be the single most important variable across the entire body of research.
The Augmentation and Atrophy Framework
A comprehensive 2025 review published in the American Journal of Education and Information Technology introduced a useful conceptual framework worth adopting directly: the Problem-Solving Trade-Off Hypothesis, which proposes that AI’s cognitive impact splits cleanly into augmentation effects and atrophy effects, often for the very same tool, depending entirely on how it is deployed. When used as a research partner, actively engaged with and questioned, AI can genuinely augment critical inquiry. When accepted uncritically as a finished answer, the same tool promotes intellectual passivity.
This framework helps explain an otherwise confusing pattern in the research literature, where some studies find AI making us smarter while others find the opposite, often examining superficially similar AI tools. A comprehensive 2026 review of the cognitive literature reached a similarly nuanced conclusion: moderate AI usage shows minimal cognitive impact, while excessive reliance correlates with decreased critical thinking abilities. The relationship is not linear, and it is not simply about the amount of AI use but the structure and intentionality of that use.
What USC’s New Research Is Actually Testing
The most rigorous ongoing effort to move beyond speculation on the AI making us smarter question is a study launched in July 2026 by USC Viterbi, funded by the National Science Foundation, examining doctors, journalists, and software engineers to determine whether structured AI use can strengthen creativity and critical thinking rather than erode it. The study’s design is explicitly informed by earlier findings, including Stadler et al.’s 2024 research showing that AI use eases mental load but often at the expense of depth of understanding, precisely the tension this two-part series has traced throughout.
What makes the USC research significant is its second phase, which moves beyond simply measuring whether harm occurs and instead attempts to redesign how humans and AI interact specifically to optimise for better creativity and critical thinking outcomes. This reflects a genuine and important shift in the research community’s framing, from asking whether AI making us smarter or dumber is happening as a fixed, inevitable outcome, toward asking how interaction design itself determines which outcome occurs.
The Original Sin of Bad Comparisons
A significant portion of the alarm in this debate traces back to comparing AI-assisted outcomes against an idealised, effortful baseline that most people were never actually achieving in the first place. Before generative AI, the realistic alternative to using ChatGPT for a first draft was often not deep, effortful, independent composition. It was frequently a rushed, low-effort draft produced under time pressure, or simply not producing the work at all. The relevant comparison for AI making us smarter is not AI use versus an idealised deep thinker with unlimited time. It is AI use versus the actual behaviour people were engaging in before AI existed, which was frequently far from ideal itself.
This reframing does not excuse genuine skill atrophy in domains, such as medical diagnosis, where the underlying skill is safety-critical and must be actively maintained regardless of convenience. But for a great deal of everyday writing, brainstorming, and problem solving, the honest comparison group was never a maximally engaged human mind. It was often a tired, distracted, or simply absent one, and against that realistic baseline, AI assistance frequently represents a genuine net gain in both output quality and, when used interactively rather than passively, in the thinking that produces it.
Fostering Collaboration Rather Than Replacement
A 2025 paper in the Journal of Student Research at Indiana University East reviewed the competing evidence directly and reached a conclusion that deserves to anchor any balanced verdict on this topic: AI fosters collaboration and efficiency, and in some cases may enhance critical thinking skills, while overuse without deliberate structure can deplete those same skills. The word collaboration is doing important work in that sentence. It suggests the healthiest relationship with AI tools treats them as a genuine thinking partner, one whose output is questioned, challenged, and integrated actively, rather than either a replacement for thought or a threat to be avoided entirely.
The CHI 2025 Tools for Thought workshop, convening 56 researchers across cognitive science, human-computer interaction, and education, framed the challenge in exactly these terms: the goal is not merely to protect human cognition from AI’s potential negative impacts, but to actively design AI tools and interaction patterns that augment thinking, in the same way that older external tools, including writing itself, have historically extended and strengthened human cognitive capacity rather than simply replacing it.
A Genuinely Balanced Verdict
Bringing both parts of this series together, the fairest conclusion is neither AI making us dumber nor AI making us smarter as a fixed, universal outcome. It is that AI is a cognitive amplifier whose effect depends almost entirely on the structure of the interaction. Passive, unstructured use, accepting AI output wholesale without engagement, reliably correlates with skill atrophy and a documented illusion of competence. Active, structured use, treating AI as a partner to question, challenge, and learn from rather than simply defer to, shows genuine evidence of strengthening rather than weakening independent capability afterward.
For the specific creative activities that motivated this series, writing and image generation, the practical implication is clear. Using an LLM to produce a finished piece of writing with minimal engagement likely does erode the specific compositional and reasoning skills that writing itself has always cultivated as a side effect of the struggle to express an idea clearly. Using an LLM as an interactive collaborator, one whose suggestions are evaluated, revised, and pushed back against, appears considerably more likely to leave those same skills intact or even strengthened, closer to how a skilled writer benefits from an editor’s feedback than a diminishment of their own capability.
Conclusion: The Choice Is Still Ours
The question of whether AI is making humanity dumber or AI is making us smarter turns out not to be a question about the technology at all. It is a question about human choices, individual and institutional, about how deeply we engage with tools that are, for the first time in history, capable of doing so much of our thinking for us if we allow them to. The calculator did not make humanity worse at abstract mathematical reasoning, because the deeper reasoning skills calculators freed us from tedious computation to pursue turned out to matter more than the arithmetic itself.
Whether generative AI follows a similar trajectory, freeing humans for a more valuable kind of thinking, or instead erodes capacities more central to what makes thinking meaningful in the first place, remains genuinely undetermined, and will likely be decided differently across different domains, different age groups, and different patterns of use.
What the evidence assembled across both parts of this series makes clear is that the outcome is not predetermined by the technology itself. It is being determined, right now, by millions of individual decisions about how deeply to engage with the tools already in nearly everyone’s hands. That is, in the end, a more demanding and more hopeful conclusion than either a simple story of decline or a simple story of progress would offer. The evidence suggests humanity retains meaningful agency in this outcome. Whether AI making us smarter becomes the dominant story or the exception may depend less on further research and more on whether that agency is actually exercised.
This concludes our two-part series on AI and human cognition. Explore Part 1 for the full evidence on cognitive offloading
-
Is AI Making Us Dumber? Part 1: The Alarming Evidence Behind Cognitive Offloading
This is Part 1 of a two-part series examining whether outsourcing creative and cognitive work to AI is degrading human thinking. Part 1 reviews the scientific evidence on cognitive offloading and skill decay. Part 2 will examine the counter-evidence, the nuance researchers have found, and what a genuinely balanced position looks like.
A Question That Refuses to Go Away
Every generation of new technology has provoked the same anxious question. Socrates worried that writing would destroy memory. Calculators sparked fears that children would forget arithmetic. Search engines were accused of hollowing out our capacity to retain knowledge.
The question of whether AI making us dumber is a genuine phenomenon or merely the latest iteration of an old cultural panic deserves to be taken seriously rather than dismissed reflexively, precisely because this time there is a growing body of controlled scientific evidence to examine rather than speculation alone.
The honest starting point is that something measurable is happening. Whether it amounts to humanity becoming dumber, in any meaningful sense of that phrase, is a harder and more contested question, one this two-part series will examine from both directions. Part 1 takes the evidence for genuine cognitive harm seriously and presents it in full.
The Concept That Explains the Mechanism
The scientific literature converges on a specific mechanism to explain how and why AI making us dumber might actually occur: cognitive offloading, the act of delegating mental tasks to an external system, reducing one’s own cognitive engagement with the problem. This is not a new concept. Humans have used calculators to support arithmetic, GPS systems to support navigation, and the internet to support memory for decades. What distinguishes AI is the breadth and depth of tasks it can now absorb, extending well beyond simple retrieval into reasoning, synthesis, and even creative composition itself.
The International AI Safety Report 2026, a major government-commissioned review of AI risks, addressed this directly, noting that cognitive offloading can free up cognitive resources and improve efficiency, but that research also indicates potential long-term effects on the development and maintenance of cognitive skills.
That report cited one particularly striking finding: three months after clinicians began using AI support for detecting tumours, their ability to detect them without AI assistance had dropped by 6 percent. This is not a hypothetical worry. It is a documented erosion of a trained medical skill, in a domain where the stakes of that erosion are genuinely serious.
What the MIT Study Actually Found
The most widely cited piece of evidence in the AI making us dumber debate comes from MIT’s Media Lab, in a 2025 study titled “Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task.” Researchers used electroencephalography, EEG, to measure brain activity in participants writing essays under three conditions: using an LLM, using a search engine, and using no external tools at all.
The results were striking. Participants who wrote essays using an LLM showed weaker neural connectivity during the task compared to those using a search engine or working unassisted. Over repeated sessions, brain activity in the LLM-assisted group declined further, a pattern the researchers described using the phrase cognitive debt, a metaphor suggesting that reliance on AI accumulates a kind of deficit in genuine engagement that compounds over time rather than remaining a one-time convenience.
While this specific study has not yet completed peer review, its findings have been influential precisely because they align with a broader pattern found across multiple independent research groups.
The 666-Participant Study and the Critical Thinking Correlation
Perhaps the most methodologically robust evidence for AI making us dumber comes from Michael Gerlich, a professor at the Swiss Business School in Zurich, who published a 2025 study in the journal Societies examining AI tool use and critical thinking across 666 participants. Gerlich found a significant negative correlation between frequent AI usage and critical thinking abilities, with cognitive offloading identified as the specific mediating mechanism. Individuals who relied heavily on AI tools for problem solving demonstrated measurably reduced independent reasoning capacity compared to lighter users. That raises the question: is AI making us dumber?
The age dimension of Gerlich’s findings deserves particular attention. Younger participants demonstrated stronger dependence on AI tools and scored lower on critical thinking assessments than older participants, a pattern replicated across several subsequent studies. This raises a specific and pressing concern about AI making us dumber that differs from earlier technology panics: if the effect is concentrated most heavily in developing minds still building foundational cognitive skills, the long-term societal consequences could be considerably more significant than a simple across-the-board decline distributed evenly across all age groups. AI making us dumber?
The Illusion of Competence
One of the more unsettling findings in the recent literature is what researchers at the University of Technology Sydney termed the illusion of competence in a March 2026 report. Participants who used AI in an unstructured way, letting it reason and synthesise on their behalf, rated their own understanding of the material as high, because the AI’s output was fluent and confident. They believed they had genuinely grasped the underlying material. When subsequently asked to reproduce the reasoning without AI assistance, they could not.
This gap between perceived competence and actual competence is arguably the most concerning specific mechanism within the broader AI making us dumber debate, because it is significantly harder to detect and correct than a simple wrong answer would be. A student who gets a maths problem wrong knows they need to study further.
A student who has an AI solve the problem, reads a fluent explanation, and feels they understand it, has no internal signal telling them their actual competence has not changed at all. The Federal University of Rio de Janeiro’s preregistered randomised controlled trial in 2025 quantified this gap directly, finding an 11 percentage point retention deficit 45 days later between AI-assisted learners and those who worked through material independently. AI making us dumber?
National Security Takes the Question Seriously
The AI making us dumber debate has moved beyond academic psychology into genuine institutional concern at the highest levels of government. The Council on Strategic Risks, an organisation that formally advises the United States government on national security matters, launched a dedicated 2026 debate series specifically examining whether AI is degrading critical thinking within the national security workforce itself.
The concern is direct and consequential: the Pentagon and State Department have rapidly deployed AI tools across their workforce in the name of efficiency, but if cognitive offloading genuinely degrades critical thinking capacity, and national security work fundamentally depends on clear, independent human judgement under pressure, efficiency gains in the short term could be quietly purchasing a less capable, less resilient institution over the longer term. AI making us dumber?
This is a genuinely significant marker for how seriously the underlying concern is being taken outside of academic circles. Governments do not typically convene formal debate series about cultural anxieties they consider unfounded. The fact that this question has reached the level of national security policy discussion suggests the evidence base, while still developing, has crossed a threshold that institutional decision makers consider worth taking seriously.
The Creative Dimension: Writing and Image Generation Specifically
The question posed at the start of this series concerned specifically creative activities, writing and image generation, rather than cognitive tasks broadly. The evidence here is somewhat more limited than for skills like arithmetic or medical diagnosis, but the mechanism identified across the wider literature applies with particular force to creative work. Writing, in particular, is not merely a output-production task.
The act of composing a sentence, revising it, and wrestling with how to express a specific idea precisely is itself a form of thinking, not merely a transcription of thoughts that already existed fully formed. When that generative struggle is outsourced entirely to an LLM, what is lost is not simply the final text but potentially the cognitive process of clarifying one’s own thinking that writing has always served, for writers, as a byproduct of the act itself.
The Google Effect research, which predates the LLM era and examined how search engines changed memory patterns, found that people who expect to have future access to information are less likely to remember the information itself, but more likely to remember where to find it. Whether an equivalent shift is occurring with creative composition, where people increasingly remember how to prompt an AI to produce writing or images rather than how to produce the work themselves, is an open and urgent research question that the field has only begun to address directly.
Conclusion
The evidence assembled in this first part of the series is genuinely substantial. Peer-reviewed studies in respected journals, a major government safety report, EEG data from MIT, and a formal national security debate series all point in a consistent direction: outsourcing cognitive and creative work to AI carries a measurable cost to the specific skills being offloaded, mediated by a documented mechanism, cognitive offloading, that researchers can observe and quantify. The illusion of competence finding is particularly troubling, because it suggests the erosion may be largely invisible to the people experiencing it until the underlying skill is tested directly.
None of this, on its own, definitively proves that AI is making humanity dumber in some broad, irreversible sense. It proves something narrower and still significant: that specific skills, when specifically offloaded to AI, tend to atrophy, and that younger users appear more vulnerable to this effect than older ones.
Whether this constitutes a genuine crisis, a manageable trade-off, or something considerably more nuanced than either extreme is where Part 2 of this series turns next, examining the counter-evidence, the conditions under which AI use appears to strengthen rather than weaken thinking, and what a genuinely balanced verdict on this question actually requires.
Part 2: The Counter-Evidence and a Balanced Verdict, coming next.
-
The Critical Wave of AI Copyright Lawsuits Reshaping the Future of Generative Models
Six Million Pirated Books and a Billion Dollar Question
On July 20, 2026, a federal judge in San Francisco approved the largest copyright class action settlement in United States history. Anthropic agreed to pay 1.5 billion dollars to a class of authors and publishers, roughly 3,000 dollars for each of an estimated 500,000 works, after admitting it had downloaded as many as seven million pirated books to train its Claude models.
That single ruling has become the anchor point for understanding the entire wave of AI copyright lawsuits now working through American courts, lawsuits that touch every major AI lab and that will, collectively, determine whether the current generation of large language models was built on a legally sound foundation or a legally precarious one.
The scale of the problem is genuinely industry wide. Dozens of cases are pending against OpenAI, Google, Microsoft, Meta, Midjourney, and Stability AI, filed by novelists, journalists, musicians, visual artists, and news organisations, all making some version of the same core allegation: that these companies trained their models on copyrighted material without consent, without a licence, and in many cases without even paying for the content in the first place.
How the Training Actually Happened
To understand why AI copyright lawsuits have multiplied so quickly, it helps to understand what actually happened during the early training runs of today’s frontier models. Building a capable large language model requires ingesting enormous volumes of text, historically hundreds of billions to trillions of words. Assembling a dataset at that scale through licensed content alone would have been prohibitively slow and expensive in the early 2020s, when the race to build the first genuinely capable chatbots was at its most intense.
Court filings across multiple AI copyright lawsuits reveal that several major labs took shortcuts. In the case against Anthropic, court records showed the company downloaded books directly from known pirate library websites, essentially the same category of site used for illegal book sharing, and stored them in a centralised internal dataset used to train Claude.
The court drew a sharp legal distinction that has become central to nearly every subsequent case: training a model on lawfully acquired copyrighted books can plausibly be fair use, because the training process transforms the material into statistical patterns rather than reproducing it. But acquiring the books through piracy in the first place is a separate, independently unlawful act, regardless of what happens to the data afterward.
The Anthropic Settlement: A Landmark With Limited Reach
The Anthropic settlement deserves close examination because it will shape negotiating positions across every other pending case. The underlying case, Bartz v. Anthropic, was filed in August 2024 by authors Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson. In June 2025, Judge William Alsup issued a pivotal ruling: training an AI model on lawfully acquired copyrighted books was fair use because the process was sufficiently transformative, but Anthropic’s use of pirated copies to build its library was not protected, and that narrower claim would proceed to trial.
Rather than face trial, Anthropic settled for 1.5 billion dollars, an amount its own lawyers described as the largest publicly reported copyright recovery in history. The settlement required Anthropic to destroy its pirated dataset entirely. Crucially, legal experts covering the wave of AI copyright lawsuits have been careful to note what the settlement does not do. It does not establish binding legal precedent, because a settled case never reaches an appeals court.
It does not grant Anthropic a licence for any future training. And it does not resolve the central industry wide question, whether training AI models on copyrighted material is lawful, since that question was never actually decided at trial. As one law professor put it, appeals courts still need to weigh in on the larger question of if and how AI companies can use copyrighted works, and they will have plenty of opportunities ahead, because dozens of similar AI copyright lawsuits remain active against other companies right now.
Meta: A Parallel Case Still in Progress
Kadrey v. Meta Platforms, filed by a similar group of authors, follows an almost identical fact pattern to the Anthropic case, and its diverging trajectory illustrates how unpredictable this legal landscape remains. The court granted Meta a partial win on fair use grounds regarding the actual training of its Llama models. But a separate and more damaging allegation survives: that Meta engaged in what is known as seeding during the torrenting process.
Meaning that in addition to downloading pirated books, Meta’s systems may have simultaneously redistributed those pirated files to other users on the same file sharing network, a potentially more serious violation than simple downloading. That claim remains active in the Northern District of California, and unlike the Anthropic case, Meta has not settled, meaning this particular thread of AI copyright lawsuits is still heading toward further discovery and potentially trial.
OpenAI and The New York Times: The Fight Over Memorisation
The most closely watched of all the active AI copyright lawsuits is The New York Times v. OpenAI and Microsoft, filed in December 2023 after nine months of failed licensing negotiations. Unlike the Anthropic and Meta cases, which centre on how training data was acquired, the Times case centres on a different and arguably more consequential legal question: whether ChatGPT can reproduce, or regurgitate, substantial portions of Times articles verbatim when prompted in specific ways.
The Times alleges that its journalists’ work does not merely inform the model statistically but can, in certain circumstances, be extracted from it nearly word for word, a claim that goes to the heart of whether training itself was transformative or whether the model functions, in part, as an unlicensed distribution mechanism for the underlying copyrighted text.
The case has become unusually contentious on discovery grounds. A magistrate judge ordered OpenAI to produce twenty million de-identified ChatGPT conversation logs, a demand OpenAI fought vigorously on user privacy grounds before ultimately complying under court order. OpenAI has publicly argued that its use of Times articles is a transformative, non-expressive analytical use protected by fair use, pointing to the Anthropic and Meta rulings on training as precedent in its favour.
As of mid-2026, the case remains in the Southern District of New York, consolidated with similar suits from other news organisations into a multidistrict litigation, with expert reports completed in late 2025 and summary judgment briefing concluding in April 2026. No trial date has yet been set, but this case, more than any other among the current AI copyright lawsuits, is widely viewed as the one most likely to produce a binding, tested legal precedent, precisely because OpenAI has shown far less inclination to settle than Anthropic did.
The Broader Pattern: Every Creative Industry Is Suing
The scope of AI copyright lawsuits extends well beyond books and news journalism. Disney and other major studios have sued Midjourney over image generation trained on copyrighted characters and artwork. Getty Images sued Stability AI over the training of its image generation models. Recording labels have pursued similar claims against AI music generation platforms. As one legal expert tracking the litigation observed, the lawsuits are coming from essentially every sector of human creativity, newspapers, recording labels, movie studios, and more, reflecting a consistent grievance across creative industries that their work was used as raw material for a multi billion dollar commercial product without consent or compensation.
The Licensing Shift: Settling Before the Courtroom
One of the more significant developments to emerge from this wave of AI copyright lawsuits is a visible shift in industry behaviour, away from litigation risk and toward proactive licensing. Facing the reputational and financial exposure demonstrated by the Anthropic settlement, several AI companies have begun negotiating licensing agreements with publishers directly rather than waiting to be sued. The New York Times itself, even while actively suing OpenAI, separately reached a multiyear licensing agreement with Amazon in 2025, reportedly worth twenty to twenty five million dollars, for use of its content in Amazon’s AI products, demonstrating that litigation and licensing are not mutually exclusive strategies for the same publisher.
A Columbia University law professor who studies literary property rights described receiving a steady stream of requests from her own publishers asking her to authorise licensing of her books to AI companies, a pattern that suggests the publishing industry as a whole is moving toward a licensing market for AI training data, driven directly by the legal exposure that the current AI copyright lawsuits have made unmistakably clear.
What the Outcome Will Determine
The stakes in this litigation extend well beyond the specific dollar amounts at issue. If courts ultimately rule broadly in favour of AI companies on fair use grounds for lawfully acquired training data, as the preliminary Anthropic and Meta rulings suggest they might, the legal foundation for training future models on publicly available text becomes considerably more secure, provided companies avoid the piracy shortcuts that triggered these specific lawsuits.
If courts rule against AI companies, particularly on the regurgitation question at the centre of the New York Times case, the entire industry may need to rebuild significant portions of its training pipelines around licensed content, a shift that would fundamentally alter the economics of building frontier AI models and could meaningfully slow the pace of model development across the industry.
Conclusion
The current wave of AI copyright lawsuits represents one of the most consequential bodies of litigation in the technology industry’s history, not because any single case will resolve every open question, but because each ruling and each settlement is incrementally shaping the legal architecture within which every future AI model must be built. The Anthropic settlement demonstrated the scale of financial exposure that piracy based training data creates.
The Meta case demonstrates that even companies who win on the core fair use question can remain exposed on narrower, related claims. And the OpenAI case, still without a trial date but approaching a decisive summary judgment ruling, may ultimately decide whether training itself, done properly and without piracy, is legally sound at all. For publishers, authors, and AI companies alike, the outcome of these AI copyright lawsuits will define the ground rules for how intelligence itself gets built for years to come.