Scope: This report examines the distinction between AI-enabled economic value and measured production, using United States evidence available by 5 October 2026, while identifying measurement implications for other economies without extrapolating American welfare estimates to them.
Executive Summary / BLUF
“Dark output” is defensible as a label for identifiable measurement gaps, provided that it does not become an unspecified monetary residual added to GDP.
The decisive distinction separates economic value, consumer surplus, firm productivity and margins, and measured GDP and official productivity, because these objects can move in different directions without any accounting contradiction.
Lower prices can reduce nominal output while preserving real output; lower labour requirements can raise productivity and profits while leaving final expenditure unchanged; better services can increase welfare without a recorded quality adjustment.
The strongest case concerns unpriced household capabilities, inadequately measured service quality, and some complementary intangible investment, whereas provider revenue, retained profits, and infrastructure expenditure already have accounting counterparts.
Current welfare surveys, production-based imputations, task-exposure estimates, and industry productivity statistics measure different objects, so their magnitudes cannot be combined into a single estimate of AI’s economic contribution.
The appropriate response is a coordinated set of production, quality, time-use, distributional, and welfare accounts, with assumptions and geographic coverage disclosed separately.
AI’s Accounting Gap Is Becoming a Policy Risk
The consequential choice in artificial intelligence is whether governments and institutions will judge production, investment and public benefit through the same statistic. On 17 November 2025, the US Census Bureau broadened its business survey question on AI, creating a break in the adoption series; on 2 October 2026, a Bureau of Economic Analysis research spotlight described an association between AI use and stronger output while preserving the distinction between correlation and causation. Those decisions expose the governing problem: measures designed for different purposes are being treated as interchangeable evidence. The fiscal and industrial risk is that expenditure will be mistaken for productive deployment, profits for consumer benefit, and unpriced capabilities for missing GDP. The remedy is complementary accounts that establish what changes, who benefits and which gains become recorded production.
The largest estimates answer different questions
The US evidence becomes more informative when its units remain separate. Census research using the Business Trends and Outlook Survey supplement collected during November 2025–January 2026 reported firm AI use at 17.9% on a firm-weighted basis and 31.2% when weighted by employment, with a reference window of the previous two weeks. Reported worker-task AI use reached 22.6% and 40.6%, respectively, over the previous six months. The employment-weighted figures describe employment located in reporting businesses; they do not describe the proportion of individual workers personally using AI. The different reference windows also prevent the gap from being interpreted as a direct measure of informal adoption. The Microstructure of AI Diffusion: Evidence from Firms, Business Functions, and Worker Tasks — US Census Bureau research working paper — April 2026.
The valuation evidence measures a different object. The April 2026 Stanford Digital Economy Lab paper What is Generative AI Worth? reports mean monthly willingness to accept compensation for losing access at $98.00 in its July 2025 wave and $124.50 in March 2026; Table 2 gives corresponding medians of $3.39 and $11.48. Its annualised, mean-based welfare aggregates rise from $116.2 billion to $172.3 billion, using estimated US user populations of 98.78 million and 115.33 million within an ages 18–64 population expansion. Those totals depend on stated choices, external user counts and annualisation. The distance between the mean and median matters because the typical respondent and the aggregate valuation answer different questions; neither is a measure of production. What is Generative AI Worth? — Stanford Digital Economy Lab — April 2026.
The production-oriented adjustment is different again. BEA research published in June 2026 estimates that including advertising- and marketing-supported free content would increase average annual US real GDP growth by 0.22 percentage point during 2022–25, under its barter-transaction methodology. The estimate includes AI but is not AI-only, and it is an experimental calculation rather than an adopted revision to headline GDP. SemiAnalysis’s approximately $1.5 trillion labour-cost exposure estimate, published in May 2026, likewise identifies potentially affected tasks rather than output already missing from the accounts. Adding these figures would combine adoption, welfare, production adjustments and exposure into a total with no coherent accounting meaning. The Progression of “Free” Digital Content to AI: Impacts on US Economic Growth and Productivity — BEA — June 2026; AI Dark Output: The Visible Cost of Invisible Output — SemiAnalysis — May 2026.
Infrastructure spending establishes inputs, not successful use
The System of National Accounts records qualifying investment because assets are produced or acquired, while downstream productivity depends on what those assets subsequently enable. A data centre, a software investment and a redesigned workflow occupy different positions in that transmission mechanism. Institutional analysis therefore needs to establish whether assets enter service, whether complementary inputs are operational and whether deployment produces accepted outcomes. The distinction is consequential for procurement: a system generating more documents may increase activity while leaving the cost of verified, usable documents unchanged.
The productivity J-curve literature supplies a mechanism for delayed gains without guaranteeing their arrival. Complementary investment in training, data and organisational processes can consume resources before mature deployment produces observable improvements, while inadequate measurement of those intangibles can distort the timing of recorded productivity. That mechanism supports collecting better evidence; it does not authorise institutions to treat every disappointing productivity observation as deferred success. The relevant test is whether spending becomes persistent use and whether persistent use improves completed production. The Productivity J-Curve: How Intangibles Complement General Purpose Technologies — American Economic Journal: Macroeconomics — January 2021.
The Bureau of Labor Statistics productivity framework makes the operational distinction explicit: labour productivity concerns real output per labour hour, while multifactor productivity also accounts for other productive inputs. Generation time alone supplies neither numerator nor denominator. Preparation, review, correction and integration belong in the workflow, and time released from a task can support additional output, better quality or unused capacity rather than lower payroll. Procurement and corporate governance need that reconciliation before converting reported task savings into a claim about financial returns or economy-wide efficiency. Labor Productivity and Costs: Concepts — US Bureau of Labor Statistics.
Quality determines whether a real gain becomes visible
The System of National Accounts can accommodate a falling price without recording a fall in real output when the deflator captures the price change correctly. For an unchanged service, lower receipts are a nominal result, not proof of less production. A quality improvement at a constant price presents a different problem: the real gain depends on whether the output measure recognises the change. “Dark output” is useful at this specific point of failure, where an inadequate service index obscures volume or quality; it becomes rhetorical when lower prices are declared invisible by definition. Measuring the Economy: A Primer on GDP and the NIPAs — Bureau of Economic Analysis.
The UN guidance on improving AI visibility identifies the difficulty of separating capabilities embedded in software and broader products. Within-firm use intensifies that difficulty because there may be no separate transaction for the task being improved. Yet internal use does not automatically escape GDP: it can support recorded final output, reduce input requirements or contribute to recognised own-account production. Statistical reform must identify which of those mechanisms operates before assigning an unobserved value to the activity. Improving the visibility of Artificial Intelligence in the National Accounts — United Nations national accounts guidance.
The Office for National Statistics offers a practical precedent for maintaining measures with different purposes. Its public-service productivity methodology incorporates quality adjustments while explaining differences from national accounts measures, including healthcare output. The institutional lesson is that an additional measure can illuminate outcomes without silently redefining the core production series. An AI service index should follow the same discipline: specify the accepted output, publish the quality variables and expose the sensitivity of the result to alternative adjustments. Public service productivity: total, UK QMI — Office for National Statistics — May 2026.
Welfare cannot finance a budget until income changes
GDP-B addresses benefits that transaction prices do not fully reveal, but a welfare estimate is not an equivalent addition to taxable income. Households may gain access to useful capabilities while market receipts remain unchanged or decline; producers may retain cost savings while buyers receive no immediate price reduction. Fiscal analysis must trace recorded income, employment and public-service expenditure rather than translate a willingness-to-accept aggregate into presumed budget capacity. The accounting boundary is a constraint on the inference, not a denial of the benefit.
The System of National Accounts also prevents visible financial gains from being relabelled as invisible production. Higher profits funded by lower unit costs are recorded outcomes, while model-provider receipts can represent intermediate purchases whose value must not be counted again alongside the customer’s final output. The distributional question is therefore separate from the production question: who receives the saving, whether prices change, how compensation and hours respond, and whether access expands. A larger margin and a larger consumer surplus can each be economically significant without describing the same gain. Measuring the Economy: A Primer on GDP and the NIPAs — Bureau of Economic Analysis.
International comparison fails when the boundaries differ
The OECD Handbook on Compiling Digital Supply and Use Tables provides a route to additional detail within the national accounts framework, distinguishing products, producers and how transactions occur. That structure can support more visible AI production, but it does not make a household valuation survey interchangeable with a business adoption survey. International comparison requires common definitions of the reporting unit, population, reference window and treatment of embedded software; transferring US willingness-to-accept estimates to another economy would bypass those requirements. OECD Handbook on Compiling Digital Supply and Use Tables — November 2023.
The UN guidance on free digital products develops satellite-account options while preserving the distinction from the central framework. A production-oriented free-content account and a GDP-B welfare account can therefore coexist without sharing a valuation method or being added together. The useful international objective is a reproducible crosswalk showing what is already recorded, what is reclassified, what improves real-volume measurement and what remains a separately estimated benefit. Without that reconciliation, more statistical detail can enlarge the scope for incompatible comparisons. Recording and Valuing “Free” Digital Products in an SNA Satellite Account — United Nations national accounts guidance — September 2022.
The next two years will test whether institutions measure outcomes
Over the next 12–24 months, the cost of retaining a single-metric interpretation will fall on identifiable participants: taxpayers where public procurement rewards generated activity rather than accepted services; adopting firms where gross time savings conceal verification and integration costs; workers where released task capacity is confused with realised employment displacement; and households where improved access remains absent from assessments of benefit. The Business Trends and Outlook Survey’s wording change demonstrates how even an adoption series can shift when the measurement boundary changes. Institutions need equivalent transparency when moving from expenditure to claims about performance.
The proposed response is a connected set of production, deployment and welfare accounts, with task-level time observations, quality-adjusted service indices and firm disclosures that reconcile AI expenditure to accepted output and labour requirements. The OECD digital supply-and-use framework supplies an institutional foundation, but implementation still requires validated definitions and observations. If institutions leave that work undone, infrastructure spending will remain easier to establish than the productive use it supports, and user benefits easier to claim than to compare. The medium-term consequence is weaker procurement, weaker fiscal inference and a distribution of gains that remains unresolved even where the transactions themselves are visible.
Navigational Index
| Thematic pillar | Required sections |
|---|---|
| Concepts and accounting mechanisms | Introduction; A short intellectual history; Formal framework |
| Evidence and the limits of measurement claims | What the current estimates actually measure; Where darkness is real and where it is rhetorical |
| Institutional interpretation and statistical reform | Implications for investors and policymakers; A measurement agenda; Conclusion; Appendix |
Master Abstract
Artificial intelligence creates a measurement problem when the relationship between expenditure and useful capability changes more rapidly than statistical systems can identify comparable services, establish quality-adjusted prices, and distinguish market production from household activity. Nevertheless, a widening difference between what users gain and what they pay does not establish that GDP is incorrectly recording transactions, because production accounts and welfare measures answer different questions. The defensible central judgment is therefore conditional: AI generates economic value that conventional production statistics capture incompletely, but the missing component must be classified before its magnitude or policy significance can be assessed.
This report separates transaction valuation, real service volume, production efficiency, and consumer benefit through an explicit accounting framework and four numerical cases. Its principal implication is that apparently contradictory outcomes—falling nominal service expenditure, rising profits, unchanged measured output, and increasing household welfare—can coexist. The empirical assessment consequently treats willingness-to-accept surveys as welfare evidence, free-content imputations as experimental production accounting, corporate expenditure as investment evidence, and productivity releases as measured outcomes requiring separate causal attribution. The assessment would change if representative task-level evidence established that quality-adjusted service indices systematically omit large, persistent gains, or if revised production accounts demonstrated that existing estimates materially understate domestic value added; neither conclusion follows merely from high token consumption or rapid technological improvement.
Introduction
The measurement problem arises when AI makes useful cognitive services cheaper, better, or newly feasible, while the observable economic record continues to consist mainly of payments, receipts, labour inputs, and imperfect service-price indices. A household that obtains previously unaffordable tutoring, an employee who produces a more reliable analysis within unchanged working hours, and a professional practice that retains an automation saving all acquire different kinds of benefit, although only some generate additional recorded final production. GDP measures production within a defined boundary rather than the entire improvement in human welfare, and its distinction between market activity and unpaid household services therefore remains essential to interpreting these cases. Production account — System of National Accounts update documentation, United Nations Statistics Division — 2025. unstats.un.org
The paper defends “dark output” as a useful descriptive category only when an identifiable omission or distortion can be demonstrated, while rejecting its interpretation as a newly discovered aggregate that can be added to GDP. Economic value encompasses useful outcomes and capabilities; consumer surplus concerns benefits exceeding the price paid; firm productivity and margins concern production efficiency and the distribution of receipts; and measured GDP and official productivity concern recorded production and its relation to measured inputs. Their relationship must be analysed rather than assumed, particularly because industry gross output includes intermediate transactions that cannot simply be summed into national value added. Measuring the Nation’s Economy: An Industry Perspective—A Primer on BEA’s Industry Accounts — U.S. Bureau of Economic Analysis — May 2011. bea.gov
A short intellectual history
The historical analogue establishes a possibility, rather than a forecast
The Solow paradox identified a discrepancy between the visible diffusion of computing and its initially disappointing appearance in productivity statistics, while the subsequent American ICT productivity acceleration demonstrated that implementation, investment, and organisational change can eventually produce measurable aggregate gains. Reviewing this history in 2006, Federal Reserve Chairman Ben Bernanke connected the post-1995 resurgence with technological progress in ICT production and improvements in industries using ICT, while emphasising the complementary investments required to turn equipment purchases into effective production systems. This experience supports the plausibility of delayed gains from a general-purpose technology, although it does not establish that AI will reproduce the same trajectory, geographic distribution, or magnitude. Productivity — Board of Governors of the Federal Reserve System — Aug 2006. Federal Reserve Board
Brynjolfsson, Rock, and Syverson’s productivity J-curve supplies a more precise mechanism: firms divert resources towards process redesign, organisational knowledge, and other complementary assets whose creation is incompletely recognised as investment, depressing initially measured productivity relative to a framework that records those assets. When the assets subsequently support market output, measured productivity can accelerate because the earlier investment and its subsequent capital services were not fully recorded. Consequently, the hypothesis concerns the timing and classification of production, rather than an assurance that every implementation expenditure will eventually earn a return. The Productivity J-Curve: How Intangibles Complement General Purpose Technologies — Brynjolfsson, Rock and Syverson, MIT-hosted NBER working paper — Oct 2018. ide.mit.edu
Free digital goods widened the separation between expenditure and benefit
Hulten and Nakamura distinguish resource-saving innovation in production from output-saving innovation in consumption, through which information enables households to obtain better outcomes from a given expenditure. Their framework explains why improved decisions and free information can increase well-being without a proportionate increase in production valued at transaction prices, while their expanded welfare measure remains conceptually distinct from conventional GDP. The relevance to AI is that a user can acquire a capability—understanding a document, planning an activity, or accessing personalised instruction—whose benefit substantially exceeds the transaction that supplied the computational input. Accounting for Growth in the Age of the Internet: The Importance of Output-Saving Technical Change — Federal Reserve Bank of Philadelphia — Aug 2018. philadelphiafed.org
The GDP-B literature formalises a complementary welfare perspective through valuation experiments, including incentive-compatible choices concerning access to goods such as Facebook and smartphone cameras. Incentive compatibility requires consequences that make truthful choices relevant to respondents, rather than merely attaching monetary amounts to hypothetical questions; accordingly, the existence of incentive-compatible experiments in this literature does not establish that every subsequent generative-AI survey uses that design. GDP-B: Accounting for the Value of New and Free Goods — Stanford Digital Economy Lab — Mar 2025. Stanford Digital Economy Lab
Cognitive services present a harder comparability problem
Generative AI is difficult to measure because the computational unit purchased is not necessarily the economically relevant service delivered: identical token expenditure can support tasks with different reliability requirements, review burdens, and consequences. Unlike a hardware characteristic that can be observed relatively directly, a cognitive service must often be evaluated jointly with the user’s expertise, the task’s difficulty, and the institutional responsibility attached to the result. The statistical challenge therefore concerns the identification of comparable outputs, rather than simply the observation of cheaper computational inputs.
SemiAnalysis’s May 2026 contribution introduced the relevant distinction between substitution of previously performed work and creation of previously uneconomic activity, while its approximately $1.5 trillion estimate concerns tasks that current-generation AI could substantially augment or automate. That figure is task exposure, rather than realised savings, incremental GDP, or measured dark output. AI Dark Output: The Visible Cost of Invisible Output — SemiAnalysis — May 2026. newsletter.semianalysis.com
Earlier measurement debates constrain the interpretation
Griliches’s hedonic work established the importance of separating product characteristics from pure price movements, while the Boskin Commission subsequently distinguished substitution, outlet, quality-change, and new-product biases in consumer-price measurement. The Commission’s historical estimate cannot be transferred to AI, but its taxonomy remains useful because missing quality improvements and delayed product entry are specific statistical mechanisms rather than evidence that all technological benefits escape the accounts. Cross-validation of quality-adjustment methods for price indexes — U.S. Bureau of Labor Statistics — Dec 2018; Toward A More Accurate Measure Of The Cost Of Living — Advisory Commission to Study the Consumer Price Index, reproduced by the Social Security Administration — Dec 1996. bls.gov
The counterweight is equally important: Byrne, Fernald, and Reinsdorf found little evidence that growing IT mismeasurement explained the post-2004 American productivity slowdown, while Syverson tested the proposed explanation against cross-country patterns, the scale of plausible missing benefits, industry arithmetic, and income-side evidence. These findings do not impose a numerical ceiling on future AI omissions, but they establish that an aggregate mismeasurement claim requires evidence commensurate with the discrepancy it seeks to explain. Does the United States have a Productivity Slowdown or a Measurement Problem — Federal Reserve Board — Mar 2016; Challenges to Mismeasurement Explanations for the U.S. Productivity Slowdown — Chad Syverson, University of Chicago author manuscript — Jun 2016. federalreserve.gov
Formal framework
Transaction value, service volume, and welfare require separate identities
For a comparable market service, define the transaction price pₜ, the number of completed services qₜ, and a quality multiplier aₜ normalised to one in the reference period:
Nominal gross output: Gₜ = pₜ × qₜ
Quality-adjusted service volume: Vₜ = aₜ × qₜ
Effective price per quality-adjusted service: dₜ = pₜ ÷ aₜ
Therefore: Gₜ = dₜ × Vₜ
These expressions are an illustrative decomposition rather than a substitute for official chain-index methods, because the quality multiplier must represent defensibly comparable service characteristics rather than an analyst’s impression of usefulness. Official industry real-value-added measures also deflate gross output and intermediate inputs separately, which prevents a simple revenue calculation from becoming a complete productivity account. Guide to the Interactive GDP-by-Industry Accounts Tables — U.S. Bureau of Economic Analysis — accessed Oct 2026. U.S. Bureau of Economic Analysis (BEA)
For an individual firm, simplify the corresponding income relationships as follows:
Value added: VAₜ = Gₜ − intermediate purchasesₜ
Operating profit: πₜ = Gₜ − intermediate purchasesₜ − labour costsₜ − other operating costsₜ
Labour productivity: quality-adjusted outputₜ ÷ labour hoursₜ
Labour costs are part of value added rather than intermediate purchases, so replacing employee work with purchased AI services changes both the firm’s input composition and the distribution of production income. GDP aggregates domestic value added, meaning that the AI supplier’s receipts cannot be added again to the purchasing firm’s final sales without subtracting the corresponding intermediate transaction. What are intermediate inputs? — U.S. Bureau of Economic Analysis — Mar 2006; What is industry value added? — U.S. Bureau of Economic Analysis — Mar 2006. U.S. Bureau of Economic Analysis (BEA)
Consumer surplus remains a separate relationship:
Consumer surplus for a purchased unit = willingness to pay for that unit − price paid
For newly feasible household activity, the useful outcome can be represented as a non-market capability Hₜ, but neither Hₜ nor its monetary welfare valuation enters the GDP identity merely because it exists. Moreover, user effort, verification time, and other burdens must be considered when moving from gross benefit to net welfare, while willingness to accept compensation for losing access is not automatically identical to willingness to pay for acquiring it.
A worked example with explicit assumptions
Consider a legal memorandum purchased by a household, using hypothetical US dollars rather than empirical estimates of legal-service prices. Before AI adoption, one memorandum costs $1,000, requires $800 of labour expenditure, and produces $200 of operating profit; the household’s willingness to pay is $1,500, giving consumer surplus of $500.
Assume that AI-assisted production reduces operating cost to $200, comprising $180 of labour and $20 paid to a domestic AI supplier, while taxes, other inputs, capital costs, and all unrelated economic activity remain unchanged. The supplier’s $20 is domestic value added in this deliberately simplified example, so the combined production chain’s nominal GDP contribution equals the household’s final payment rather than that payment plus the token bill.
The real-output calculations below use reference-period prices and an explicitly assumed quality adjustment, while the GDP changes concern only this isolated production chain; economy-wide changes would additionally depend on displaced workers, expenditure reallocation, imports, and subsequent demand.
| Case | Final price | Quality-adjusted volume | Operating cost | Legal firm profit | Nominal GDP contribution | Real final output at reference prices | Consumer surplus |
|---|---|---|---|---|---|---|---|
| Baseline | $1,000 | 1.0 | $800 | $200 | $1,000 | $1,000 | $500 |
| Cost saving fully passed into price | $400 | 1.0 | $200 | $200 | $400 | $1,000 | $1,100 |
| Cost saving retained as margin | $1,000 | 1.0 | $200 | $800 | $1,000 | $1,000 | $500 |
| Quality rises at constant price | $1,000 | 1.5 | $200 | $800 | $1,000 | $1,500 | $800 |
| Newly feasible free household activity | $0 | Additional household capability | No incremental market expenditure | No assumed profit change | $0 incremental | $0 incremental within GDP boundary | $250 additional |
The final row concerns a separate activity that previously did not occur, with assumed willingness to pay of $250; it is not a free replacement of the baseline memorandum.
Cost saving fully passed into price
The reduction in cost from $800 to $200 permits the final price to fall from $1,000 to $400 while preserving the firm’s $200 profit, so nominal GDP associated with the isolated final service declines by $600 while consumer surplus rises by the same amount under unchanged willingness to pay. Properly measured real final output remains unchanged because the household receives the same memorandum, although labour productivity rises if fewer total production hours are required for that unchanged output.
If the deflator fails to recognise the lower price of the comparable service and remains fixed, the measured real service value instead falls from $1,000 to $400, producing a spurious 60% decline in volume. This outcome requires the specified deflator failure; falling prices alone do not make real output invisible.
Cost saving retained as margin
With the final price held at $1,000, the firm’s profit rises from $200 to $800 while nominal final expenditure, real service volume, and consumer surplus remain unchanged under the stipulated assumptions. The legal firm’s own value added falls from $1,000 to $980 because it now purchases a $20 intermediate service, but the domestic supplier contributes the corresponding $20, leaving combined nominal GDP unchanged.
The efficiency gain is therefore real, and its distribution towards profit is visible, although it does not require an increase in final output. Official labour productivity can also rise without a price or quantity change when measured output remains constant and measured hours fall; consequently, lower input requirements are another route through which productivity gains enter the statistics.
Quality rises at constant price
Assume that an independently defensible service-quality measure rises from 1.0 to 1.5, and that willingness to pay rises from $1,500 to $1,800 rather than mechanically following the quality multiplier. Nominal GDP remains $1,000, reference-price real final output rises to $1,500, profit rises to $800 under the common cost assumption, and consumer surplus increases from $500 to $800.
The 50% real-output increase depends entirely on the assumed quality adjustment, while the $300 surplus increase depends on the separate valuation assumption. Without a valid quality adjustment, recorded real output remains unchanged, even though the stipulated capability improvement has occurred.
Free household use and employee use require different treatments
For the newly feasible household activity, assume zero monetary payment, unchanged provider monetisation and costs during the comparison interval, and willingness to pay of $250. The activity produces additional gross consumer surplus but no incremental recorded household service production; if participation and verification consume resources worth $30 to the household, net benefit is $220 rather than $250.
Employee use cannot be assigned the same automatic zero-output treatment, because an internal AI-assisted task can improve a marketed product, reduce paid labour requirements, increase final throughput, or create an own-account asset. Its contribution is therefore assessed through the employer’s final output and input accounts, while the absence of a separately billed internal document does not establish that its economic effect is absent from GDP.
Substitution reduces GDP only under specified conditions
Substitution lowers the isolated nominal GDP contribution when a household replaces an expensive final purchase with a cheaper equivalent, while its effect on real GDP depends on comparability, quality measurement, and whether activity crosses into household production. By contrast, replacing a domestic business’s purchased intermediate service with internal AI-assisted work can redistribute value added between suppliers and the purchasing firm without reducing the value of the final product.
Demand expansion provides another countervailing mechanism: lower prices can induce sufficient additional purchases to stabilise or increase nominal expenditure, while displaced resources can produce other goods and services. Creation likewise enters measured output when the newly feasible activity is sold, embodied in a marketed product, or recorded as eligible own-account investment, whereas household capabilities and unmeasured intermediate quality improvements require complementary evidence.
What the current estimates actually measure
Key evidence and estimate comparison
The following table preserves differences in coverage and method, while distinguishing statistical observations from experimental estimates and company-defined indicators.
| Evidence | Concept measured | Geography and population | Reference period | Magnitude | Method and issuer | What it is not |
|---|---|---|---|---|---|---|
| Token expenditure | Purchased computational services | Coverage depends on provider and customer accounts | Transaction period | No harmonised aggregate established here | Invoices, subscriptions, and usage records | Completed-task value or welfare |
| Provider monetisation example | Company-defined AI revenue run rate | Microsoft’s business; no US-only attribution | Announced January 2025, concerning FY2025 Q2 | Above $13bn annual run rate | Company statement filed with SEC | Realised annual revenue, token-only revenue, or GDP |
| Generative-AI WTA | Annualised access valuation | US users aged 18–64 in aggregation | July 2025 / March 2026 | Mean-based $116.2bn / $172.3bn | Brynjolfsson, Collis, Eggers, Kazinnik and Nguyen | Market output |
| WTA inputs | Monthly compensation to surrender access | Same aggregation; 98.78m / 115.33m users | 2025 / 2026 | Means $98 / $124.50; medians $3.39 / $11.48 | Paper’s Table 2 | Interchangeable mean and median estimates |
| Experimental free-content adjustment | Additional measured production and volume | US advertising- and marketing-supported content, including AI | 2022–2025 | +0.22 percentage point average annual real GDP growth | Nakamura, Samuels and Soloveichik | AI-only contribution or consumer surplus |
| Official industry productivity | Real sectoral output per hour | US software publishers | 2025 | +12.9% | BLS industry accounts | Causally identified AI effect |
| Infrastructure expenditure example | Cash additions to property and equipment | Microsoft consolidated operations | Fiscal year ending June 2025 | $64.551bn | Audited cash-flow statement | AI-only investment or US domestic value added |
Exact supporting records: Microsoft Cloud and AI Strength Drives Second Quarter Results — Microsoft, SEC-filed Exhibit 99.1 — Jan 2025; What is Generative AI Worth? — Stanford Digital Economy Lab — Apr 2026, Table 2; The Progression of “Free” Digital Content to AI: Impacts on U.S. Economic Growth and Productivity — BEA Working Paper WP2026-13 — Jun 2026; Productivity and Costs by Industry: Selected Service-Providing Industries—2025 — U.S. Bureau of Labor Statistics — Aug 2026; Microsoft 2025 Annual Report — Microsoft — Jul 2025. sec.gov
Welfare valuation is sensitive to the estimator and denominator
The 2026 Stanford paper uses stated-preference binary choices, acknowledges hypothetical bias, and annualises estimated monthly WTA using external adoption counts; the surveys and adoption observations are not exactly contemporaneous, while Table 2 aggregates ages 18–64 rather than all adults. Calculated median-based annualisations using its table inputs are approximately $4.02 billion for 2025 and $15.89 billion for 2026, which illustrate distributional sensitivity rather than provide statistically equivalent substitutes for the mean-based totals. What is Generative AI Worth? — Stanford Digital Economy Lab — Apr 2026, pp. 4–7 and 9. digitaleconomy.stanford.edu
The earlier $97 billion appears in Collis and Brynjolfsson’s Wall Street Journal headline, but the accessible article does not expose the underlying estimation table needed to revalidate its reference-year population, mean-versus-median basis, and aggregation assumptions. Consequently, it is acknowledged as an earlier published communication rather than inserted into a harmonised time series with the subsequently documented survey estimates. AI’s Overlooked $97 Billion Contribution to the Economy — Collis and Brynjolfsson, Wall Street Journal. WSJ
The free-content adjustment concerns production, rather than willingness to pay
Nakamura, Samuels, and Soloveichik model commercially supported free content as a barter transaction, excluding content already paid for directly or bundled into paid products. Their June 2026 working paper identifies growth-pattern breaks around 1995 and 2022, but the reported adjustment covers free content beyond AI and depends on production-cost valuation and strongly assumed deflators; it is neither an adopted BEA revision nor a welfare estimate. Its relevance is that a specified accounting extension changes measured growth, rather than that a consumer-surplus number has been added to production. The Progression of “Free” Digital Content to AI: Impacts on U.S. Economic Growth and Productivity — BEA Working Paper WP2026-13 — Jun 2026, pp. 1–3 and 13–16. bea.gov
Productivity and investment observations require separate attribution
BLS reports that productivity increased in 15 of its 30 selected service-providing industries during 2025, illustrating that the official record contains heterogeneous outcomes rather than a uniform service-sector response. Such releases measure outcomes but do not isolate AI from capital intensity, demand, workforce composition, or other changes, so even a strong industry productivity result cannot identify the technology’s causal contribution without further evidence. Productivity and Costs by Industry: Selected Service-Providing Industries—2025 — U.S. Bureau of Labor Statistics — Aug 2026. 2025 A01 Results
Corporate infrastructure spending is observable, but worldwide property-and-equipment expenditure does not map directly into domestic GDP because geographic production, imports, timing, and asset coverage must be reconciled. Accordingly, AI-related construction and equipment investment can support measured growth before widespread downstream efficiency gains appear, while this report assigns no numerical share of 2025 GDP growth to AI investment without a corresponding decomposition.
Where darkness is real and where it is rhetorical
Observable profits are not an omitted output category
Retained cost savings become visible through operating profit and production-income distributions, although company accounting profit requires reconciliation with national-accounts operating surplus. The resulting capability improvement can remain incompletely measured, but describing the increased margin itself as invisible confuses the distribution of income with the measurement of service quality.
The same distinction applies to provider revenue and infrastructure expenditure: their economic interpretation can be incomplete, yet their transactions are observable. Conversely, a task’s exposure to automation does not establish that substitution occurred, that quality was preserved, or that displaced labour generated additional production elsewhere.
Service deflators can misidentify a changing production process
A defensible service-measurement failure occurs when receipts decline after cheap tasks move to an AI supplier or into household production, while the sampled prices continue to describe the more difficult tasks retained by incumbent professionals. The remaining sample can then indicate stable or increasing prices even though comparable simple services became cheaper, creating a coverage and composition problem rather than proving universal incapacity to measure service productivity.
SemiAnalysis supplies this mechanism through its legal-service discussion, but its broader suggestion that receipts-and-price accounting cannot accommodate productivity gains is too strong: price-adjusted output can remain constant while measured hours fall, and valid deflators can register cheaper comparable services. AI Dark Output: The Visible Cost of Invisible Output — SemiAnalysis — May 2026. newsletter.semianalysis.com
Quality adjustment therefore requires evidence about characteristics that users actually receive, rather than a presumption that a newer model always produces a better service. Existing hedonic methods demonstrate how characteristics can be separated from price, but their extension to cognitive services requires validated measures of task completion, reliability, review burden, and comparability. Frequently Asked Questions about Hedonic Quality Adjustment in the CPI — U.S. Bureau of Labor Statistics — accessed Oct 2026. bls.gov
Household capabilities create the clearest boundary issue
Unpaid household tutoring, translation, and administrative assistance can increase capability without becoming additional household service production inside GDP, whereas an internal business improvement can affect measured final output or recognised assets even without a separate sale. The strongest defensible claim is therefore that some capabilities lack a separate statistical representation, rather than that every unpriced AI-assisted task lies outside the accounts.
The analogy with “dark matter” is limited to incomplete observation, because these economic gaps contain heterogeneous accounting and welfare objects rather than one homogeneous residual.
The critical argument corrects exaggeration but does not remove the boundary problem
James Broughel’s June 2026 critique locates much of the problem in price indices, which correctly challenges the implication that falling prices automatically make production invisible. Nevertheless, improving deflators cannot by itself bring unpaid household services inside the production boundary or transform consumer surplus into production income, so the appropriate remedy depends on whether the omission concerns price measurement, output coverage, or welfare valuation. The Real Reason AI Doesn’t Show Up In The GDP Statistics — James Broughel, Forbes — Jun 2026. forbes.com
Income can also decline while physical productivity rises, as the passed-through numerical case demonstrates, because producing an unchanged service with fewer resources does not guarantee preservation of its selling price or the previous distribution of earnings. Likewise, resource savings are not automatically net social gains when verification costs, transition losses, and displaced workers’ subsequent outcomes remain unresolved.
Implications for investors and policymakers
Four questions require four bodies of evidence
The analytical error is using one statistic to answer questions about demand, implementation, surplus capture, and aggregate production, because a welfare survey cannot establish supplier cash generation and a capital-expenditure figure cannot establish successful deployment.
| Decision question | Evidence capable of answering it | Evidence insufficient on its own |
|---|---|---|
| Is demand for the technology real? | Sustained usage, completed tasks, renewals, paid demand, and consequential access valuations | Infrastructure spending or aggregate GDP |
| Are complements being built? | Implemented workflow redesign, integration, training, and identifiable organisational assets | Announced hardware budgets or token counts |
| Who captures the surplus? | Price pass-through, wages, margins, supplier economics, and user valuations | Productivity growth or provider revenue alone |
| Is aggregate productivity rising? | Quality-adjusted domestic output relative to measured labour and capital inputs | WTA aggregates, task benchmarks, or corporate capex |
For institutional analysis, the relevant discipline is to require an observable transmission mechanism between capability and the outcome under examination. Supplier revenue establishes monetisation but does not reveal the user’s return; lower unit cost establishes an operating improvement but does not establish market expansion; higher consumer surplus establishes a benefit under the valuation method but does not establish additional taxable income.
For policymakers, the corresponding distinction separates welfare improvement from fiscal capacity and market-production growth. Free household assistance can improve access to capabilities, while policy still requires separate evidence about employment income, public revenue, and resource allocation; assuming that a welfare gain automatically finances adjustment costs would erase precisely the distinctions the measurement framework is intended to preserve.
Cross-country application requires local observation
The conceptual architecture applies beyond the United States, but adoption rates, willingness to accept, employment arrangements, language performance, and service-production boundaries cannot be assumed identical. The following proposals identify possible national measurement priorities rather than assert comparative adoption levels or estimate national dark output.
| Jurisdiction | Illustrative measurement priority | Evidence needed before inference |
|---|---|---|
| Italy | AI-assisted administrative and professional tasks | Representative task samples, actual review hours, and comparable completed services |
| France | Public-service and professional-service quality | Outcome measures separated from production costs and household welfare |
| Germany | Technical documentation and industrial-service integration | Allocation between intermediate activity, marketed output, and capitalised assets |
| United Kingdom | Professional services and digitally enabled household activity | Consistent service units and household time-use evidence |
| European Union collectively | Comparable digital supply-use and quality classifications | Common definitions reconciled with national compilation methods |
International coordination has an existing institutional foundation: the 2025 SNA update recommends a suite of indicators covering AI and other digital activities, although adoption of a conceptual standard does not establish that every country has implemented comparable estimates. New Standards for Economic Data Aim to Sharpen View of Global Economy — International Monetary Fund — Jul 2025. imf.org
A measurement agenda
Task-level time-use should measure completion and redeployment
Statistical agencies should collect repeated observations of comparable tasks, recording time spent producing, checking, correcting, and approving outputs, alongside whether the activity substitutes for prior work or represents a newly feasible service. The relevant denominator is total effective labour input rather than prompting time alone, while saved time should be classified according to redeployment into production, leisure, reduced paid hours, or transition between jobs.
Representative sampling is essential because provider records overrepresent users of the observed platform and task benchmarks do not necessarily represent normal working conditions. Privacy-preserving collection should distinguish economic activity and output characteristics without requiring disclosure of confidential document contents.
Service indices should follow equivalent completed outcomes
Experimental service indices should begin with categories for which completion and quality can be observed consistently, such as a translated document with a specified error tolerance or a software feature passing a stable acceptance test. Comparing an unreviewed AI draft with an accountable professional service would otherwise count the removal of review or responsibility as a quality-preserving price decline.
Pilot indices should run alongside existing series, disclose their treatment of changing task difficulty, and publish sensitivity to alternative quality weights before becoming official deflators. This approach applies the established distinction between pure price movements and changing characteristics without assuming that laboratory performance supplies the relevant economic valuation. Hedonic Models in the Producer Price Index — U.S. Bureau of Labor Statistics — Jun 2011. bls.gov
A GDP-B satellite account should preserve its welfare identity
A supplementary welfare account should disclose the access restriction being valued, the experimental consequences, the valuation distribution, the eligible population, and assumptions behind annualisation. Mean and median results should be presented together, while dependence on alternative tools, income, workplace access, and household circumstances should be examined rather than concealed inside a headline aggregate.
Production extensions and welfare accounts should remain distinguishable: an imputed exchange valued consistently with production accounting differs from a monetary valuation of benefits from access. The United Nations’ endorsed guidance on free digital products in satellite accounts provides an institutional basis for supplementary treatment without pretending that all such benefits belong in the core production measure. Recording and Valuing “Free” Digital Products in an SNA Satellite Account — United Nations Statistics Division, endorsed guidance — 2023. unstats.un.org
Firm disclosure should reconcile AI expenditure with labour and output
Voluntary statistical modules could distinguish AI operating expenditure, capitalised software, complementary integration costs, realised labour-hour changes, throughput, quality, and review burdens. Claimed hours displaced should be reported separately from actual reductions in total paid hours, because faster execution of selected tasks does not establish an equivalent change in firm-wide inputs.
Such disclosure would impose collection and confidentiality costs, so implementation should begin with standardised confidential surveys and carefully selected pilots. The public output should aggregate comparable measures rather than encourage firms to publish unverifiable monetary valuations of every internal AI task.
Growth decomposition should separate building AI capacity from using it
Agencies should identify infrastructure production and investment separately from downstream AI adoption, while reconciling imported equipment, domestic construction, capital services, and complementary intangible assets. This would distinguish growth generated by building capacity from productivity gains obtained through its use, without treating the former as either proof or disproof of the latter.
The measurement programme should also avoid classifying every intangible as currently missing, because software and other recognised assets already enter investment accounts. Its purpose is to locate specific residual gaps in coverage and valuation, consistent with the SNA digitalisation work’s treatment of AI systems and encouragement of more granular reporting. Digitalisation — System of National Accounts update documentation, United Nations Statistics Division — 2025. 2025_SNA_CH22_BPM7_CH16_Digitalisation_V11
Principal gaps, watch indicators, and decision thresholds
| Gap capable of changing the assessment | Observation needed | Interpretation threshold |
|---|---|---|
| Comparable AI-assisted service prices | Repeated matched-service prices including review and responsibility | Persistent divergence from the existing deflator |
| Conversion of task savings into production | Firm-level hours, output, and redeployment over time | Sustained gains beyond pilot tasks and implementation costs |
| Reliability-adjusted quality | Independent completion and error measures | Improvement preserved after checking and correction |
| Welfare valuation robustness | Consequential experiments and alternative population estimates | Benefits robust to design, distribution, and user-count assumptions |
| Investment-versus-use contribution | Domestic production and capital-service decomposition | Separately identified downstream productivity effects |
These thresholds concern the evidence required for interpretation rather than predetermined numerical targets, because assigning arbitrary percentages would reproduce the false precision the agenda is intended to correct.
Conclusion
The defensible claim is that AI shifts some economic value from paid services towards capabilities whose usefulness is only partially represented by transaction prices, while other gains already appear through lower input requirements, profits, investment, or marketed output. Production accounts inherited from twentieth-century frameworks will record parts of this transition late, incompletely, and sometimes with the wrong measured sign when coverage or quality-adjusted deflators fail, although the nominal decline associated with a genuinely cheaper final service is not itself an accounting error.
A rigorous assessment therefore preserves the distinction between a price decline, a productivity improvement, a redistribution of income, and a welfare gain. The remedy is a set of complementary accounts that identifies where each occurs and who benefits, rather than a single corrected GDP enlarged by token expenditure, task exposure, consumer-surplus estimates, and investment figures that measure different economic objects.
Appendix — One-page conceptual and numerical reference
Glossary
| Term | Consistent definition |
|---|---|
| Measured output | Production within the national-accounts boundary, valued at transaction prices or recognised imputations; GDP measures domestic value added, whereas industry gross output includes intermediate production. |
| Real output and productivity | Quantity- and quality-adjusted production, with productivity relating output to measured inputs; gains can appear through better output measurement or lower input requirements. |
| Firm productivity and margins | Task efficiency, throughput, error rates, labour requirements, and the share of receipts retained after costs; observable margin gains are not inherently dark output. |
| Consumer surplus | Willingness to pay minus price paid; access WTA supplies a related welfare valuation whose equivalence depends on method and assumptions, and neither is automatically added to GDP. |
| Newly feasible activity | Activity not previously undertaken because its cost exceeded the relevant benefit or budget; its treatment depends on whether it becomes market output, eligible own-account production, or household activity. |
| Substitution and creation dark output | Substitution replaces previous work, whereas creation enables additional activity; either can be visible, partly measured, or outside the production boundary. |
Accounting definitions follow the BEA Industry Accounts Primer — May 2011, while the welfare distinction follows GDP-B: Accounting for the Value of New and Free Goods — Stanford Digital Economy Lab — Mar 2025. bea.gov
Four-case numerical reference
Assumptions: The baseline final service costs $1,000, entails $800 labour expenditure, yields $200 profit, and has willingness to pay of $1,500; AI-assisted operating cost becomes $200, including a $20 domestic intermediate input, with unrelated activity held constant.
| Case | Nominal GDP change | Real final-output change | Firm-profit change | Consumer-surplus change |
|---|---|---|---|---|
| Saving fully passed through: price $400 | −$600 | $0 if comparable price decline is captured | $0 | +$600 |
| Saving retained: price $1,000 | $0 | $0 | +$600 | $0 |
| Quality multiplier 1.5; price $1,000; WTP $1,800 | $0 | +$500 if quality adjustment is valid | +$600 | +$300 |
| Separate newly feasible free household task; WTP $250 | $0 incremental under stated assumptions | $0 incremental inside GDP | $0 assumed | +$250 gross benefit |
Interpretation: The first three rows compare the same marketed final service, while the fourth adds a separate household capability; employee use instead requires tracing effects through the employer’s output, inputs, or assets. The calculations demonstrate conditional accounting mechanisms rather than estimate AI’s aggregate economic contribution.
Open-source analytical assessment · Visual scheme
Dark Output
What national accounts miss when artificial intelligence shifts value from prices to capabilities
The conceptual architecture
Introduction: six objects, distinct boundaries
The economic significance of an AI-assisted activity depends on what changes: its price, completed volume, quality, production inputs, or the set of activities that users can perform. A change in one object does not establish an equivalent change in the others.
Measured output
Market production and recognised imputations; GDP measures domestic value added, while gross output also includes intermediate production.
Real output and productivity
Quantity- and quality-adjusted production relative to labour or other inputs; lower measured hours can reveal gains even when final output stays constant.
Productivity and margins
Unit costs, throughput, error rates and operating profits; retained margins are observable rather than an inherently missing output category.
Consumer surplus
Benefit above the price paid; willingness to accept loss of access supplies a related valuation, with method-dependent assumptions.
Newly feasible activity
Tasks not previously undertaken because their cost exceeded the relevant benefit or budget; household and market activity have different treatments.
Substitution and creation
Substitution replaces prior work, while creation enables additional activity; either can be recorded, partly captured, or outside the production boundary.
Source: Measuring the Nation’s Economy: An Industry Perspective — BEA — May 2011.
Source: AI Dark Output: The Visible Cost of Invisible Output — SemiAnalysis — May 2026.
A short intellectual history
From the Solow paradox to ICT
The late-1990s American productivity acceleration shows that widespread use can precede measurable aggregate gains, but that historical sequence does not establish AI’s future magnitude or timing.
The productivity J-curve
Integration, training and organisational redesign require complementary investment whose creation can be incompletely recorded, while implementation costs arrive before productive returns.
Source: The Productivity J-Curve — Brynjolfsson, Rock and Syverson — October 2018.
Free goods and GDP-B
Welfare valuation complements production accounts; incentive-compatible experiments in the GDP-B literature must be distinguished from stated-preference surveys with hypothetical bias.
Source: GDP-B: Accounting for the Value of New and Free Goods — Stanford Digital Economy Lab — March 2025.
The empirical constraint
Earlier research challenged the claim that IT mismeasurement explains the post-2004 productivity slowdown, so any large AI residual requires evidence matched to its claimed scale.
Source: Challenges to Mismeasurement Explanations for the U.S. Productivity Slowdown — June 2016.
Illustrative accounting · Not empirical legal-service data
Formal framework: identical technology, different accounts
Nominal gross output: price × number of services.
Quality-adjusted volume: quality multiplier × number of services.
Value added: gross output − intermediate purchases.
Consumer surplus: willingness to pay − price paid.
Reference-price examples simplify official chain-index methods. Summing industry receipts double-counts intermediate purchases; domestic value added is the relevant GDP aggregate.
Source: Measuring the Nation’s Economy: An Industry Perspective — BEA — May 2011.
Baseline: one household legal memorandum, price $1,000, cost $800, profit $200 and willingness to pay $1,500. AI-assisted cost is assumed to fall to $200, including $20 of domestic supplier value added; other activity is held constant. The quality case assumes a multiplier of 1.5 and willingness to pay of $1,800.
Graph: four accounts, four market-service cases
Levels in illustrative US dollars; common 0–1,500 scale across all panels. Real output uses baseline service prices.
Nominal GDP contribution
Real final output
Legal firm operating profit
Consumer surplus
| Market-service case | Nominal GDP | Real output | Profit | Consumer surplus |
|---|---|---|---|---|
| Baseline | 1000 | 1000 | 200 | 500 |
| Passed through | 400 | 1000 | 200 | 1100 |
| Retained margin | 1000 | 1000 | 800 | 500 |
| Higher quality | 1000 | 1500 | 800 | 800 |
The separate free-household case
A newly feasible activity with zero payment and assumed willingness to pay of $250 adds $250 gross consumer benefit but no incremental GDP under unchanged provider monetisation and costs. A $30 participation burden reduces net benefit to $220. Employee use must instead be traced through employer output, inputs or assets.
This is a separate activity, not a free replacement of the $1,000 memorandum; it is excluded from the market-service graph to avoid comparing incompatible baselines.
Interactive accounting simulator
Adjust assumptions to inspect the isolated market-service mechanism. Quality-adjusted real output assumes a valid deflator; willingness to pay is illustrative, not inferred from token expenditure.
Sourced estimates · Not additive
What the current estimates actually measure
Graph: mean versus median sensitivity
Annualised US access valuations, USD billions; linear scale 0–180. Median annualisations are calculated sensitivity illustrations, not equivalent aggregate-surplus estimators.
Source: What is Generative AI Worth? — Stanford Digital Economy Lab — April 2026, Table 2.
Table 2 population: US users aged 18–64. Monthly means: $98 / $124.50; medians: $3.39 / $11.48. Users: 98.78m / 115.33m. Assumptions include external adoption counts, annualisation, and stated preferences; hypothetical bias and unmatched observation dates remain relevant.
Graph: experimental free-content adjustment
Additional average annual US real GDP growth, percentage points per year; interval averages, not annual observations. Linear scale 0–0.25.
Source: The Progression of “Free” Digital Content to AI — BEA WP2026-13 — June 2026.
Advertising- and marketing-supported free content, including AI; a research adjustment rather than an adopted BEA revision. It is production-based, not consumer surplus, and depends on imputation and deflator assumptions.
| Period | Mean-based aggregate, $bn | Median sensitivity, $bn | Users, millions |
|---|---|---|---|
| 2025 | 116.2 | 4.02 | 98.78 |
| 2026 | 172.3 | 15.89 | 115.33 |
| Interval | Added average annual real GDP growth |
|---|---|
| 1929–1995 | 0.04 percentage point/year |
| 1995–2022 | 0.09 percentage point/year |
| 2022–2025 | 0.22 percentage point/year |
| Measure | Verified observation | Interpretation boundary |
|---|---|---|
| Task exposure | Approximately $1.5 trillion, SemiAnalysis, May 2026 | Potential augmentation or automation; not realised output |
| Industry productivity | US software publishers: +12.9%, 2025, BLS | Measured result; no causal AI attribution |
| Infrastructure expenditure | Microsoft: $64.551bn cash property/equipment additions, FY2025 | Consolidated worldwide expenditure; not AI-only or domestic GDP |
Source: AI Dark Output: The Visible Cost of Invisible Output — SemiAnalysis — May 2026.
Source: Selected Service-Providing Industries—2025 — BLS — August 2026.
Source: Microsoft 2025 Annual Report — July 2025.
Where darkness is real—and where it is rhetorical
Observable transactions and income
Retained profits, purchased AI services and infrastructure investment already have accounting counterparts; their distribution and AI-specific attribution require reconciliation, rather than a presumption of invisibility.
Specific gaps and boundaries
Household capabilities can lie outside the production boundary, while unmatched service prices and missed quality changes can distort real output; internal firm activity must be followed through marketed outputs or assets.
Price-index reform addresses comparability, but it does not transform welfare into production income. Lower receipts can coexist with unchanged service volume, and physical productivity can rise while nominal income falls when prices decline.
Source: The Real Reason AI Doesn’t Show Up In The GDP Statistics — James Broughel, Forbes — June 2026.
Source: Frequently Asked Questions about Hedonic Quality Adjustment in the CPI — BLS.
Implications for investors and policymakers
| Question | Relevant evidence | Insufficient on its own |
|---|---|---|
| Is demand real? | Sustained use, renewals, completed tasks, paid demand | Capex or aggregate GDP |
| Are complements being built? | Integration, training and implemented workflow changes | Hardware budgets or token counts |
| Who captures the surplus? | Pass-through, wages, margins and user valuations | Productivity growth alone |
| Is productivity rising? | Quality-adjusted domestic output relative to inputs | WTA surveys or corporate capex |
Italy, France, Germany and the United Kingdom require local task samples, comparable service measures and household evidence; the EU coordination priority is common definitions and consistent digital supply-use classifications. American willingness-to-accept totals are not estimates for those economies.
Source: New Standards for Economic Data Aim to Sharpen View of Global Economy — IMF — July 2025.
A measurement agenda
Task-level time-use
Include completion, review, correction and the actual redeployment of saved hours.
Quality-adjusted services
Compare equivalent completed outcomes with consistent responsibility and review requirements.
Welfare satellite account
Disclose mean, median, experiment design, population and annualisation assumptions.
Firm input disclosure
Separate AI opex, capitalised assets, implementation expenditure and realised labour-hour changes.
Investment versus use
Separate infrastructure production from capital services and downstream deployment gains.
Watch indicators
Require sustained matched-service gains, verified quality and robust valuation before revising the aggregate assessment.
Source: Recording and Valuing “Free” Digital Products in an SNA Satellite Account — UN endorsed guidance.
Conclusion
Appendix: the four numerical cases and their assumptions
| Case | Δ nominal GDP | Δ real output | Δ firm profit | Δ consumer surplus |
|---|---|---|---|---|
| Full price pass-through | −$600 | $0 | $0 | +$600 |
| Retained margin | $0 | $0 | +$600 | $0 |
| Higher quality | $0 | +$500 | +$600 | +$300 |
| Separate free household activity | $0 incremental | $0 within GDP | $0 assumed | +$250 gross |
The first three cases concern one marketed service; the fourth adds a separate household activity. Real output requires the assumed quality adjustment; the household case assumes unchanged provider monetisation and costs. Broader GDP effects depend on demand, imports, expenditure reallocation and displaced workers’ subsequent production.
Dark Output: Concepts and Accounting Mechanisms
The central accounting question is where an AI gain appears after production is reorganised. A reduction in professional-service revenue may represent a lower price for unchanged output, a transfer of value added to another industry, an import substitution, or a genuine contraction. These mechanisms require different interpretations. None establishes, by itself, an amount of economic value missing from GDP.
Index
| Section | Focus |
|---|---|
| Introduction | Accounting boundaries and the evidence needed to identify missing output |
| A short intellectual history | The transition from measuring production inputs to measuring service outcomes |
| Formal framework | Value-added transfers, demand responses, billing units, quality, and internal production |
Introduction
The next step in the dark-output argument is to examine the complete production chain rather than the revenue of the activity being replaced. When AI replaces a service purchased by a household, lower expenditure can reduce nominal final consumption. When it replaces a service purchased by a business, the reduction in intermediate expenditure can increase the purchasing business’s value added. A disappearing invoice therefore has different aggregate consequences depending on who previously paid it.
GDP aggregates domestic value added, whereas gross output also records transactions between producers. Consequently, declining gross output in one service industry need not imply an equivalent decline in GDP. BEA, Measuring the Economy: A Primer on GDP and the NIPAs. bea.gov
A useful empirical starting point is the scale of the accounts already recording production and income.
Table — The recorded US economy before any proposed dark-output adjustment
| Measure | Reference year and coverage | Magnitude | Interpretation |
|---|---|---|---|
| Nominal GDP | 2025; US domestic production | $30,861.3 billion | Recorded production at current prices |
| Nominal gross domestic income | 2025; income generated by US domestic production | $30,907.7 billion | Independently estimated income counterpart |
| GDP minus GDI | 2025; calculation from the two BEA estimates | −$46.4 billion | Statistical discrepancy, approximately −0.15% of GDP |
| Private nonresidential software fixed investment | 2025; US investment expenditure | $751.3 billion | Recorded software investment, including purchased and own-account software; not an AI-only measure |
Estimator and vintage: BEA, September 30, 2026 annual-update methodology summary. These are national-account aggregates, not survey means or medians. The discrepancy and its percentage are calculations from the published totals. BEA, Updated Summary of NIPA Methodologies, September 2026. apps.bea.gov
The GDP–GDI discrepancy is not a candidate estimate of dark output. The two aggregates should coincide conceptually, but their independently collected source data produce differences in practice. An AI capability that generates neither a recorded transaction nor recorded income can be absent from both sides. Conversely, a discrepancy can arise without any AI measurement problem. BEA, Why do GDP and GDI differ, and what does that imply?. U.S. Bureau of Economic Analysis (BEA)
The analytical burden is therefore more demanding than identifying a gap between technological enthusiasm and aggregate statistics. A defensible claim must specify:
| Required identification | Question to resolve | Consequence of leaving it unresolved |
|---|---|---|
| Purchaser | Was the former service bought by a household, business, or government? | Final expenditure and intermediate expenditure become conflated |
| Domestic production boundary | Where is the replacement service produced? | Foreign value added may be attributed to domestic GDP |
| Service unit | Is output an hour, a document, a resolved issue, or a completed engagement? | Input reductions may be interpreted as output reductions |
| Quality standard | Is the replacement equally accurate, reliable, and complete? | More generated material may be mistaken for more useful output |
| Destination of saved resources | Are hours redeployed, paid but idle, or eliminated? | Task efficiency may be mistaken for realised firm productivity |
| Accounting treatment | Is the activity current production, intermediate consumption, or capital formation? | Internally produced assets may be incorrectly classified as wholly invisible |
These conditions turn “dark output” into a testable measurement proposition. They prevent it from becoming a general explanation for every instance in which AI appears useful but recorded growth remains modest.
A short intellectual history
The unresolved issue is the unit of service output
One lesson from earlier technological transitions is that a production input can remain easy to observe after it becomes a poor proxy for output. Professional hours are relatively straightforward to count. The quantity and quality of completed advice are harder to standardise.
This difficulty predates generative AI. The Eurostat–OECD methodological guide for service producer prices distinguishes time-based pricing from methods intended to follow specified services. In its discussion of architectural services, it explains that hourly charge rates can fail to capture productivity changes when the time required to deliver a service changes. Representative-service and model-pricing approaches offer alternatives, although maintaining comparable specifications is demanding. Eurostat–OECD, Methodological Guide for Developing Producer Price Indices for Services, second edition, 2014. oecd.org
AI intensifies this existing problem because it can change the relationship between paid time and delivered work rapidly. A professional may charge a higher hourly rate while requiring far fewer hours to complete an engagement. An hourly price index can rise even while the price of the completed service falls.
The historical record also cautions against treating service industries as organisationally fixed. A BLS examination of legal services after the financial crisis describes changes in outsourcing, staffing, technology use, and reliance on in-house lawyers. Such changes complicate the interpretation of industry receipts even before AI enters the analysis. Joseph Valentine, BLS, Producer prices in the legal services industry after the Great Recession, November 29, 2019. bls.gov
Table — How organisational change alters the measurement question
| Production arrangement | Observable transaction | Additional information needed |
|---|---|---|
| Specialist firm provides the entire service | Fee paid to the specialist | Completed-service quantity and quality |
| Client purchases the service but automates part of its preparation | Specialist fee plus AI expenditure | Changes in task scope and duplication of work |
| Client brings the activity in-house | Employee compensation and purchased inputs | Internal output and its contribution to marketed production |
| Household performs the activity using AI | Subscription expenditure, if any | Household time, useful outcomes, and welfare benefit |
| Firm develops a reusable internal software asset | Development expenditure and eligible own-account investment | Asset boundary, development inputs, and subsequent capital services |
This history supports a narrower proposition than “new technologies evade national accounts.” Production accounts already accommodate organisational change and some forms of internal production. Their vulnerability increases when the observable transaction ceases to represent a stable service.
Direct task evidence improves identification, but does not complete aggregation
The study Generative AI at Work provides a useful example of measuring an outcome rather than merely counting AI use. Brynjolfsson, Li, and Raymond examine the staggered introduction of an assistant among 5,172 customer-support agents. The authors’ updated summary reports an average 15% increase in issues resolved per hour, with substantial differences across workers. This is an average estimated workplace effect, not a median effect, a nationally representative US estimate, or an aggregate GDP contribution. The paper was subsequently published in the Quarterly Journal of Economics in 2025. Stanford Digital Economy Lab, Generative AI at Work; published article, 2025. Stanford Digital Economy Lab
The measurement advantage is the denominator: the study relates useful service outcomes to labour time. The remaining aggregation problem concerns what happens to those gains.
| Finding at task level | What it establishes | What remains unidentified |
|---|---|---|
| More issues resolved per hour | Higher measured throughput in the studied setting | Whether firm revenue or total output increases |
| Faster completion | Lower time required for the measured task | Whether saved time is redeployed |
| Different effects across workers | Heterogeneous benefits from assistance | Economy-wide workforce-weighted gains |
| Improved service performance | A potentially valuable quality change | Its treatment in official output and price indices |
| Adoption within a particular firm | Feasibility in that production environment | Transferability across industries and economies |
The historical advance is thus methodological: AI effects can increasingly be studied at the level of completed tasks. National-account interpretation still requires a bridge from task outcomes to establishments, industries, and domestic value added.
Formal framework
Follow the value added through the production chain
Consider a manufacturer selling $10,000 of final products. Initially, it purchases $1,000 of domestic consulting services. It subsequently replaces that purchase with $100 of domestic AI services.
Assumptions: final sales and product quality remain unchanged; there are no other intermediate inputs; both service suppliers’ receipts equal their value added; all production occurs within the same period. These are constructed accounting examples, not empirical estimates.
Table — Business-service substitution can leave nominal GDP unchanged
| Accounting item | Before substitution | Domestic AI replacement | Change |
|---|---|---|---|
| Manufacturer’s final sales | $10,000 | $10,000 | $0 |
| Manufacturer’s purchased intermediate service | $1,000 | $100 | −$900 |
| Manufacturer’s value added | $9,000 | $9,900 | +$900 |
| Consultant’s value added | $1,000 | $0 | −$1,000 |
| AI supplier’s value added | $0 | $100 | +$100 |
| Total domestic value added | $10,000 | $10,000 | $0 |
| Combined gross output across establishments | $11,000 | $10,100 | −$900 |
The consulting industry contracts, and combined gross output falls. Aggregate nominal GDP does not fall because the manufacturer’s higher value added offsets the change elsewhere.
This example establishes neither a welfare gain nor its distribution. Those depend on service equivalence, implementation costs, displaced workers, and the use of released resources. It establishes that lost intermediate-service revenue cannot be treated as an equal loss of GDP. The accounting basis is BEA’s distinction between gross output and value added. BEA, Measuring the Economy. bea.gov
Now change the AI supplier’s location.
| Accounting item | Domestic AI supplier | Foreign AI supplier |
|---|---|---|
| Final expenditure on manufactured products | $10,000 | $10,000 |
| Manufacturer’s domestic value added | $9,900 | $9,900 |
| Replacement supplier’s domestic value added | $100 | $0 |
| Imported AI service | $0 | $100 |
| Domestic GDP | $10,000 | $9,900 |
Under these assumptions, imported replacement services reduce domestic value added by $100 relative to domestic supply. That is a production-location effect, not evidence that an additional $900 of domestic output has become invisible. Imports are deducted in the expenditure approach to exclude foreign production from domestic GDP. BEA, The Expenditures Approach to Measuring GDP, June 3, 2025. U.S. Bureau of Economic Analysis (BEA)
Price pass-through does not determine revenue without a demand response
Let the initial price be $200 and unit cost $120. AI reduces unit cost to $60. If the entire $60 saving is passed through, the new price is $140.
Revenue then depends on how demand responds. Under an illustrative constant-elasticity demand specification,
Here, ε is the absolute price elasticity of demand.
Table — The same cost reduction can produce falling, unchanged, or rising receipts
| Assumed demand elasticity | Quantity after price reduction | Receipts after reduction | Change from initial $20,000 | Operating profit after reduction |
|---|---|---|---|---|
| 0 | 100.00 services | $14,000 | −30.0% | $8,000 |
| 0.5 | 119.52 services | $16,733 | −16.3% | $9,562 |
| 1.0 | 142.86 services | $20,000 | 0.0% | $11,429 |
| 1.5 | 170.75 services | $23,905 | +19.5% | $13,660 |
Constructed sensitivity analysis: initial quantity is 100; initial operating profit is $8,000. The calculations assume full pass-through, unchanged service quality, constant unit costs, no fixed or transition costs, and the specified demand response. Fractional quantities represent aggregate service volume.
The results qualify the substitution-dark-output claim. Lower prices do not mechanically reduce nominal receipts. They do so when the quantity response is insufficient to offset the price decline. Nor do the receipts establish GDP effects: a final service, an intermediate service, and an imported service occupy different accounting positions.
Even when receipts fall, correctly measured real output can rise. In the elasticity-0.5 case, nominal receipts decline by 16.3%, while service quantity rises by 19.5%. That combination is consistent with an ordinary price-and-volume decomposition. It becomes a measurement problem when the deflator fails to recognise the relevant service price.
Hourly billing can produce a double measurement failure
Suppose an engagement initially requires ten hours at $200 per hour. After AI adoption, the same engagement requires four hours at $220 per hour.
The hourly rate increases by 10%, but the price per completed engagement falls by 56%.
Assumptions: 100 comparable engagements are completed in both periods; quality and complexity are unchanged; hours include all relevant production time; there are no omitted review hours.
Table — Input-hour pricing versus completed-service pricing
| Measure | Before AI | After AI | Change |
|---|---|---|---|
| Completed engagements | 100 | 100 | 0% |
| Hours per engagement | 10 | 4 | −60% |
| Total labour hours | 1,000 | 400 | −60% |
| Hourly billing rate | $200 | $220 | +10% |
| Price per engagement | $2,000 | $880 | −56% |
| Nominal receipts | $200,000 | $88,000 | −56% |
| Real gross output using a matched-engagement deflator | $200,000 | $200,000 | 0% |
| Apparent real gross output using only the hourly-rate deflator | $200,000 | $80,000 | −60% |
| Completed engagements per labour hour | 0.10 | 0.25 | +150% |
| Hourly-deflated receipts per labour hour | $200 | $200 | 0% |
The same mistaken deflator can therefore conceal both unchanged service volume and a substantial labour-productivity gain. Apparent real output falls in proportion to hours, leaving the resulting output-per-hour measure unchanged.
This is a demonstration of a possible gross-output measurement failure. It is not an estimate of bias in a current official legal-services index or aggregate GDP. Moving from gross output to real value added also requires appropriate measurement of intermediate inputs. The underlying concern about hourly pricing and productivity is documented in the Eurostat–OECD service-price methodology guide. oecd.org
WordPress graph: dark-output-billing-graph.html
The file contains only the graph fragment, its assumptions, source links, and an expandable data table. Paste its contents into a WordPress Custom HTML block at this point in the document.
The relevant unit cost includes verification and rework
A generated draft is an intermediate step. The relevant economic unit is the accepted service after necessary checks.
Consider a task requiring three human hours at $50 per hour before AI. Its initial labour cost is $150.
Table — A constructed cost calculation per accepted task
| Cost component | Assumption | Cost |
|---|---|---|
| Preparation and prompting | 0.10 hour × $50 | $5 |
| Professional review | 1.20 hours × $50 | $60 |
| Rework | 0.30 hour × $50 | $15 |
| Expected escalation | 10% probability × two additional hours × $50 | $10 |
| Model expenditure | Assumed expenditure per task | $2 |
| Total AI-assisted cost | $92 | |
| Initial human-only cost | Three hours × $50 | $150 |
| Saving per accepted task | $150 − $92 | $58, or 38.7% |
A calculation using only preparation and model expenditure would report a $7 cost and a 95.3% saving. Including the stipulated review, rework, and escalation requirements reduces the saving to 38.7%.
These figures are illustrative assumptions, not findings about any profession. They show why token prices cannot establish the cost of completed output.
Beyond that threshold, the AI-assisted workflow is more expensive in this example. If errors affect service quality, further adjustment is required before comparing real output.
Quality measurement needs more than a count of generated items
A single quality multiplier is convenient for exposition but demanding in practice. Two services can differ in accuracy, timeliness, coverage, and reliability. Aggregating those differences requires explicit weights or a clearly specified comparable service.
| Quality dimension | Candidate observable | Identification problem |
|---|---|---|
| Accuracy | Independently verified error rate | Averages can conceal consequential errors |
| Completeness | Required components successfully covered | Longer output need not be more complete |
| Timeliness | Time to an accepted result | Faster delivery may have different value across users |
| Reliability | Failure rates across repeated comparable tasks | Selected demonstrations omit unsuccessful attempts |
| Complexity | Performance within task-difficulty categories | A shift towards easier tasks inflates apparent productivity |
| Downstream effectiveness | Resolution, implementation, or sustained use | Outcomes also depend on the user and surrounding organisation |
For measurement, the appropriate starting point is accepted, comparable output, with unsuccessful attempts and review costs recorded alongside it. If the task mix changes, a before-and-after count is insufficient.
This also limits claims about creation dark output. A newly feasible activity demonstrates expanded capability. Its economic value still depends on usefulness, reliability, opportunity cost, and whether the activity replaces something already counted.
Internal use has several accounting destinations
The claim that AI used inside a firm “never enters GDP” is too broad. Its contribution may appear through marketed output, profits, lower input requirements, or eligible own-account investment.
BEA defines own-account investment as fixed assets produced by an establishment for its own use, including software. Internal production therefore does not automatically lie outside the production boundary. BEA, Own-account investment. U.S. Bureau of Economic Analysis (BEA)
Table — Where internal AI activity can enter the accounts
| Internal activity | Potential accounting destination | What can remain poorly measured |
|---|---|---|
| AI-assisted production of goods sold to customers | Market output and intermediate expenditure | Quality improvement at unchanged prices |
| Faster routine administration | Lower inputs, higher margins, or redeployed capacity | Useful work performed within fixed paid hours |
| Development of eligible reusable software | Own-account fixed investment | Capability improvement relative to development cost |
| Experimental workflow redesign | Current expenses or qualifying investment, depending on the activity | Organisational benefits without a separate transaction |
| Household tutoring or planning | Purchased subscription, where applicable | Unpriced household outcomes and time savings |
Own-account software presents a particularly important measurement issue. BEA’s current methodology estimates its production using cost information, including occupational employment data. If AI reduces development inputs for a comparable asset, current-dollar production costs can fall. Whether real investment captures unchanged or improved capability depends on the volume and deflation methods. This is a conditional measurement concern, not proof that existing software investment is systematically understated. BEA, Updated Summary of NIPA Methodologies, September 2026. apps.bea.gov
Saved task time is not automatically realised productivity
Finally, distinguish three denominators:
BLS defines labour productivity as real output per hour. Its total-factor-productivity framework also accounts for relevant non-labour inputs and labour composition; labour productivity alone does not identify an efficiency gain after all inputs are considered. BLS, Labour Productivity and Costs: Concepts. bls.gov
Table — Different uses of saved time produce different outcomes
| Response to faster task completion | Task efficiency | Firm-level implication |
|---|---|---|
| Employees produce more accepted output within existing hours | Rises | Labour productivity can rise |
| Total hours fall while comparable output is maintained | Rises | Labour productivity rises |
| Employees remain on payroll with paid idle time | Rises | Firm output per total hour may remain unchanged |
| Time is redirected to training or redesign | Rises locally | Current marketed output may change little |
| Verification absorbs the initial time saving | Drafting efficiency rises | End-to-end productivity may not rise |
| Labour is replaced by additional computing and purchased services | Labour efficiency can rise | Total-factor-productivity effects require broader input measurement |
The strongest version of the dark-output claim consequently requires evidence at successive levels: accepted tasks improve; gains survive verification and implementation costs; saved resources are productively used or genuinely released; and official quantities or deflators fail to reflect the resulting production change.
The defensible conclusion for this pillar is conditional. AI can move value away from billed hours and towards completed capabilities. Some of that change is already visible as income, value added, or investment. Some is household welfare outside GDP’s purpose. Some can be missed because service units and quality adjustments remain inadequate. Identifying which mechanism operates is the necessary step before assigning either a macroeconomic magnitude or a measurement correction.
When hourly billing conceals a service productivity gain
Constructed example, not observed data. Every series starts at 100. The same 100 completed services require 1,000 hours before AI and 400 hours afterwards. Quality, task complexity and acceptance criteria remain constant.
Mechanism: the hourly rate rises from $200 to $220 (+10%), but hours per service fall from 10 to 4. The completed-service price falls from $2,000 to $880 (−56%). Deflating $88,000 of receipts by 1.10 gives $80,000 of apparent real output, against $200,000 initially. Deflating by the matched-service price ratio, 0.44, preserves the original service volume.
Scope: this illustrates a possible gross-output measurement failure. It does not estimate an error in any current official index or in aggregate GDP; industry value added also requires measurement of intermediate inputs.
Methodological reference: Eurostat–OECD, Guide for Developing Producer Price Indices for Services, second edition (2014). Productivity definition: BLS, Labour Productivity and Costs: Concepts.
Exact graph data and assumptions
| Metric | Before AI | After AI |
|---|---|---|
| Completed services | 100 | 100 |
| Total hours | 1,000 | 400 |
| Receipts | $200,000 | $88,000 |
| Matched-service output index | 100 | 100 |
| Hourly-deflated output index | 100 | 40 |
| Completed-service productivity index | 100 | 250 |
| Hourly-deflated productivity index | 100 | 100 |
Dark Output: Evidence and the Limits of Measurement Claims
The available evidence supports a substantial gap between AI’s value to users and the revenue earned by its suppliers, while providing much weaker support for a quantified gap between actual production and measured GDP. The decisive distinction concerns the object being estimated: consumer valuation, exposed labour expenditure, model spending, infrastructure investment, and quality-adjusted production cannot be combined into a single total of invisible economic output.
Reference date: October 5, 2026. US welfare and productivity evidence is distinguished throughout from worldwide commercial expenditure and company financial reporting.
Index
| Section | Analytical focus |
|---|---|
| What the current estimates actually measure | Estimator comparison, distributional sensitivity, production adjustments, and observable expenditure |
| Where darkness is real and where it is rhetorical | Identification standards, counterarguments, historical bounds, and evidence that would change the assessment |
What the current estimates actually measure
The estimates form separate evidence streams
The most consequential mistake in the current debate is to place figures with different denominators beside one another and interpret their relative size as evidence of an accounting failure. A willingness-to-accept estimate concerns compensation for losing access; an exposure estimate concerns labour expenditure associated with potentially affected tasks; an expenditure forecast concerns purchases within a specified market; and an infrastructure figure concerns resources committed to productive capacity. Their differences reveal something about production, distribution, and valuation, but do not constitute an unexplained national-account residual.
Table — Comparison of the principal quantitative evidence
| Evidence | Concept measured | Geography | Reference period | Magnitude | Method | What it is not |
|---|---|---|---|---|---|---|
| SemiAnalysis task exposure | Labour expenditure associated with potentially augmentable or automatable tasks | US labour data underlying the exercise | April 2026 mapping; published May 2026 | Approximately $1.5 trillion | Task mapping and classified evidence of substitution potential | Realised displacement, missing GDP, or measured welfare |
| Generative-AI willingness to accept | Consumer valuation of continued access | US; aggregation uses adults aged 18–64 | July 2025 and March 2026 survey waves | $116.2 billion and $172.3 billion annualised, using means | Choice responses, fitted valuation curve, estimated users, and annualisation | Recorded production or realised full-year income |
| GenAI model expenditure | End-user spending in the specified model market | Worldwide | 2025 forecast issued July 2025 | $14.2 billion | Gartner market forecast | Observed US consumer revenue or total AI expenditure |
| Free-content production adjustment | Additional growth under experimental barter accounting | United States | 2022–2025 | +0.22 percentage point annually in real GDP growth | Imputed production, expenditure allocation, and constructed price indices | AI-only consumer surplus or an adopted official GDP revision |
| Software-publisher productivity | Measured industry output per labour hour | United States | Calendar 2025 | +12.9% | BLS industry output and hours estimates | A causal estimate of AI’s contribution |
| Microsoft property-and-equipment additions | Company cash expenditure on fixed assets | Worldwide consolidated operations | Fiscal year ended June 30, 2026 | $115.948 billion | Company cash-flow reporting | US AI-only investment or benefits from using AI |
The original estimators and publication vintages are documented in AI Dark Output: The Visible Cost of Invisible Output — SemiAnalysis — May 2026, What is Generative AI Worth? — Stanford Digital Economy Lab — April 2026, Gartner Forecasts Worldwide End-User Spending on GenAI Models to Total $14.2 Billion in 2025 — Gartner — July 2025, The Progression of “Free” Digital Content to AI: Impacts on U.S. Economic Growth and Productivity — BEA — June 2026, Productivity and Costs by Industry: Selected Service-Providing Industries—2025 — BLS — August 2026, and Earnings Release FY26 Q4 — Microsoft — July 2026. newsletter.semianalysis.com
The table deliberately preserves incompatible concepts rather than converting them into a common dollar total, because such a conversion would require assumptions about adoption, substitution, valuation, domestic production, and quality that the underlying estimates do not jointly establish.
Consumer valuation is highly sensitive to the distribution
Brynjolfsson, Collis, Eggers, Kazinnik, and Nguyen’s April 2026 working paper reports materially different means and medians, with its national aggregates calculated from the means rather than the valuation of the typical user.
Table — Distributional sensitivity using the working paper’s Table 2
| Metric | July 2025 | March 2026 |
|---|---|---|
| Mean monthly willingness to accept | $98.00 | $124.50 |
| Median monthly willingness to accept | $3.39 | $11.48 |
| Estimated US users, millions | 98.78 | 115.33 |
| Published annual aggregate using the mean | $116.2 billion | $172.3 billion |
| Calculated annual equivalent using the median | $4.02 billion | $15.89 billion |
| Calculated mean-to-median ratio | 28.91 | 10.84 |
The last two rows are calculations from the cited inputs, not alternative totals endorsed by the authors. The PDF’s narrative gives a different 2025 median, $2.27, from Table 2’s $3.39; this assessment consistently uses Table 2. The authors integrate the fitted valuation curve up to $500 and acknowledge hypothetical bias. What is Generative AI Worth? — Stanford Digital Economy Lab — April 2026. digitaleconomy.stanford.edu
The mean–median difference should not be presented as a confidence interval around a single quantity, because the two statistics answer different questions. Multiplying the median by the user population describes an economy in which every user receives the typical valuation; it does not sum heterogeneous individual benefits. Nevertheless, the difference matters for interpretation: a large aggregate based on the mean can coexist with a substantially smaller valuation for the middle user, which makes claims about broadly shared benefits dependent on the underlying distribution.
A second distinction concerns the experimental design. Earlier digital-goods research included incentive-compatible experiments in which compensation and the loss of access were implemented, whereas the generative-AI working paper describes stated choices and explicitly discusses hypothetical bias. Random assignment of an offered compensation amount improves identification of the response curve, but does not by itself establish that respondents would honour the same choice when access is actually withdrawn. Using massive online choice experiments to measure changes in well-being — Brynjolfsson, Collis and Eggers, PNAS — April 2019. PMC
Table — Conditions governing the interpretation of a valuation aggregate
| Condition | Interpretation if satisfied | Risk if untested |
|---|---|---|
| Choices correspond to actual access restrictions | Responses approximate behaviour under the stated counterfactual | Hypothetical answers differ from implemented choices |
| The sampled valuation distribution represents the target users | Population aggregation is defensible | Digitally engaged respondents receive excessive influence |
| The upper tail is measured adequately | The mean reflects high-value users reliably | A small group strongly influences the total |
| User counts match the valuation population | Valuations and adoption rates share a denominator | Coverage differences distort aggregation |
| Monthly access values can be annualised | A monthly estimate supports an annual equivalent | Learning, substitution, and adaptation alter longer-term value |
| External effects are assessed separately | Access benefits remain distinguishable from wider social welfare | Private benefits are mistaken for net social benefits |
These conditions do not imply that the estimates should be discarded; they establish why a consumer-access valuation is evidence of welfare benefits rather than an independently measured quantity of omitted market production.
The commercial denominator also requires correction
The $14.2 billion figure sometimes used to compare supplier revenue with consumer benefits is especially sensitive to its label. Gartner’s original July 2025 release identifies it as a forecast of worldwide end-user spending on GenAI models, comprising foundation and specialised models, while discussing purchases by organisations; the release does not establish realised US consumer revenue for a specified group of suppliers. Gartner Forecasts Worldwide End-User Spending on GenAI Models to Total $14.2 Billion in 2025 — Gartner — July 2025. gartner.com
This qualification also corrects the earlier description of that figure as consumer revenue: its primary source supports a spending forecast with broader institutional coverage, so a ratio using it cannot be interpreted as a measured division of returns between US consumers and AI producers.
Table — Requirements for a defensible benefits-to-revenue comparison
| Dimension | Required alignment | Consequence of mismatch |
|---|---|---|
| Geography | Domestic benefits compared with domestic supplier sales | Worldwide expenditure changes the denominator |
| Population | Household users distinguished from business purchasers | Consumer welfare and enterprise expenditure are mixed |
| Period | Comparable calendar periods and observation dates | Survey snapshots are compared with forecasts or run rates |
| Product boundary | Equivalent coverage of tools and services | Model purchases omit or overlap application expenditure |
| Accounting basis | Realised revenue distinguished from projected expenditure | A forecast acquires the appearance of an observed transaction |
| Valuation object | User benefit distinguished from supplier profit | Revenue is mistaken for producer surplus |
Even a carefully aligned ratio would describe the relationship between user valuation and monetisation rather than the amount missing from GDP, because revenue is neither producer surplus nor a comprehensive measure of the social cost of providing the service.
The BEA exercise changes production accounting rather than inserting consumer surplus
Nakamura, Samuels, and Soloveichik’s June 2026 exercise has a different purpose from the willingness-to-accept studies, using barter accounting to recognise commercially supported content and corresponding user services.
Table — Additional average US real GDP growth in the experimental free-content account
| Period | Estimated addition to annual real GDP growth |
|---|---|
| 1929–1995 | 0.04 percentage point |
| 1995–2022 | 0.09 percentage point |
| 2022–2025 | 0.22 percentage point |
These are changes in growth rates for advertising- and marketing-supported content collectively, including AI, rather than AI-only estimates or percentages of consumer welfare. The authors associate the later trend break with AI development, but the timing does not isolate a causal contribution. The Progression of “Free” Digital Content to AI: Impacts on U.S. Economic Growth and Productivity — BEA — June 2026. U.S. Bureau of Economic Analysis (BEA)
Table — Assumptions that carry the experimental production adjustment
| Component | Construction in the paper | Principal sensitivity |
|---|---|---|
| Content valuation | Production-cost approach | Costs need not track useful capabilities |
| Household/business allocation | Usage evidence and extrapolation | Allocation changes final consumption |
| AI content prices | Software and cloud proxies | Proxy improvements need not match accepted outputs |
| Software capability proxy | Training-dataset expansion | Data volume does not directly establish service quality |
| TFP aggregation | Production-based weighting | Time-use weights would estimate a different object |
The authors’ constructed AI-content price index declines by approximately 41% annually during 2022–2025, while their experimental account raises private-business TFP growth by approximately 0.18 percentage point annually over that period. Both results depend on the specified price and allocation methods, and neither represents an adopted official revision. The Progression of “Free” Digital Content to AI: Impacts on U.S. Economic Growth and Productivity, sections 3.5–5.3 — BEA — June 2026. bea.gov
The important implication is methodological: even after expanding the production treatment, real growth remains sensitive to how capabilities are translated into price indices. Recognising an additional transaction does not eliminate the need to establish the quantity and quality of what that transaction represents.
Industry statistics contain observable gains without identifying their cause
The 2025 BLS service-industry results provide a more informative test than a single software-publisher productivity figure, because the separate movements in output and hours reveal different routes to higher output per worker hour.
Table — Selected US industry movements, calendar 2025
| Industry | Real output change | Labour-hours change | Diagnostic interpretation |
|---|---|---|---|
| Software publishers | +12.1% | −0.7% | Output expansion accompanied by slightly fewer hours |
| Line-haul railroads | +2.3% | −8.0% | A substantial hours reduction drives much of the productivity gain |
| Wireless telecommunications carriers | +4.1% | −4.3% | Higher output and lower hours contribute together |
| Commercial banking | +2.9% | −0.3% | Output growth with broadly stable hours |
| Engineering services | +2.6% | +2.8% | Output growth does not outpace hours |
| Wired telecommunications carriers | −8.3% | +2.8% | Lower output accompanied by higher hours |
Estimator and coverage: BLS industry aggregates, published in its 2025 service-industry chart package; these are annual changes, not survey means or medians. 1-year percent change in hours and output, and current employment level — BLS — 2026 release. bls.gov
The variation demonstrates why an industry productivity print cannot identify an AI effect without adoption and counterfactual evidence. A strong result might reflect AI, other software, capital deepening, demand changes, workforce composition, or changes in the measured output mix; similarly, a weak result can coexist with useful AI adoption if other forces offset its contribution.
Investment is visible, but its productive payoff remains a separate observation
Microsoft’s latest full-year disclosure materially extends the expenditure evidence available in the preceding discussion, with cash additions to property and equipment increasing from $64.551 billion in FY2025 to $115.948 billion in FY2026, an increase of $51.397 billion, or approximately 79.6%, calculated from the reported values. These are worldwide company totals rather than US AI-only investment, and the same release reports more than 30 million paid Microsoft 365 Copilot seats, which represents a commercial adoption count rather than active usage or measured productivity. Earnings Release FY26 Q4 — Microsoft — July 2026. Microsoft
Table — What observable commercial growth establishes
| Observation | Defensible inference | Additional evidence needed |
|---|---|---|
| Higher fixed-asset expenditure | More resources committed to capacity | Domestic location, AI allocation, utilisation, and returns |
| More paid seats | More contracted access | Active use, renewal, and accepted output |
| Higher token expenditure | More purchased model services | Task purpose, failure rates, and useful outcomes |
| More provider revenue | Greater monetised demand | Value added, profitability, and customer benefits |
| Lower customer labour expenditure | Changed production costs | Comparable output, transition costs, and redeployment |
Infrastructure construction can contribute to recorded production before users achieve better end-to-end outcomes, which means that an investment-led contribution to growth and a productivity gain from AI use must remain separate empirical claims.
Where darkness is real and where it is rhetorical
Missing welfare, missing quality, and delayed production are different diagnoses
The strongest interpretation of dark output concerns valuable outcomes that existing statistics do not measure adequately, but the remedy depends on whether the gap arises outside the production boundary, inside a defective volume measure, or before the economic benefit has actually been realised.
Table — Diagnostic classification of the apparent gap
| Observed pattern | Most relevant diagnosis | Evidence needed before calling it missing output |
|---|---|---|
| A household receives valuable assistance without additional expenditure | Welfare outside the core production measure | Valuation, time-use, and outcome evidence |
| A service improves while its price remains unchanged | Potentially omitted quality growth | Comparable tasks and validated quality adjustment |
| A firm saves preparation time but retains extensive verification | Partial task improvement | Full workflow time and accepted output |
| A firm spends on training and redesign before output improves | Implementation lag | Subsequent output, cost, and organisational evidence |
| Higher margins follow lower costs | Recorded distributional change | Separation of efficiency from other causes of profit growth |
| Supplier receipts fall following internalisation | Changed production organisation | Complete value-added reconciliation |
| An experimental satellite account records more production | Alternative statistical treatment | Boundary consistency and sensitivity analysis |
A household welfare gain need not be described as an error in GDP, because a production statistic can perform its intended function while remaining incomplete as a measure of living standards. Hulten and Nakamura formalise this possibility through a separate consumption technology that allows households to obtain more utility from a given amount of production, distinguishing that mechanism from conventional resource-saving productivity growth. Accounting for Growth in the Age of the Internet: The Importance of Output-Saving Technical Change — Federal Reserve Bank of Philadelphia — August 2018 revision. philadelphiafed.org
Their expanded measure is a theoretical alternative with an additional welfare component, which should be interpreted alongside conventional production accounts rather than relabelled as an unexplained correction to national income.
Exposure is a map of potential change, not evidence of its realised magnitude
SemiAnalysis explicitly describes its approximately $1.5 trillion figure as exposed labour expenditure rather than missing output, while reporting that much of its collected evidence concerns augmentation rather than replacement. Its monitor therefore identifies economically relevant task categories, but does not establish how much labour has disappeared, how much accepted output has increased, or how much real production official accounts have missed. AI Dark Output: The Visible Cost of Invisible Output — SemiAnalysis — May 2026. newsletter.semianalysis.com
The conversion from exposure to realised production gains requires several additional observations, each capable of reducing or changing the estimated effect.
Table — The evidence required to move from exposure to realised gains
| Stage | Required observation | Why exposure alone is insufficient |
|---|---|---|
| Technical suitability | Comparable tasks can be completed | Capability under selected conditions does not establish routine performance |
| Deployment | Firms use the system in production | Available technology can remain unused |
| Acceptance | Outputs meet operational requirements | Generated material can require rejection or substantial revision |
| Net efficiency | Gains survive review and integration costs | Faster drafting can coexist with slower completion |
| Economic adjustment | Firms expand output or release resources | Local time savings can become idle capacity |
| Statistical omission | Official measurement misses the identified gain | A realised gain can already appear in output or income |
These stages should not be converted into numerical probabilities without evidence, because multiplying an exposure figure by unsupported deployment or acceptance rates would manufacture precision rather than improve measurement.
End-to-end results can contradict perceptions, and newer evidence can become harder to identify
METR’s early-2025 randomised experiment illustrates the difference between perceived assistance and measured completion, finding that 16 experienced open-source developers took approximately 19% longer to complete 246 tasks when AI tools were allowed, in the studied repositories and with the tools available during February–June 2025. The result concerns a specific experienced-worker setting rather than software development generally, and should not be presented as an estimate of current 2026 tools. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity — METR — July 2025. METR
METR’s February 2026 update is equally important for interpretation, because the organisation judged its later experiment an unreliable indicator of the current effect, citing participation and task-selection problems, together with difficulties recording time when developers used agents concurrently. The researchers considered greater speedups plausible, but did not regard their later data as strong evidence of the magnitude. We are Changing our Developer Productivity Experiment Design — METR — February 2026. METR
The resulting lesson is that evidence quality does not necessarily improve merely because a study is newer: as adoption changes who participates, which tasks enter an experiment, and how work is organised, the counterfactual can become more difficult to maintain.
Price declines require better deflators, not an automatic addition to GDP
Broughel’s June 2026 intervention usefully separates nominal transactions from the price indices needed to interpret their real content, while arguing that welfare analyses should remain distinct from national-income measurement. Its strongest contribution is the insistence that falling prices are not automatically invisible; the issue is whether the relevant deflator captures comparable quantity and quality. The Real Reason AI Doesn’t Show Up In The GDP Statistics — James Broughel, Forbes — June 2026. forbes.com
That argument does not resolve household production or every new-service boundary problem, but it correctly limits a broad claim that lower receipts must represent an omission. Nominal income can decline while physical productivity improves, and a lower nominal value is not itself a statistical error; the error arises when the real-volume calculation fails to recognise the identified production change.
The Griliches tradition of hedonic measurement provides a disciplined response by relating prices to product characteristics, although characteristics must remain economically meaningful and sufficiently comparable. A token count, benchmark score, or training-data measure cannot become an output index merely because it is observable. Handbook on Hedonic Indexes and Quality Adjustments in Price Indexes — OECD — 2006. oecd.org
Historical measurement research imposes a scale test
The Boskin Commission’s December 1996 estimate of approximately 1.1 percentage points of annual upward CPI bias concerned a historical US cost-of-living measurement problem, rather than a transferable correction to current GDP growth. Applying that figure to AI would require evidence that the relevant products, expenditure weights, and sources of bias were comparable, which the historical estimate does not establish. Toward A More Accurate Measure Of The Cost Of Living — Advisory Commission to Study the Consumer Price Index — December 1996. ssa.gov
BLS’s subsequent response also emphasised that the Commission’s quality and new-product estimates included rough component calculations, illustrating why recognition of a genuine measurement mechanism does not establish its aggregate magnitude. The BLS Response to the Boskin Commission Report — BLS — Spring 2006. bls.gov
Byrne, Fernald, and Reinsdorf add a second constraint: measurement problems existed before the post-2004 productivity slowdown, and their quantified IT adjustments did not explain that slowdown; in their analysis, some adjustments increased the earlier period’s productivity growth more than the later period’s. Consequently, the existence of bias is insufficient without evidence of a change in bias large enough to explain the growth pattern. Does the United States have a Productivity Slowdown or a Measurement Problem? — Federal Reserve Bank of San Francisco — March 2016. San Francisco Fed
Syverson’s published 2017 assessment supplies a complementary scale test, examining cross-country patterns, the size of digital benefits, implied ICT-industry growth, and the GDP–GDI relationship. These tests challenge attempts to explain the earlier slowdown predominantly through measurement, without imposing a permanent ceiling on future AI gains. Challenges to Mismeasurement Explanations for the US Productivity Slowdown — Journal of Economic Perspectives — Spring 2017. American Economic Association
Table — Constraints that a large dark-output claim must satisfy
| Test | Required demonstration | Claim that does not satisfy it |
|---|---|---|
| Scale | Omitted production is large relative to the aggregate being explained | A striking isolated example |
| Change over time | Measurement bias has increased during the relevant period | Longstanding statistical difficulty |
| Domestic weighting | The affected production has sufficient domestic weight | Worldwide technology spending |
| Counterfactual | Output would differ without AI | Adoption coincides with higher performance |
| Quality | Gains survive consistent acceptance standards | More generated items |
| Accounting reconciliation | The gain is absent after transactions and value added are reconciled | Lost supplier receipts |
| Welfare separation | Production and consumer benefits remain distinct | A surplus estimate added to GDP |
Key judgments and evidence that would change the assessment
The evidence is strongest for heterogeneous consumer benefits, visible commercial expenditure, and specific measurement vulnerabilities; it is weaker for the magnitude of omitted market production, because the available studies do not jointly identify adoption, accepted output, quality, domestic value added, and the corresponding official measurement error.
Retained margins, infrastructure expenditure, and paid model services are observable economic objects whose interpretation requires attribution rather than a new accounting residual, whereas household outcomes and unrecognised quality improvements require complementary measures that preserve the distinction between production and welfare.
Table — Observations capable of strengthening or weakening the assessment
| Evidence to obtain | Would strengthen a missing-output claim | Would weaken that claim |
|---|---|---|
| Matched service outcomes and official deflators | Accepted output improves while the deflator fails to reflect comparable prices or quality | Official volume measures capture the improvement |
| Full workflow records | Gains remain after verification, integration, and rework | Initial savings disappear at completion |
| Linked adoption and establishment data | Adopting establishments improve relative to credible controls | Comparable improvements occur without AI |
| Firm accounts reconciled with industry statistics | Identified production remains absent after reconciliation | Benefits appear in recorded value added |
| Implemented access-choice experiments | Stated valuations persist under actual restrictions | Valuations change substantially when choices are enforced |
| Domestic investment and utilisation data | Capacity supports sustained increases in accepted output | Expenditure precedes unused or weakly productive capacity |
The principal unresolved record is a linked dataset combining AI expenditure, accepted task outcomes, labour time, prices, quality, and establishment accounts across a representative set of adopting firms. Until that bridge exists, the defensible conclusion is that current measures identify different parts of the transformation, while none provides a comprehensive total of dark output or a valid single adjustment to GDP.
Dark Output: Institutional Interpretation and Statistical Reform
The institutional task is to identify which economic object each indicator measures, and then establish how those objects connect. Expenditure demonstrates a transaction; adoption demonstrates use; a reduction in unit cost demonstrates an operational improvement; consumer surplus demonstrates a benefit relative to the price paid. None, by itself, establishes the contribution of AI to aggregate productivity.
This final part develops that distinction into a framework for institutional decisions and statistical reform. It proposes complementary production, deployment, and welfare accounts, while preserving the accounting discipline that prevents intermediate expenditure, investment, profits, and consumer surplus from being combined into an artificial measure of total AI output
Implications for investors and policymakers
Match the question to the evidence
The principal institutional error is to ask one statistic to answer several different questions. A company can demonstrate sustained demand for an AI service without demonstrating higher productivity among its customers. An adopting business can improve its margins without reducing final prices. Households can receive substantial benefits without generating proportionate market expenditure. Infrastructure investment can contribute to current production before the infrastructure produces an identifiable improvement in downstream services.
These are different stages and distributions of economic activity. An assessment should therefore begin with the question being asked, rather than with whichever headline estimate is largest.
| Institutional question | Evidence capable of answering it | Evidence that cannot answer it alone | Condition for a defensible interpretation |
|---|---|---|---|
| Is demand for the technology real? | Repeated paid use, renewals, utilisation, accepted tasks, and customer retention | Aggregate GDP growth, registrations, benchmark performance, or announced capacity | Separate experimentation from sustained use; identify subsidies, bundled access, and changes in price |
| Are productive complements being built? | Workflow integration, training, software investment, data preparation, and infrastructure entering service | Capital expenditure totals or installed computing capacity alone | Establish whether complements are operational and whether they support completed production |
| Who captures the surplus? | Matched evidence on prices, unit costs, compensation, profits, and user benefits | Provider revenue, consumer willingness to accept compensation, or a productivity statistic in isolation | Trace pass-through and distinguish transfers between participants from changes in total benefits |
| Is aggregate productivity rising? | Consistently deflated value added, labour hours, capital services, and multifactor productivity | Adoption rates, reported time savings, task benchmarks, or willingness-to-accept aggregates | Account for industry composition, complementary inputs, quality adjustment, and competing explanations |
The table is an analytical decision framework, not an investment recommendation. Official productivity concepts distinguish real output per labour hour from multifactor productivity, which also accounts for other productive inputs. Labor Productivity and Costs: Concepts — U.S. Bureau of Labor Statistics. bls.gov
The distinction between demand and productivity is particularly consequential. A service can attract paying customers because it provides convenience, flexibility, or access to a previously unavailable capability. Those benefits can support demand even when the service produces no measurable reduction in labour hours. Conversely, an embedded AI feature may improve an existing product without generating a separately identifiable AI transaction.
Institutional analysis should therefore examine the conversion of use into completed outcomes. For a document-review service, the relevant operational denominator might be accepted reviews of comparable complexity, including verification and correction. For software development, it might be deployed functionality meeting specified requirements, including maintenance and rework. Neither token consumption nor drafts generated provides that denominator.
Adoption statistics require explicit denominators
Recent US evidence demonstrates why adoption cannot be represented by a single percentage. Bonney and co-authors report the following estimates in Table C.7 of their April 2026 Census working paper, using the Business Trends and Outlook Survey supplement collected during November 2025–January 2026.
| Measurement layer | Reference window | Firm-weighted share | Employment-weighted share |
|---|---|---|---|
| Firm reports AI use | Previous two weeks | 17.9% | 31.2% |
| Firm reports worker-task AI use | Previous six months | 22.6% | 40.6% |
| Firm reports worker-task generative AI use | Previous six months | 20.8% | 38.9% |
| Firm reports AI use in a business function | Previous six months | 27.7% | 37.1% |
Estimator and population: Bonney, Breaux, Dinlersoz, Foster, Haltiwanger, and Pande; US private non-agricultural employer firms; survey collection November 2025–January 2026. These are weighted proportions, not mean or median valuations. Employment-weighted percentages describe employment located in reporting firms; they are not percentages of individual workers personally using AI. Worker-task use is reported by firm respondents. Different reference windows prevent interpreting the differences between rows as a pure measure of informal adoption. The Microstructure of AI Diffusion: Evidence from Firms, Business Functions, and Worker Tasks — U.S. Census Bureau research working paper — Apr 2026. www2.census.gov
Question wording also matters. Census broadened its core question in November 2025 from AI used in producing goods or services to AI used in any business function, and identified a break in the series. A higher reading across that transition cannot automatically be attributed to faster diffusion. BTOS AI Core Question Updates — U.S. Census Bureau — Dec 2025. census.gov
For international comparisons, harmonisation must cover the sampling population, reporting unit, definition of AI, reference period, weighting, and treatment of embedded software. A US employer survey and another country’s household survey can both be informative without estimating the same object. Neither provides a basis for transferring US consumer valuations to another population.
Association is an intermediate finding
A BEA research spotlight published on 2 October 2026 reports that state–industry cells with higher worker-reported AI use experienced stronger real-output trajectories after 2020. Employment differences were generally positive but less precisely estimated; comparable regressions relying only on industry-level BTOS variation produced null results. The authors identify reverse causality and thin samples among the limitations. These findings concern associations and do not isolate AI’s causal contribution. AI Utilization and Changes in Economic Performance — Bureau of Economic Analysis — Oct 2026. apps.bea.gov
The institutional implication is that statistical aggregation can remove informative variation, while greater disaggregation can introduce additional sampling uncertainty. Neither problem is resolved by choosing the estimate most favourable to a particular interpretation.
A stronger research design would observe adoption dates, pre-adoption performance, complementary investment, and comparable non-adopting units. It would also investigate whether higher-performing firms are more likely to adopt. This is a proposed evidential standard: correlations remain useful for locating diffusion and generating hypotheses, but should not be converted directly into a national productivity dividend.
Different institutions require different outputs
| Institutional setting | Decision-relevant evidence | Interpretation to avoid |
|---|---|---|
| Monetary policy | Real output, hours, compensation, unit labour costs, prices, and evidence of productive capacity | Treating expected future AI gains as an already realised increase in supply |
| Fiscal analysis | Recorded production, taxable income, employment income, and identifiable changes in service costs | Treating estimated consumer surplus as an equivalent increase in the tax base |
| Public procurement | Accepted outcomes, total workflow cost, turnaround time, and service quality | Counting automated activity as a successful public-service outcome |
| Corporate governance and institutional analysis | Persistent use, utilisation, implementation costs, net labour requirements, and reconciliation to financial accounts | Treating expenditure, theoretical time savings, or operating margins as interchangeable evidence |
| Competition analysis | Changes in prices, access, switching, service quality, and the distribution of gains | Inferring consumer benefits solely from lower provider costs |
These are proposed interpretative rules, not descriptions of newly adopted institutional requirements.
A decline in transaction prices can expand access while reducing receipts in an incumbent industry. An increase in margins can improve the income of producers without increasing consumer surplus. A rise in service quality can improve outcomes while remaining weakly represented in the deflator. Institutional analysis needs to preserve those possibilities until evidence establishes which mechanism operates.
A measurement agenda
Build connected accounts with different purposes
Statistical reform should extend existing accounts through explicit classifications and supplementary measures. It should preserve a distinction between recording production and evaluating welfare.
There is already an institutional foundation for this approach. The OECD’s digital supply-and-use framework adds detail about products, producers, and how transactions occur while remaining consistent with conventional national accounts. BEA research separately examines the definitions and data required to identify AI production across existing industries. Neither approach treats all benefits from AI use as additional GDP. OECD Handbook on Compiling Digital Supply and Use Tables — OECD — Nov 2023; Concepts and Challenges of Measuring Production of Artificial Intelligence in the U.S. Economy — Bureau of Economic Analysis — Jan 2025. OECD
The proposed architecture has three components.
| Account | Principal question | Contents | Relationship to headline GDP |
|---|---|---|---|
| AI production account | Where is AI-related production occurring? | Identifiable products, investment, intermediate inputs, imports, exports, and domestic value added | Reclassifies and details recorded production; any additions require an explicit accounting justification |
| AI deployment and performance account | What changes when organisations and households use AI? | Tasks, hours, accepted outcomes, quality, implementation costs, and newly feasible activities | Supplies evidence for productivity analysis and improved measurement; does not automatically add task valuations |
| Welfare satellite account | What benefits exceed the prices users pay? | Carefully specified valuation experiments, population estimates, access, and distribution | Remains separate from production totals |
This separation also prevents different satellite-account proposals from being conflated. The UN guidance on free digital products considers an SNA satellite account; a GDP-B-style valuation exercise addresses willingness to pay or accept compensation. Their boundaries and valuation methods need not coincide. Recording and Valuing “Free” Digital Products in an SNA Satellite Account — United Nations national accounts guidance — Sep 2022; Using massive online choice experiments to measure changes in well-being — Brynjolfsson, Collis and Eggers — Apr 2019. unstats.un.org
Introduce five measurement modules
The following programme is proposed rather than reported as an existing statistical obligation. The US institutional examples indicate plausible expertise; other economies would allocate responsibilities according to their statistical systems.
| Module | Additional observations | Published result | Essential safeguard |
|---|---|---|---|
| Task-level time use | Task type, active labour time, checking, correction, simultaneous activity, and accepted completion | Time requirements for comparable completed tasks | Separate reported counterfactual savings from observed changes |
| Quality-adjusted service indices | Transaction price, complexity, timeliness, accuracy, and relevant service outcomes | Experimental price and volume indices | Publish quality assumptions and alternative specifications |
| GDP-B welfare satellite | Implemented choices where feasible, valuation distributions, active-user counts, and access conditions | Population-specific welfare estimates | Report means, medians, uncertainty, and user-count sensitivity separately |
| Firm deployment disclosure | AI operating expenditure, implementation costs, labour hours, throughput, and rework | Reconciled cost and performance panels | Avoid counting gross task savings before additional checking and integration |
| AI investment and use decomposition | Investment, assets entering service, capital services, utilisation, and downstream output | Separate infrastructure and user-industry contributions | Keep purchases, productive services, and output gains distinct |
Task time should measure the complete workflow
The American Time Use Survey provides an established diary-based framework for observing activities. Its principal-activity structure, however, would need additional task and tool information to distinguish AI-assisted work from the broader activity in which it occurs. That extension is a proposal, not an existing AI productivity measure. American Time Use Survey: 2025 Results — U.S. Bureau of Labor Statistics — Jun 2026. 2025 A01 Results
The measurement unit should be a completed task with a specified acceptance standard. Recording only generation time omits preparation, verification, correction, integration, and unsuccessful attempts. Recording only elapsed time can also misstate labour input when an employee performs another task while a system runs.
Four observations should remain separate:
- Active labour time: human effort devoted to the task.
- Elapsed completion time: the interval before an acceptable result becomes available.
- Computing time: machine resources consumed.
- Released capacity: human time actually available for another activity after the full workflow changes.
Released capacity is not automatically a reduction in paid hours. It can support additional output, improve quality, reduce work intensity, or remain unused. A task panel should observe those subsequent uses rather than assigning all reported savings a wage-based value.
Service indices need an auditable quality model
A useful service-price index must distinguish a lower price for the same service from a change in the service purchased. For AI-intensive services, the statistical challenge is to identify features that are observable, consequential, and sufficiently stable for comparison.
| Service | Candidate output unit | Candidate quality observations | Main measurement difficulty |
|---|---|---|---|
| Document review | Accepted review of a defined document set | Material omissions, corrections, coverage, and turnaround | Complexity and error detection vary across assignments |
| Tutoring | Instruction delivered within a defined learning objective | Validated learning outcomes and continuity of support | Outcomes depend on prior attainment and other inputs |
| Software development | Deployed functionality meeting specified requirements | Reliability, maintenance burden, and rework | Faster initial completion may conceal later costs |
| Administrative services | Completed case or resolved application | Accuracy, timeliness, and subsequent reopening | Faster processing may accompany changes in service scope |
These are proposed measurement designs, not validated official AI deflators.
A quality model should publish the variables used, their treatment over time, and the sensitivity of results to alternative specifications. A statistically convenient feature should not receive a quality weight merely because it is easy to count.
The UK provides a relevant institutional precedent for maintaining complementary measures: ONS public-service productivity estimates incorporate quality adjustments, while explaining differences from national accounts measures, including the treatment of healthcare quality. This demonstrates the possibility of publishing distinct measures with explicit purposes; it does not supply a transferable numerical estimate of AI benefits. Public service productivity: total, UK QMI — Office for National Statistics — May 2026. Office for National Statistics
Firm disclosure should reconcile expenditure with performance
Deployment reporting should connect financial records to operational observations. Separate disclosure of AI expenditure is useful, but cannot establish productivity without information about output and the other inputs required to produce it.
| Disclosure field | Recommended treatment | Analytical purpose |
|---|---|---|
| Model and cloud expenditure | Separate identifiable AI use; disclose allocation rules for shared services | Establish operating cost without implying equivalent final value |
| Embedded software expenditure | Identify the incremental AI component where defensible; retain an unallocated category otherwise | Avoid assigning the entire software bill to AI |
| Training and integration | Record labour and purchased services; distinguish expenditure from eligible investment | Capture complementary inputs |
| Verification and rework | Include checking, correction, rejected results, and remediation | Measure net workflow performance |
| Accepted output and labour hours | Use consistent task definitions and acceptance criteria | Connect costs to completed production |
| Asset treatment | Identify ownership, capitalisation, depreciation, and assets entering service | Reconcile operational reporting with financial and national accounts |
Internally produced software may qualify as own-account investment; internal spending cannot therefore be classified uniformly as current operating expense. Own-account investment — Bureau of Economic Analysis. U.S. Bureau of Economic Analysis (BEA)
A reduction in task hours should also be distinguished from a reduction in payroll. The former is an operational result; the latter depends on employment, compensation, redeployment, and demand. Reporting the connection between them would make distributional analysis more informative.
Separate investment flows from productive use
Infrastructure measurement should distinguish expenditure during a period from the services supplied by the accumulated assets.
| Object | What it records | What remains to be established |
|---|---|---|
| Investment expenditure | Acquisition or production of qualifying assets | Domestic production content and the asset’s eventual use |
| Productive asset stock | Assets available to support production | Capacity actually used and the services supplied |
| Capital services | Productive input supplied by assets during the period | Output attributable to the combination of capital and other inputs |
| Downstream performance | Output, quality, costs, and labour requirements in user industries | AI’s contribution relative to other changes |
In growth-accounting terms, real-output growth can be decomposed into contributions from capital services, labour input, and measured multifactor productivity. The residual is not a pure measure of AI: it can contain other innovation, organisational change, and measurement error. Labor Productivity and Costs: Concepts — U.S. Bureau of Labor Statistics.
The proposed AI extension would tag relevant assets and deployment rather than insert an additional “dark output” term into this decomposition. It would also distinguish domestic production from imports and expenditure abroad. National accounts aggregate domestic value added, not every receipt along an international supply chain. Measuring the Economy: A Primer on GDP and the NIPAs — Bureau of Economic Analysis. bea.gov
Welfare estimates require their own publication standards
A welfare satellite should publish its valuation distribution rather than rely on a single aggregate. The mean and median answer different questions: the mean is relevant to aggregation but sensitive to the upper tail; the median describes the central respondent but is not a substitute for summing heterogeneous valuations.
Population expansion requires an independently justified count of eligible users. Estimates should specify geography, age coverage, active-use criteria, access conditions, and survey period. They should disclose whether choices were implemented and whether compensation involved a credible restriction on access.
For multiple AI services, a welfare programme should also examine joint access. Adding separate valuations can overstate total benefits when services are substitutes. These proposed publication requirements extend the experimental welfare-measurement approach; they do not convert consumer surplus into production. Using massive online choice experiments to measure changes in well-being — Brynjolfsson, Collis and Eggers — Apr 2019.
Publish an explicit reconciliation
A credible programme needs a bridge between experimental results and existing statistics. Each release should identify what is already recorded, what is reclassified, what improves the measurement of volume, and what remains outside the production boundary.
UN guidance on AI visibility provides a foundation for greater detail, while recognising difficulties in defining and identifying AI embedded in broader products. It should be treated as guidance for measurement development, rather than evidence that all countries have already implemented a common AI account. Improving the visibility of Artificial Intelligence in the National Accounts — United Nations national accounts guidance. unstats.un.org
| Proposed release stage | Required deliverable | Criterion for proceeding |
|---|---|---|
| Definitions and crosswalks | Product, task, industry, and asset classifications | Boundaries are reproducible and overlap is documented |
| Pilot collection | Linked expenditure, hours, task, and outcome observations | Respondents can report the variables consistently |
| Experimental publication | Separate production, deployment, and welfare results | Assumptions, uncertainty, revisions, and reconciliation are public |
| Integration into core measures | Validated improvements to classifications or deflators | Results meet the statistical standards applicable to official accounts |
The central publication discipline is to identify the status of every result. A welfare estimate should remain labelled as welfare; an experimental deflator should remain labelled as experimental; a causal estimate should identify its research design. Their proximity in a report should not erase their differences.
Conclusion
“Dark output” is useful when it identifies a specific failure of observation: an unpriced benefit, a poorly measured quality change, a newly feasible activity outside market production, or a volume gain obscured by an inadequate service deflator. It becomes misleading when it relabels recorded revenue, profits, investment, or intermediate expenditure as invisible production.
The institutional consequence is a need for multiple connected measures. Production accounts establish what is produced within their boundary. Deployment measures establish how tasks and input requirements change. Distributional evidence identifies who receives gains. Welfare accounts examine benefits beyond transaction prices. Their results can inform one another without being summed into a single corrected GDP.
Value is shifting from prices toward capabilities. Production accounts can record that shift late, partially, and sometimes with the wrong sign when quality and volume are inadequately measured. That conclusion concerns particular measurement mechanisms, not every disappointing productivity observation. Contemporary statistical standards are evolving; the reform task is to implement and validate complementary observations rather than abandon production accounting. System of National Accounts — United Nations Statistics Division.
The appropriate remedy is a set of complementary accounts with explicit boundaries, reliable denominators, and published reconciliation. A strong AI assessment should be able to report substantial user benefits, uneven producer gains, and modest measured productivity simultaneously, then investigate the evidence connecting them.
Appendix
Compact reference sheet. All numerical amounts below are hypothetical US-dollar illustrations, not empirical estimates.
Glossary
| Term | Definition |
|---|---|
| Measured output | GDP records value added within the production boundary at transaction values, including established imputations. Industry gross output also includes production subsequently used as intermediate inputs. |
| Real output and productivity | Real output is quantity and quality-adjusted volume. Labour productivity is real output per labour hour; multifactor productivity accounts for additional productive inputs. |
| Firm productivity and margins | Operational performance includes throughput, errors, unit costs, and labour requirements. Margins concern financial returns; a higher margin is not itself invisible output. |
| Consumer surplus | Willingness to pay minus price paid. It is a welfare object and must not be added to GDP. |
| Newly feasible activity | A task made possible by lower costs or improved capabilities. Its accounting treatment depends on whether it becomes market production, recognised own-account production, or non-market activity. |
| Substitution versus creation dark output | Substitution replaces previously billed work; creation enables previously unavailable activity. Neither category determines the GDP effect without prices, quantities, quality, and production-boundary information. |
Four-case numerical example
For A–C, the baseline is one final legal memo: price $1,000; cost $800; profit $200; willingness to pay $1,500; consumer surplus $500. AI reduces cost to $200. All production and inputs are domestic; final expenditure represents combined value added across the production chain, with no intermediate double counting. Other expenditure is held constant.
| Case | New transaction and quality | Δ nominal GDP | Δ real GDP, at baseline prices | Δ firm profit | Δ consumer surplus |
|---|---|---|---|---|---|
| A. Full cost pass-through | Price $400; unchanged quantity and quality | −$600 | $0 | $0 | +$600 |
| B. Saving retained | Price $1,000; unchanged quantity and quality | $0 | $0 | +$600 | $0 |
| C. Quality improvement | Price $1,000; quality multiplier 1.5; willingness to pay $1,800 | $0 | +$500 | +$600 | +$300 |
| D. Newly feasible free household task | Previously absent; price $0; willingness to pay $250 | $0* | $0* | $0* | +$250 gross |
Assumptions: A assumes complete pass-through and successful deflation. B assumes no pass-through. C assumes an accepted 1.5 quality adjustment; willingness to pay is independently assumed. Without that adjustment, reported real growth could remain zero. D assumes unchanged provider production and costs; zero changes are conditional, not a general rule for free services. A $30 household time burden would reduce D’s net benefit to $220. Employee use can affect recorded costs, output, or eligible investment.
These are controlled accounting examples, not predictions of economy-wide effects after spending and employment adjust. Measuring the Economy: A Primer on GDP and the NIPAs — Bureau of Economic Analysis; Labor Productivity and Costs: Concepts — U.S. Bureau of Labor Statistics; Using massive online choice experiments to measure changes in well-being — Brynjolfsson, Collis and Eggers — Apr 2019.



















