Claims about AI-led growth often compress several different propositions into one sentence: a tool improves a task, a company produces more, investment creates jobs, and society becomes more prosperous. Each step may be plausible without proving the next. A serious evaluation separates them, names the affected population, and asks what evidence would distinguish a durable gain from a shifted cost.
The need for that discipline is visible in the G20's 2026 AI agenda. Ministers connected research, commercialization, technical education, digital infrastructure, and industrial investment through voluntary objectives. The framework established a direction, but it did not supply a binding funding mechanism, common timetable, or complete system for measuring labor outcomes. That gap is useful beyond any one meeting: policy commitments, company announcements, and market forecasts are inputs to an evaluation, not proof of results.
This guide provides a five-part framework for examining productivity, exposure, transition costs, infrastructure, and distribution. It can be used for a government program, corporate AI rollout, investment proposal, training initiative, or public claim about employment.
Begin by converting the claim into a test
Before examining a statistic, rewrite the claim in operational terms. Identify the intervention, comparison, time period, affected group, and proposed outcome. “AI will create jobs” is too broad to test. “A funded data-center program will produce a stated number of local construction jobs over three years and a smaller number of permanent operating roles” is testable. So is “AI assistance will reduce the time required for a defined workflow without increasing error rates or unpaid rework.”
Keep forecasts, commitments, activity, and outcomes in separate columns. Announced investment is not contracted spending; contracted spending is not installed capacity; installed capacity is not productive use. Course enrollment is not completion, and completion is not placement. A model's ability to perform part of a task is not evidence that an occupation has disappeared.
Also identify the counterfactual. Productivity may improve because of AI, process redesign, stronger demand, new staff, or several changes at once. Employment at an infrastructure site may rise even as entry-level office hiring falls elsewhere. Without a comparison—an earlier baseline, a similar unit, or an agreed target—the evaluator cannot attribute the change confidently.
Test 1: Does task productivity become organizational value?
Start at the narrowest supported level. Measure a defined task before making claims about a role, company, industry, or economy. Useful task measures include completion time, throughput, quality, correction effort, and the share of outputs requiring human escalation. A faster first draft is not a productivity gain if verification and repair consume the saved time.
Next, trace how task-level savings move through the organization. Management can use saved time to increase output, improve service, reduce backlogs, shorten working hours, or reduce staffing. These choices produce different economic and workforce effects even when the same tool performs equally well. Record which pathway leaders intend to pursue and which outcome actually follows.
Company-level measures should pair output with quality and cost. Track completed work, error or incident rates, customer outcomes, labor hours, AI service costs, and the cost of oversight. When a claim extends to the economy, require additional evidence such as sustained business formation, investment, wages, and output—not a collection of isolated demonstrations.
The key question is causal and practical: did AI help the organization create more useful value with the resources consumed, or did it merely move work into review, infrastructure, or another department?
Test 2: Is the workforce claim about exposure, transformation, or displacement?
These terms are not interchangeable. Exposure means AI can perform some activities within an occupation. Transformation means the mix of tasks changes. Displacement means demand for workers declines or workers lose positions. Job creation concerns a different flow and may occur in other occupations or places.
The International Labour Organization's refined global index illustrates the distinction. Its 2025 analysis found that one in four workers was in an occupation with some generative AI exposure, while 3.3 percent of global employment was in the highest exposure category. It concluded that transformation was more likely than complete replacement because most occupations still contain tasks requiring human involvement. Those findings are a baseline for inquiry, not a forecast that every exposed employer will adopt AI in the same way.
Segment the evidence by occupation, task, income group, gender, region, seniority, and contract type. The ILO found exposure of 34 percent in high-income countries and 11 percent in low-income countries, as well as higher representation of women in the highest-exposure category. An aggregate percentage can therefore conceal materially different risks and opportunities.
Watch flows rather than a single employment total: hiring, vacancies, promotions, hours, wages, layoffs, retention, and movement between occupations. Pay special attention to entry-level roles. An organization can avoid layoffs while reducing junior hiring, weakening the pathway through which future workers gain experience.
Test 3: Are transition costs included in the calculation?
Training is often presented as the bridge between exposure and opportunity, but a course is only one part of a transition system. Workers may also need recognized credentials, prerequisites, time, childcare, transport, equipment, apprenticeships, and employers with real vacancies. A program designed for data-center technicians will not automatically serve an administrative worker whose tasks are changing.
Evaluate training as a sequence: targeted occupation, access, enrollment, completion, credential recognition, placement, wage progression, and retention. Report where participants exit the sequence. Completion rates alone cannot show whether the program connects people to durable demand.
Include geography and timing. Infrastructure work can appear near manufacturing plants, data-center clusters, power systems, or universities, while software-related displacement may occur elsewhere. Construction roles may be temporary; operating roles may be fewer and more specialized. Retraining has limited value when a new role arrives in another region or after a worker's income support ends.
Assign transition responsibilities explicitly. Employers can disclose changing skill requirements and provide paid learning time. Education providers can align credentials with named vacancies. Governments can fund access and track mobility. If a proposal promises adaptation without naming who pays, who hires, and when support becomes available, its labor case is incomplete.
Test 4: Can the infrastructure support the projected growth?
AI is not only software. Growth claims may depend on chips, servers, data centers, cooling, electrical equipment, grid connections, construction labor, financing, and permits. Capacity at one layer does not remove bottlenecks at another.
Energy deserves its own ledger. The International Energy Agency projects global data-center electricity consumption of about 945 terawatt-hours in 2030, more than double the current level and just under 3 percent of global electricity demand. It projects roughly 15 percent annual growth from 2024 through 2030 and identifies AI as the most important driver alongside other digital services. These are global projections; local effects depend on where demand concentrates and what generation and transmission are available.
For a project or policy, document expected computing capacity, total electricity demand, connection dates, construction stages, cooling or water needs, and responsibility for upgrades. Separate efficiency per server or model request from total demand. Lower unit consumption can coexist with higher aggregate use when deployment expands.
Then examine local effects. Who pays for grid or public-infrastructure upgrades? What happens if connections or permits are delayed? Are construction jobs counted separately from permanent operations? A credible growth estimate includes the physical schedule and community costs rather than treating compute as instantly available.
Test 5: Who receives the gains and who carries the risk?
Aggregate productivity can rise while its benefits remain concentrated. Evaluate at least four recipients: workers, employers, customers, and communities hosting infrastructure. For workers, examine wages, hours, job quality, bargaining power, and career mobility. For employers, measure output, cost, resilience, and new revenue. For customers, examine price, access, quality, and recourse. For communities, include employment, tax benefits, energy or water pressure, and public costs.
Distribution also applies to failure. In employment, education, health care, or public services, an AI-assisted decision can affect people who did not choose the system. The G20 objectives call for secure, reliable, and trustworthy adoption while allowing flexible national approaches. An evaluator should translate those principles into named controls: access rules, source retention, human review, responsibility for errors, and a path to challenge consequential decisions.
Do not net unlike effects into a single jobs figure. New construction work does not cancel the loss of an office role for a particular worker. A national gain does not prove a host community is better off. Report benefits and costs by group first; aggregate only after the distribution is visible.
Build an evidence scorecard
For every major claim, record the baseline, target, data owner, reporting interval, and confidence level. Label evidence as forecast, commitment, activity, intermediate result, or outcome. This simple taxonomy prevents a press release from being compared directly with measured employment or productivity.
Use paired indicators to expose trade-offs: output with error rates, time saved with review effort, training completion with placement, investment with operational capacity, jobs created with job duration, and efficiency per unit with total energy demand. Note missing data rather than filling gaps with assumptions.
Review the scorecard at predetermined intervals. A short operational pilot may be reviewed monthly, while occupational mobility and wage effects require longer observation. Update the evaluation when the use case, deployment scale, workforce plan, or infrastructure schedule changes.
Practical evaluation checklist
Before accepting or implementing an AI growth proposal, confirm that you can answer these questions:
- Is the claim written with a defined intervention, population, comparison, period, and measurable outcome?
- Are forecasts, funding commitments, deployed activity, and verified outcomes reported separately?
- Does the productivity measure include quality, correction work, oversight, and AI operating costs?
- Does workforce analysis distinguish task exposure, role transformation, displacement, and job creation?
- Are results segmented by occupation, region, gender, income group, seniority, and contract type where relevant?
- Does training reporting extend from access and completion to recognized credentials, placement, wages, and retention?
- Are temporary construction roles separated from permanent operating roles and from jobs transformed elsewhere?
- Does the infrastructure plan name computing, electricity, grid, cooling, schedule, and local cost constraints?
- Are gains and risks reported separately for workers, employers, customers, and host communities?
- Are human review, source access, responsibility for errors, and appeal routes defined for consequential uses?
- Is each indicator assigned to a data owner and a review date?
- What evidence would cause leaders to pause, redesign, or stop the program?
The purpose of this framework is not to assume that AI growth is beneficial or harmful. It is to make broad claims auditable. Productivity, workforce change, training, infrastructure, and distribution interact, but they should not be blended before each has been measured. When decision-makers preserve those distinctions, they can identify genuine gains, expose shifted costs, and revise a program before optimistic language hardens into policy or investment commitments.
AI Tools Radar separates product facts, editorial judgment, and commercial placement. Updated facts retain their verification date.
