From Paper to Practice

Helping governments build the teams that turn AI strategies into working public services — and a low-cost way to test whether it works

Partnership for the Americas · Research Division · research@partnershipamericas.org

The Argument in Brief

The development community has done the diagnostic work on artificial intelligence in government with unusual speed and candor. We now know, from the multilateral institutions' own research, that more than seventy countries have national AI strategies; that translating them into operational reality is where many stall; that public administrations face widespread AI literacy gaps and acute shortages of advanced technical expertise; that this shortage produces over-reliance on external vendors; and that AI projects fail at high rates — on cross-sector estimates, not public-sector-specific ones — without rigorous business cases and evaluation.2 The problem has been named precisely.

The 2026 World Development Report, published in August, sets out the destination: developing economies should adopt existing tools, adapt them to local conditions, and only then advance toward building their own — supported by clearer procurement frameworks, institutional quality and skilled public workforces.1 What is less settled is how an administration with few technical staff climbs the first two rungs. Development institutions can fund an AI project, assess readiness, publish standards and train officials. What we could not find is an instrument that puts experienced practitioners inside a ministry to build something real alongside the people who will run it — the step at which a great deal of the value is realized or lost.

This paper argues that the missing instrument is a small one, and that it can be tested cheaply. It also argues that the objective should be understood correctly from the outset. The point is not to deliver AI systems to governments, and not to hand one over and depart. It is to help build the permanent internal team an administration needs to select, adapt, govern and sustain these systems — and, in the interim, to leave officials able to specify, commission and supervise the work. On present trends, the organizations selling AI to a finance ministry are frequently larger than the state buying it. Whether governments meet that asymmetry as informed clients or as customers is an open question that development institutions are well placed to influence.

The diagnosis is widely shared. What is thin is delivery capacity — and the judgment inside government that hands-on delivery appears to build faster than training alone.

We suggest a modest first move: that an interested institution commission a short feasibility and design engagement — on the order of three months — to test demand with real ministries, work out the procurement and financing path, and return with either a costed pilot design or an honest recommendation not to proceed. This paper is offered as a contribution to that discussion rather than as a proposal, and its authors would welcome disagreement with any part of it.

I. What the Evidence Already Says

It is worth beginning with what the multilateral institutions have themselves established, because the case for delivery capacity does not require any new claims. It follows from findings already published.

The gap between strategy and operation

The World Bank's assessment of public institutions in the age of AI puts the central problem plainly: while more than seventy countries have adopted national AI strategies, converting them into operational reality remains the difficulty. The countries that manage it, it observes, tend to pair aspiration with enforceable directives, mandatory risk assessments, budget allocation and performance indicators — practices associated with success rather than proven causes of it.2 The same pattern appears in the Bank's GovTech Maturity Index, which now covers 197 economies: when governments were asked whether they track the actual usage and uptake of the systems they have built, the responses revealed a significant need to improve — many countries have policies and platforms without the means to tell whether they work. The index relies substantially on self-reporting and is a measure of maturity, not performance.3

Figure 1

Where public-sector AI stalls

Attrition from AI strategy to production Four stages — national strategy, standards and guidance, pilot projects, and systems in production — with each stage narrower than the last, and the final stage marked as the point where most efforts stop. National strategy 70+ countries Standards & guidance Widely published Pilots Funded and run In production Where efforts thin out fewest instruments operate here THE ATTRITION
Illustrative; widths are not to scale. Author synthesis of the pattern described in World Bank, Public Institutions in the Age of AI and the GovTech Maturity Index 2025 — not a measured attrition rate.

The capacity constraint underneath it

The reason is not a shortage of intent. It is a shortage of people. The Bank's own language is direct: public administrations face widespread AI literacy gaps and acute shortages of advanced technical expertise, and without structured capability governments risk data leaks and over-reliance on external vendors. It describes public servants already using AI tools without guidance or training as a shadow AI workforce, carrying the risks that implies.2 Harvard's Belfer Center documents the same constraint from the hiring side in Latin America: a thin technical pipeline, competition from the private sector, and public hiring processes that struggle to keep pace.4

A gap that is not closing

The distributional picture makes this urgent rather than merely important. The IMF's AI Preparedness Index, covering 174 economies on a zero-to-one scale, averaged 0.68 for advanced economies, 0.46 for emerging markets and 0.32 for low-income countries in 2023.5 That is a snapshot rather than a trend, but the World Bank's GovTech data does track movement, and points the same way: higher-income economies have generally advanced while low-income ones have regressed.3 Whatever AI does to public administration, it is currently on course to do it unevenly.

Figure 2

AI preparedness by income group, 2023

IMF AI Preparedness Index by income group Advanced economies average 0.68, emerging markets 0.46, and low-income countries 0.32 on a zero to one scale. 0.68 0.46 0.32 Advanced economies Emerging markets Low-income countries 0.0 0.25 0.50 0.75 1.00 AI Preparedness Index score THE GAP
Empirical data. IMF AI Preparedness Index, 174 economies, 0–1 scale, group averages for 2023 — a single-year cross-section, not a trend. The World Bank's GovTech Maturity Index does track movement over time and points the same way.

And a prize large enough to justify the effort

The upside is not speculative either. The United Kingdom's trial of Microsoft Copilot across twenty thousand civil servants reported average time savings of twenty-six minutes per working day. The figure is self-reported, the trial had no control group, and the evaluators noted they could not establish how the saved time was actually used.6 It is an indication of promise, not proof of service improvement — and that distinction is precisely the point. The question is not whether the technology can help. It is whether governments can get it into production and then tell whether it worked.

Figure 3

The most-cited government trial, and its limits

Measured productivity effect of an AI assistant in the UK civil service A trial across twenty thousand civil servants found average savings of twenty-six minutes per working day, equivalent to roughly two working weeks per official each year. 20,000 civil servants in the trial 26 min saved per working day self-reported no control group THE EVIDENCE, HONESTLY
Empirical data, with important caveats. UK cross-government Microsoft 365 Copilot experiment, 20,000 participants across 12 organisations, Sept–Dec 2024. Time savings were self-reported, there was no control group, and the evaluators stated they could not identify how the saved time was spent. An indication of promise, not evidence of service improvement.

II. The Missing Instrument

When large enterprises hit this same wall, their answer was to pair the software with people. Forward-deployed engineers embed inside the client organization, work with its staff on its actual systems, and configure and integrate in place — alongside the product, not instead of it. The model is most closely associated with Palantir and has spread quickly through enterprise AI. OpenAI has announced an initial investment of more than four billion dollars in a separate deployment business and an agreement to acquire the consultancy Tomoro, bringing roughly 150 forward-deployed engineers; Accenture describes a practice numbering in the thousands, and Deloitte has launched one of its own.7, 8, 9

The analogy should be handled with care, and we would rather state its limits than oversell it. These are, in part, distribution strategies for platforms, and enterprises differ from ministries in ways that matter enormously: statutory authority, procurement law, due process, transparency obligations, and continuity across changes of government. A model designed to make a corporation adopt a vendor's product cannot be transplanted whole into a public institution.

The underlying observation is not only a commercial one. Government digital delivery arrived at a similar conclusion by a different route. The structure developed in the United Kingdom and adopted widely since — a discovery phase that investigates the problem before anyone builds, an alpha that tests ideas in prototype code explicitly expected to be thrown away, then a beta, live operation, and an assessment at each boundary where the work may proceed, repeat or stop — exists because remote specification does not work for public services either.10 Two different traditions, one shared conclusion: someone has to be in the room.

Figure 4

Two traditions, one conclusion

Commercial and public-sector delivery models converge Enterprise AI adopted forward-deployed engineers; government digital services adopted discovery and alpha phases with assessments. Both place practitioners inside the organization before building. VENDOR-EMBEDDED DEPLOYMENT Forward-deployed engineers Associated with Palantir; now widespread PUBLIC DIGITAL-SERVICE DELIVERY Discovery, alpha, assessment GOV.UK Service Manual and its descendants Someone has to be in the room
Author synthesis. Vendor-embedded deployment and public digital-service delivery developed for different reasons and differ in important ways — the first configures a product, the second builds a service. Both concluded that specification at a distance does not work.

Why AI is harder than ordinary government software

Much of this would be true of any government technology project. Three things make AI different, and all three argue for permanent capability rather than a one-time build. Outputs are probabilistic rather than deterministic, so a system cannot be tested once and declared correct. Performance drifts as data and models change underneath it, which makes monitoring a standing operating function rather than a closing task. And the vendor and model landscape moves fast enough that portability — changing supplier without rebuilding — is a live design requirement. A government that cannot evaluate, monitor and switch is exposed in a way it would not be with a payroll system. The corollary is that AI should not be used where a rule change, a process redesign or conventional software would be safer or cheaper, and part of a competent team's job is saying so.

III. The Intelligent Client

There is a second reason to send people, and for governments it may be the more important one.

Even a government that intends to buy rather than build needs enough internal capability to specify what it wants, evaluate what it is offered, and supervise what is delivered. British government guidance calls this the intelligent client function — an in-house capability to translate policy into outcomes and to monitor, challenge and verify what providers deliver. That guidance was written for the property and estates function rather than for technology, and we borrow the term deliberately, because the problem it describes transfers exactly.11 Where the capability is absent in digital procurement, the United Kingdom's own audit office has documented the results plainly: repeated delays and cost overruns, commercial teams lacking the technical depth to manage digital contracts, and growing dependence on suppliers that, in its words, are bigger than governments themselves.12

If that is the position of a wealthy government with a mature digital service, the position of a ministry with two or three technical staff is harder still. Capability varies widely and some administrations are formidable buyers; but where no official can interrogate a model's error rate or tell a sound proposal from an expensive one, the ministry is not really negotiating. The World Bank's own recommendation on this point — that governments make careful decisions about internal expertise versus external procurement — presupposes exactly the capability that is often missing.2 A careful decision about what to outsource requires someone who understands what is being outsourced.

Capability of this kind does not appear to come from training courses alone. It seems to come from having built something once, alongside someone who had done it before — a proposition this paper treats as a hypothesis worth testing, not a settled finding.
Figure 5

The asymmetry a ministry negotiates against

Capability asymmetry between technology suppliers and government buyers Suppliers bring deep technical capability and dedicated commercial teams; ministries often have neither, and the intelligent client function is what closes the gap. THE SUPPLIER Deep technical capability Dedicated commercial teams Global scale and pricing power A product to sell THE MINISTRY Often generalist procurement staff Few who can interrogate an architecture Limited baseline to test claims against A problem to solve negotiation The intelligent client function is what makes this a negotiation rather than a queue.
Author synthesis; capability varies widely between administrations and many are formidable buyers. The UK's National Audit Office describes government dependence on suppliers “bigger than governments themselves.” The intelligent client term comes from UK property guidance and is borrowed here by analogy.

IV. Who Does What Today

It would be wrong to suggest this field is empty. It is not. A review of published documentation, current to August 2026, shows a well-populated landscape — and one specific combination we could not find in it.

  • The World Bank, through its GovTech program, benchmarks digital government across 197 economies, convenes an AI working group spanning 35 governments and institutions, runs AI bootcamps for public administrations, publishes the most useful synthesis available of how public institutions are adopting AI, and has just devoted its flagship World Development Report to AI for development.1, 2, 3, 13
  • The Inter-American Development Bank, through fAIr LAC, has built the reference tools for applying ethical principles across the phases of an AI project and supports pilots developed with its partners and regional hubs.14
  • UNDP supports partners to deploy senior strategic technical talent embedded in government — in its own description, to make decisions on vendor selection and management, IT architecture and recruitment — alongside local capability building and open-source work.15
  • GovStack supplies the specifications, building blocks, playbooks and training that let countries design, prototype and scale services, delivered through its partner institutions.16
  • CAF, with UNESCO, has taken AI readiness assessment and ethics into member countries, part of a wider programme of support to more than fifty countries on ethical AI policy.17

Each of these is necessary and several are excellent, and there is genuine overlap with what we describe below — embedded advisers, prototyping support, pilot financing and capability building all exist today. What the published record does not describe is a standardized, repeatable facility combining four things at once: platform-neutral embedded delivery, co-production with the ministry's own staff, explicit transfer of the work to a permanent internal team, and measurement of whether that team is still running the system a year later. We offer that as a preliminary finding from a documentation review rather than a settled conclusion. Interviews could change it, and we would rather be corrected early than late.

Others assess readiness, set standards, train officials, fund pilots and advise on what to buy. What we could not find is a repeatable way of building alongside a ministry and leaving a permanent team behind to run it.
Figure 6

A well-populated landscape, and the combination we could not find

Layers of AI support to governments and who occupies them Principles, standards, benchmarking, training and pilot financing are each occupied by named institutions. The build-and-hand-over layer has no standing occupant. THE STACK Principles & ethics IDB fAIr LAC · CAF–UNESCO · UNESCO Standards & building blocks GovStack Benchmarking & diagnosis World Bank GTMI · IMF AIPI Training & advisory World Bank AI bootcamps · UNDP embedded advisors Pilot financing MDB lending and technical cooperation Build it together, then hand it to a permanent team not found
Author synthesis from published programme documentation, current August 2026; not based on interviews. Institutions are listed where their published work concentrates, and several operate across more than one layer. Adjacent provision is extensive — the bottom row marks a combination we did not find described as a standing, repeatable facility.

V. What Such Teams Would Do

What follows is a proposed shape rather than a settled method. Anyone claiming to know the correct duration of each phase before running one is guessing, and we would rather say so than present false precision. The structure borrows deliberately from the discovery-and-assessment model already established in government digital delivery.10

A short visit, and an honest list

A small team spends roughly two weeks inside an agency that has asked for help. It examines real systems, real data and real workflows rather than presentations. It leaves behind four things: a diagnosis of where AI would and would not help; a shortlist of candidate use cases scored for feasibility and value; a costed roadmap written so the ministry's finance officials can read it; and a list of what the government must have in place before anyone can responsibly build anything — data access authority, a named product owner, security clearances, infrastructure, and staff assigned rather than borrowed.

Return when it is ready

The team comes back to build once the government has completed that list. This is the design decision that carries the most weight. It means nothing is promised that cannot be delivered, and it means the government's own follow-through determines what happens next rather than an external judgment of its readiness. A ministry that assembles the prerequisites has demonstrated commitment in a way no external assessment can; a ministry that does not was unlikely to succeed against a fixed calendar anyway. The list should carry a validity window, so that a scoping exercise does not sit indefinitely against changed circumstances.

The build phase runs longer than the visit. It should be understood as a route to production rather than production itself: a first working version tested with real users, then the security, legal and rights assurance a public service requires before it carries real cases. Nothing should go live without a named permanent owner, an operating budget, a monitoring plan and an exit route. Whether the right build length is six weeks or twelve, and whether a team works best at four people or six, are questions a first set of engagements should answer rather than a paper should assert. What is already clear is that software engineering alone is insufficient: public service work also needs product judgment, service design, security and privacy, and someone who understands both the sector and the language.

Figure 7

The engagement, and the gate in the middle of it

The proposed engagement cycle with a readiness gate A two-week discovery visit produces a diagnosis, a roadmap and a prerequisites list. A readiness gate follows: the build phase begins only once the government has completed the list. The build ends in handover. THE CYCLE STEP 1 · ~2 WEEKS Discovery visit Real systems, real data, real workflows LEAVE-BEHIND Diagnosis · scored use cases Costed roadmap Prerequisites checklist THE GATE Government completes the list STEP 2 · WEEKS Build together Then assurance, production, and a permanent team not ready — the list keeps, within a validity window Nothing is promised that cannot be delivered. A ministry that assembles the prerequisites has demonstrated commitment in a way no external readiness rating can.
Author synthesis; a proposed shape, not a settled method. The gate is the design decision that carries the most weight: the government's own follow-through determines whether the build happens, and when. Note that a build phase reaches a tested first version — production additionally requires security, legal and rights assurance, a permanent owner, an operating budget and an exit route.

VI. What Success Would Look Like

Because the objective determines the design, it is worth stating precisely. A successful engagement produces three things. They are listed in the order they appear, not in order of importance — the public value and the capability matter equally, and an engagement that delivered one without the other would have half failed.

  • Something that works. One bounded, useful system running on government infrastructure, addressing a problem the agency's own staff selected.
  • People who can run it. Named counterparts who built alongside the team, hold the documentation, and can operate and extend the system unaided.
  • A team that outlasts the engagement. A small permanent group inside the administration with the authority, budget and standing to select, adapt, govern and sustain these systems — able to specify, commission and supervise the next project whether they build it, buy it, or simply ask a supplier a harder question than they could have asked before.
Figure 8

What an engagement should leave behind

Three outcomes of an engagement in ascending order of importance A working system, people who can run it, and people who can judge the next project. The third compounds and is the intended end state. ALL THREE, NOT A RANKING 1 · Something that works One bounded system, running on government infrastructure, on a problem its own staff chose 2 · People who can run it Named counterparts who built alongside the team and can operate and extend it unaided 3 · A team that outlasts the engagement A small permanent group with the authority, budget and standing to select, adapt, govern and sustain these systems after the external team has gone
Author synthesis. All three are necessary; the order is the sequence in which they appear, not a ranking. Public value and lasting capability matter equally, and both should be measured at six and twelve months alongside service quality, cost and safety.

The third is the outcome that compounds, and it is what makes this a development proposition rather than a procurement of services. It is also where the model departs from a conventional build-and-handover: the external team is a transitional mechanism whose job is to help stand up a permanent internal one, not a supplier that finishes and leaves. This is the direction the World Development Report already points — adopt, then adapt, then advance — and the unresolved question it leaves open is how an administration with few technical staff gets onto that ladder at all.1 Evaluation should follow accordingly. Alongside service quality, cost and safety, the test at six and twelve months is whether the agency still runs the system, still has the people, and still has the budget.

VII. A Feasibility and Design Phase

The honest state of the evidence does not support a large commitment, and we are not suggesting one. It is not established whether ministries will commit staff and data access to this, what it should cost, or which financing instrument fits. Those questions are answerable — by a combination of going and asking, and of costing what the answers imply.

An interested institution could commission a short feasibility and design phase — roughly three months, the size of an ordinary technical cooperation — producing five things. It can test demand, institutional fit, readiness, procurement feasibility and delivery economics. It cannot establish whether capability endures, which is what a subsequent pilot would be for.

  • A demand test. Structured conversations with a defined set of governments and country offices, aimed at specifics rather than generalities: which agency, which problem, which official would sponsor it, and what the government itself would commit.
  • A partnership map. An honest account of what the institution's peers already provide, and where this capability would complement rather than duplicate them — including where a government should simply be referred elsewhere.
  • A tested delivery model. The engagement design worked out in practice, ideally including one or two real discovery visits with willing agencies, so the model is described from experience rather than invented.
  • A financing and procurement path. Worked through with the institution's own legal and procurement staff: how a build phase would be tendered so that conducting a discovery visit confers no unfair advantage, how continuation would be financed, and what the rules actually permit. This cannot be assumed from outside.
  • A costed pilot design and a recommendation — unit costs, staffing, travel, procurement lead times, expected country contributions, and the recurring cost of operating a system after handover — including, explicitly, a recommendation not to proceed if the evidence does not support it.

This commits an institution to learning something rather than to building anything, and its output is what is actually needed to make a decision later: evidence about demand, a validated delivery model, a legal pathway, and a defined go-or-no-go point. On the question of whether such a capability ultimately belongs inside the institution or with an external operator, there are serious arguments both ways, and they turn on facts about each institution's own hiring and procurement rules that are better established from the inside than assumed from outside.

VIII. What We Do Not Know

It seems more useful to set out the open questions than to paper over them. Each is a reason the test comes before the program.

  • Will governments actually complete a prerequisites list, and fund a permanent team afterward? These are the real tests of demand, and enthusiasm in a meeting is no substitute for either.
  • What team composition works? The right mix of engineering, product, design, security and sector expertise is an empirical question.
  • How is a build phase procured so that the discovery visit does not prejudice competition? This must be settled before the first engagement, not after.
  • Which use cases are appropriate to start with, given that the highest-value applications are often those that touch citizens' rights most directly?
  • How much genuinely transfers between countries? That a pattern proven in one revenue authority accelerates the next is plausible and untested.
  • What does a government need to keep a system running two years later, and who pays for it?

IX. Risks Worth Naming

Embedded assistance can create dependency instead of capability. The answer is the objective itself, made measurable: engagements judged on whether officials can operate the system without the team, tested at six and twelve months. If they cannot, the engagement failed regardless of what shipped.

The rights risk is more serious and deserves more than a sentence. The World Bank's synthesis documents what failure looks like at scale in AI and automated decision systems — the Dutch childcare benefits scandal, which wrongly accused more than twenty-six thousand families through algorithmic risk selection; contested identification errors in Canadian immigration processing; the United Kingdom's pandemic-era exam grading algorithm, a statistical allocation model rather than machine learning — and concludes that such failures can violate rights at scale and fall hardest on marginalized populations.2 The implication is a formal risk assessment and a presumption that early engagements begin away from automated decisions on benefits, health, enforcement, policing or justice. That narrows what can be attempted first, which is a real cost and the correct trade.

Figure 9

Where to start — and why category alone is not enough

Risk tiering for early engagements Bounded internal workflows are appropriate to start with. Decisions affecting benefits, health, enforcement, policing and justice should be deferred until assurance is proven. RISK TIERING START HERE Bounded internal workflows Document processing and summarization Case triage with the decision retained by staff Internal search across administrative records Drafting support and translation Still requires assessment — none of these is risk-free DEFER UNTIL ASSURANCE IS PROVEN Automated decisions affecting rights Benefit eligibility and entitlement Clinical triage and treatment Policing, surveillance and immigration Tax liability, enforcement, justice Where documented failures have caused harm at scale The Dutch childcare benefits scandal wrongly accused more than 26,000 families. Similar failures are documented in Canadian immigration processing and UK exam grading. Public-sector AI failure violates rights at scale.
Author synthesis; failure cases documented in the World Bank's Public Institutions in the Age of AI. The left column is a starting presumption, not a safe-list: risk depends on consequence, data sensitivity, degree of automation, scale, reversibility, and whether appeal and manual fallback exist. Case triage and summarization can affect rights even inside a back office.

Risk tiering should be done properly rather than by category. A bounded internal workflow is not automatically safe: case triage shapes who gets seen first, internal search can surface information a user should not see, and summarization can quietly distort the record a decision rests on. What matters is the consequence for the person affected, the sensitivity of the data, how much of the decision is automated, the scale, whether it can be reversed, and whether an appeal and a manual fallback exist. Human review is a necessary safeguard but not a sufficient one, since reviewers tend to defer to the machine.

Teams rotating across finance, tax, health and security institutions in several countries also concentrate sensitive knowledge in a small group, which requires country-level compartmentalization, personnel vetting, and clear rules about what leaves the building. The written record from each engagement should separate patterns safe to publish as a public good from specifics — fraud-detection logic, architecture, known weaknesses — that plainly are not.

Finally, there is the risk of duplicating what already exists. That is why the partnership map belongs in the first three months rather than in a footnote. If the conclusion is that another institution is better placed, that is a useful finding reached at very low cost.

In Closing

Governments are being handed a technology capable of augmenting a great deal of public work, at a moment when few people can help them use it well and many are positioned to sell it to them. The World Development Report has set out where developing economies need to get to. The open question is how administrations with thin technical benches take the first step. Teams that build alongside public servants — and leave behind a permanent group more capable than the one they found — are one answer worth testing, and development institutions are among the few actors with the reach, the relationships and the standing to test it.

We do not ask anyone to accept that on the strength of a paper. We suggest it is worth three months to find out, and we would welcome the argument either way.

Notes and Sources

  1. 1World Bank, World Development Report 2026: The Promise of Artificial Intelligence (August 2026), and the accompanying release, "AI Offers Lifeline to Developing Economies in an Era of Weak Growth" — source for the adopt–adapt–advance sequence and for the emphasis on procurement frameworks, institutional quality and skilled public workforces.
  2. 2World Bank, "Public Institutions in the Age of AI: Emerging Practices" — source for the seventy-plus national AI strategies, the AI literacy and expertise gaps, over-reliance on external vendors, the "shadow AI workforce," the internal-expertise-versus-procurement recommendation, and the failure cases in Section IX. The high project-failure figure it reports is a cross-sector estimate drawn from RAND, not a public-sector-specific rate.
  3. 3World Bank, GovTech Maturity Index 2025 Update: 197 economies assessed; monitoring of actual usage identified as a significant gap; higher-income economies advancing while low-income economies regress. The index measures maturity rather than performance and relies substantially on self-reported survey responses.
  4. 4Belfer Center for Science and International Affairs, Harvard Kennedy School, "Closing the Government Tech Talent Gap: Lessons for Latin America".
  5. 5IMF, AI Preparedness Index, covering 174 economies on a 0–1 scale. Group averages for 2023: 0.68 advanced economies, 0.46 emerging markets, 0.32 low-income countries. A single-year cross-section, not a time series.
  6. 6UK Government, "Microsoft 365 Copilot Experiment: Cross-Government Findings Report": 20,000 participants across 12 organisations, September–December 2024. Time savings were self-reported, there was no control group, and the report states that it was not possible to identify how the saved time was spent.
  7. 7OpenAI, "OpenAI Launches the Deployment Company" (May 2026): an initial investment of more than US$4 billion and an agreement to acquire Tomoro, bringing approximately 150 forward-deployed engineers. Cited as evidence of announced investment, not of results.
  8. 8Accenture, "Accenture Launches Microsoft Forward Deployed Engineering Practice" (March 2026).
  9. 9Deloitte, "Announcing Forward Deployed Engineering" (December 2025). The announcement does not state the practice's scale.
  10. 10UK Government Digital Service, GOV.UK Service Manual. Note that alpha explicitly produces prototype code teams should "expect to throw away" — it is not a production build, and the phases after it (beta, live, and continuous operation) are where a service becomes real.
  11. 11UK Government, "Intelligent Client Roles: Functional Guidance". This guidance sits under Government Functional Standard GovS 004: Property; the term is borrowed here by analogy to digital and AI procurement, where no equivalent named function exists.
  12. 12UK National Audit Office, "Government's Approach to Technology Suppliers: Addressing the Challenges".
  13. 13World Bank, Global Program on GovTech and Public Sector Innovation, including the GovTech Maturity Index, the AI working group and the AI bootcamps for public administrations.
  14. 14IDB, fAIr LAC.
  15. 15UNDP, "UNDP Highlights Successful Digital Transformation Strategies at Inaugural GovTech Congress".
  16. 16GovStack, Country Collaborator.
  17. 17UNESCO, "UNESCO to Support More Than 50 Countries in Designing Ethical AI Policy", the programme under which CAF and UNESCO have worked with member countries in Latin America.