Editorial NoteThis report was commissioned by Coefficient Giving (CG) and produced by Rethink Priorities from November 2025 to January 2026. We lightly edited the report for publication. Coefficient Giving does not necessarily endorse our conclusions, nor do the experts we interviewed or the organizations with which they are affiliated. The report evaluates the potential of “fieldbuilding” activities in the “AI for Good” (AI4G) space: efforts to create the enabling conditions for many organizations to develop and deploy AI tools in low- and middle-income countries (LMICs), as distinct from funding individual implementers directly. We take up four questions: what barriers stand in the way of launching AI4G interventions in LMICs; what the likely impact of alleviating those barriers would be, and what kinds of bets or organizations CG might fund to do so; who the other major funders in the space are and what they have supported; and which organizations might serve as AI4G “fieldbuilders.” The report is structured accordingly: we first map the lifecycle of a typical AI4G intervention and catalog the barriers that arise along it; we then examine four priority barriers in depth and assess candidate solutions for each; and we close with a review of the current funding landscape. Supporting detail on our methods and analysis is provided in the appendix. To produce this report, we drew on a scan of the published and gray literature and conducted 13 interviews with experts across health, agriculture, education, and disaster response, including both practitioners deploying AI tools and commentators who take a more critical or holistic view of the field. Eight of these experts agreed to be named; the remainder are referred to anonymously. We assessed the severity of barriers and the promise of candidate solutions qualitatively, supplemented by a structured Bayesian aggregation of our own team’s beliefs; the resulting scores reflect our subjective judgments rather than definitive estimates. We have no conflicts of interest to declare. We have tried to flag major sources of uncertainty throughout, and we remain open to revising our views in light of new evidence or further research. |
Executive Summary
This work was commissioned by Coefficient Giving (CG) to evaluate the potential of “fieldbuilding” activities in the Artificial Intelligence for Good (AI4G) space in low- and middle-income countries (LMICs). In this report, we identify launch bottlenecks, survey the existing funding landscape, and highlight promising solutions to enable more interventions.
What We Did
- We outlined the lifecycle of the typical AI4G intervention, drawing from first principles and implementation research.
- We cataloged 60 possible barriers to the implementation of AI4G interventions across the lifecycle, drawing from our own intuitions, a literature scan, and 13 interviews with key experts. (We hold high confidence (90–99%) that this catalog represents the primary obstacles currently discussed in the field, despite some limitations regarding survivorship bias in our interview sample.)
- We used this list to qualitatively assess the top barriers, identifying 15 in total, of which a third (5) were identified as top barriers within the scope of fieldbuilding.
- We examined four of these in-scope barriers in detail.
- We identified 23 potential solutions to these four barriers and used Bayesian Aggregation Methods to quantify our views on whether these were promising.
- We also conducted a brief investigation into the state of funding for AI4G.
Key Findings
Barriers
- Regarding the overall distribution of barriers, we find that they concentrate most heavily at the beginning and end of the development pipeline: Solution Design (Stage 2 in our pipeline) and Scaling/Sustainability (Stage 8 in our pipeline).
- Because so many barriers appear at the early implementation stages, we believe that unlocking these challenges might have a large return because doing so will allow more interventions to proceed to testing, unlocking more solutions for development challenges.
- Our list of top barriers within scope (challenges that can be addressed through fieldbuilding activities and those that cannot) includes:
- Lack of context-specific training data
- Tech and telecommunications infrastructure, equipment, and costs
- Misalignment between AI builders and decision makers
- Local regulatory barriers
- User economics (we did not explore this barrier or its solutions in detail).
Solutions
- In examining 23 solutions for the first four barriers, we found that only eight met our threshold to be classified as “Likely Promising,” though we only have medium to high confidence in one of these. The two highest-scoring solutions specifically address the lack of context-specific training data, suggesting this is a promising area to consider targeting.
Overview of Proposed Solutions
Below, we provide high-level summaries of the four top barriers, and some detail with regard to the fieldbuilding solutions we find most promising for each.
- Lack of context-specific training data: AI models trained on datasets from high-income countries frequently underperform in LMIC settings due to critical gaps in context data (e.g., local crop diseases) and language data (e.g., under-resourced dialects). These gaps persist due to structural misalignment regarding what data is collected versus what is useful to local communities, competitive market pressures that encourage data hoarding, and regulatory uncertainty.
- Context Data: We view Direct Data Generation (funding local teams to create datasets via crowdsourcing or service delivery campaigns) as the most promising solution.
- Language Data: Our highest-leverage recommendation is to support NGOs in fine-tuning existing frontier models for moderately resourced languages (like Swahili), while investing in ethical raw data collection for under-resourced languages. We are reasonably confident this is a promising solution.
- Tech and telecommunications infrastructure, equipment, and costs: A lack of reliable, affordable high-speed internet in LMICs creates a hard infrastructure barrier for AI adoption. This threatens data-heavy applications (e.g., video/audio recognition for medical diagnosis), posing fewer risks to text-based tools.
- Offline First: Rather than attempting to subsidize connectivity or build new telecommunications infrastructure, we believe the highest-leverage solution is enabling offline capabilities (Edge AI and tinyML) that allow AI interventions to function without constant internet access.
- AI in a Box: Because existing software is often sufficient to improve upon the counterfactual (no doctor/expert available), we see a plausible funding opportunity in subsidizing mid-range “AI in a Box” servers for AI4G interventions, though we are uncertain regarding specific costs and demand.
- Misalignment between AI builders and decision makers: The disconnect between builders and decision makers leads to three consequences: builders often misunderstand the context; decision makers lack frameworks to assess AI risks/benefits; and trust is eroded.
- For Builders: We believe funding advisory support for AI builders offers a potentially promising route to addressing context gaps. This can be part of incubator/accelerator programs, which are effective but expensive; short-term consultancy services offer a less costly model, although we are uncertain about impact.
- For Policy Makers: While advisory support for policy makers is promising in the medium-term (~5 years), there is no good evidence this type of intervention works.
- Local regulatory barriers: Regulatory barriers can be passive (e.g., lack of harmonization) or active (e.g., strict data sharing laws). We view passive barriers, particularly for health AI, as the most tractable focus area.
- Legislation & Hubs: The most promising interventions are supporting the drafting of new AI-specific legislation for health and establishing “regional hubs” to diffuse regulatory best practices from LMICs with robust AI infrastructure to those with less capacity.
- Risks: We remain unsure about how promising these interventions might be, given uncertainty over costs, the actual demand from countries to adopt legislative models from other settings, and the difficulty of measuring success in this area.
Funding
- We estimate very roughly that ~$302 million of philanthropic funding or official development assistance has been committed to AI for Good in LMICs since 2017 (~$33 million/year). We are uncertain about the trajectory of this funding: official development assistance is likely to decrease in the next five years, but this may be (massively) offset by commitments by technology companies.
- Current funding prioritizes direct grants to individual implementers over broader ecosystem enablers. Fieldbuilding activities receive relatively little attention, with the exception of very general and non-sector-specific convenings/conferences.
Introduction
We can think of two possible compelling approaches to funding in the AI4G sector: 1) identify and fund effective AI4G implementors, and 2) specify pain points in the launch phase of AI4G endeavors. In the report, we focus on the second theory of change—referred to as “fieldbuilding” AI4G. Specifically, we ask:
- What are the barriers[1] in the process of launching AI4G interventions in LMICs, whether adapting an existing product from High-Income Countries (HICs) or developing a new intervention from scratch?
- In identifying a comprehensive list of barriers, we distinguish between those that act as fundamental roadblocks versus those that are hurdles that are easier to overcome. We label the first group Top Barriers. Within this set, we differentiate between those that can be addressed through fieldbuilding activities and those that cannot. We label the first subgroup Top Barriers Within Scope.
- What is the likely impact of alleviating these barriers, and what types of bets or organizations could CG fund to achieve this?
- Who are the other major funders in the AI4G space, and what have they invested in?
- What are the organizations that could be considered AI4G fieldbuilders?[2]
We answer these questions roughly in turn. Our approach included a short literature scan, as well as 13 conversations with experts from across the field of AI4G, including practitioners in health, agriculture, and disaster response, as well as commentators who take a wider view of AI interventions as a whole, both from a positive and critical perspective. Details can be found in the table below.
Table 1: List of experts interviewed for this project that agreed to attribution
| Name(s) | Affiliation | Role |
|---|---|---|
| David Yanagizawa-Drott | Poverty Action Lab (J-PAL) | Co-Chair, Partnership for AI Evidence (PAIE) |
| Edwin Klinkenberg | Delft Imaging | Clinical Application Specialist |
| Evert Bopp | Crisis Cognition | Co-Founders |
| Jay Patel | Jacaranda Health | Director of Technology |
| Maëlle-Marie Corso | PATH | Health AI Policy Researcher |
| Markus Goldstein | CGD | Vice President |
| Rikin Gandhi | Digital Green | CEO |
| Shabnam Aggarwal | Rising Academies | CTO |
| 5 other anonymized experts |
Identifying Challenges
To begin, we developed from first principles the conceptual steps required to build or deploy an AI4G intervention in an LMIC. We then used this framework to identify a full list of possible challenges, and to group them conceptually by stage, severity, and other characteristics. We did so to help us identify top barriers.
Process Mapping
Over a period of one working week, we developed a comprehensive process map to understand how AI4G interventions might be created and scaled up, and then identify where in the process they might encounter obstacles. The initial goal was to create a framework granular enough to identify specific intervention points, while remaining generalizable across health, agriculture, education, and other sectors. We ultimately organized the process into an eight-stage linear framework with an additional category for cross-cutting issues:
- Problem Identification & Scoping
- Solution Design & Development
- Validation & Testing
- Regulatory & Policy Navigation
- Implementation & Deployment
- User Adoption & Training
- Monitoring & Evaluation
- Scaling & Sustainability
To ensure our process map was comprehensive and aligned with established implementation science, we compared our framework against 11 existing models from technology development, diffusion theory, health systems, and development practice. The comparison revealed that no single framework covers all aspects of AI deployment in LMICs; we believe our eight-stage process map is more comprehensive than any other individual framework. We list the relevant frameworks and discuss the findings in slightly more detail in the appendix.
List of Challenges
Approach
We compiled a comprehensive catalog of obstacles that AI4G interventions encounter in the process of building tools and scaling in LMICs through a literature review (discussed further in the appendix), expert interviews,[3] and our own brainstorming.
The final list contains 60 barriers. Each entry includes a concise name for reference, detailed description of the challenge, three to four concrete examples from different sectors, source attribution (including if the primary source is our own brainstorming), the primary implementation stage(s) where the barrier manifests, classification by barrier type, and other helpful criteria that we used to anchor our research.[4]
We acknowledge various limitations to our approach. With regard to the literature, it remains difficult to assess whether our (or any) scan/review could adequately capture the full range of challenges.[5] Moreover, no single resource did a good job of mapping the severity of any specific challenge, which was important to our analysis of identifying the top issues. Still, we believe our scan has identified the majority of obstacles related to AI4G ventures that are currently being discussed in print, and that any gaps would have likely surfaced in conversations with experts.
We also acknowledge that our interview selection process may be biased toward relevant, positive outcomes: organizations that successfully launched interventions managed to overcome existing barriers. This survivorship bias in our sample suggests we may be overstating the impact of barriers that these particular organizations faced (and overcame), and we risk missing “terminal” barriers that stop other projects entirely. To protect against this, we have also engaged experts with a more holistic view of the field who could speak to failed ventures, but there is no way to know if we have fully avoided this bias.
Ultimately, despite these issues, we feel fairly confident that our approach has been successful at identifying the types of barriers people are currently discussing or writing about (90–99% certain).
Mapping Barriers to Implementation Stages and Categorizing by AI-Specificity
We assigned each barrier to one or more of the eight implementation stages based on where it most commonly blocks or slows progress.[6] Some barriers appear in multiple stages when they manifest differently at different points in the deployment journey. Figure 1 is a diagram of the barriers by stage.[7]
Figure 1: Diagram of key barriers by various dimensions[8]
We find that the distribution across stages reveals patterns about where organizations typically encounter obstacles and what types. The concentration of barriers in solution design (stage 2) and scaling/sustainability (stage 8) suggests that many potential failure points concentrate at the end points of the development process: in this case, in building contextually appropriate AI systems, and in sustaining solutions over time.
We reason that unlocking these early-stage challenges is likely to have the highest payoff, both because they will allow more interventions to move through the development funnel but also because they are possibly more likely to be neglected relative to other types of challenges.
Top Barriers and Solutions
We approached the identification of top barriers holistically. We drew from interviews, literature, and a quantitative exercise[9] to score the severity of each barrier. Our method was qualitative in nature, selecting the barriers most likely to be fundamental roadblocks to overcome without support. In importance, tractability, and neglectedness (ITN) terms, our focus was on the I (importance). From this list, we tagged those within the scope of the current exercise. For example, a frequently cited challenge was the lack of funding for AI-specific interventions. Because we don’t think of grants as fieldbuilding, we tagged this as a top barrier falling outside our scope and did not pursue further. The full list of top barriers can be found in this tab (columns I and J). We reproduce a copy of these data with only the top barriers within scope in Table 2, below.
The next section covers deep dives into four barriers. We selected these four in agreement with CG because we had gained confidence there might be tractable solutions to explore. For each barrier, therefore, we describe the issues, the conceptual challenges to overcome, and solutions tied to these challenges. We also focused research time on exploring what possible funding opportunities would be like in this space. Finally, we provide a view on whether each solution seems promising to us, given the information we have collected, our general impression of whether this is something that CG would fund, and remaining uncertainties.[10]
Table 2: Top barriers within scope
| Barrier/Difficulty | Description | Type | Stage |
|---|---|---|---|
| Lack of context-specific training data | Models require data reflecting local languages, disease presentations, crops, etc., not available in standard datasets | AI-specific | 2. Solution Design & Development, 3. Validation & Testing |
| Misalignment between technical AI builders and actual decision makers | Disconnect between what technologists build and what program managers/officials actually need | General | 1. Problem Identification & Scoping, 2. Solution Design & Development, 6. User Adoption & Training |
| Tech and telecommunications infrastructure and equipment | Insufficient internet connectivity or mobile network coverage for AI deployment | Tech | 4. Regulatory & Policy Navigation |
| Passive and active barriers from local governance and regulation | Local policymakers lack familiarity with new tools or procurement processes; slow or poorly adapted permitting processes; lengthy bureaucratic approvals delay or block implementation, despite no clear policy benefit in doing so. Passive blocking by local government. | General | 4. Regulatory & Policy Navigation, 5. Implementation & Deployment |
| User economics | End-users cannot afford to pay for services, limiting financial sustainability | General | 2. Solution Design & Development, 6. User Adoption & Training |
To quantify whether something seemed promising or not, after having conducted and read our joint work, our core team of three researchers deployed a Bayesian Belief Aggregation Model. We used this model to provide a sense of not only the direction (promising or not) but also the strength, uncertainty, and consensus regarding beliefs. To do so, each researcher provided a range of values reflecting their beliefs on how promising a given intervention would be. So, for instance, a range of 0.5 to 0.55 would imply a coin toss slightly in favor of the solution, whereas values of 0.9 to 0.99 would imply a near certainty that the solution was promising. Details on our aggregation methods, simulation procedures, and other notes can be found in the appendix.
We display our results in a series of density plots representing the distribution of our posterior collective beliefs regarding how promising a given solution is. We label solutions whose distribution is fully above the 0.5 threshold as Likely Promising. The height of the curve indicates the strength of consensus: tall peaks suggest strong alignment, while wide curves indicate more disagreement. We would like to emphasize that this is not a definitive take on all solutions, but our subjective beliefs after conducting a limited review of the evidence; we are open to re-evaluating our views on the basis of new evidence. We track and provide takes on 23 possible solutions.
Figure 2: Distribution of aggregate beliefs over solution promise

Before examining the 23 solutions across four barriers in detail, here are a few takeaways from Figure 2:
- Based on our current research, we believe that eight interventions appear promising (where the densities encompassing our posterior beliefs are mostly above the 0.5 threshold). However, we have significant uncertainty over even these; there is only one solution for which we have medium confidence that is promising: fine-tuning frontier models.
- Our top two scoring solutions relate to context-specific training data, suggesting various opportunities to engage with this barrier. While we identify at least one promising solution for each of the remaining barriers, our confidence varies. We are least confident about opportunities to address tech and infrastructure challenges. Further, we have particularly mixed views regarding the misalignment between AI builders and decision makers, suggesting that the potential of these solutions is unknown but potentially promising.
We remain, of course, open to updating our views on any and all of these solutions, and have tried to describe the main reasons for our assessments (and what would change our minds about them) in the main text.
Barrier: Lack of Context-Specific Training Data
Description
Artificial intelligence models require training data that reflects the specific contexts in which they will be deployed. Because AI models are trained on datasets from high-income countries, they frequently underperform when deployed in low- and middle-income settings. Broadly speaking, there are two dimensions to this issue: (1) sourcing better data to improve model knowledge (context data), (2) improving models’ ability to communicate with users (language data). In this section, we describe what these dimensions entail and provide an examination of possible solutions for each.
Before moving on, we provide a brief review of the problem. Context-specific data is important because various types of AI capabilities depend on representative training data to understand what they are being asked about. With regard to Recognition, for instance, if a crop disease algorithm is trained mostly on Western diseases, it will likely fail to accurately diagnose local variants in an LMIC context. Similarly, Edwin Klinkenberg from Delft Imaging noted that diagnostic AI performance can vary across patient populations, including for comorbidities and older infections.[11] Specifically, lower accuracy can happen because these clinical presentations are underrepresented in training datasets.[12]
Representative language data is important for Interaction and Generation. A chatbot trained on English or Mandarin might falter at Swahili or Yoruba interactions, and cultural norms around communication (directness, formality, appropriate topics) require culturally specific training data. Goldstein suggested that Google and Meta do not benchmark against these types of languages as key indicators, meaning voice applications are likely to limit effectiveness in settings where these languages are prevalent.[13]
Why do such gaps in context and language data persist? Our sense is that this pattern stems in large part from the following structural issues, which are addressed in turn below:
- A mismatch between the data that is available to train AI models versus what is needed
- Competitive market pressures
- Regulatory uncertainty
Regarding data mismatching, the data that is collected and available to AI developers is often quite different from the data required to give users the models that they would find useful. In our conversations, organizations emphasized that broad datasets often fail to capture the specific contexts needed to drive impact. For context data, Rikin Gandhi (Digital Green) noted that while agricultural R&D funding heavily prioritizes staple crops like maize, their farmer chat service has received millions of queries focused on high-value crops such as coffee, avocado, mango, and chili. As Gandhi put it, “1% of queries are about maize…99 [percent] on this other stuff,” yet the “AI model stuff so far has all been on maize.” Another example comes from healthcare: Klinkenberg noted critical gaps in pediatric TB X-rays, where model performance degrades due to the underrepresentation of specific patient subgroups.[14] With language data, users will not be able to use models if they cannot talk to them, an issue that arises because there is a lack of training data on under-resourced languages. Jacaranda Health’s experience when expanding to Ghana illustrates this problem. While text data for the Twi language was already scarce, Jay Patel noted that “voice data is even…more scarce.” This distinction is vital because, unlike in Kenya, where Jacaranda operates successfully, Patel explained that “literacy is…a much bigger problem” in Ghana. As a result, the intended users require voice-based interaction to engage with the tool at all. Without the proper language setting, models can sit unused.
Regarding competitive market pressures, organizations may lack incentives to share their valuable language datasets. Expert H describes the current landscape as a “Hunger Games” environment, where organizations leverage their proprietary data for financial gain.[15] This tends to make rare data, whether it is pediatric TB images or dialect-specific voice logs, quite expensive and valuable. Klinkenberg shared that “if [pediatric TB data] is publicly available, immediately lots of companies will respond.” As a result, implementers will tend to keep data to themselves rather than contribute to a digital commons.
Finally, regulatory uncertainty, such as structural and legal barriers, might prevent organizations from sharing their data even if they wanted to. Expert H points out that data ownership is often “messy”, with multiple stakeholders, noting that “governments mistakenly believe they own this data,” which complicates release permissions. Furthermore, many organizations lack the specific budget to release data safely (e.g., anonymization costs). Organizations may also feel like they may receive no credit or attribution for sharing data. While exceptions exist (e.g., Digital Green is already making their data public), it is our impression that regulatory uncertainty, cost, and competitive pressures keep data siloed.
Together, this means that those who stand to benefit the most from AI applications are the hardest to serve.
Potential Fieldbuilding Solutions
Table 3 provides an overview of our solution sets organized by whether they relate to context or language data. But these solutions also map onto the three structural issues that create this barrier in the first place. When the primary challenge is data mismatch, our solutions focus on creating data, like Direct Data Generation for context data and Investing in Raw Data Collection for language data. When the barrier is competitive market pressure, we focus on improving incentives to engage, such as Building Shared Responsibilities or Data for Access Partnerships. Because we address regulatory issues elsewhere, we do not spend time on them here. We expand on all these solutions and more after the table.
Table 3: Evaluation of solutions for lack of context-specific training data
| Approach | Structural Barrier | Promising? | Primary Reason |
|---|---|---|---|
| Context Data Solutions | |||
| Direct data generation | Data mismatch | Perhaps | Bypasses competitive market pressures; various examples of it working, though some uncertainty |
| Building shared data repositories | Competitive market pressure | No | Crowded field; unsure how willing organizations will be to share/sell data |
| Data-for-access partnerships | Competitive market pressure | No | Unclear what role CG could play; possibility it creates new data silos |
| Language Data | |||
| Fine-tuning frontier models | Data mismatch | Yes | For moderately resourced languages: support NGOs to fine-tune existing frontier models using language incubators/ embedded experts tFor under-resourced languages: support investment in raw data collection with a focus on developing gold standards for the ethics of this process |
| Purpose-building small models | Data mismatch | No | Will be less powerful and more expensive per language than working with frontier models |
| Translation layers | Data mismatch | No | Will always be inferior in computation needs and cost compared to working directly with frontier models |
Context Data Solutions
Direct Data Generation
We view direct data generation as the most promising solution in the context data category because it bypasses competitive incentives entirely. Moreover, there appear to be some reasonable approaches that have been tested and appear scalable, but we have not conducted sufficient research to assess costs.
The premise of this solution is that most existing data remains in analog formats (paper records, unrecorded conversations, undocumented local knowledge), or exists but hasn’t been digitized for machine learning purposes.[16] CG could fund local teams to systematically create new datasets through several approaches, in order of our rough sense of tractability below:
- Crowdsourced labeling. One option is to set up a crowdsourcing platform where local experts (agronomists, radiologists, teachers) review and label data, paid per task. Gandhi[17] and Expert H[18] have both used versions of this model. The approach has precedent: ImageNet’s (Deng et al., 2009) early strategy relied on Amazon Mechanical Turk to label millions of images, and Sama uses SMS-based crowdsourcing in LMICs.
- Screening campaigns generating data through service delivery. Another option is to fund health screenings that reduce cost barriers for participants while simultaneously building training datasets. Participants would consent to de-identified data use at the time of screening. What is appealing about this approach is the ethical and logistical advantage of providing immediate health benefits, with the ability to target specific data gaps through focused campaigns (e.g., pediatric screenings only). The downsides are operational: this process requires equipment protection, staff training, and quality assurance. There appears to be precedent for these types of programs, though after a few minutes, we could not find credible sources with sufficient detail to verify this.[19]
- Digitizing existing records. Another option is to digitize existing medical or government records. The MIMIC database, which digitized over 60,000 ICU admissions, provides a precedent (Johnson et al., 2016), though we think doing something like this in an LMIC setting would be considerably more expensive than this exercise was. A key challenge is that data ownership is often unclear, leading to complex negotiations over permitted uses and revenue sharing if commercial applications result.
- Recording voice data. A final option is to recruit native speakers to record scripted conversations covering medical, agricultural, or educational content in target dialects, with linguists hired to transcribe and validate. The challenge here is scale. Expert G told us that community-led efforts to collect voice data for very low-resource languages can run for years and still yield only a few hours of usable audio, because the data must be gathered from scratch.
We found evidence of various organizations pursuing direct data generation and believe they could support in doing so on contract, though we have not contacted them:
- Masakhane – A grassroots volunteer community creating African natural language processing (NLP) datasets through distributed data collection. They have created MasakhaNER (21 languages) and various translation datasets.
- Karya – Pays workers in India to create training datasets, including speech, image labeling, and translation.
- Translators without Borders/CLEAR Global – Generates language datasets through their translation work and releases some openly, with a focus on crisis-affected languages.
Building Shared Data Repositories
We view this solution as having relatively low potential, mostly because there already seems to be significant activity in this space, and because we believe that competitive market pressures will be hard to overcome. Both of these things seem challenging.
The premise of this solution is to get valuable private datasets into public or shared repositories where they can be used to train and fine-tune models. We first explored whether voluntary sharing could work at scale. Some ad-hoc sharing is already happening: Digital Green publishes agricultural assets on Hugging Face, Masakhane releases African-language datasets through volunteer efforts, and Zindi runs challenges encouraging data contribution.[20] However, we are skeptical that voluntary initiatives will scale.
We therefore explored funding the purchase of datasets (i.e., paying organizations to release datasets they would not otherwise share). In our research, we found multiple instances of valuable data held by implementing organizations. In healthcare, Klinkenberg noted that Delft Imaging holds TB chest X-rays from CAD4TB deployments and ultrasound data from BabyChecker. He estimated that for pediatric TB, a release of ~10,000 X-rays would trigger an immediate market response: “if this is publicly available, immediately lots of companies will respond.”[21]
Whether or not this solution is viable hinges at least in part on whether or not resources are the main reason why organizations do not release data. Many organizations often want to share data for the sake of fieldbuilding but lack the budget for proper anonymization and legal compliance. For example, Shabnam Aggarwal seemed at least interested in this solution. In these instances, purchasing data could resolve the bottleneck by covering these costs.[22] We remain relatively skeptical that organizations collecting data at a high cost will want to share it at all, but it may be possible in some areas. In our research, we identified various organizations already pursuing this approach, suggesting a crowded field and, consequently, limited opportunities for CG to fund this by themselves:
- The Lacuna Fund has awarded ~$30M to create and release labeled training datasets for LMIC applications in health, agriculture, and education.
- AI4D (Artificial Intelligence for Development Africa) coordinates the creation and release of African-language datasets.
- Mozilla Common Voice purchases or commissions voice recordings in 100+ languages, serving as a potential model for systematic acquisition.
Data-for-Access Partnership
Our view is that it would be in general difficult for medium to large sized philanthropic funders to encourage these types of arrangements, both because we don’t know how one would incentivize large tech organizations to do so, but also it might unintentionally create new data silos. It also does not solve for competitive pressures.
The general idea of this solution is an exchange whereby data contribution gives providers direct access to in-kind benefits such as tools, discounts, or pilot access. In healthcare, for example, hospitals or clinics would provide data (mammography images, TB X-rays) and receive AI tools in exchange that they couldn’t otherwise afford. What is appealing about this solution is that it solves multiple problems simultaneously: hospitals gain tools while developers gain diverse validation data.[23] However, the feasibility of this approach depends heavily on who is offering the tools, a point which makes us discount this solution.
For those organizations building models, one can envision multiple ways of structuring such arrangements, with different levels of partnership and integration. At one end, a single stream of data might be uploaded in exchange for access to one AI model; at the other, a whole suite of data could be shared in return for a host of services. A key crux is finding an AI tool or intervention of high enough value (either because it is expensive or because it is known to be highly effective) to motivate data sharing. There is also the question of how these data would ultimately be used. Several concerns arise:
- Power imbalance concerns: Okafor and Nyarko (2025) warn that data “may be extracted from vulnerable communities under the pretense of healthcare improvement but later used for commercial purposes”
- Long-term sustainability is unclear because it may be the case that the data required is collected relatively quickly, while the organization providing it may require the AI tools it is receiving for a much longer period of time;
Organizations currently pursuing this approach include:
- PATH: Maëlle-Marie Corso mentioned that PATH works on identifying how AI can increase access and quality of healthcare in LMICs, for example by conducting one of the first clinical trials of LLM-based clinical decision support in Sub-Saharan Africa.[24]
- D-tree International – Develops digital health tools in Tanzania and other LMICs, with data partnerships trading tool access for data contribution.
- Microsoft AI for Good Lab – Provided compute credits to Lelapa AI for InkubaLM training in exchange for access to African language data and validation.
Language Data Solutions
Language barriers are a key data and functionality gap that came up repeatedly in interviews because they can exclude intended beneficiaries from direct engagement across a wide range of programs. For example, while the most cutting-edge cloud-based closed-source models have the potential to perform quite well on some African languages, current performance varies dramatically by language. An internal OpenAI assessment found that Swahili reached 60–70% of English performance (OpenAI, 2024, Tables 8-9). Hausa jumped from 6.1% (GPT-3.5) to 71.4% (GPT-4o) on science questions, and Yoruba went from 28.3% to 51.1% on questions testing factual accuracy. Lower-resource languages like Twi and Wolof achieve only 30–45% of English performance, and an external review found that translation tasks to Tigrinya produced “gibberish outputs” in an early version of ChatGPT (Tsanni, 2023).
A further benchmarking attempt, Irokobench, uses model performance on three tasks to assess AI model performance in African languages:
- AfriXNLI tests whether models understand logical relationships between sentences (e.g., “The woman is pregnant” entails “The woman will have a baby”)
- AfriMGSM requires solving grade-school math word problems involving multi-step arithmetic
- AfriMMLU asks university-level multiple choice questions across subjects from medicine to history (Adelani et al., 2024)
The results indicated that African languages face a significant penalty in performance, with GPT-4o achieving 48% accuracy across African language tasks (zero-shot solving in-language) versus 72% in English.[25]
We considered four options to bring improved AI functionality to underserved language communities:
- Improving general language competencies of multilingual frontier models
- Using dedicated translation layers on frontier model outputs
- Purpose-building small models for underrepresented languages
- Fine-tuning the language capabilities of frontier models on specific subject matter
The following sections discuss the merits and drawbacks of each of these in turn, ultimately identifying fine-tuning the language capabilities of frontier models as the most promising idea. Two sets of solutions are then explored, dependent on the level of linguistic accuracy for a given language in current frontier models:
- For moderately resourced languages: support NGOs to fine-tune existing frontier models through embedded experts, language-specific incubator programs, funding the sharing of learnings/weighting between organizations, and/or developing a repository of best-practice guidance.
- For under-resourced languages: support investment in raw data collection, focusing on languages with a high ratio of native speakers to digital representation. Ensure that fundamental ethical questions about data protection, extraction, and power imbalances are integral to any initiative.
Improving General Language Competencies of Multilingual Frontier Models
Our impression is that, without dedicated philanthropic investment, some higher-resource languages like Swahili and Hausa will continue to improve, but languages with minimal web presence will remain poorly supported because no economic incentive drives training data creation for markets generating negligible revenue. Aside from encouraging anchor organizations to gather, digitize, and publish training data, we have not found any ideas that might change the trajectory of low-resource languages in frontier large language models (LLMs). We therefore spent no further time exploring this area.
Using Dedicated Translation Layers on Frontier Model Outputs
In addition to raw model performance in-language, we are also interested in the potential for a machine translation layer that could mediate between low-resource languages and high performance in a higher-resource language. Meta’s NLLB-200 is one open-source model that claims to improve on the Bilingual Evaluation Understudy (BLEU) scores[26] of previous state-of-the-art models “by an average of 44 percent across all 10k directions of the FLORES-101[27] benchmark,” including on 55 African languages (Akula et al., 2022). However, the IrokoBench scores for translated task completion are roughly comparable to in-language tasks (Adelani et al., 2024, Table 2), and we anticipate that both approaches will carry a considerable quality penalty as long as the data gap persists.
We don’t recommend funders consider investing in translation models, because we expect that translation gaps and cultural barriers would make AI translation a permanently inferior solution compared to in-language processing. In addition, such a layer would introduce additional compute and transmission costs, reducing the potential for cost-effectiveness and speed.
Organizations currently pursuing this approach:
- Translators without Borders/CLEAR Global – Working on translation tools for crisis-affected and underserved languages.
Purpose-Building Small Models for Underrepresented Languages
Purpose-built models take a fundamentally different approach to the above. While exact training costs are not publicly disclosed, the investment required for pre-training such models typically ranges from hundreds of thousands to millions of dollars—much less than frontier models, but orders of magnitude more than fine-tuning approaches.
Several organizations are pursuing this as a route to impact. Masakhane, Lelapa AI, and Lesan AI are all developing AI tools applicable across more languages in Africa and the Global South. Lelapa AI’s InkubaLM, described as Africa’s first multilingual large language model, was trained from the ground up on 1.9 billion tokens of data collected specifically for five African languages: Swahili, Yoruba, isiXhosa, Hausa, and isiZulu (Tonja et al., 2024). The training corpus combined open-source datasets from repositories, including Hugging Face, GitHub, and Zenodo, supplemented with work alongside linguists and local communities to collect, annotate, and validate data offline as well as online (Tsanni, 2023). Lelapa AI credits Microsoft’s AI for Good Lab with providing the substantial compute credits required for training the model (Mzekandaba, 2024).
The conventional wisdom, articulated by Expert J, holds that NGOs should not attempt to solve fundamental language problems. If a model does not adequately cover a particular language, organizations should wait for governments or large technology providers to address the gap, or alternatively pivot to problems where existing models already function adequately.[28]
We don’t recommend funders prioritize this route. Any small language models that are developed will likely be far less powerful and more expensive per language than either general development or fine-tuning of frontier models.
Fine-Tuning the Language Capabilities of Frontier Models on Specific Subject Matter
Fine-tuning existing multilingual foundation models requires substantially less data and resources than building language understanding from scratch, and there seems to be significant promise in fine-tuning general models for both specific use cases and for local languages and dialects. According to Patel, even languages that share a common base, such as Swahili, spoken in Kenya versus Tanzania, often require separate fine-tuning to incorporate local knowledge.
For languages where multilingual base models demonstrate partial comprehension, even when generative outputs remain unreliable, fine-tuning of multilingual models can support narrow but operationally useful tasks such as intent classification, urgency triage, and template routing. This approach does not create general language capability, but it is often sufficient for automated systems that rely on prewritten responses and fixed templates rather than free-text generation (often termed PROMPTS-style systems).
Jacaranda Health found this route very promising; after an initial $50,000 learning investment, they have been able to create fine-tuned models with the ability to answer many common questions for about $3,000 per additional language.[29] Reflecting on how the landscape has shifted since, Patel cautioned that he would not take this build-and-fine-tune route if starting today.[30] The cost differential between Jacaranda’s initial model development versus subsequent languages suggests that codifying and transferring implementation knowledge could reduce barriers dramatically for follower organizations. A “language model cookbook” documenting the process, from base multilingual model selection through domain-specific fine-tuning, evaluation methods, and cost-effective inference deployment, would capture learnings that took years and significant resources to develop. Similarly, Expert J recommended embedding AI teams within organizations as a fieldbuilding intervention to allow more implementers to successfully fine-tune and deploy models.[31]
It is worth noting here that fine-tuned models may perform poorly on benchmarking exercises such as Irokobench despite being at least adequate for narrowly defined practical applications. For example, Jacaranda has had considerable operational success with Meta/Llama 3, which performed at only 34.8% in English on Irokobench (Adelani et al., 2024; conversation with Patel). Even at relatively low performance on sophisticated metrics, such a system reduces manual sorting workload while maintaining safety. Similar tolerance likely applies to agricultural pest identification, supporting extension workers or educational content recommendation, where errors cause minimal harm.
Exploring Solutions Within the Fine-Tuning Option
Ultimately, the most promising interventions we identified focus on fine-tuning the language capabilities of frontier models. The specifics of the solutions depend on the current level of support for a given language within the frontier models, as follows:
- For moderately resourced languages, fine-tune existing frontier models: Interview data suggests that existing multilingual frontier models (e.g., GPT-4, Claude, Gemini) work adequately for moderately resourced languages such as Swahili or Filipino, and will likely progress rapidly in accuracy in the near future, without the need for philanthropic support. Moreover, the baseline functionality of frontier models in these languages means that fine-tuning using domain-specific vocabularies (i.e., agricultural terminology for a weather forecasting app) is feasible.[32] The example from Jacaranda both supports these assertions and further indicates that fine-tuning models on domain-specific language terms is not cost-prohibitive for NGOs serving users at scale.[33] However, fine-tuning costs could form a significant barrier to entry for smaller start-ups. Moreover, Tech for Dev indicated that LMIC-based NGOs may struggle to find machine learning local talent to do the fine-tuning, stating that “organizations might not be able to hire a resource even though they have the money.”
One solution pathway is therefore to upskill priority organizations through focused incubator programs and/or embedded experts, potentially dropping the costs closer to $5–10K per org.[34] Alternatively, one could fund organizations to fine-tune the models themselves, or purchase the tailored model weights (if not open source) from organizations that have already done the work. The latter solution assumes that more experienced organizations are willing to participate and share their outputs. However, commercial considerations or competitive dynamics might limit participation, because organizations viewing language models as proprietary competitive advantages would require either higher compensation or structured arrangements that preserve their market position while enabling broader field access to underlying technology. Finally, developing a repository on how to fine-tune the frontier models in particular languages seems like a positive first-principles idea. With more time, we would have explored possible examples or corollaries of this idea in the literature.
Overall, funding-focused language incubators/embedded experts seem the most tractable option for moderately resourced languages such as Swahili or Filipino.
- For minority/under-resourced languages, invest in raw data collection: For minority languages/those with minimal digital presence, we recommend investing in raw data collection, and with the aim of creating shareable datasets or open models. Ideally, this would focus on those languages with a high ratio of primary language speakers to the existing level of digital representation. A number of organisations are currently pursuing this approach:
- Lacuna Fund – Primary funder of LMIC training dataset creation. Has awarded grants to 80+ projects across 29 countries.
- AI4D Africa – Coordinates African language dataset efforts. Provides small grants for data collection and model development.
- Mozilla Common Voice – Crowdsourced voice data collection in 100+ languages. Community volunteers contribute recordings.
- Masakhane – Community-driven approach to African language data collection through distributed volunteer networks.
However, language data collection also stalls in ways that illustrate the tension between typical funder expectations and community requirements. For very low-resource languages, the data often does not exist anywhere (not even within the largest AI labs) so it must be collected from scratch, which is slow and expensive. Progress is frequently further constrained by communities’ need for full control and autonomy over the data they collect, including a reluctance to open-source it.[35][36]
This creates a direct conflict between community priorities and standard fieldbuilding and other philanthropic funding models. Echoing the community’s concerns, the Sharma et al. (2020) report on AI startups in LMICs identifies “many fundamental questions about data protection, ingrained bias as a result of poor data collection methods, social inclusion and the responsible use of AI” (p. 2) and recommends that “AI solutions should be guided by sound privacy and ethical principles” (p. 2). As such, any raw data collection solution must address the ethical concerns Okafor and Nyarko (2025) raise around data extraction and power imbalances, which we have heard echoed in several expert conversations so far, including with Expert G and Expert H.
Overall, investing in neglected languages with large numbers of native speakers could represent a sound investment opportunity, particularly if the process of doing so created best-practice standards and cross-cutting learnings for other similar initiatives.
Barrier: Tech and Telecommunications Infrastructure, Equipment, and Costs[37]
Description
Lack of access to high-speed internet is one of the most obvious challenges to the deployment of AI for Good in LMICs. As the United Nations Development Programme (UNDP) notes, hard infrastructure is the “tangible backbone needed” to leverage AI in low-income settings (UNDP, 2025, p. 90). However, one report suggested that across a sample of nine LMICs,[38] only 10% of the total population, and 5% in rural areas, had access to a “meaningful internet connection” (A4AI, 2022, p. 10).[39] In our research, we identified two dimensions of this problem:
- Connectivity (coverage and quality): Millions of people in developing countries simply lack a clear option to connect to the internet, whether by computers or smartphones. The African continent, for example, “accounts for nearly half of the 350 million people around the world who do not live in areas covered by mobile broadband networks” (GSMA, 2024, p. 10). As a result, only 34% (World Bank Group, 2024) of people in sub-Saharan Africa report using the internet in the last three months. There is also a significant urban-rural gap. In Nigeria, for instance, less than a quarter (Umeh, 2025) of rural communities have any type of internet access.
- Affordability: Assuming connections exist, cost is a non-trivial issue. A 2023 ITU (2023) report notes that in the poorest low-income countries, fixed broadband costs rise to 31% of monthly GNI per capita, compared to just 1% in high-income nations (p. 7). If internet access is prohibitively expensive, AI for Good tools must either subsidize connectivity or find ways to function without it. Beyond connectivity itself, legacy telecommunications costs can dwarf the cost of the AI.[40]
These issues pose a threat at both ends of the development pipeline: developers may find it hard to upload training data, and end-users might not reliably access cloud-based products. Regarding connectivity, historically, AI models consume large amounts of computing power at data centers typically located in developed economies. Roughly 40% (Statista, 2026) of data centers are located in the USA. Regardless of where data centers for AI4G organizations are located, providing these centers with training data from LMICs requires internet transmission.[41] For example, an AI-supported program that uses images to aid in screening or diagnosis requires broadband access to transmit heavy training data. In short, developers building products in remote areas may simply be unable to train the relevant models.
With regard to affordability, interacting with LLMs or other types of AI models as users typically requires internet access. As Okafor & Nyarko (2025) put it in their working paper, “reliable internet connectivity is a foundational requirement for cloud-based AI applications” (p. 3). Users of ChatGPT, Gemini, and Claude essentially stream access via chatbots. Innovation in how we interact with these LLMs (via video or audio) requires even stronger connections. Speaking at the 2024 AI for Good Summit, Aminu Maida of the Nigerian Communications Commission (NCC) noted that it would be challenging for users on 2G or 3G connections to effectively use these tools. For example, long lags and sudden disconnects might turn people off from using LLMs.
Our current view is that the challenge of high-speed internet infrastructure is general across sectors/verticals; however, the severity of the bottleneck probably hinges on the specific type of AI application. It seems to us that applications relying on recognition (“the ability to improve identification of patterns or trends in data”) are the most likely to depend on high-speed internet to operate well, as they need to transmit heavy images and audio, rather than simple text, to be useful. Similarly, applications focused on generation, specifically those producing video or audio content, are likely to face significant latency hurdles compared to text-only tools. Conversely, the applications least likely to be affected are those deploying interaction via text chatbots; a close second might be those using prediction on new data (e.g., weather forecasting), provided the models are run locally, and training has occurred elsewhere.
Potential Fieldbuilding Solutions
We explored five solutions, finding that supporting offline AI efforts is likely to be promising. A summary of our takes is available in the following table.
Table 4: Evaluation of promising solutions for internet connectivity
| Approach | Promising | Primary Reason |
|---|---|---|
| Offline Hardware | Yes; we highlight some possible pathways | AI in a Box” might allow cheaper scaling; a gap exists for mid-range hardware. |
| Satellite Internet | Maybe | Per-user costs are prohibitive for scale. |
| Community Networks | No | Relies on existing infrastructure (does not fix the root cause). |
| Regulation | No | Hard for philanthropy to influence market prices. |
| Offline AI Models | NA | Tech is “good enough” vs. counterfactual; big players are already fixing accuracy. |
Our current sense is that there are two approaches to helping solve this issue. We discuss them below in order of our sense of tractability.
Offline Hardware That Eliminates the Need for Internet Access
The ability to bypass traditional internet infrastructure is potentially one of the highest-leverage opportunities in this landscape. Our thinking is that improving access to models capable of processing data locally, without relying on good or consistent connectivity, could enable more AI4G interventions. In theory, the technology exists:
- Edge AI: Defined by Singh & Gill (2023) as “the practice of doing AI computations near the users at the network’s edge, instead of a centralized location like a cloud service” (p. 71).
- tinyML: A subset of Edge AI focused on running machine learning models on low-power hardware. As noted by Professor Marcelo Rova, speaking at an AI for Good event, this involves intelligence running “inside the physical world in very, very small devices . . . working all the time.” (AI for Good, 2021, 13:59)
- Offline-First Tech: Systems designed to prioritize local processing to mitigate issues like “latency, power consumption, and also security.” (AI for Good, 2021, 21:17)
We spent roughly four hours investigating the key cruxes to making offline AI widely available, and identifying investment opportunities if relevant.[42] A summary of our takes is available in the table below, followed by more detailed answers.
Table 5: Issues in deploying offline AI
| Issue | Answer | Confidence | Investment Recommendations |
|---|---|---|---|
| Is the tech good enough for real-time AI interventions (e.g., triage)? | Yes, we think it is good enough. While offline models may lag behind state-of-the-art cloud models, they are at least sufficiently capable to improve upon the counterfactual (i.e., no doctor/diagnosis at all). We also expect rapid improvement via external players (Google DeepMind, Fastagger), meaning the models themselves are not the primary bottleneck. | Medium to High (Current capability) Very High (future improvement) | No action. The model landscape is crowded with well-funded entities (Google, etc.) rapidly solving the “accuracy” problem. CG should assume the software will exist and instead focus on the delivery mechanisms (see below). |
| Is hardware a bottleneck? What would we need to buy to deploy at scale? | Hardware can be a bottleneck; we are betting that “AI in a box” can be transformative. Providing high-end user devices (laptops/phones) at scale is too expensive ($500-$2,000/user) and risky (theft/maintenance). The catalytic intervention is increasing access to local edge servers (“AI in a box”) for AI4G implementers. This moves the compute cost to one shared node (hosting the AI/LLM/database via local Wi-Fi), allowing users to connect via cheap, low-end smartphones. | Medium | Promising investment opportunity Investigate “Mid-Range” Edge Server Hardware. There is a massive market gap between the low-end tools (Crisis Cognition’s “0-LA” @ ~$500) and corporate tools (AWS Snowball @ ~$55K). We found three possible paths for investment, which we explore below. |
Is the tech good enough for AI4G-style interventions? For example, real-time AI-supported triage? If not, will they improve soon? Who is working on this?
We believe that offline AI is sufficiently “good enough” to support real-time medical triage. The people we spoke to expressed confidence in offline tools for medical diagnostics (Klinkenberg[43]), agricultural diagnostics (Expert H[44]), and information retrieval (Bopp[45]), noting their ability to improve over humans in remote areas. We slightly downweight these takes by the fact that these individuals are themselves developers of these tools.
Moreover, for particular applications, the technology may not be there yet. Expert N argued that offline models will “always be worse and trailing”[46] the best cloud models, and Shabnam Aggarwal suggested accuracy rates (in the context of education applications) fall well below 40–50% for Global South use cases, a “pretty poor” performance that our research confirms is common wisdom.[47] (In this case, we would think the crux is whether or not the models are trained appropriately, not whether they are capable in the abstract.)
Despite this poor performance, our take is that while high accuracy is important, models must be judged on whether they improve over the counterfactual (e.g., a scenario where a person might receive no diagnosis at all[48]), and, therefore, we believe these models are likely good enough for triage in low-resource settings (medium to high confidence). Furthermore, we are very confident that these models will improve in short order, as evidenced by organizations like Fastagger helping to shrink models for edge devices in LMICs, and successful collaborations like Crane AI Labs and Google DeepMind building powerful, fully offline models. Ultimately, we do not think the models themselves are a clear bottleneck and, in any case, we do not see a straightforward investment case for CG.
Is hardware a bottleneck to running the best offline models? What would we need to buy to deploy these types of products at scale?
There are two categories of infrastructure to address: 1) “AI in a box” and edge servers, and 2) end-user devices (smartphones, tablets, or laptops capable of running diagnostic AI models offline).
Our medium-confidence take is that the former is the more catalytic. Producing an accessible AI hardware/software bundle at scale without internet access could provide communities or hospitals with high-accuracy diagnostic tools. However, before exploring that solution, a brief review of end-user devices is in order. Our research suggests that existing consumer hardware is already capable of supporting good open-source models locally (e.g., Expert H notes that standard MacBooks can effectively run large-parameter models[49]). The cost of providing them at scale would be prohibitive, with capable devices ranging from $500 to $2,000 for a single user.[50] Moreover, the effective cost increases when considering potential theft[51] and maintenance (including updating devices to new models). The solutions to this cost challenge do not seem obvious or particularly tractable to us.[52]
Returning to the first category, the appeal of “AI in a box” solutions is that they create a local network, allowing connected devices to access powerful LLMs without needing the internet. Under such infrastructure, the constraint is no longer the user device, which could be an inexpensive smartphone or tablet, but the server itself and the Wi-Fi connection. This server would be capable of handling a powerful model and multiple simultaneous queries, while the user device would only need to connect to the local server.
We have evidence of these types of devices operating in practice. In our work, we found various instances of interventions that rely on this type of infrastructure. For instance, Crisis Cognition’s “0-LA” device creates a local “mini internet” powered by solar batteries to host an LLM and database for use in refugee camps, completely independent of global internet access.[53] Shabnam Aggarwal (Rising Academies) suggests a similar but potentially more cost-effective approach for schools, using a Raspberry Pi that devices can ping when connectivity is unavailable (though this would depend on AI capability).[54] Expert N expressed significant skepticism regarding the potential of offline models to work (or outperform proprietary models),[55] and Aggarwal noted that there were limited examples in practice.[56] Still, we think this is technically possible for some subset of AI4G applications.
Our sense is that while these devices exist, access to them is limited, and there is an opportunity to build and market AI in a Box for AI4G interventions. The question is whether we can increase access to these products at a reasonable cost. Evert Bopp shared that the 0-LA device will market for approximately $500, which he highlights is significantly cheaper than corporate alternatives like Amazon’s AWS Snowball Edge, which carries a license fee of $55,000 a month.[57] However, the 0-LA product is a chatbot meant to provide fairly straightforward answers and, in our understanding, doesn’t process inputs like images. We suspect, therefore, that building a product for more complex applications would be much more expensive than $500.
Still, we think that building these at scale is feasible and worth pursuing at some level. We spent roughly one hour researching who could build this technology at scale. We imagine a scenario where CG is funding or co-funding a series of “AI in a box” systems with different capabilities, but we want to make these systems sufficiently developed to make use of the full range of capabilities. These would be sold at market price (or subsidized) so that organizations interested in deploying their solutions in remote areas are able to do so without relying on constant internet connectivity. Three possible pathways:
- First, identify a candidate to build. One such candidate might be Crisis Cognition. Given their experience with 0-LA, it is likely they have the know-how and connections to build different types of devices for various ranges of capabilities. As mentioned before, 0-LA does not currently deploy particularly demanding AI capabilities, so the technical challenge here is not to be discounted. We are unsure about their interest.
- Second, we attempted to identify other organizations doing similar work. Our most promising hit was World Possible’s RACHEL, an offline server deployed in schools, correctional facilities, and health centers in up to 50 countries. RACHEL does not currently deploy “AI in a box”; it simply hosts interactive, static material in local servers. As such, implementing AI would require upgrading their hardware, but we think this is a natural next step that would build on their existing infrastructure and know-how. In terms of cost, their current prices seem in line with devices like 0-LA (roughly $5,000 for one server and 10 laptops, or ~$550 per connection). While the necessary hardware upgrades would likely increase this cost, we also believe it would likely represent significant savings over our estimates for other interventions (e.g., providing satellite internet) We have not contacted them, though a partnership seems possible.
- A final pathway could entail funding an entirely new startup to provide these capabilities at scale. This would require researching appropriate technical requirements and contacting industrial providers, including companies already building rugged AI servers like Eurotech (ReliaCOR series) and OnLogic (Karbon Series). We think this is likely a good idea; it could require only a small amount of seed funding to pilot a self-sustaining enterprise, and could be combined with other investment bets that CG makes to get AI solutions to the most remote places.
Improving Internet or Making it Cheaper
We also looked at various solutions aimed at improving the internet infrastructure itself or making it cheaper. Our tentative conclusion, after a few hours of research, is that broad investment in physical internet infrastructure, while an important direct way to address these challenges, is likely outside of scope. Specifically, this area does not appear to meet basic criteria for neglectedness or tractability, though we acknowledge that one-off, geographically contained interventions could be exceptions (but this, of course, runs counter to the notion of fieldbuilding outside of our specific context). Our brief investigation identified three primary avenues: satellite technology, community networks, and regulatory reform.
First, we looked at satellite-to-cellphone technology as an alternative to traditional fiber optic or cell tower construction. While potentially reducing infrastructure capital expenditure, satellite internet seems too costly on a per-user basis. The Center for Global Development (CGD) notes, for example, that user costs can reach ~$600 upfront with ongoing monthly fees of ~$100 (Croshier, 2022, p. 5). This would mean that this technology is likely to exclude large swaths of populations we would otherwise be interested in reaching if the user base is high, though it might be acceptable if the intervention is focused on a limited number of users. For example, if we were to equip all community health workers in Nigeria with access to smartphones, the yearly cost might be something around $130M (which seems high to us).[58]
Second, we examined community networks or cooperative-owned ISPs as a mechanism for reducing costs and solving “last mile” delivery challenges. The Zenzeleni network in South Africa shows that this concept can work. However, these community networks rely on, rather than replace, the foundational infrastructure required to bring internet connections to remote communities. (So funding would not solve the first-order barrier).
Finally, we considered the impact of regulatory and market shifts. A World Bank blog attributed substantial market price drops in India and Cambodia to intensifying competition (Foster et al., 2021). While plausible, we find that these market dynamics might be intractable for CG because it is not clear how effective philanthropic dollars can be in influencing such a large, profitable market.
Barrier: Misalignment Between AI Builders and Decision Makers
Description
Applications of artificial intelligence may be impeded by the often-wide gap between those developing AI systems and decision makers involved in approving and implementing them.[59] This gap is common to many development interventions, where the developers of interventions are often removed from the contexts in which they intend to work. But the recent development of AI tools means that many decision makers also lack an understanding or framework by which to evaluate their functionality, as well as limitations, widening this gap.
AI builders here include both private companies and smaller NGOs that develop an AI system, or a team within a large NGO or implementer that builds or adapts an AI tool. Decision makers may be government policymakers or NGO leaders who approve the integration of these tools into their work. Our overall impression from the examples below is that the most important misalignments are between AI builders and government policymakers. It is also the case that to achieve scale, many interventions require the involvement of government policymakers, making them a particularly important decision-maker to consider.
This misalignment has (at least) three potential dimensions:
- AI builders lack an understanding of the context of the intervention, resulting in effort being spent on the wrong problems or tools being developed that are not fit for purpose. A RAND report on AI applications in business (Ryseff et al., 2024) found this to be the most commonly cited primary cause of failure for AI projects. They stress that misunderstandings between AI builders and decision makers can exist even within (large) organizations, e.g., data science teams may struggle to understand what outputs would be most helpful for the organization. Expert J also highlighted that a lack of focus on the user frequently stymied tech applications in general, arguing that developers often skip the question of “who are you doing this for and why?’.[60]
- Decision makers lack an adequate framework with which to judge the potential risks and requirements of proposed AI applications, resulting in missed opportunities for beneficial applications, while potentially exposing themselves to risks of inappropriate applications. Goldstein (CGD) told us that government officials had approached them for advice on how to decide whether AI systems that were being proposed to them should be piloted; the government had requested CGD’s help to ascertain which proposals for AI applications were most promising.[61]
- AI developers and program implementers may need established relationships with decision makers in order to carry out pilots as well as programs at scale. Okafor and Nyarko (2025, p. 9) highlight that past experiences with failed technology implementations have sometimes resulted in healthcare workers and policymakers becoming disillusioned with digital solutions. To establish the trust required for successful implementation, they recommend that stakeholders be engaged in every stage of AI project planning. David Yanagizawa-Drott told us that the capability to access data and scale their work was to build relationships with senior stakeholders in government.[62] Similarly, Shabnam Aggarwal highlighted that Rising Academies’ work with government schools in Ghana depended on the involvement of an academic who had relationships with a subset of schools.[63] Rikin Gandhi (Digital Green) also stressed that the requirement to build relationships with distribution partners, including government channels, before being able to distribute an AI application was a key potential barrier.[64]
Potential Fieldbuilding Solutions
We identified three approaches to addressing this misalignment: providing advisory support for builders; providing advisory support for decision makers; and bridging the gap between the two parties with convenings and spaces for connection. Based on our landscaping analysis, we xfconcluded that the third of these approaches was not neglected, and therefore did not explore it further. We conducted deep dives on the first two approaches. A summary of our overall takes is shown in Table 6.
Advisory support for builders addresses the first dimension of this issue (AI builders lacking an understanding of the context of their AI solution), while advisory support for decision makers addresses the second dimension (decision makers lacking an adequate framework to judge interventions).[65]
Table 6: Evaluation of solutions for misalignment
| Approach | Solving Which Dimension? | Explored In-Depth? | Promising? | Primary Reason |
|---|---|---|---|---|
| Advisory support for builders | AI builders lack an understanding of context | Yes | Possibly | We think that advisory support for organizations building AI tools can help to tailor their tools to the context they will be working in and reduce mistakes at an early stage. |
| Advisory support for decision makers | Decision makers lack an adequate framework to judge interventions | Yes | Likely no, at least in the short term | There is currently very little evidence regarding how to integrate AI solutions into the work of governments or large NGOs, on which technical advice can be based. We have more confidence that this could be an important solution in the medium term (~5 years), once more evidence has been gathered on applying AI to government work. |
| Bridging gap between parties with convenings and spaces for connection | AI developers and program builders need established relationships with decision makers | No | No | This solution seems non-neglected according to our landscaping analysis. |
Advisory support for builders can help organizations to understand the context in which their tools will be applied (especially through a focus on evaluation and user experience), addressing the first dimension of the issue. We have explored this approach below. Our takeaway is that this is worth exploring further. In particular, incubator/accelerator models offer an intensive version of this support with additional features, and could be a strong model for funding promising individual implementers. Funding short-term tech consultancy services could present a less expensive solution, but we are highly uncertain about impact.
A second set of interventions provides advisory support for decision makers, building capacity to help them evaluate and implement proposed AI solutions. This could theoretically be a highly leveraged solution, as decisions about what AI systems to take up—especially decisions by senior government officials—could have very wide-ranging impacts. We explore this class of solution below. Our takeaway is that there is currently a lack of evidence on integrating AI solutions to inform decision making; we are therefore moderately confident that this is not a promising solution in the short term, although it may be promising in the medium term (~5 years).
A final approach is interventions that bridge the gap between parties with convenings and spaces for connection. David Yanagizawa-Drott, for example, told us that J-PAL was attending (and holding a side-event) at the AI for Good Summit (organized by UN-ITU), with the primary goal of connecting with policymakers who may be willing to collaborate on the deployment of an AI system.[66] However, our impression from our landscaping of current funder activities is that this solution is not neglected, with several such initiatives already in existence, and limited value to an additional initiative. We therefore deprioritized this solution and do not discuss it further below.
Advisory Support for Builders
Many of the organizations we spoke to that had successfully rolled out an AI intervention had also received some kind of advisory support.[67] Expert J suggested that a key focus of this advisory support is on helping participants to understand precisely what their tech solution is for, and how it should be designed at a high level: this includes a focus on evaluation, including on user experience, at an early stage.[68] This kind of consulting is therefore helpful in supporting organizations to tailor their tool to the context they will work in, addressing the first dimension of this barrier (AI builders lacking context).
The fact that many of the key barriers to AI implementation outlined in our original scoping are common to tech solutions, and not specific to AI, is also suggestive that this kind of support can be beneficial, as this kind of expert consulting can draw on lessons learned from technology implementations over many years. We suspect that this kind of support can result in fewer mistakes made by organizations at an early stage, resulting in better-designed AI tools. This kind of technical support can come in different forms; we have provided a framework for these models in Table 7.
Table 7: Interventions that provide advisory support to builders
| Intervention | Features | Examples | Costs (per participant) | Overall Take on Whether Promising |
|---|---|---|---|---|
| ccelerator/incubator programs | Intensive technical and management advice; peer network; seed funding; (sometimes) free/subsidized access to compute, or access to AI engineers | LEVI Math program; CGD/Agency Fund’s AI for Global Development Accelerator, Google’s GenAI accelerator | Broadly, this can be very high. AI for Global Development Accelerator: $500K per org in funding, $100K in credits, plus advisory and administration time However, we think cheaper models are possible, especially in partnership with providers of compute (Google, OpenAI, Amazon) | Potentially promising as a model for grants to individual implementers. |
| Short-term consulting | Intensive technical and management advice | Tech4Dev’s Fractional CxO; Amazon’s Now Go Build CTO Fellowship; in-house technical consultants at grantmakers | Fractional CxO program: estimated[69] $36K per organization. | Potentially promising, although highly uncertain |
| Community of practice | Light-touch technical and management advice; peer network | Tech4Dev’s Community of Practice | Likely low | No; unlikely to be impactful |
Incubators/Accelerators
Incubators or accelerators present a common form of advisory support to builders. Examples include the Learning Engineering Virtual Institute (LEVI) Math program and Google.org’s Growth Academy. Beyond advisory support, these programs provide a peer network for organizations to learn from one another. Shabnam Aggarwal highlighted this as a benefit of Rising Academies’ participation in the LEVI Math program, although she also stressed the importance of bringing together organizations that are not in direct competition.[70]
They also often include grant financing, as well as access to engineering support, addressing both the misalignment barrier and issues related to human capital. Many NGOs and social enterprises struggle to recruit AI engineering talent, as typical salaries for these workers far outstrip typical salary structures. Expert L, at an AI startup in an LMIC, told us that engineers working with them accepted salaries around a third as much as what they might receive for an equivalent Silicon Valley role, but that this was still multiples of the salaries for most of their employees. Beyond the strict financial burden, Expert L suggested that this kind of gulf in pay would cause discontent among existing employees, making hiring difficult to make such hires even if they seemed cost-effective.[71] This is an advantage of programs that provide short-term engineering talent to organizations in-kind, where salary costs are taken on by the funder.
The primary issue with this solution is that the marginal costs of supporting one organization are often high, as many of these programs also provide other resources, such as computation credits or grant financing. The provision of computation credits is easier for organizations such as Google and OpenAI that are able to provide their own services in-kind. It is also not necessarily neglected—several of these programs already exist, as shown in our collation of existing funding activities.
However, partnership models, with organizations like Google or OpenAI, that can provide computation credits and engineering support in-kind, may provide a lower-cost model for CG. An example of this arrangement is the AI for Global Development Accelerator, where the Agency Fund provided advisory support, while OpenAI provided computation credits in-kind.
Given the role of these kinds of programs in supporting successful examples of AI solutions, we think this kind of partnership model is worth exploring further as a promising opportunity for CG, at least as a model for funding individual implementers that they identify as promising.
Short-Term Consulting
Several models exist for builders to receive short-term consulting to support them rolling out an AI solution, or integrating it into their existing work:
- Tech4Dev has a Fractional CxO program seconding experts in integrating tech solutions into implementing organizations in India. These last 3–9 months and are at a subsidised cost to the organization, with the difference being made up by donors.
- Amazon runs a Now Go Build CTO fellowship, which trains existing NGO staff in developing technology tools.
- Grantmakers may have in-house technical expertise that can be used to support grantees: our interviewees gave us cases of Google.org providing this kind of support. Expert J also mentioned that the Patrick McGovern Foundation has in-house technical support for grantees.[72] This is likely a more flexible model, with the degree of support depending on the grantee’s needs.
This model will tend to be significantly less expensive per organization than the incubator/accelerator model described above, as it focuses solely on providing advisory support. This offers greater potential scalability, and it is therefore plausibly quite leveraged. But it is difficult to assess the effectiveness of this kind of support on its own: the organizations we spoke to had generally been through the more intensive accelerator/incubator model, and benefited from its additional features. After searching for ~45 minutes, we also struggled to find compelling evidence on the effectiveness of technology consulting in general. Discussions with recipients of this support, such as alumni of Tech4Dev’s Fractional CxO program, would be informative to understand how important this support was to their overall success.
We are therefore highly uncertain about whether this solution may be impactful, but we think it may be worth exploring further. Discussions with recipients of this support would be especially helpful for confirming this view.
Community of Practice
A type of support with lower costs and greater potential reach than incubators or consulting might be a model like Tech4Dev’s Community of Practice: this provides monthly webinars, quarterly in-person convenings, and a Discord channel to foster collaboration and knowledge-sharing between builders and provide low-intensity advisory support.
While costs for this solution are low, we are pessimistic that it would make a difference in enabling organizations to successfully roll out an AI solution. Our impression from discussions with successful builders is that issues which they encounter in building and deploying AI solutions are highly specific to the type of AI application and the sector they work in. This suggests to us that more intensive, tailored advice as part of the models listed above would be much more likely to be effective than highly generalized, non-targeted advice through broadly aimed webinars.
Additionally, no interviewees mentioned peer support on its own as important to their success, suggesting that this only has a minor role in helping organizations to build and deploy AI solutions. For these reasons, we do not view this as a promising solution for medium to large funders.
Advisory Support for Decision Makers
A potential route to addressing the second dimension of this barrier would be to support decision makers, including government officials, in making effective decisions regarding integrating AI solutions.[73] Markus Goldstein told us that the Indian government had requested CGD’s help to ascertain which proposals for AI applications were most promising.[74]
We have a limited understanding of what specifically decision makers struggle with in making these decisions. To narrow down what kinds of support might be most beneficial, it would be useful to speak to key decision makers: this could include senior civil servants/ministers in health, education, or agricultural ministries in LMICs, or the leaders of large implementers of health, education, or agriculture programs in LMICs. What follows is a description of several possible interventions to support decision makers generally, but our analysis would benefit from further conversations. A summary of our takes is available in Table 8.
Table 8: Interventions that provide advisory support to decision makers
| Intervention | Examples | Costs (per participant) | Overall Take on Whether Promising |
|---|---|---|---|
| Advisory materials | UN-International Telecommunications Union AI Skills Coalition; Agency Fund’s AI Evaluation in the Social Sector | Theoretically, very low (as materials are aimed broadly). But potentially very low readership or take-up of evidence. | No; materials are unlikely to be read or acted upon. |
| Embedded advisors in government | Digital Impact Alliance’s Country Engagements[75] | >Moderate, as advisors directly engage with just one government at a time. | In the medium term (~5 years), possibly. In the short term, impact is limited by the lack of evidence/experience in applying AI to government work. |
| General training for government officials | Asia Foundation and Stanford Institute for Human-Centered AI’s (HAI) AI Perspectives from Asia program | Low. Training could reach many policymakers at once. | Highly uncertain, but tentatively no. The current types of programs available may be too broad to be valuable. |
| Support collaboration between governments | World Bank’s Artificial Intelligence Working Group | Low. Convenings could theoretically involve many country participants. | No; non-neglected, limited value of additional work by CG. |
Advisory Materials
Unlike direct advisory relationships, producing advisory materials offers a more indirect route to influencing how decision makers evaluate and implement AI solutions. One example is the UN-International Telecommunications Union AI Skills Coalition, which is aimed at decision makers and produces an extensive range of online courses on AI and its implementation. Another example is the playbook for AI Evaluation in the Social Sector, produced by the Agency Fund and drawing on work with CGD and J-PAL. Other evaluation frameworks for AI deployment have been developed by Microsoft (Zhang, 2024) and Precision Development (Kamra et al., 2025). None of these has yet emerged as the gold standard, potentially creating an opportunity for CG to support a framework they consider superior.
However, there is evidence that materials aimed at policymakers in developing countries that do not involve direct engagement with policymakers generally have limited impact. Doemeland and Trevino (2014) found that 31% of World Bank policy reports are never downloaded, and that 87% are never cited (p. 4). Work by Alix Bonargent with the International Growth Center (Bonargent, 2024) found that only 3% of non-partnership research projects result in evidence uptake, while the involvement of policymakers increases this to around 20% (p. 3).
Overall, we are not convinced that these materials will be widely used by governments in their decision-making, and therefore do not see them as a promising solution for funders.
Embedded Advisors in Governments
Embedding advisors with expertise in applying and integrating AI into governments could improve decision-maker efficacy. This could be designed along the lines of the Digital Impact Alliance’s Country Engagements, which work directly with the governments of Sierra Leone, Ethiopia, Nigeria, and Benin to support their digital government strategies and data. Another potential partner might be the Zambia Evidence Lab, which is a joint initiative of the Ministry of Finance and National Planning initiative and the International Growth Center, and which has also published previously (Wani et al., 2025) on how governments can make the best use of AI. The UNDP also conducts country assessments (UNDP, 2025) on readiness for AI, which involve government policymakers.
The study by Bonargent (2024) above aligns with our strong prior that policymakers are substantially (according to Bonargent’s data, 7x) more likely to be influenced by assessments that they actively participate in. In terms of the pure outcome of influencing policymakers, we therefore think that this is the most promising solution listed here.
However, we are uncertain that it will significantly improve government decision-making in the short term, due to the current lack of evidence on which technical advice can be based. This is discussed further below.
General Training for Government Officials
Directly training government officials in AI skills could help them better understand how to integrate AI into government work. The African Development Bank is currently collaborating with Intel on a project in Africa (AFDB, 2024) to train 30,000 government officials (and three million citizens) in AI skills. Similarly, the Asia Foundation and Stanford Institute for Human-Centered AI (HAI) founded an AI Perspectives from Asia (The Asia Foundation, 2023) program to provide policymakers and civil society groups with briefings on the potential and risks of AI, and allow them to discuss this material.
We are highly uncertain about the effectiveness of these programs. We have a weak prior that this kind of general training may be too broad to be useful, relative to sector-specific training. As with embedded advisors, the effectiveness of policy-maker training may also be limited by the lack of evidence on AI integration into government. We do not currently view it as a promising solution for CG, but conversations with policymakers would be helpful to confirm this view.
Supporting Collaboration Between Governments
A final model is to support collaboration between governments, as many will be confronted with similar questions about how to best implement AI. One example of this is the World Bank’s Working Group on GovTech and Public Sector Innovation (World Bank Group, 2025). However, we do not think that philanthropies are well-suited to this role, compared to UN agencies that have a mandate to convene governments. Our landscaping analysis also suggests that this is not a neglected area. We therefore do not see it as a promising solution for CG.
Drawbacks to Advisory Support for Decision Makers
There is currently a lack of evidence and, consequently, reliable technical advice for how AI can be integrated into government processes. For example, a blog (Wani et al., 2025) by the International Growth Center on making AI work for LMIC governments cited a US public inventory of AI use cases as a primary resource on past attempts to use AI for government work. The absence of examples in low and middle-income countries potentially limits the usefulness of advisory support of this nature in general. We think this kind of solution may have greater benefits in the medium term (~5 years’ time), once there are more examples of integrating AI into government work.
Barrier: Local Regulatory Barriers
Description
Unlike many high-income countries that have established or emerging frameworks for artificial intelligence regulation, most low- and middle-income countries (LMICs) lack robust regulatory structures to address the use of AI tools. The absence of clear guidance on how to safely and effectively use and monitor AI tools creates uncertainty for developers, deployers, and data providers about how to design a product that operates according to existing local laws (e.g., on patient data sharing) and how to bring a product to market, given longer timelines for approval compared to high-income countries. Local regulatory barriers are common in LMICs, particularly in the health sector, where any product must necessarily pass stringent benchmarks for safety and effectiveness to be approved for use in the population.[76]
We decided to focus primarily on potential solutions to regulatory barriers for “AI for health” interventions, since barriers are likely to be more common in this sector, and unblocking them would lead to high impact. While we believe barriers to educational AI are also of high importance, due to the potential for misuse when involving minors, we believe that most of the specific barriers in this sector are likely to be active barriers and not solvable by fieldbuilding. Regulatory solutions in agriculture are likely to be of lesser importance because the anticipated risk and impact are lower than in other sectors.
Absent, conflicting, or highly complex regulations for AI use can be further broken down into passive and active barriers. Passive barriers arise from ambiguous or absent elements of the regulatory system (where solutions may involve developing new policies or streamlining existing processes), while active barriers are explicit laws or community prohibitions related to the deployment of AI tools. Active barriers (e.g., cross-border data sharing laws) are often necessarily in place due to the high risk of using patient data and ethical questions around the ability to monitor and link tools to specific outcomes. In contrast, passive barriers frequently originate from a lack of resources or cohesion across countries to develop and streamline processes, and in our view, are therefore likely to be more tractable. Passive barriers include:
- A fragmented regulatory framework for AI tools: There are currently major gaps in AI regulation in LMICs, where no guidance exists on how to safely and effectively implement new technologies. In an interview, Maëlle-Marie Corso emphasized that some AI tools may fit neatly into pre-existing regulatory frameworks (e.g., medical image diagnostics), while others do not (e.g., LLMs, which do not have a specific intended use as medical devices, can nevertheless be used to provide medical advice).[77] Medical devices are traditionally assessed as hardware devices, and formal regulatory frameworks for “Software as a Medical Device” (SaMD) only exist in a select number of LMICs.[78] Products themselves may be so novel that there are also unforeseen risks of regulating them through existing frameworks, creating low stakeholder and community trust for tools in high-risk sectors that do not have robust frameworks for approval or post-hoc monitoring and evaluation.[79]
- Extended timelines for approval: Low financial and staff resources, as well as underlying workflow challenges, can mean that national regulatory approval processes in LMICs have substantially longer timelines compared to high-income countries. This can create a lag between the fast-paced development of new AI tools and their adoption and dissemination, which means that LMICs cannot benefit from helpful products on the same timescale as high-income countries.
- Low market incentives for high-regulation sectors:
- Corso emphasized that the specific regulatory requirements to obtain market authorization for each country may require collecting additional evidence beyond WHO pre-qualification, such that “the markets themselves may not be big enough to justify… extra trials and clinical proof.”[80]
Active barriers include:
- Data governance laws: Stringent cross-border data sharing and privacy laws create significant barriers by forcing “data residency,” which requires countries to store and process information on physical servers within their own borders. This can prevent local innovators from accessing the high-performance global cloud infrastructure and diverse international datasets necessary to train accurate, unbiased models. Evert Bopp from Crisis Cognition identified several overlapping constraints from their work deploying AI systems in emergency contexts, including restrictions on data export or prohibitions on cloud use, low institutional trust in externally hosted systems, concerns around handling sensitive medical, protection, or population-level data, and operational risk from cyber intrusion or infrastructure disruption, particularly where connectivity is unreliable.[81] Organizations considering AI deployment must navigate these regulatory and institutional constraints regardless of technical capabilities.
Potential Fieldbuilding Solutions
We explored five solutions, finding that support for drafting or adapting new AI-specific regulations and establishing regional hubs to transfer regulatory “best-practices” to other countries are the most promising. Our specific takes are presented in the following table:
Table 9: Evaluation of solutions to regulatory issues
| Approach | Promising | Primary Reason |
|---|---|---|
| Support drafting or adaptation of new AI-specific regulations | es | Solves a primary passive barrier of gaps in AI-specific regulation |
| Regional hubs | Maybe | Creates a local reliance framework for regulatory harmonization, but would be hard to measure success |
| Regulatory sandboxes | Maybe | Strengthens evidence base for tools while exploring best-practice regulations, but could have limited impact |
| Direct support for governments and advocacy | No | Could weaken state regulatory capacity development, and best-practice policy objectives are not clear. |
| Subsidized country-specific validation fund | No | Not an obvious organization for funding or analogous example of success from another field |
The solution sets we explore below are focused on addressing passive regulatory barriers because active barriers seem less tractable for health.[82] We are relatively uncertain about the solutions we explore below for three reasons: (1) we did not interview policymakers; (2) we are not confident about how to measure success well in this area, which affects how we think about tractability; and (3) we may be relying on implementation fictions about regulatory processes and the state capacity of LMICs to translate policy into outcomes. Our sense is that potential solutions are divided into those that would promote the development of AI-specific regulations (where none exist or are currently ambiguous) and those that would streamline regulatory approval (to reduce friction). We are moderately confident that interventions addressing development (by filling gaps in legislation) are more likely to be tractable at this stage than those streamlining regulatory submission processes; we think that regulatory frameworks for AI are not robust in many LMICs due to the novelty of the field, and that streamlining processes only becomes more feasible once regulations have been solidified.
Developing New Regulations for AI Tools
Supporting the development of new regulations for the safe and effective use of AI in LMICs would address the primary barrier of substantial gaps in policy. We think some promising interventions in this area are to:
Support drafting or adaptation of regulations to address gaps
- For example, the Gates Foundation provided initial seed funding ($12.5M; Baeza, 2011) to establish the African Medicines Regulatory Harmonization initiative, which drafted the 2016 AU Model Law on Medical Products Regulation. This provided a template for African nations to harmonize their own domestic laws, ensuring that a “joint approval” in one region is legally recognized in another. If there is work being done to adapt the AU Model Law to include SaMD, there could be a tractable opportunity for CG to support this work, but we are not sure if this is the case, and inconsistent adoption could be a barrier to success.[83]
- One other option would be to fund a company like HealthAI, a nonprofit that focuses on building in-country, government-led regulatory capacity-building for health AI, including in five LMIC “pioneer countries.”[84] We think this could be a promising funding opportunity, but we have done limited research on this specific organization.
To determine if this solution would be a good candidate for investment, we would want to know (1) where specifically more legislation is most needed (i.e., countries with significant health AI activity); (2) if introducing new regulatory frameworks measurably reduces potential AI-related harms or unlocks beneficial deployment that would otherwise be blocked; and (3) how much capacity existing organizations have to absorb additional funding. Our loosely held view is that this space is likely to be moderately neglected since most funds are directed towards AI innovation itself, not regulation, but there are multiple organizations working on the topic, and we are not sure of the total costs required to construct new regulatory frameworks.
Deploy regulatory sandboxes
- Corso said that her team at PATH would be supporting an African National Regulatory Authority[85] to deploy a regulatory sandbox. A regulatory sandbox enables companies and regulators to work closely together to 1) “define best practices for testing products”, and 2) help regulators identify where regulatory guidance would be needed in the industry more broadly. Ideally, learnings from the sandbox then “enter the guidance that [the regulatory authority] publishes for other companies to leverage,” therefore creating a “feedback loop.” In this specific sandbox, companies will also get a decision from the regulator around whether or not they can enter the market.
- The African Union Development Agency’s New Partnership for Africa’s Development (AUDA-NEPAD) also advocated for AU member states to develop a “Pan-African AI Regulatory Sandbox” (AUDA-NEPAD, 2024) in December 2024, though we are unsure if any steps have been taken to implement this in practice.
We think regulatory sandboxes could be an investment opportunity for CG because they would strengthen the evidence base for AI tools alongside the development of safe regulations, but it would only be valuable insofar as its results would be accepted by external regulators and stakeholders. We are also unsure of the costs required, and we would want to know the number of companies that could be unblocked by a large-scale initiative; our sense is that some regulations would need to be relatively niche and that there could be limited adoption of findings by individual countries.
Direct support for governments and advocacy
A grant could also fund technical advice or advocacy directly supporting national governments to adopt regulations enabling the approval of new AI tools.
- The UN-International Telecommunications Agency produces guidelines and holds convenings on the topics of AI standards and governance (AI for Good, 2025), although their impact on LMIC government policy is unclear.
- Another model would be funding advisors within target countries to advocate more directly for the adoption of best practice policies, similar to work on lead paint elimination done by the Partnership for a Lead-Free Future, though the specific “best practice” policy objectives are less clear for AI tools.
We are unsure if funding a government directly would be seen as using foreign influence over domestic laws or weakening state capacity to develop independent regulations, so our sense is that this is not a good funding opportunity for CG.
Streamlining Regional Regulations
Regulatory harmonization efforts would aim to reduce burdens on AI developers and streamline approval processes, primarily by (1) implementing joint submission and review procedures, where agencies from different settings assess applications together, and (2) using reliance, where one authority defers to the rigorous deliberations of another. We previously researched reproductive health regulatory harmonization efforts in East Africa and Francophone West Africa (Gosnell et al., 2025) and found that some programs had successfully reduced registration timelines by ~50%,[86] but there was inconsistent adoption of harmonization efforts across national agencies, and resource constraints limited uptake in some settings.
Some examples of initiatives to streamline regulations that we think could be promising are:
Regional hubs
Supporting AI systems to establish in regional hub countries with relatively strong regulatory frameworks, then subsidizing expansion to neighboring countries, could create templates for cross-border deployment while building implementation capacity incrementally.
One recently-launched example of this is the AI Hub for Sustainable Development, under which 14 priority countries in Africa are being positioned as “regional hub countries” for AI development across six key sectors[87] due to their relatively advanced tech sectors and existing national AI strategies. The AU’s Continental AI Strategy creates a unified framework under which the hub countries can pilot specific applications, while Italy’s Mattei Plan (backed by G7) provides the financial support to help transfer these technologies and regulatory best practices from the hubs to neighboring countries.[88]
There could be an opportunity for CG to support this program, but we are unsure how much money has been specifically allocated for regulatory purposes or their capacity to use additional funding.[89] We think success would be difficult to measure convincingly, but could include the percentage of countries adopting the hub’s unified framework by the end of the initial three-year phase (though the counterfactual would also be challenging to estimate in the context of a current regulation vacuum).
Subsidized country-specific validation
A fund supporting the incremental costs of country-specific compliance could enable organizations to expand beyond their initial markets. This differs from direct implementation grants by specifically targeting the regulatory friction costs that block otherwise viable expansions. Commercial investors may be wary of the binary risk of regulatory approval in new markets, and targeting these one-time, high-friction costs could allow access to new markets without developers absorbing the risk directly. We spent a short time (~15 minutes) searching for prior examples of this and did not find specific instances of a successful fund to subsidize expansion to new markets for existing products, so we are not confident that this would be a tractable solution without a specific funding opportunity.
We also considered a shared service[90] to address hurdles like regulatory navigation, evidence generation, and regional approval processes[91] which seemed like a plausible theoretical solution, but we discontinued further research in favor of the more promising ex-ante solutions that are identified above.
In short, we think there are several promising funding opportunities in this area, but have limited confidence that any one is an outstanding candidate, given concerns about tractability and measuring impact in the regulatory environment. While we think funding an organization to address regulatory barriers is highly uncertain, affecting regulation is a leveraged intervention and has several promising funding avenues for CG.
Funders and Fieldbuilders
Total Volume of Funding in This Area
We evaluated total funding in the AI for Good space , including activities that we do not consider fieldbuilding. After conducting a short (15-minute) search, we found only one source (Lin et al., 2025) with an estimate of total funding for the area we were interested in (p. 33-36). They conducted a review of funding documents towards “AI for social good”: this consisted of a keyword search conducted in August 2025, followed by snowball sampling. While this was not a systematic review, their selection of keywords appears to be relatively comprehensive; we therefore believe that they are highly likely to have captured the bulk of funding towards AI for social good up to 2025.
Their approach identified $410 million in investments globally since 2017. To tailor their findings to our interest in AI for Good in LMICs specifically, we conduct further analysis here. We exclude initiatives they find that focus on high-income countries, and count only 50% of the amount of funding for initiatives eligible for applications in high-income and low and middle-income countries (assuming roughly half go to relevant countries). This yields a figure of $302 million of funding committed or disbursed over nine years, or ~$33 million annually. This figure may be inflated by the inclusive approach that Lin et al. take to investments in AI, including grants for “data science”. On the other hand, for several Requests for Proposals, the total amount committed across grantees is not discernible from funding documents, so these are excluded. Overall, we think it is highly unlikely (<5%) that the true figure is more than an order of magnitude away from our estimate.
One notable caveat, however, is that this excludes a recent announcement (Taylor, 2025) by OpenAI that its nonprofit wing would commit $25 billion to innovation in health and AI resilience.
Current Funders
To identify organizations actively funding or supporting this space, we first drew on our literature review and the interviews we conducted with experts. We supplemented this by spending around 1.5 hours searching for programs for AI for Good by other established philanthropies and bilateral and multilateral agencies, and collected results in this sheet. It is possible that with more time to search, we might identify additional initiatives, especially in cases where AI only plays a minor role (and which may therefore not show up using a keyword search). However, we think that our approach should capture major actions by key donors.
The most active bilateral donors include the UK Foreign, Commonwealth, and Development Office (FCDO), Canada’s International Development Research Centre (IDRC), and the Swedish International Development Cooperation Agency (SIDA). Notably, the FCDO and SIDA have supported the GSMA Innovation Fund since 2013 (GSMA, 2021), which has disbursed more than £28 million to support digital innovation in LMICs. The most active philanthropic organization has been the Gates Foundation, which, for example, issued $5M of grants (Brey, 2023) for LLM-enabled development interventions in 2023, and has also invested $15M to support development and deployment of AI technologies through “AI scaling hubs” in Rwanda (Gates Foundation, 2025) and Nigeria (Bilali, 2025). Alongside traditional funders, tech companies involved in developing frontier models have supported AI for Good applications through their philanthropic arms, notably Microsoft and Google.org.
Several key funders coordinate investments through the AI for Development (AI4D) Funders Collaborative, which supports structured collaboration, joint funding, and the synthesis of evidence gathered by its members. Partners currently include the Gates Foundation, the German Agency for International Cooperation (GIZ), IDRC, Japan International Cooperation Agency, FCDO, SIDA, and Community Jameel. Collectively, these organizations have committed more than $100 million, of which FCDO and IDRC are the largest contributors, to support infrastructure, governance, skills and talent, and the development and scale-up of applications of AI for development programs. AI4D has so far launched the AI Evidence Alliance for Social Impact (AEASI), a £2.75 million research and evidence initiative to support AI for Good.
Funding Activities
We quickly examined and loosely categorized the types of activities funded by these donors and international organizations. A key interest was the relative share of attention devoted to fieldbuilding, compared to a focus on making grants to individual implementers to support their activities. Our definition of fieldbuilding is broad, going beyond the classic sense of community building to include within it all activities that provide the conditions for applications of AI for Good;[92] this excludes only grants to individual implementing organizations.[93]
A summary of these activities is available in Table 10. Overall, the most prominent activities were grants to individual implementers. But we found 16 unique initiatives with some kind of fieldbuilding function.[94] Several organizations position themselves in a coordinating or convening role: the AI4D funders collaborative, Google.org, the World Bank, WHO, and the UN-International Telecommunication Agency, the latter of which has hosted an annual ‘AI for Good’ summit since 2017. We also identified five programs aimed at expanding training to support implementers in developing AI applications; five programs to develop research and evidence on applying AI for Good; four aimed at establishing standards or policies for responsible and effective AI use; three for supporting enabling infrastructure, especially access to datasets for model training; and three with incubator or accelerator models, where early-stage organizations or initiatives are given access to funding, expertise, and connections.
Table 10: Funder activities
| Activities | Programs Identified | Examples |
|---|---|---|
| Grants to implementors | 25 | Gates Foundation “Grand Challenges” AI grants for organizations using LLMs to solve problems in Global Health and Development (Brey, 2023) |
| Coordination/convening | 7 | AI for Good summit, hosted annually by the UN International Telecommunications Union |
| Research/ evidence | 5 | AI Evidence Alliance for Social Impact (AEASI), £2.75 initiative to provide grants for research to support evidence-informed deployment of AI for social good in Africa |
| Training | 5 | >African Development Bank project with Intel to train 3 million Africans and 30,000 government officials in general AI skills |
| Policy/ Standards | 4 | Work as part of the AI for Development (AI4D) Collaborative to support locally relevant policy frameworks for innovative and safe AI |
| Incubators/ accelerators | 3 | AI for Global Development Accelerator, with OpenAI and CGD, offers several organizations with established health and development interventions training and support to integrate an AI solution into their work |
| Infrastructure | 3 | Masakhane Language Foundation is collating datasets in low-resource languages for model training |
Note. We treat all activities as fieldbuilding, with the exception of grants to implementers.
Most grantmaking and other activities were not restricted to any particular sector or region, with 28 out of 39 initiatives being cross-cutting, and only five being restricted regionally (Tables 13 and 14 in the appendix).
Likely Important Funders in the Future
- Given announced cuts, it is likely that FCDO’s role in the space will become less prominent in the next few years.
- OpenAI has also signaled a growing interest in supporting AI for Good applications. So far, this has focused on applications in the United States, but they recently announced $25 billion of funding for innovation in health and AI resilience (Taylor, 2025).
We find it likely (~70%) that Official Development Assistance funding in this area will decrease in the next 5–10 years, both relatively and absolutely (which may be more stark, as announced aid cuts by several bilateral funders take hold).[95] This was the view espoused by one interviewee, Markus Goldstein, who suggested that AI’s status as the “new shiny thing” in the development space means that it currently receives a disproportionate share of funding.[96] However, this may be more than offset by commitments by technology companies, including OpenAI and Anthropic, which have shown interest in this space. We are currently highly uncertain how likely this interest is to translate into funding commitments and, if so, where this funding will be directed.
Contributions and AcknowledgmentsRuby Emerson, Britta Jewell, Rory Todd, and Thomas R Vargas researched and wrote this report. Vargas served as project lead. Aisling Leow, Abbie Clare, and John Firth managed the project; Clare also contributed to the research. Special thanks to our managers for helpful comments on drafts, and to Deena Mousa and Oliver Kim at Coefficient Giving for their input. Thanks also to Shane Coburn and Thais Jacomassi for copyediting and to [comms/publishing staff] for assistance with publishing the report online. Further thanks to the experts who took time to speak with us: Shabnam Aggarwal, Evert Bopp, Rikin Gandhi, Markus Goldstein, Edwin Klinkenberg, Jay Patel, Maëlle-Marie Corso, Samantha Wolhuter, and David Yanagizawa-Drott, and unnamed anonymous experts. Coefficient Giving provided funding for this report, but it does not necessarily endorse our conclusions. |
Appendix
What We Would Do With More Time (WWWDWMT)
- We would attend the Devex webinar on “Why AI for Good isn’t scaling,” involving Kanika Bahl, who is leading the new AI Access Initiative, or alternatively speak with Bahl directly (Cheney, 2025).
- We would speak to policymakers involved in decisions about taking up AI tools. We would focus on speaking to senior civil servants—or ideally, ministers—within ministries of health, education, or agriculture, especially those currently making decisions about whether and how to take up AI tools.
- We would speak to policymakers involved with regulating health products and devices to better understand where AI tools might require new forms of regulation and what preferences would be for developing and adopting new laws.
- We would speak to someone at the AI Hub for Sustainable Development about specific plans for distributing regulatory best practices from regional hubs to surrounding countries and ask about costs of supporting regulatory harmonization (we discovered the organization late in our research process).
- We would listen to this podcast interview between Maëlle-Marie Corso and Dr. Hugh Harvey about health AI regulation (Troadec, 2025).
- We were suggested this resource because it speaks to capital as being a big barrier to AI4G, but we deprioritized it due to our perception that it is not a top barrier within the scope of this work. We would scan with more time.
- Maëlle-Marie Corso proposed the creation of regulator-endorsed national testing datasets to overcome local data training gaps. Rather than requiring each AI developer to independently source local data for validation, regulators could establish (or support the creation) of standardized datasets that represent actual population characteristics. However, we did not have the time to evaluate this solution in the current study.
Comparison to Other Frameworks
Technology development frameworks:
- NASA Technology Readiness Levels (TRL 1-9): Strong technical focus, but weak on adoption and sustainability dimensions
- Stage-Gate Process: Business-focused with relatively complete coverage, but limited attention to LMIC-specific challenges
- Biotech/Medical Device Development: Similar structure to our framework, with heavy emphasis on regulatory pathways
Adoption and diffusion models:
- Rogers’ Diffusion of Innovations: Individual psychology focus; weak on technical development and institutional barriers
- Lean Startup: Iterative emphasis valuable, but missing regulatory complexity and sustainability planning
- Greenhalgh et al. Implementation Model: Distinguishes assimilation from adoption; emphasizes non-linearity that our framework does not fully capture
Health and development-specific frameworks:
- WHO Digital Health Guidelines: Most comprehensive for health in LMICs; explicitly addresses governance, policy, and infrastructure—closest analog to our framework
- Consolidated Framework for Implementation Research (CFIR): Five domains (Intervention Characteristics, Outer Setting, Inner Setting, Characteristics of Individuals, Process) provide systematic barrier taxonomy that we integrated
- USAID Journey to Self-Reliance: Three dimensions (Commitment, Capacity, Financing) address sustainability gaps in technology-focused frameworks
- Digital Public Infrastructure (DPI) Framework: Three layers (Foundational, Services, Governance) distinguish infrastructure from applications—relevant for AI systems requiring compute/data infrastructure
- Market Systems Development: Models market dynamics particularly relevant for agriculture AI; highlights difference between direct service delivery and market-enabling interventions
Interested readers can refer to our spreadsheet to compare the various process maps and frameworks across the stages.
Caveats to Our Approach
Our goal of making a broadly applicable process map has necessitated that the end product is more generic and less specifically diagnostic than some other approaches might be. For example, some frameworks would split our Stage 6 (User Adoption & Training) into at least two phases or workstreams, with distinctions between individual trial, organizational routinization, and systemic integration. Our single stage may conflate these mechanisms. In addition, our approach to baseline infrastructure needs may be insufficient to analyze some cases, and some AI interventions may require parallel tracks—one for foundational infrastructure (compute, connectivity, data systems) and another for AI applications built atop that infrastructure. Our framework also does not cleanly account for nonlinearities in process, such as feedback loops or agile development processes.
Despite these refinements, we retained the eight-stage linear structure for analytical clarity to move on to mapping process challenges, while acknowledging that real-world implementation is likely much messier and more iterative.
Detail on Literature Scan for Barrier Identification
Drawing on existing literature, technology deployment case studies, and preliminary understanding of the AI landscape, we generated an initial brainstorming list covering technical, institutional, and contextual challenges. To validate and refine this initial catalog, we conducted structured interviews with practitioners deploying AI solutions in LMICs and with experts across the space. We also cross-referenced our barrier list against established implementation science frameworks, WHO Digital Health guidelines, and published case studies of AI deployment failures and successes in LMICs.
Regarding the literature, we spent a couple of hours scanning for published articles, working papers, reports, blog posts, and other non-technical publications. We identified 27 resources related to AI for Good in LMICs, listed in the corresponding tab of our spreadsheet with corresponding search terms and other metatags, and selected roughly half (13) for deeper examination, given our sense of how useful they would be to our work. Across these resources, we identified 26 unique challenges and quantified how frequently these challenges were mentioned. The top 10 challenges by mentions are:
Table 1A: Top 10 barriers by mention in literature
| Challenge | Count |
|---|---|
| Local training data | 9 |
| High-speed internet access | 8 |
| Culture/norms | 7 |
| Reliable power grid | 6 |
| Capital | 6 |
| AI talent | 6 |
| Compute | 5 |
| AI Hardware | 5 |
| UX | 4 |
| Regulatory environment | 4 |
Note. Drawn from 13 sources, which are listed here.
Detailed Findings on Categorizing Barriers by AI-Specificity
After assigning a rating to each barrier, AI-specific problems account for six barriers, or 10% of the total catalog. These are challenges unique to artificial intelligence systems that would not apply to non-AI interventions. The most prominent AI-specific barriers include context-specific training data needs, such as the lack of local language corpora, regional disease presentations in medical imaging, or crop varieties specific to smallholder agriculture in particular geographies. Furthermore, we found potential issues with model performance degradation in populations underrepresented in training data, dependence on external APIs and foundation models, and the black-box opacity of many AI systems.
General technology problems represent 20 barriers, or 33% of the catalog. These challenges apply to any digital technology deployment, not just AI systems. For example, we heard about infrastructure limitations around internet connectivity, mobile network coverage, and electricity access, constraining what is possible in many LMIC contexts. The shortage of technical staff with skills to build or maintain digital systems affects all technology interventions, though AI skills are particularly scarce.
General development or implementation problems comprise 34 barriers, or 57% of the catalog—the largest category. These are institutional, political economy, and capacity challenges that affect any development intervention, regardless of whether technology is involved. For example, donor-funded pilots might not transition well to sustainable government budgets, political transitions lead to abandonment of a predecessor’s initiatives, as new leaders want to launch their own flagship programs, or misalignment between external builders and local decision makers means what gets built doesn’t match what’s actually needed.
The distribution of barriers according to this categorization across implementation stages reveals important patterns about what kinds of challenges emerge at different points in the deployment journey. The earliest stage emphasizes general challenges, with problem identification being similar to many other types of programs. Beginning with solution design, validation, and regulatory navigation, the process shows higher concentrations of AI-specific barriers, centered on training data availability, computer access, specialist workforce skills, and validation methodologies appropriate for AI systems. Middle stages show more balanced barrier distributions. Late stages become dominated by institutional challenges, such as sustained funding, political economy, capacity to maintain systems, and adaptation to changing contexts—not technical AI problems. By the time an AI system is being scaled and sustained, the technical questions have largely been answered; what remains are the eternal challenges of development work in resource-constrained, politically complex environments. We expect that the most high-leverage AI for Good interventions may therefore need to focus on the earlier stages of product testing and introduction, or take a broader approach than simply AI enablement.
Bayesian Aggregation Methods
For each solution, we elicited assessments regarding:
- Directional Vote: A binary assessment of whether the solution is promising (Yes/No).
- Confidence Interval: How confident the evaluators are in their assessment.
The data structure and interpretation guide are shown in the table below:
Table 2A: Translation of Promisingness to Scores in our Model
| Rough Interpretation | Vote | Min | Max |
|---|---|---|---|
| “It’s a coin flip.” | Yes or No | 0.5 | 0.55 |
| “I think it WON’T work, but I’m unsure.” | No | 0.6 | 0.75 |
| “I think it WON’T work, and I’m certain.” | No | 0.9 | 0.99 |
| “I think it WILL work, and I’m certain.” | Yes | 0.9 | 0.99 |
| “It’s promising, but I have doubts.” | Yes | 0.6 | 0.75 |
Our assessments can be found here.
We use a Weighted Geometric Mean of Odds to take individual assessments and turn them into a subjective, collective team belief; this method rewards consensus and penalizes unsupported outliers. There are other nice features as well: In small groups, Weighted Geometric Mean of Odds ensures that skepticism from any single evaluator (under our weights, particularly the domain expert) reduces the certainty that something is labeled promising. It also produces a consensus bonus, boosting confidence when evaluators align. We tried just weighting them equally as well, but the results did not change much. Since we think this version of the results is more consistent with our preferences, we decided to keep this aggregation method.
We assign a larger weight (w=0.5) to whoever was the principal researcher for a given barrier, reflecting the idea that they possess deeper context, and divide the remaining weight among the two peer evaluators (w=0.25 each). To model the uncertainty associated with our confidence intervals, we deploy a Monte Carlo simulation (n=5,000 draws per solution). We sample random confidence values from each evaluator’s range, calculate the Bayesian update for each iteration, and generate a probability density function (PDF) for every solution.
Additional Funding Analysis
Table 13 shows our breakdown of funder activities by sector.
Table 3A: Sector focus
| Sector | Programs Identified |
|---|---|
| Cross-cutting | 28 |
| Health | 5 |
| Education | 2 |
| Agriculture | 2 |
| Humanitarian | 1 |
Table 14 shows our breakdown of funder activities by region.
Table 4A: Region focus
| Region | Programs identified |
|---|---|
| LMICs | 34 |
| Africa | 4 |
| Southeast Asia | 1 |
- We also call such obstacles “barriers,” “challenges” and “issues” interchangeably. ↑
- We adopt a broad definition of fieldbuilding, including within it all activities which provide the enabling conditions for (multiple) organizations or actors to develop and implement AI for Good interventions in LMICs; this excludes grants to individual implementers. More detail is provided in the ‘Funders and Fieldbuilders’ section. ↑
- In identifying experts for interviews, we sought to arrange discussions with funders, researchers, and implementers who are either active in or familiar with the AI4G space. These include at least one representative from each of our predetermined verticals (agriculture, education, and health) speaking to their specific experiences, research, and challenges. We also sought at least a few critical perspectives. ↑
- This included a severity rating on a one-to-five scale (discussed here). ↑
- Two reasons for this: For one, the field is nascent, so there may be many important problems or barriers whose cumulative effect might not be obvious today. For another, even if organizations are experiencing the most important barriers today, some unknown selection process might determine what issues are written about, which may not fully map onto the former. ↑
- To note, the fact that many barriers counted in multiple stages mean the sum of barriers in each category exceeds the 60 total barriers. ↑
- To understand whether barriers are uniquely related to AI technology or reflect broader challenges in LMIC development, we also classified each barrier into one of three categories:
- AI-specific problems (barriers unique to the implementation of AI, for example, lack of adequate training data),
- General technology problems (barriers related to the implementation of tech-for-good interventions; for example, access to appropriate hardware or internet connectivity), and,
- General development problems (barriers related to standard development challenges; for example; questions related to state capacity.)
These tags can be found in the spreadsheet. We are not entirely sure this categorization is helpful, however. While conceptually interesting, the response to any given problem could, in principle, still be AI-specific (e.g., targeted funding). For this reason, we have chosen to remove much of the analysis pertaining to this categorization and placed it in the appendix, where it has not been fully updated. ↑
- Thanks to Jamie Elsey for producing this diagram. ↑
- At different points early on, we rated each on a one-to-five severity scale based on how completely the barrier blocks progress and how difficult it would be to overcome. (A severity rating of 1 indicates a minor inconvenience easily addressed, while five represents a near-absolute blocker that could kill or minimize the positive impact of many projects. Ratings of 2 suggest notable challenges with multiple workarounds, 3 indicates significant barriers requiring substantial effort, and 4 represents major obstacles that block many implementations.) We each worked through the full list of barriers and assigned our personal assessment to each barrier, then calculated average severity scores based on our independent ratings from three team members. The average score also assigns greater weight the more a problem is mentioned by experts or literature, and is normalized on a 1–5 scale. After additional expert interviews and discussions with Coefficient Giving, we further adjusted the scores to be more in line with our understanding of which barriers are likely to be important and solvable with fieldbuilding interventions within the scope of this project. The highest-severity barriers we identified (including data access, infrastructure, and human capital) were independently mentioned by multiple experts without prompting. This convergence between our analytical framework and expert experience increases confidence in the severity rankings. We used these lists to inform our holistic evaluation of top barriers (and indeed there is much overlap), but we think our quantitative approach created the impression of precision and at times might be double-counting inputs. ↑
- Because we explore so many solutions, we prioritized research into some over others and have attempted to convey how we did this in the write-ups below. ↑
- Klinkenberg: “Performance is decreased due to older infections leaving cavities and chronic parenchymal destruction behind in the lung which can look like active TB, human readers also show decreased performance in this group. This is an inherent limitation of the population and x-ray. Comorbidity alters the presentation of TB, such as in HIV making it more subtle in an x-ray for example. These limitations do not necessarily arise from a limited training dataset but are inherent to the modality used” ↑
- Secondary sources we consulted point in the same direction. Mismatches between the populations, disease presentations and training data can lead directly to poor performance in low-resource healthcare settings (Okafor & Nyarko, 2024, p. 4). Schwalbe and Wahl (2020) warn that applying models trained on one population to another creates a “threat that false associations could be identified and integrated into new AI-driven health interventions” (p. 1582), with unknown outcomes. Giacomelli (2025) agrees, noting that “AI models trained on global datasets often fail to capture local nuances,” creating a fundamental barrier to relevance in LMIC contexts. ↑
- Goldstein: “But anyhow but yeah but they don’t benchmark on this right. So, so back to the constraints like they don’t this is not a primary indicator they don’t really care about serving you know people of Gundi who speak Kurundi . . . .they had been explicit in their to the AI companies about this and they’ve been explicit that this is not a primary they don’t care about Swahili and Swahili worries me more right or you know just on the number.” ↑
- Klinkenberg: “Data is critical . . . .Making data publicly available is very important—fast-tracks AI development” “Need 10,000 pediatric images . . . .Within X-ray space, data on silicosis etc. is critical.” ↑
- Expert H: “Neither FAO [nor] CG [share data] because everybody’s using knowledge is power and the power is to get money. People are using knowledge in order to get more money . . . . Um, but this like we’re in a hunger games and this is because the people you work for have created this. The philanthropic organizations have created the Hunger Games that everybody’s scrabbling around there.” ↑
- It occurs to us that many of these solutions could be deployed in combination with initiatives to make these data public and accessible to many, but we think there is also a universe in which these initiatives provide data directly to LLMs to improve these capabilities themselves. ↑
- Gandhi: “So like what we do is we also have like this reinforcement learning process where we’ve engaged local agronomists in different countries that we operate in to basically go in and then annotate . . . .This is the system not for the farmers. This is the system for the local agronomists who get paid basically gig work . . . . to like have this local agronomist, you know, review something like 20,000 odd things, then yeah, you can get this up to like 80 kind of like percent.” “It doesn’t cost that much. So like, you know, it’s answering these questions generally for them, they’re like experts . . . . on average answer about like one question takes them about like 30 seconds something like that . . . . It’s mostly clicking the buttons that takes the time.” ↑
- Expert H: “But then we pull that information into the back end . . . . So you have the record coming in and you have also multiple types of information uh whereby you can see was the AI correct or not. So humans can update this . . . . if a human expert at, say, the CG center says that’s not correct, we can improve it. So you got that human in the loop part.” “AI can be a job creator in Africa where it’s a job destroyer in the global north. And so as a job creator, you want to take these unemployed young people and give them a job . . . . We employ a bunch of young people who are running around helping farmers use the tools . . . . They’ve got degrees in agronomy and stuff and they’re ongoing training.” ↑
- An example seems to be Google’s partnership with others in the context of eye health (Sawhney, 2024). ↑
- We also found but did not examine this competition to release datasets (Tools Competition, 2025). ↑
- It is possible that such a reaction might lower costs to end users as well, if it allows more entrants into the relevant market. ↑
- Another possibility is that organizations may simply not want to provide data because of concerns over competition or the regulatory issues we highlighted above. Of these, the most tractable to overcome seems to be resources, so we focus efforts on this dimension. ↑
- Thank you to Maëlle-Marie Corso for floating this possibility. ↑
- In her previous research at MIT, she worked in a lab who developed breast and lung cancer risk prediction models to support more personalized screening strategies, and models to more accurately diagnose tuberculosis. See Mateen et al. (2025). ↑
- As part of this report, we considered the fieldbuilding option of funding the establishment of practical benchmarks for language evaluation and the continued assessment of different models on a variety of priority languages. However, our impression is that several organizations are already working on this for various sub-fields, and it does not appear to be neglected. (Gandhi: “there’s an initiative from University of Illinois that’s doing this thing called AI AgriBench (Univ of Illinois). . . . There’s another one which is from the UAE government, the MBZUAI . . . . with the Gates Foundation support are also trying to do this benchmarking thing.”) This is our own conclusion, loosely held.Organizations currently supporting this benchmarking:
- Masakhane – Created IrokoBench, MasakhaNER, and other benchmarks specifically for African languages. Active ongoing benchmarking work
- GEM (Generation, Evaluation & Metrics) – Multilingual generation benchmark, including some African languages. Could expand coverage with funding
- AfriCLEVER/AfricaNLP – Research community developing benchmarks and evaluation datasets for African languages
- BLEU is the standard metric used to evaluate the quality of text generated by a machine (like an LLM or a translation system) by comparing it to one or more human-written reference translations. ↑
- FLORES-101 is a benchmark dataset created by Meta AI (Facebook). If BLEU is the ruler used to measure performance, FLORES-101 is the exam paper the models are taking. It is particularly famous because it shifted the industry’s focus from “rich” languages (like French, German, and Chinese) to low-resource languages (like Swahili, Urdu, and Nepali). ↑
- Expert J: “I don’t think we should be in the game of let’s solve the language because the Indian government is putting so much money [in] . . . . So while language is a critical thing to me it’s more like a yes or no, right, if there’s no adequate language support, then pivot . . . like there might not be adequate language support in audio, [but] there might be pretty good language support in text, because typically that’s the funnel, then then do a text intervention for now if you think you can pivot to that.” ↑
- Patel: “If you know how to do it, you can go from a base model through to language fine-tuning and add domain-specificity, all for $3,000.” ↑
- Patel: “Frontier models have gotten good enough that building your own is often no longer worth it for moderately-resourced languages. The recent Gemini and GPT releases handle them really well. If I were starting today, I would not train my own model. I would take a frontier model, engineer the prompts, and focus instead on the safety, evaluation, and localization scaffolding around it. That scaffolding is what turns a general model into something safe to put in front of a mother asking about a danger sign in her own language.” ↑
- Expert J: “Ideally, if funders can put in small AI teams within the organization and just use them in [an] advisory capacity or [say] go work with another organization and build an AI team that can help advise.” ↑
- Expert J: “Saying like hey, Swahili does not work, let’s get swahili to work, is a much much harder undertaking than hey, let’s try to fine-tune this model to understand the agricultural terminology in Swahili.” ↑
- Jacaranda serves 650K mothers at $0.74/mother; $3K = 4,000 users’ annual cost ↑
- Some members of our team expressed skepticism of costs going this low. ↑
- Expert G: “One of the most important elements here is for the community to own data. This is data that doesn’t exist anywhere, even OpenAI doesn’t have this kind of data.” ↑
- Expert G: “The community want full control and autonomy over the data. So they don’t want to open source it.” ↑
- There are also non-trivial costs related to supporting AI hardware outside of actual internet access. For instance, Delft Imaging’s CAD4TB deployment in Cameroon faced equipment failure due to heavy rains (“In Cameroon, rain ruined it.”), and Klinkenberg highlighted theft and misuse as frequent concerns. Specifically, he mentioned that in contexts where “people aren’t used to equipment or don’t know how to use it properly” sometimes equipment “gets abused or stolen.” While these are important challenges, they are site- or problem-specific (e.g., require localized security and maintenance strategies), and would not be easily solved though broad, field-wide interventions of the type we explore here. As such, we focus here on the issue of access. ↑
- The countries include Colombia, Ghana, India, Indonesia, Kenya, Mozambique, Nigeria, Rwanda, and South Africa. ↑
- A4AI (2022) defines meaningful connectivity as a user having daily access via a smartphone with at least 4G speeds, alongside an unlimited broadband connection at their home, work, or place of study. ↑
- Interviewees warn that the costs are punishing on all sides. Goldstein points to the exorbitant data rates paid by “dumb phone” users (“If you have a dumb phone, [you] pay 85,000 times more for a byte of data than somebody with a smartphone,” – though he later mentioned this obviously varies by country) while Patel and Expert I argue that legacy telecom costs (like SMS and voice) and volatile cloud pricing can actually dwarf the cost of the AI itself. (Patel: “A large part of our cost is SMS, not AI training or inference, compute or database costs. Simply the cost of sending the message. Even though it’s relatively low in Kenya, the costs add up at scale. In other countries across Africa, it’s much higher.”; Expert I: “The challenge is that even though AI hosting costs are artificially low at the moment, they do go up and variants of models disappear quite quickly.”) ↑
- Of course, if one is thinking about building a domestically-hosted LLM, the challenge is even larger because of the power and compute infrastructure required. ↑
- Offline capabilities appear to be particularly well-developed for medical applications. Delft’s CAD4TB program works offline on dedicated hardware that travels alongside the portable X-Ray. ↑
- Klinkenberg: “[Delft’s offline product CAD4TB’s] performance has been validated extensively. Head-to-head comparisons with radiologists show it does as well as senior radiologists with 10 years of experience.” ↑
- Expert H: “It’s as good as human experts . . . . We have AI for knowledge 24/7 offline. Then we have it in the cloud where researchers can improve it and then we also have it as public institutions … the government of Malawi, and then iteratively we can just kind of think of ways in which we can work collaboratively on this to improve it going forward.” ↑
- Bopp: “[0-LA] doesn’t search the whole internet. So it doesn’t search basically a lot of trash that you don’t need. It just uses your specific information. And because this device is offline or works offline, it’s not hackable . . . [and] one of the things that we hardcoded into this LLM is that when it doesn’t know the answer, it tells you. It doesn’t make up the answer, which is one of the issues with a lot of LLMs.” ↑
- Expert N: “There is some innovation happening with small models and on edge but they’re always going to be worse and trailing. I just think it’s going to be a harder path forward.” ↑
- Aggarwal: “If you look at the accuracy rate on low resource languages and use cases in the Global South, it’s like 40-50%, even on the proprietary models . . . accuracy rates are pretty poor.” ↑
- We thank Expert N for this helpful point. Specifically, she mentioned that in contrast to a patient in the US who might have to wait a week to see a human doctor, in some of the settings we care about, many might—in the absence of an AI diagnosis tool—never see a doctor at all, or they might speak to someone who is “deeply misinformed.” ↑
- Expert H: “You can run I think it’s about 30 million parameter models quite easily on the MacBook [Pro].” ↑
- However, perhaps such a model looks promising on a back-of-the-envelope calculation (BOTEC) basis in settings, where users operate in isolation in remote locations (such as community health workers [CHWs] in certain contexts). ↑
- Expert H brought up the question of theft in the context of solar panels used to power servers, but we think this is likely to be an even larger problem for phones, which have an infamous black market “We have to weld this to the floor . . . . We have to get the solar panels. We have to get the roofers. And we have to put it on the wall with welding so people don’t steal it.” ↑
- For example, one option here is to improve market access through initiatives that lower the cost of hardware, either through bulk purchases or by lobbying to reduce taxes, which according to some sources comprise up to a third of the cost (Namanya, 2025). We found one example of GSMA (Warwick, 2025) undertaking these types of activities, but we do not know how successful these initiatives might be, nor a good way to judge them. ↑
- Bopp: “You create your own mini internet . . . . We made a power module, which is basically [a] relative[ly] small-scale power battery with a solar panel ” Bopp clarified that while they might have something like 100 users, “the capacity is scalable far past that.” (Correspondence on June 15, 2026). ↑
- Aggarwal: “I think for equality in access to AI, [offline access] is really critical . . . . So, for example, a Raspberry Pi that you can set up inside the school that all of the devices can ping whenever there is no connectivity.” ↑
- Miles: “There is some innovation happening with small models and on edge, but they’re always going to be worse and trailing and I just think it’s going to be a harder path forward.” ↑
- Aggarwal: “We looked for examples of actual use cases where edge AI was being used in these contexts in education and it’s pretty limited right now.” ↑
- Bopp: “[0-LA] device will market for $500 . . . a project that came at Amazon AWS that they did was called Snowball Edge . . . but the cost the operating cost alone if you wanted to use this the monthly user fee was $55,000 a month, that’s without the hardware purchase.” ↑
- The number of CHW in Nigeria is around ~71K (Oseni et al., 2024; we round down). The annual cost of connection is $1860 (600 + 12*100). ↑
- We treat the misalignment between AI builders and decision makers as distinct from misalignments between AI builders and end-users, which includes frontline workers and beneficiaries. ↑
- Expert: “Very rarely do [Indian NGOs] get the end user into the equation . . . one of the first things which I do when I get on a call with someone trying to build technology [is ask] who are you doing this for and why?” ↑
- Goldstein: “The government of India reached out to us because somebody’s trying to sell them an AI product and they don’t know if it’s good or not.” ↑
- Yanigazawa-Drott: “[A key barrier is] access to data…the only way to do something at scale is you often need to build relationships with the C suite people, or the people that own an organization or a minister or a president” ↑
- Aggarwal: “Coming into…these countries where it’s just so much more about relationships also . . . we’re in a southern part of Ghana where we’re working with government schools. It was all about the professor who already had relationships with a subset of schools so he could get us in.” ↑
- Gandhi: “One thing which is related to the government policy stuff around distribution. So how do you actually get this stuff out to these people at scale? And that’s where for us the telco partnerships have been good, but you could expand that to even the government in-person networks, the satellite people, [or] the IVR [Interactive Voice Response] networks.” ↑
- It is notable that there are few clearly strong solutions to the third dimension of this barrier (the need to build relationships of trust with decision makers). Our sense is that there is no easy shortcut to this issue, with these relationships taking time to build. Among organizations we spoke to that had successfully deployed AI applications in LMICs, many had run established interventions before integrating an AI system. We expect one reason is that many had existing relationships with decision makers. This suggests that programs directly supporting builders may have more success when selecting beneficiaries with these kinds of relationships. On the other hand, favoring these organizations may encourage a feedback loop resulting in established organizations capturing the bulk of resources and making it difficult for new organizations to break through. ↑
- Yanigazawa-Drott: “We’re co-organizing an event [the AI Impact summit]. That’s for networking, that’s for building this context and finding the policymakers that want to be champions.” ↑
- This includes Full Fact, Crisis Cognition, Jacaranda Health, and Rising Academies. ↑
- Expert J: “One of the first things which I do when I get on a call with someone trying to build technology [is ask] who are you doing this for and why? . . . . The third thing we are pushing on our end is getting a product person as soon as possible.” ↑
- Taking the typical salary of CTO in India = ~$60,000 per annum; 3–9 month secondment (taking 6 months as midpoint); adding 20% administrative costs. This comes to $36,000 for one placement. ↑
- Aggarwal: “We were also sharing a lot of learnings between the seven of us [partners in LEVI program] . . . a lot of that knowledge sharing I think in building Rori came from this program and from the funding but also the really strong kind of sharing of resources…as well as knowledge . . . as a philanthropy it’s important to kind of like if you’re funding an accelerator I think it’s really important to identify organizations that are not directly competitive.” ↑
- Expert L: “We struggle to find the right technical expertise while maintaining internal pay equity. Because we operate in regions with lower compensation norms, the global market rate for an AI engineer vastly exceeds what we pay our local staff. It creates a massive parity problem. There were many times I wanted to hire someone, but knew our maximum offer would fall so far below their market value that neither party would be satisfied.” ↑
- Expert J: “Having either these [advisory] entities reside within the foundations, which is what Patrick McGovern has done, where they have a fair bit of in-house AI capability, is one way of addressing [grantee support].” ↑
- One key challenge for governments concerns the effective regulation of AI applications. We address this challenge in the section on local regulatory barriers, and therefore do not discuss this here. ↑
- Goldstein: “The government of India reached out to us because somebody’s trying to sell them an AI product and they don’t know if it’s good or not.” ↑
- This is an analogous example, as it focuses more on broader government data strategies than the applications of AI specifically. ↑
- Some existing processes have been introduced to speed up regulatory approval in these contexts: for example, new health products (in vitro diagnostics, medicines, vaccines and immunization devices, vector control products, and inspection services) can go through the process of WHO pre-qualification to reduce assessment time by roughly two-thirds compared to a full approval pathway (Hodges et al., 2022, p. 13). Once pre-approved, products may choose to go through the WHO Collaborative Registration Procedure (WHO, 2024), which aims to issue a decision on national approval within 90 days. ↑
- Corso: “If you take tools that kind of fit the regulatory framework today—diagnostic devices, image based—these fit very easily within the framework of software as a medical device . . . . the things that don’t fit the existing regulatory framework are those tools which have not which don’t have a single intended use, so [for example] LLM chatbots.” ↑
- For example, the South African Health Products Regulatory Authority (SAHPRA) introduced specific regulatory processes for AI/ML-enabled devices (SAHPRA, 2025) and software lifecycle processes in September 2025, one of the first African countries to do so. ↑
- For example, one new challenge with AI tools is that models are adaptive by design, and thus require constant monitoring to ensure that algorithmic drift does not materially change the product from the approved version. ↑
- Corso: “One of the big challenges as well is linked to the market incentives because when a company wants to enter a market with a vaccine [or] a drug in sub-Saharan Africa . . . . I’ve spoken to regulatory experts who say that even if they obtain WHO pre-qualification, they still have to do things within each regulatory body or for each country. And so the markets themselves aren’t big enough to justify them doing all these extra trials and extra proof.” ↑
- Bopp: “Across our work with Crisis Cognition and Disaster Tech Lab, we consistently see a combination of regulatory, operational, and trust-related constraints that slow or block adoption, [including]: restrictions on data export or prohibitions on cloud use in many countries and jurisdictions; low institutional trust in externally hosted systems during high-risk situations; concerns around handling sensitive medical, protection, or population-level data; operational risk from cyber intrusion or infrastructure disruption, particularly in environments where connectivity is unreliable. For these reasons, offline and edge-based AI solutions provide a fundamentally different deployment profile. They allow organizations to process information locally without transmitting raw data outside the operational environment, which materially reduces perceived and actual privacy exposure and substantially accelerates approval processes.” ↑
- One way to sidestep potentially complex data governance and sovereignty concerns is to use offline AI systems that “provide a fundamentally different deployment profile,” because they “materially reduce perceived and actual privacy exposure and substantially accelerate approval processes,” according to Evert Bopp. See more about our take on this here. ↑
- This has been a major barrier to the success of the AU Model Law, and the target of 25 countries adopting the law by 2020 was missed (Ncube et al., 2023, p. 12). ↑
- Specifically, India, Zambia, Indonesia, Brazil, and Vietnam. ↑
- In this case, by “identifying ten companies that have innovative products”—specifically, AI mental health tools—but “don’t have a clear pathway to approval.” ↑
- Specifically, the East African Community Medicines Regulatory Harmonization (EAC-MRH) Programme. ↑
- Energy, agriculture, health, water, education and training, and infrastructure. ↑
- Italy has committed to an annual budget of €5 million through 2028 (Decode39, 2025). ↑
- If we had more time, we would have liked to have spoken to someone at the Hub to gain a better understanding of progress and potential blockers. ↑
- A technical assistance organization helping AI deployers with: (1) regulatory approval navigation (both for WHO pre-qualification and national approval), (2) evaluation design (conducting RCTs or using quasi-experimental methods to generate robust, peer-reviewed evidence), (3) privacy and intellectual property (IP) compliance frameworks, and (4) evidence synthesis for policymakers. ↑
- Rather than individual country approvals, regional bodies (e.g., the African Union, East African Community) could establish a single validation study and regulatory submission to access multiple markets simultaneously, improving economics and market incentives for companies looking to expand their product outside of an existing market. ↑
- Prize competitions, such as Renaissance Philanthropies’ “Tools Competition” for AI tools for education, sit in a gray area in this framework (Renaissance Philanthropy, 2026). These provide financial reward for winners, and in this sense resemble traditional grantmaking. However, they also incentivize innovation among organizations other than the eventual winner, and so can have a fieldbuilding function. While we don’t discuss it in our discussions elsewhere of fieldbuilding activities, we think that it can be a strong model for grantmaking to individual implementers. ↑
- We have some uncertainty over whether this definition is most helpful, however. ↑
- Note that the total for the fieldbuilding activity types in Table 10 exceeds 16, since several initiatives involve multiple fieldbuilding functions. ↑
- Expert I told us that the civic tech space in particular is already experiencing a “total crisis” in funding due to its political nature. ↑
- Goldstein: “I think right now there’s a lot we’re in a declining aid budget. So if we think about official spending, I do think there’s a disproportionate amount going to AI in the sense that everybody thinks it’s the new shiny thing. And I think the shine will wear off very quickly.” ↑
