Search

Where Better Data Matters Most. Prioritizing Disease Burden Data Gaps in Sub-Saharan Africa

Share with
Or share with link

Executive Summary

Disease burden estimates in Sub-Saharan Africa rely on weak empirical foundations

Every year, billions of dollars in global health funding are allocated based on disease burden estimates that, in much of Sub-Saharan Africa (SSA), rest on little or no local data. For many conditions, what looks like a precise figure is largely statistical modeling in the absence of real empirical inputs. That uncertainty means resources may be flowing to the wrong places and the wrong problems. This report asks: where would investments in better data matter most?

We built a framework to prioritize data gaps across countries and conditions

We scored every combination of 49 countries and 24 high-burden conditions across three dimensions: disease burden, evidence strength, and feasibility of improvement. Our approach identifies priorities: first by using a priority filter to flag the biggest problems we can realistically address, and then by applying a regression-based approach to identify the low-hanging fruit—where existing capacity is not being fully used.

The findings point to both regional and targeted opportunities

Fifty-seven country-condition pairs meet our criteria for priority data investment. Across both approaches, a consistent picture emerges:

  • Diarrheal diseases and neonatal disorders stand out above everything else—high burdens and weak evidence across regions, and flagged by every analytical approach.
  • Priority conditions cluster into two types—event-based conditions that happen outside health facilities (e.g., road injuries and drowning), and conditions that lack dedicated surveillance infrastructure (e.g., neonatal disorders and diarrheal diseases). Both tend to fall through the cracks of routine data collection.
  • Road injuries are a regional problem, not a country one—consistently flagged as high-priority but with uniformly weak data everywhere. No single country stands out as the place to start, pointing to a need for coordinated regional action.
  • Burundi, Rwanda, and Benin show systemic gaps across multiple conditions—suggesting that broad data system strengthening in these countries could have returns across many conditions simultaneously.
  • East Africa is the clearest low-hanging fruit—Burundi and Rwanda have relatively high capacity but surprisingly weak data, suggesting targeted investment could yield quick returns across multiple conditions.

A starting point for funders and policymakers

This analysis is designed to generate hypotheses and support early-stage priority-setting, helping narrow down a large landscape of potential data gaps to a manageable shortlist of contexts worth investigating further. Results should be treated as indicative rather than definitive. They are most useful when scoping where to focus, comparing opportunities across countries or conditions, or checking whether a specific context of interest has particularly weak data relative to its burden. We recommend supplementing findings with local knowledge and expert input before directing funds.

Several caveats are worth keeping in mind

First, disease burden estimates rest on weak underlying data in many SSA settings, which means our prioritization could shift if true burden differs substantially from estimated burden. Second, our “Feasibility score” is a rough, exploratory measure that lacks external validation. Third, this analysis identifies where data gaps are largest and most addressable, but does not prescribe specific data collection approaches; determining the most appropriate intervention in each context requires further investigation.

High Stakes, Weak Data: Where Should We Invest to Improve Disease Burden Data in Sub-Saharan Africa?

Why Prioritizing Data Gaps Matters for Health Decision-Making

Disease burden estimates, such as those produced by IHME’s Global Burden of Disease (GBD) study, are widely used to guide global health priorities (see Appendix A for a brief overview of how GBD estimates are constructed). Funders, researchers, policymakers, and implementers rely on these estimates to compare health problems and decide where to allocate resources.[1] However, the strength of these estimates depends heavily on the availability and quality of underlying empirical data.

In Sub-Saharan Africa (SSA), even burden estimates for comparatively well-funded diseases such as malaria[2] remain highly uncertain, with substantial variation across models. For example, GBD’s 2023 estimate of malaria mortality in SSA spans a wide 95% uncertainty interval, ranging from roughly 250,000 to 1.2 million deaths.[3] Substantial differences also arise across commonly used modeling approaches. As highlighted by Noor (2018, p. 3), malaria case estimates based on adjusted routine surveillance data are in some settings about four times higher than estimates derived from household surveys. This level of uncertainty complicates prioritization and increases the risk of misallocating scarce resources.

A key reason for this uncertainty is that disease burden estimates are only as strong as the underlying data. Improved modeling cannot fully compensate for sparse, outdated, or poor-quality empirical inputs. In SSA, these data limitations reflect persistent structural constraints, including incomplete civil registration and vital statistics systems, limited diagnostic and laboratory capacity, and infrequent or non-representative population-based surveys (Mbondji et al., 2014; Alegana et al., 2020; Kudymowa et al., 2026). Because these constraints differ in severity across countries and conditions, the strength of evidence supporting burden estimates is uneven rather than uniformly weak.

These challenges are unlikely to resolve in the near term. Progress in strengthening routine data systems has been slow, and recent reductions in external funding for surveys and data infrastructure, including from major bilateral donors,[4] have further strained already fragile information systems (Kudymowa et al., 2025). Given limited resources, strengthening data systems everywhere at once is not feasible. A central question for funders and decision-makers is therefore not simply where data gaps exist, but where investments in better data collection should be prioritized to strengthen health decision-making.

This report systematically compares data gaps, disease burden, and improvement feasibility across Sub-Saharan African countries and conditions[5] to answer that question. It enables funders to identify which country-condition combinations offer the best return on data investments, helps national health ministries see where their evidence base is weakest relative to burden, and reveals whether data gaps are isolated problems or systematic challenges affecting multiple countries.

Moving from Documenting Data Gaps to Prioritizing Them

Existing research has documented many challenges in health data availability and quality across SSA. Some focus on data gaps for specific conditions such as malaria (Alegana et al., 2020), cancer (Bray et al., 2022), and cardiovascular diseases (Kariuki et al., 2015). Others examine particular components of health information systems, including civil registration and vital statistics (Yokobori et al., 2021), routine facility data (Hoxha et al., 2020), and household surveys (Seidler et al., 2025). Additional studies take a country-level perspective, describing data constraints within individual national contexts (Rumisha et al., 2020; Chekol et al., 2023). While these analyses provide important insights, they rarely compare data gaps across both countries and conditions, or provide guidance on how limited resources for data improvement should be allocated.

In practice, decisions about investing in better data must weigh several considerations at once: the scale of the health problem, the strength of existing data, and the feasibility of improvement. Focusing solely on disease burden would direct resources toward the largest health problems regardless of whether good data already exists or whether improvement is feasible in that context. Conversely, focusing only on data gaps would prioritize conditions with weak evidence even if their health impact is minimal. Emphasizing feasibility alone would favor easy wins over high-impact opportunities. A framework that balances these considerations helps identify contexts where data investments can be both high-impact and achievable.

This report addresses that gap by systematically examining three key dimensions:[6]

  • Burden: the scale of disease burden
  • Evidence Strength: the strength of the empirical evidence base underlying burden estimates, assessed through indicators of data availability, recency, and quality
  • Feasibility: the capacity to improve the evidence base, based on country context and diagnostic complexity

Focusing on the highest-burden conditions in SSA, we identify country-condition pairs where all three dimensions align: substantial health problems, weak existing data, and realistic prospects for improvement. The aim is not to produce definitive rankings, but to offer a transparent starting point for guiding targeted investments in better health data.

Data and Approach

The GBD is among the most comprehensive compilations of health burden data across low- and middle-income countries (LMICs) and applies consistent definitions, methodologies, and metrics across all regions. This makes it possible to compare disease burden and the availability and quality of data across countries and conditions within a region. We focus on SSA, covering 49 countries, and include the top 30 conditions by disability-adjusted life year (DALY)[7] burden based on GBD 2023 estimates. Together, these conditions account for roughly 75% of the total DALY burden in the region (see Appendix B).[8] Six conditions in this set are excluded from the evidence strength assessment for methodological reasons described below. The remaining 24 conditions form the basis of our analysis.

To operationalize our prioritization framework, we combine disease burden estimates with metadata on empirical input sources, as well as a set of country- and condition-level indicators capturing the feasibility of improving evidence on disease burden.

Measuring Burden, Evidence Strength, and Feasibility

Burden

In our primary analyses, disease burden is measured using total DALYs from GBD 2023 to enable prioritization of conditions where the absolute number of affected individuals—and thus the potential scope for intervention—is greatest. We also conducted all analyses using DALY rates per 100,000 population to enable comparison across countries of different population sizes. Focusing on total burden shifts priorities to countries with greater population size, but this is consistent with our analytical focus on the absolute scope of disease impact; a condition affecting several million at a modest rate may offer a better opportunity for improved evidence collection than one with a high rate in a small population.[9]

Evidence Strength

Evidence Strength reflects the quality and availability of data underlying disease burden estimates. We combine information from two sources: metadata from the GBD 2023 Sources Tool, which documents the empirical studies informing each burden estimate, and IHME’s cause-of-death data quality ratings, based on the star rating system reported in the GBD 2023 Causes of Death analysis (GBD, 2025; GBD, 2023, Fig. S5).

GBD burden estimates are based on two types of empirical data: cause-of-death data and non-fatal outcome data, which serve to estimate the two components of DALYs: years of life lost (YLLs) and years lived with disability (YLDs), respectively. For this analysis, we restrict the evidence strength assessment to cause-of-death input data. Two considerations motivated this choice. First, IHME provides standardized star ratings only for cause-of-death data quality, allowing a consistent measure of empirical reliability across countries and conditions. Comparable quality indicators are not available for non-fatal outcome inputs, which would require additional subjective assessment. Second, YLLs account for 84% of total DALY burden across the top 30 conditions in SSA, with YLDs accounting for only 16%.[10] Focusing on cause-of-death data, therefore, captures the vast majority of overall burden while maintaining a transparent and comparable assessment framework.[11]

From these sources, we extract three variables capturing different aspects of the empirical foundation behind burden estimates:

  • Number of input data sources, derived from the GBD Sources Tool and used as a proxy for the volume of empirical evidence.
  • Recency of available data, based on the most recent year of underlying data collection recorded in the GBD Sources Tool, used as a proxy for how up-to-date the empirical evidence base is.
  • Cause-of-death data quality, based on IHME’s country-level star ratings, a scale from zero to five reflecting the completeness of vital registration systems and the proportion of deaths assigned to well-defined conditions.[12]

Each variable is converted into a 0-5 point scale as shown in Table 1 below:

 

Table 1: Scoring system for Evidence Strength components

ScoreNumber of sources[13]Recency[14]Data quality
0No dataNo data0 stars
11–2 sources≤ 20001 star
23–5 sources2001–20092 stars
36–11 sources2010–20173 stars
412–24 sources2018–20214 stars
5≥ 25 sources≥ 20225 stars

These three components are then summed to produce a single Evidence Strength score ranging from 0 to 15, with higher values indicating a stronger empirical evidence base and fewer data gaps.

We exclude subnational entries to avoid double-counting, as spot checks indicate that these typically represent disaggregations of national datasets rather than distinct empirical sources.[15]

Feasibility

Feasibility captures the relative difficulty of improving the evidence base on the burden of a particular condition in a particular country. A higher score reflects greater feasibility or ease of additional data collection across multiple dimensions. In our search of the literature, we did not come across an existing attempt to generate this kind of score. We therefore constructed our own score, based on country-level and condition-level characteristics. Our process for selecting country-level indicators is described in Appendix C.

Our Feasibility indicator is made up of four equally weighted components. Three of these four components are based on country-level data: Government Effectiveness (part of the World Governance Indicators), Healthcare Access and Quality Index (an analysis of the IHME GBD 2019 study), and the Rural Access Index (designed by the World Bank).[16] The fourth component is a five-point indicator reflecting the requirements for diagnosing each condition.[17] Conditions with a possible “epidemiological diagnosis”—those in which death results from a specific, discrete external occurrence (e.g., road injuries or drowning), rather than an internal disease process or gradual physiological decline—were considered easiest to diagnose.[18] Those in which clinical diagnosis—i.e., via assessment by a trained healthcare professional—is sufficient (without any form of follow-up testing) were considered next easiest to diagnose, followed by those in which additional testing is useful, but not required for diagnosis. The conditions with the lowest diagnostic ease are those requiring a test, which is further subdivided into those with the availability of an acceptable point-of-care (POC) test and those that require a laboratory- or hospital-based test (considered the hardest to diagnose). Details of the specific ease of diagnosis indicator assessments by condition are available in Appendix D.

 

Table 2: Indicators used to construct Feasibility score

Sub-indicatorDimensionDescriptionData yearReason for inclusion
Government effectiveness (World Governance Indicators)CountryThis captures perceptions of the quality and efficiency of a country’s government and civil service.2024Data collection will generally require the approval or active participation of government bodies.[19] Inefficiency in government can therefore inhibit or slow data collection.
Rural Access Index (World Bank)CountryMeasured as the share of the rural population living within 2 kilometers of an all-season road.2019Poor connections to rural areas make data collection substantially more difficult logistically.[20]
HealthCare Quality and Access Index (IHME)CountryCombines mortality-incidence ratios and risk-standardized death rates to reflect mortality amenable to healthcare access and quality. Higher scores reflect stronger health systems with greater capacity to prevent avoidable deaths.

We used the age-standardized index to capture healthcare system strength broadly.

2019Higher quality health system strength will reflect greater numbers of trained health workers, wider clinic availability, and better logistics and organization for healthcare delivery. These features will also help to facilitate data collection on health burden.[21]
Ease of diagnosisConditionOur assessment of the ease of accurate data collection for each condition.[22] We used a five-point scale:
  1. Epidemiological diagnosis (easiest to diagnose): sufficient event-based information for diagnosis
  2. Clinical diagnosis: clinical diagnosis is sufficient without follow-up testing
  3. Testing is useful for diagnosis: testing is useful, but not required, for diagnosis
  4. Testing is required for diagnosis: testing is required for diagnosis, with a POC test sufficient
  5. Laboratory- or hospital-based testing is required for diagnosis: facility-based testing is required for diagnosis; a POC test is either unavailable or insufficient (hardest to diagnose)

N/AConditions that require medical training to diagnose, or laboratory testing to confirm, demand greater resources for data collection and, in some cases, may be largely unavailable in LMICs.

The three country-level components were normalized from 0 to 1 using min-max scaling. Diagnostic difficulty was placed on a 0 to 1 scale, with easier-to-diagnose conditions assigned a maximum of 1 and high-difficulty causes assigned a minimum of 0; this ensured that for all indicators, a high score reflects greater feasibility. All four indicators were then combined with equal weight to produce our score, on a scale from 0 to 4.

Approach to Identifying Priority Data Gaps

We use a two-part approach to identify priorities.

First, we use a priority filter to identify absolute needs—contexts where major health issues coincide with weak data in settings where improvement is feasible. The priority filter analysis asks: “What are the biggest problems we can realistically address?”

From there, we construct relative priority scores to identify relative gaps and potentially low-hanging fruit—contexts where data is surprisingly weak given the feasibility of data collection. This analysis asks: “Where is existing capacity not being fully applied and where should we care about this the most?”

In sum, the priority filter provides a menu of country-condition pairs that warrant further investigation, from which our relative priority ranking offers one plausible way to rank these gaps in order of importance.

Priority Filter: Identifying High-Impact Investment Opportunities

We first use a filter to identify country-condition pairs representing clear priority data gaps. The framework is built around combinations of criteria: high burden, poor data, and feasible context must all be present, and no single dimension can compensate for an exceedingly low score on another. Specifically, we flag pairs that meet all three of the following criteria:

  • Burden: ≥100,000 total DALYs[23]
  • Evidence: strength ≤5 out of 15
  • Feasibility: ≥2 out of 4

This approach is visualized in Appendix E, with bubbles in gold representing country-condition pairs meeting all three criteria.

The Burden threshold of ≥100,000 total DALYs captures approximately the top half of country-condition pairs by burden. The Evidence Strength threshold of ≤5 represents the lower third of our scale. The Feasibility threshold ≥2 represents the midpoint. We also provide an interactive tool for exploring how results change under alternative filter specifications, and tested two specific alternative scenarios—one stricter, one more lenient—to assess sensitivity to threshold choice.

Rather than identifying a single “best” country-condition to focus on, this approach yields a menu of candidate priorities—all meeting minimum standards for burden, data quality, and feasibility. Funders and policymakers can select from this menu based on their specific mandates, geographic focus, or technical capacity.

We also analyze patterns by counting how frequently each condition and country appear among pairs meeting all criteria. High frequency can indicate opportunities for coordinated, multi-country interventions. For instance, a condition appearing across many countries may benefit from standardized data collection protocols that can be deployed broadly.[24] Pairs appearing less frequently are not necessarily lower priority, but may be better suited to targeted, context-specific investments.

Relative Priority Scores: Identifying Potentially Low-Hanging Fruit

To add specificity to the priority filter analysis, we then use a regression-based approach to identify contexts where evidence is weaker than expected. If we know how strong evidence typically is for a given Feasibility score, we can then flag country-condition pairs that fall short of that benchmark.[25] By regressing Evidence Strength on Feasibility, the residuals of the model reveal whether a given pair has better or worse data than we would ordinarily predict, making them a natural tool for revealing surprising data gaps rather than merely expected ones.

Second, we recognize that our Feasibility score may not capture all of the inherent variation in the ability to collect data. A country may face bureaucratic or infrastructural challenges that our score does not fully reflect, or a condition may be systematically harder to survey for reasons that go beyond our estimate of diagnostic ease. To account for these variations, we also use a regression model that includes both country and condition fixed effects.[26] The residuals from this model capture the gap between a country-condition pair’s actual Evidence Strength score and what we would expect given that country’s typical data quality and that condition’s typical surveillance challenges. This highlights pairs where data gaps continue to be surprising, even after accounting for these broader patterns. The fixed effects approach has the advantage of absorbing a wide range of factors that our Feasibility score may miss, but because so much variation is soaked up by the fixed effects themselves, the model also makes it difficult to draw general conclusions about any particular country or condition in isolation.

For pairs with worse data than expected, we then multiply residuals by burden to create relative priority scores.[27] Large scores indicate substantial disease burden combined with unexpectedly weak data. This may represent low-hanging fruit where existing capacity has not been fully applied—but only where country capacity is sufficiently high to support improved data collection. Where Feasibility scores are low, the unexpected weakness in evidence may reflect structural constraints rather than untapped potential. To address this, we apply our priority filter to capture only those pairs that meet our criteria for high burden, low evidence strength, and high feasibility. This approach is summarized in Figure 1 below, and equations for the regression models are available in Appendix G.

 

Figure 1: Methodology of constructing relative priority scores and priority data gaps[28]

Data Gaps are Widespread but Vary Across SSA

Before identifying priority data gaps, we describe patterns in Evidence Strength across SSA—gaps that are both widespread and uneven, and that motivate the systematic prioritization that follows.

First, data availability is poor across high-burden conditions in many SSA countries (see Figure 2). Equatorial Guinea, Eritrea, and South Sudan had the weakest evidence bases, with 71% of the top 24 highest-burden conditions scoring zero across all three measures of data quantity and quality, followed by the Central African Republic, Djibouti, and Lesotho (each scoring zero across 67% of all conditions). Poor Evidence Strength correlates moderately with smaller population size (r = 0.34), though this does not explain most of the variance across countries.

 

Figure 2: Evidence Strength by country in SSA, averaged across 24 top high-burden conditions

Second, data gaps are also widespread at the condition level—one third of high-burden conditions have no local data available across at least half of SSA countries. Invasive Non-typhoidal Salmonella (iNTS) had the least data available overall,[29] with 78% of SSA countries having zero local data sources, followed by ischemic heart disease (71%), idiopathic epilepsy (67%), and diarrheal diseases and stroke (both 63%).

Third, chronic conditions generally appear to have less available data than high-burden infectious diseases. HIV/AIDS, malaria, and tuberculosis benefit from dedicated international funding and monitoring programs, which we suspect contribute to their comparatively stronger evidence base. We are not aware of equivalent infrastructure for most other conditions in SSA.

Burden and Feasibility show similarly uneven distribution across SSA (see Appendix H for country-level maps of average scores). Together, these patterns illustrate why systematic prioritization is needed: data gaps are widespread, but the contexts where high burden, weak evidence, and realistic prospects for improvement align are fewer—and it is these contexts that our analysis identifies.[30]

Where to Prioritize Data Investments in Sub-Saharan Africa

The Biggest Problems We Can Realistically Address

By using our priority filter, we identified 57 country-condition pairs (out of 1,176) where three criteria align: Burden is substantial (≥100,000 DALYs accrued annually), the data supporting burden estimates is weak (Evidence Strength ≤5), and the context makes data improvement realistic (Feasibility ≥2). Figure 3 shows these pairs, with larger bubbles indicating higher burden and darker red showing weaker Evidence Strength scores.[31]

 

Figure 3: Priority-filtered country-condition pairs meeting minimum thresholds for Burden, Evidence Strength, and Feasibility (by total DALY burden)

Diarrheal Diseases, Neonatal Disorders, and Road Injuries Dominate the Priority Landscape

  • Diarrheal diseases is the condition category that appears most frequently. It appears across 10 of the 16 countries that have at least one condition meeting the priority criteria. The burden is moderately high across all flagged countries, with West and East Africa standing out as regions with particularly poor data and substantial burden.
  • Neonatal disorders and road injuries also stand out as priority conditions.[32] Cameroon and Côte d’Ivoire face particularly high neonatal burdens with weak data, suggesting that neonatal disorders represent one of the most significant unmet data needs across all high-burden conditions in SSA.[33]
  • Burundi, Côte d’Ivoire, and Rwanda stand out as countries with a broad range of flagged conditions across multiple disease categories, followed by Cameroon, Benin, and Uganda. This suggests systemic data gaps within each country, rather than condition-specific ones. For example, investment to improve the vital registry system in these countries could potentially have broad-scale effects across multiple high-burden conditions.
  • A majority of conditions appear more sporadically, flagged in only a handful of countries. This suggests specific gaps that might be filled by more targeted interventions to collect better data on a given condition in a given location.

Priority Conditions Tend to Lack Dedicated Surveillance Programs and Often Occur Outside Facilities

Several characteristics may help explain why these particular conditions emerge as priorities:

  • Most lack dedicated surveillance infrastructure. Unlike malaria, HIV/AIDS, and tuberculosis, which benefit from vertical programs with active case finding and reporting systems, we suspect that conditions like diarrheal diseases, neonatal disorders, and road injuries rely primarily on vital registration systems.[34] When these registration systems are weak, data gaps emerge.
  • Many involve events outside the formal health sector. Health facilities serve as the primary source of routine vital statistics in much of SSA, but passive surveillance systems struggle to capture events occurring outside this setting. Neonatal deaths and drowning most often occur entirely outside any facility (Price et al., 2019); diarrheal deaths frequently follow delayed or absent care-seeking (Marew et al., 2025); and road injuries and interpersonal violence may reach facilities but remain severely undercounted at the population level (Mwale et al., 2023; Van der Walt et al., 2025). In all cases, the gap between what facilities capture and what is actually happening in communities is substantial and in need of improvement.

Priority Countries are Concentrated in West and East Africa

The countries appearing most frequently—Rwanda, Burundi, Côte d’Ivoire, Uganda, Benin, and Cameroon—primarily span East and West Africa. They tend to be small-to-medium population countries rather than the region’s largest, meaning that a high absolute burden is not the only factor driving the results. Beyond meeting our Feasibility threshold by design, we have not identified any clear common characteristic. For countries like Burundi and Rwanda with gaps across diverse disease categories, strengthening core vital registration infrastructure could potentially improve data across multiple conditions simultaneously.

The Low-Hanging Fruit: Where Capacity is Not Being Used

Regressing Evidence Strength on Feasibility yields 58 country-condition pairs flagged as priority data gaps, of which a subset of 40 pairs are prioritized when accounting for country and condition characteristics.[35] Table 3 shows the prioritized ranking of the top 20 country-condition pairs from the regression model without fixed effects (visualized in Figure G.1).[36] The equivalent table with results from the fixed effects regression is shown in Table G.1. In both models, our priority filter is also applied to capture only pairs with high burden, surprisingly poor evidence, and realistic feasibility for collecting more data.

The residual data gap quantifies the magnitude of the difference between the actual Evidence Strength score and the predicted Evidence Strength score. The relative priority score itself has no units: it only represents a relative magnitude used to rank data gaps in order of priority, and is designed to scale proportionally with burden—such that an equivalent data gap in a condition with twice the burden is weighted twice as heavily. If the residual data gap is small, this suggests that data availability is not much worse than expected, but if the burden is high, we should place greater weight on resolving that data gap. Conversely, if the burden is moderate but the residual data gap is substantial, that also represents a potential opportunity to meaningfully improve the evidence base. Rankings should therefore be treated as a useful starting point for thinking about relative priorities, rather than a definitive guide about where to direct funds specifically.

 

Table 3: Top 20 ranked country-condition pairs from regressing Evidence Strength on Feasibility

CountryConditionEvidence StrengthFeasibilityResidual data gapBurdenRelative priority score
CameroonNeonatal disorders02.17-5.312.0M10.4M
Côte d’IvoireNeonatal disorders12.18-4.332.3M10.1M
BurundiNeonatal disorders02.35-5.621.1M6.2M
CameroonDiarrheal diseases02.17-5.311.1M5.7M
BeninNeonatal disorders12.35-4.62880,0004.1M
RwandaNeonatal disorders12.81-5.42640,0003.5M
UgandaDiarrheal diseases12.47-4.83560,0002.7M
Côte d’IvoireDiarrheal diseases12.18-4.33620,0002.7M
BurundiDiarrheal diseases02.35-5.62440,0002.5M
CameroonCongenital birth defects02.17-5.31440,0002.3M
TogoNeonatal disorders02.23-5.42430,0002.3M
Côte d’IvoireCongenital birth defects12.18-4.33500,0002.2M
BurundiLower respiratory infections22.10-3.19610,0001.9M
BurundiRoad injuries22.60-4.06440,0001.8M
LiberiaNeonatal disorders12.04-4.09440,0001.8M
TogoDiarrheal diseases02.23-5.42320,0001.7M
RwandaRoad injuries43.06-2.85480,0001.4M
BurundiCongenital birth defects02.35-5.62240,0001.4M
RwandaDiarrheal diseases12.81-5.42240,0001.3M
MauritaniaNeonatal disorders02.00-5.02240,0001.2M

Note. Burden = total DALY burden. Evidence Strength = score out of 15. Feasibility = score out of 4. Residual data gap = actual Evidence Strength score minus the expected Evidence Strength score. Relative priority score = absolute value of the burden-weighted residual, i.e., the residual data gap multiplied by the burden. Values are rounded to the nearest 100,000 if ≥1M and nearest 10,000 if <1M. Blue shading indicates variables used in the regression. For residual data gap, burden, and relative priority score, coloring indicates magnitude, with darker red colors indicating worse data, a higher burden, and higher overall priority, respectively. Country-condition pairs are ranked in order of the largest relative priority score.

Though very high-burden conditions (those with >1M DALYs annually) tend to rise to the top of the country-condition pair ranking, pairs with the highest relative priority scores have both surprisingly poor data and high burden.[37] When comparing pairs across both regression models, several other themes emerge:

Prioritized Rankings Highlight Neonatal Disorders and Regional Opportunities in East Africa

  • Neonatal disorders rises to the top priority condition category. Similar to results using our priority filter, neonatal disorders account for roughly a third of the country-condition pairs flagged as high priorities in both regression models. Specific highest-priority opportunities for data strengthening emerge in Cameroon, Côte d’Ivoire, Burundi, Benin, and Rwanda, each driven by a very high burden and substantially poorer evidence than expected. The consistency of neonatal disorders surfacing as a priority data gap across multiple regions in SSA—appearing in both modeling approaches and across multiple analytical methods—provides the strongest evidence that this is a genuine, robust data gap rather than an artifact of any particular methodological choice.
  • Diarrheal diseases and congenital birth defects offer several opportunities for investment. Six specific high-priority countries are identified for data strengthening for diarrheal diseases, evenly divided between East Africa (Uganda, Burundi, and Rwanda) and West Africa (Cameroon, Côte d’Ivoire, and Togo). Congenital birth defects emerge as a priority in three countries, while high-ranking conditions that only appear once or twice (e.g., lower respiratory infections in Burundi) may highlight specific data gaps rather than a consistent pattern by condition.
  • Road injuries drop down the list of priorities, particularly after accounting for country and condition characteristics. This drop is likely because road injuries do not have as much surprisingly low evidence in any particular country, once country and condition fixed-effects are accounted for. Evidence gaps for road injuries appear to be relatively uniform across countries, likely reflecting systemic gaps in data availability across SSA.
  • East Africa presents multiple opportunities for data system strengthening. By country, Rwanda, Burundi, and Côte d’Ivoire appear most frequently across both regression models (six of the top country-condition pairs). This may indicate a regional pattern in which strengthening data systems in East Africa, in particular, could have compound effects on improving the evidence base for multiple conditions. This is further supported by the relatively high Feasibility scores of both countries, suggesting that the barriers to better data are not primarily logistical, and that targeted investment in data infrastructure could yield improvements across several conditions.

Sensitivity Analyses

Burden–Evidence Strength–Feasibility Component Independence

A concern is whether the three dimensions (Burden, Evidence Strength, and Feasibility) measure distinct constructs or whether they are so highly correlated that they provide redundant information. If the latter were true, the framework would add little value over simply ranking by any individual component alone.

We calculated pairwise correlations between the three components across all 1,176 country-condition pairs.[38] Burden and Evidence Strength are very weakly correlated (r=0.15), and Burden and Feasibility are essentially uncorrelated (r=-0.02). Evidence Strength and Feasibility show a weak positive correlation (r=0.30) that aligns with our expectations: countries with stronger data systems tend to have somewhat better governance capacity and health infrastructure. However, the correlation is sufficiently weak that the dimensions capture genuinely distinct information. Data-poor contexts are not automatically infeasible to improve.

Comparison With GBD Uncertainty Intervals

We considered whether our Evidence Strength metric might be redundant with GBD’s uncertainty intervals. Although these intervals are not explicitly designed as data quality indicators, one might expect that weaker empirical data would lead to wider intervals.

We tested this by calculating the correlation between our Evidence Strength scores and the relative width of GBD uncertainty intervals.[39] The correlation is negligible (r=-0.07).[40] Closer examination reveals why: country-condition pairs with the narrowest GBD uncertainty intervals frequently have zero national-level input data.[41] For example, neonatal disorders show narrow intervals in several countries despite minimal or no local empirical data.

These narrow intervals occur because GBD uses covariate-based imputation and statistical modeling to generate estimates in data-sparse contexts.[42] In such cases, the resulting uncertainty intervals reflect model uncertainty (confidence in covariate-based predictions) rather than data availability (whether local empirical observations exist). Thus, a narrow interval may indicate either robust local data or a precisely modeled estimate derived from statistical relationships in the absence of direct empirical inputs. For investment decisions, our Evidence Strength score directly captures whether local empirical data exists, which may be more decision-relevant than model-based uncertainty intervals.

Alternative Priority Filter Scenarios

We considered two alternative priority filter scenarios for Burden, Evidence Strength, and Feasibility[43] to examine how varying these criteria affected the number and distribution of priority candidate countries and conditions.

  • A more restrictive version in which burden must be higher, evidence weaker, and feasibility higher, with the following thresholds applied:
    • Burden ≥250,000 total DALYs
    • Evidence Strength ≤3 out of 15
    • Feasibility ≥2.5 out of 4
  • A more lenient version in which burden can be lower and evidence stronger,[44] with the following thresholds applied:
    • Burden ≥50,000 total DALYs
    • Evidence Strength ≤7 out of 15
    • Feasibility ≥2 out of 4

Results are relatively consistent with our main findings. The restrictive scenario yielded only three country-condition pairs meeting all criteria: neonatal disorders (in Rwanda) and road injuries (in Benin and Burundi). In contrast, the more lenient version yielded 125 country-condition pairs, with the top conditions (in rank order), including interpersonal violence, drowning, neonatal disorders, diarrheal diseases, and road injuries. In the more lenient scenario, four out of five conditions and countries overlap with our main analysis, with only interpersonal violence jumping up the list of priorities substantially.

This sensitivity analysis suggests that our main findings are robust to reasonable changes in threshold criteria, with diarrheal diseases and neonatal disorders emerging as priority condition categories, and Benin, Burundi, and Rwanda as priority countries, regardless of the thresholds applied.

 

Limitations

This analysis was designed to enable rapid cross-condition and cross-country comparison using simple indicators and rough proxies, prioritizing breadth over precision. Findings should be treated as hypothesis-generating rather than definitive, and as a starting point for further validation before informing investment decisions. The two limitations most likely to affect our overall conclusions are:

  • Circularity in burden estimates. DALY estimates were used as a proxy for burden, but these estimates themselves rely on weak underlying data in many SSA settings—if the true burden of a condition is systematically over- or under-represented, our prioritization could shift accordingly. This circularity is difficult to resolve without the kind of improvements in evidence strength that this analysis is designed to motivate.
  • An unvalidated Feasibility score. The Feasibility score is exploratory and lacks external validation; since it determines which pairs make the final priority list, miscalibration here could systematically include or exclude certain contexts. Several specific weaknesses are worth noting: diagnostic complexity was assessed through internal judgment rather than a validated framework, and is inherently harder to classify for conditions encompassing multiple causes of death (e.g., neonatal disorders); three of the four Feasibility components are measured at the country level, meaning all conditions within a country receive the same score regardless of condition-specific data collection challenges; only the diagnostic difficulty component varies by condition; and some indices draw on 2019 data and may not fully reflect current conditions. We therefore recommend treating Feasibility scores with particular caution.

Several additional limitations are worth noting, though we believe these are less likely to affect our overall conclusions:

  • Relative rather than absolute benchmarks. Scores are based on relative comparisons across the countries and conditions included in this analysis, with no gold standard threshold for what constitutes adequate data availability or feasibility for strengthening the empirical evidence base.[45]
  • Nonfatal outcomes not captured. The assessment focuses on cause-of-death data and does not capture nonfatal health outcomes, which may underestimate available information for some conditions. Six conditions whose burden consists entirely of nonfatal outcomes—such as anxiety disorders and low back pain—were excluded from the analysis entirely, as no cause-of-death data exists for them by definition.
  • Evidence Strength components are crude proxies. Publication counts were used as a rough indicator of evidence availability, even though a single high-quality dataset may be more informative than many smaller studies. GBD star ratings similarly reflect only vital registration and verbal autopsy quality, missing other potentially informative sources such as police records for injuries or sibling history for maternal conditions.
  • Spillover effects not captured. Improvements in data infrastructure for one condition may benefit others relying on similar data pipelines, but these cross-condition gains are not incorporated in our analysis.

Conclusions

The empirical foundation for health decision-making in SSA is uneven, with critical resource allocation often relying on modeled estimates based on statistical curve-fitting, rather than local data. This analysis identifies a clear set of priority country-condition pairs where high burden, weak existing evidence, and high feasibility for improvement align. Our results show a number of recurring themes:

  • Critical disease burden estimates lack a solid empirical foundation. Data availability and quality across SSA is often poor to non-existent, even for very high-burden conditions, with many estimates based on no local evidence.
  • Diarrheal diseases and neonatal disorders are clear priorities: Across all analytical approaches, both emerge as the most significant and addressable data gaps. Neonatal disorders are frequently flagged in countries with substantial burden and minimal evidence, while diarrheal diseases appear most frequently under our priority filter criteria. As categorical conditions, subsequent analyses that are disaggregated into specific conditions could also provide more direction on specific investments that could be made.
  • “Out-of-facility” events are systematically under-evidenced: Conditions in which death frequently occurs outside formal health settings—such as drowning, road injuries, and diarrheal diseases—are consistently undercounted by passive surveillance systems.
  • Regional opportunities in East and West Africa: Results highlight Burundi and Rwanda in particular as potentially low-hanging fruit, where existing infrastructure and relatively high Feasibility scores suggest that targeted investments in core data systems rather than condition-specific programs could yield valuable cross-condition improvements in data availability. The frequent appearance of countries in West Africa (Cameroon, Côte d’Ivoire, and Benin) suggests there may also be similar opportunities for regional data strengthening there.

This framework serves primarily as a hypothesis-generating tool for funders and policymakers. While it identifies where investments are most likely to be impactful and achievable, these findings should be taken as a starting point and validated with local expert knowledge before making funding decisions.

Our results may also be of interest to specific parties: for example, philanthropic funders interested in improving the quality of health data in SSA or survey implementers. For instance, identifying countries where data on a particular condition is sparse or outdated could help funders prioritize commissioning new surveys or adding relevant modules to planned ones. Similarly, survey implementers could use our findings to advocate for expanding core questionnaires or introducing new modules in regions where gaps are most acute. More broadly, funders using burden estimates to guide philanthropic giving may find our results useful for understanding where data limitations introduce the most uncertainty into those estimates, and for factoring that uncertainty into their prioritization decisions. Beyond funders and implementers, our results may also be relevant to national governments and ministries of health seeking to understand where their country’s data landscape falls short, and to researchers and modelers who rely on survey data as inputs and need to account for uncertainty stemming from data limitations.

By shifting from documenting gaps to prioritizing them, stakeholders can ensure that limited resources are directed toward strengthening empirical evidence where it matters most, ultimately driving smarter investment in tackling specific disease burdens.

Contributions and acknowledgments


Britta Jewell, Jenny Kudymowa, Rory Todd, Abbie Clare, and John Firth jointly researched and wrote this report. Britta Jewell and Jenny Kudymowa served as project co-leads. Abbie Clare and John Firth co-supervised the project.

Thank you to Shane Coburn and Thais Jacomassi for copyediting and Elisa Autric for assistance with publishing the report.

Appendices

Appendix A: How GBD Burden Estimates are Constructed

The GBD study, coordinated by IHME, estimates mortality and morbidity across 200+ countries by pooling available empirical data and modeling across gaps. Estimate quality is directly tied to the quality of underlying data inputs.

GBD draws on multiple data source types that feed into different components of the estimates (mortality, cause of death, morbidity, and risk factors). These sources fall into two broad groups (IHME, 2021):

Sources used in all settings:

  • Vital registration (VR): considered the gold standard for mortality data; continuously records deaths with medically certified causes of death coded using the International Classification of Diseases (ICD). VR is the primary source for mortality and cause-of-death estimates, but coverage is limited across much of Sub-Saharan Africa (Mikkelsen et al., 2015).
  • Censuses and household surveys: e.g., Demographic and Health Surveys (DHS); used for population denominators, mortality, and a broad range of morbidity and risk factor indicators
  • Disease registries: centralized databases for specific conditions, notably cancer; used for incidence, prevalence, and mortality
  • Hospital and insurance claims data: inpatient and outpatient records; used for morbidity estimates only, not cause of death
  • Morbidity surveillance and published literature: disease notification systems and peer-reviewed studies; used for incidence and prevalence estimates

Sources used where vital registration is absent or incomplete:

  • Sample registration systems: VR-like systems covering a population sample rather than all deaths; found mainly in India, China, Indonesia, and Bangladesh
  • Demographic surveillance systems (DSS): continuous monitoring of geographically defined populations; provide mortality and cause-of-death data for specific localities, mainly through the International Network for the continuous Demographic Evaluation of Population Health in developing countries (INDEPTH) network in Africa and Asia
  • Verbal autopsy: post-death interviews with family members to infer cause of death; used where VR is absent

Appendix B: Top 30 Conditions by Total DALYs in Sub-Saharan Africa (GBD 2023)

Chart

Note. Calculations can be found in this spreadsheet.

Appendix C: Process for Choosing Country-Level Components of Feasibility Score

To construct the Feasibility score, we first tried to identify key country-level factors (at a conceptual level) that are likely to make conducting data collection or improving existing data systems easier or more difficult. First, brainstorming from first principles, and then spending around ~5 hours searching for literature, we found evidence that political stability and government effectiveness,[46] health system strength,[47] laboratory availability, hard infrastructure—including electrification and internet connectivity[48]—rural access,[49] and education levels[50] are important factors influencing the feasibility of data collection. However, we did not find sufficient information to discern which were the most important.

Given time constraints, we relied on existing country-level indicators, rather than trying to create new indicators from raw data sources. We did not find an existing indicator for laboratory availability, and also found that indices for data system strength used underlying indicators that were likely to be highly correlated with our Evidence Strength metric (for instance, the quality of vital registration systems).

In general, we prioritized simplicity and intuitiveness, and thus were restrictive about the number of indicators used to construct the score. We also prioritized indicators that measured distinct and independent factors influencing the difficulty of improving the evidence base on disease burden, and in one case, excluded an indicator that correlated too strongly with another.[51]

Based on these criteria, we ultimately selected three country-level indicators: Government Effectiveness (part of the World Governance Indicators), Healthcare Access and Quality Index (an analysis of the IHME GBD 2019 study), and the Rural Access Index (designed by the World Bank). Details on the indicators and an explanation for their selection are presented in Table 2 in the main text. Correlations between these indicators are moderate,[52] indicating that they contribute independent information about the feasibility of improving evidence, but likely with some overlap. Given the lack of information with which to compare indicators that were highlighted in the literature, this choice was based on our intuitions regarding the most important factors. We also have little way to validate how well the score measures the true feasibility of improving the evidence on disease burden. Our results, therefore, draw on this score with low confidence: further detail is provided in the Limitations section.

Appendix D: Ease of Diagnosis Component of the Feasibility Score

We classified the top DALY burden conditions in SSA by a custom metric of “ease of diagnosis,” including the typology of diagnosis and testing requirements, divided into the following categories:

  1. Epidemiological diagnosis: sufficient event-based information for diagnosis (e.g., drowning)
  2. Clinical diagnosis: clinical diagnosis is sufficient without follow-up testing (e.g., diarrheal diseases)
  3. Testing is useful for diagnosis: testing is useful, but not required, for diagnosis (e.g., pneumonia)
  4. Testing is required for diagnosis: testing is required for diagnosis, with a point-of-care (POC) test sufficient (e.g., malaria)
  5. Laboratory- or hospital-based testing is required for diagnosis: facility-based testing is required for diagnosis; a POC test is either unavailable or insufficient (e.g., ischemic heart disease)

Scores for each category were assigned according to equal intervals on a 0–1 scale, with a highest-possible score of 1 as follows: epidemiological diagnosis = 1; clinical diagnosis = 0.75; testing is useful for diagnosis = 0.5; testing is required for diagnosis = 0.25; laboratory- or hospital-based testing is required for diagnosis = 0.

We spent ~5 hours searching for diagnostic requirements for each condition. We prioritized guidance from the World Health Organization (WHO) where available, especially guidance targeted at Sub-Saharan Africa. In other cases, we consulted scientific literature to understand which diagnostic approaches were possible. We found several cases where both a point-of-care (POC) test and a laboratory test were available; POC tests are typically less accurate, but may still be considered acceptable diagnostics, depending on the condition. In these cases, we looked for evidence regarding whether POC tests were viewed as of acceptable accuracy for assessing burden at a population level. In ambiguous cases, we tended to be lenient, coding conditions as of medium rather than low ease of diagnosis. This is because we found that guidance for locations with weak laboratory capacity, which includes much of sub-Saharan Africa, tended to be less stringent. Given large data gaps in the region, data from POC testing is more likely to be acceptable: data from these tests are a better option than no data at all. Table D.1 below shows our assessment of conditions:

 

Table D.1: Ease of diagnosis for top DALY burden conditions

ConditionCategoryIs testing required?Score
Neonatal disorders2Sometimes; varies due to wide range of conditions[53]0.75
Malaria4Yes; rapid diagnostic blood test available[54]0.25
Lower respiratory infections3Sometimes; primarily clinical diagnosis; rapid antigen tests or chest X-ray, if possible[55]0.5
HIV/AIDS4Yes; POC rapid antibody testing and self-testing available[56]0.25
Diarrheal diseases2Sometimes; primarily clinical diagnosis[57]0.75
Road injuries1No; event-based epidemiological diagnosis1
Tuberculosis4No; POC rapid molecular diagnostic test available[58]0.25
Congenital birth defects2No; primarily clinical diagnosis based on newborn exam[59]0.75
Stroke5Yes; hospital-based MRI and CT scans recommended[60]0
Meningitis4Yes; POC CSF tests (e.g., cryptococcal antigen) or culture required[61]0.25
Measles5Yes; laboratory NAAT required[62]0
Maternal disorders2Sometimes; varies due to wide range of conditions[63]0.75
Ischemic heart disease5Yes; diagnosis via ECG, exercise stress test, angiography, or MRI[64]0
Interpersonal violence1No; event-based epidemiological diagnosis1
Diabetes mellitus4Yes; POC glucose meters and self-testing available[65]0.25
Protein-energy malnutrition2No; primarily clinical diagnosis[66]0.75
Chronic kidney disease4Yes; POC measurement possible for creatinine and urine albumin[67]0.25
Cirrhosis and other chronic liver diseases5Yes; laboratory confirmation (AST level and a platelet count) required, if not biopsy[68]0
Hemoglobinopathies and hemolytic anemias4Yes; while POC testing exists for sickle cell disease, for other types laboratory testing is required[69]0.25
Pertussis5Yes; laboratory confirmation recommended[70]0
Invasive Non-typhoidal Salmonella (iNTS)5Yes; blood culture required[71]0
Drowning1No; event-based epidemiological diagnosis1
Idiopathic epilepsy3Sometimes; clinical diagnosis or via EEG, imaging[72]0.5
COVID-194Yes; rapid antigen test available[73]0.25

Note. POC = point-of-care. MRI = magnetic resonance imaging. CT = computed tomography. CSF = cerebrospinal fluid. NAAT = Nucleic Acid Amplification Test. ECG = electrocardiogram. AST = Aspartate aminotransferase.

Appendix E: Visualization of Priority Filter Approach

 

Figure E.1: Country-condition pairs shown by whether they meet minimum thresholds for Burden (by total DALYs), Evidence Strength, and Feasibility

Appendix F: Priority Filter Approach by DALY Rate per 100,000

 

Figure F.1: Priority-filtered country-condition pairs meeting minimum thresholds for Burden, Evidence Strength, and Feasibility (by DALY rate per 100,000)

Appendix G: Regression Methods and Additional Results

Equations for Regression Models

The model without fixed effects uses Feasibility as the sole predictor of Evidence Strength, as reflected in Equation G.1, which we estimate via ordinary least squares (OLS).

 

Equation G.1: Core OLS regression

We estimate the following equation:

yic = β0 + β1xic + εic

Where:

  • xic = Feasibility score for country i, condition c
  • yᵢc = Evidence Strength score for country i, condition c

We then calculate the regression residuals:

ε̂ic = yicŷic

representing the difference between actual Evidence Strength (yic) and regression-predicted Evidence Strength (ŷic ≡ β̂0 + β̂xic).

The model with fixed effects uses country and condition fixed effects to predict Evidence Strength, absorbing systematic differences in evidence levels across countries and diseases rather than estimating a single slope coefficient. This produces a fitted value for each country-condition pair based on the average Evidence Strength score for that country and that disease, and the residual captures how far a given pair falls above or below what would be expected given those averages. The equation is shown in Equation G.2:

 

Equation G.2: OLS regression with two-way country and condition fixed effects

EvidenceStrengthi,c = α + Σj βj · Countryj + Σk γk · Conditionc + εi,c

Where:

  • α = intercept
  • βj = coefficient for country j (j = 1, … , 49)
  • γₖ = coefficient for condition k (k = 1, … , 24)
  • Countryj = dummy variable equal to 1 if observation belongs to country j, 0 otherwise
  • Conditionc = dummy variable equal to 1 if observation belongs to condition c, 0 otherwise
  • εi,c = residual (actual minus fitted Evidence Strength score)
  • i = country index, c = condition index

Visualization of Results from Both Regression Models

Figure G.1 below shows actual Evidence Strength values plotted against the values predicted by the non-fixed effects model.[74] The model here explains only 9% of variation in Evidence Strength (R² = 0.09), suggesting that Feasibility alone captures relatively little of what drives differences in evidence across country-condition pairs. Points falling below the line (shown in pink) are those where actual evidence is worse than the model predicts given the country and cause—these are the country-condition pairs driving the negative residuals and, when weighted by burden, the relative priority scores shown in our main text.

 

Figure G.1: Actual versus expected Evidence Strength scores, after regressing Evidence Strength on Feasibility

Figure G.2 below shows the same plot for the fixed effects model. This model has a much higher R² of 0.84, reflected in the relatively tight clustering around this line, which confirms that country and cause fixed-effects together explain most of the variation in Evidence Strength.

 

Figure G.2: Actual versus expected Evidence Strength scores, after regressing Evidence Strength on country and condition fixed effects

Fixed Effects Regression Model Highlights Similar Priority Data Gaps, but Only After Priority Filter is Applied

Table G.1 below shows the prioritized ranking of the top 20 country-condition pairs from the regression model with fixed effects. Our priority filter is also applied to capture only pairs with high burden, surprisingly poor evidence, and realistic feasibility for collecting more data.

 

Table G.1: Top 20 ranked country-condition pairs from regressing Evidence Strength on country and condition fixed effects

CountryConditionEvidence StrengthFeasibilityResidual data gapBurdenRelative priority score
NigeriaInterpersonal violence52.38-2.451.3M3.1M
Côte d’IvoireNeonatal disorders12.18-1.142.3M2.7M
CameroonNeonatal disorders02.17-1.302.0M2.5M
UgandaDiarrheal diseases12.47-2.59560,0001.5M
RwandaNeonatal disorders12.81-2.14640,0001.4M
BeninNeonatal disorders12.35-1.34880,0001.2M
BurundiNeonatal disorders02.35-1.011.1M1.1M
Côte d’IvoireMaternal disorders12.18-4.07270,0001.1M
Sierra LeoneDiarrheal diseases32.04-2.05400,000810,000
CameroonDiarrheal diseases02.17-0.671.1M720,000
LiberiaNeonatal disorders12.04-1.47440,000640,000
Côte d’IvoireCongenital birth defects12.18-1.28500,000640,000
CameroonCongenital birth defects02.17-1.45440,000640,000
KenyaDiarrheal diseases52.40-0.75770,000580,000
SenegalDrowning32.68-3.37110,000370,000
TogoNeonatal disorders02.23-0.84430,000360,000
RwandaDiarrheal diseases12.56-1.50240,000360,000
RwandaStroke12.06-1.50210,000310,000
Côte d’IvoireDiarrheal diseases12.18-0.50620,000310,000
BeninCongenital birth defects12.35-1.49200,000300,000

Note. Burden = total DALY burden. Evidence Strength = score out of 15. Feasibility = score out of 4. Residual data gap = actual Evidence Strength score minus expected Evidence Strength score. Relative priority score = absolute value of the burden-weighted residual, i.e., the residual data gap multiplied by the burden. Values are rounded to the nearest 100,000 if ≥1M and nearest 10,000 if <1M. Blue shading indicates variables used in the regression (here, feasibility is absorbed by the fixed effects). For residual data gap, burden, and relative priority score, coloring indicates magnitude, with darker red colors indicating worse data, a higher burden, and higher overall priority, respectively. Country-condition pairs are ranked in order of the largest relative priority score.

We also show results from the fixed effects model with no priority filter applied in Table G.2. Without the priority filter, the top country-condition pairs surfaced often appear to be poor candidates for data investment: they either already have relatively strong empirical evidence (e.g., HIV/AIDS in South Africa) or low feasibility for further data collection (e.g., invasive non-typhoidal salmonella in Nigeria).

 

Table G.2: Top 20 ranked country-condition pairs from regressing Evidence Strength on country and condition fixed effects, without priority filter applied

CountryConditionEvidence StrengthFeasibilityResidual data gapBurdenRelative priority score
South AfricaHIV/AIDS102.33-3.046.1M18.6M
MozambiqueHIV/AIDS101.32-2.672.9M7.6M
NigeriaIschemic heart disease31.38-2.802.2M6.3M
NigeriaInvasive Non-typhoidal Salmonella (iNTS)31.38-2.372.5M6.0M
KenyaHIV/AIDS71.90-2.841.7M4.8M
AngolaNeonatal disorders11.96-2.431.9M4.5M
NigeriaCirrhosis and other chronic liver diseases31.38-3.151.4M4.4M
CameroonMalaria01.67-1.961.9M3.8M
EthiopiaMeasles101.28-3.481.0M3.6M
Côte d’IvoireMalaria11.68-1.792.0M3.5M
NigerDiarrheal diseases11.70-1.921.7M3.3M
EthiopiaHIV/AIDS101.53-2.591.3M3.3M
NigeriaInterpersonal violence52.38-2.441.3M3.1M
NigeriaPertussis81.38-1.761.7M3.0M
Democratic Republic of the CongoCongenital birth defects11.63-2.201.3M2.8M
United Republic of TanzaniaHIV/AIDS91.83-1.382.0M2.8M
Côte d’IvoireNeonatal disorders12.18-1.142.3M2.7M
Democratic Republic of the CongoLower respiratory infections31.38-0.783.4M2.6M
CameroonNeonatal disorders02.17-1.302.0M2.5M
EthiopiaPertussis81.28-4.10600,0002.5M

Note. Burden = total DALY burden. Evidence Strength = score out of 15. Feasibility = score out of 4. Residual data gap = actual Evidence Strength score minus expected Evidence Strength score. Relative priority score = absolute value of the burden-weighted residual, i.e., the residual data gap multiplied by the burden. Values are rounded to the nearest 100,000 if ≥1M and nearest 10,000 if <1M. Blue shading indicates variables used in the regression (here, feasibility is absorbed by the fixed effects). For residual data gap, burden, and relative priority score, coloring indicates magnitude, with darker red colors indicating worse data, a higher burden, and higher overall priority, respectively. Country-condition pairs are ranked in order of the largest relative priority score.

Appendix H: Average Burden and Feasibility Scores Across Countries in SSA

Figure H.1: Total DALY burden by country in SSA, averaged across 24 conditions

Figure H.2: Feasibility score by country in SSA, averaged across 24 conditions

  1. For example, each country’s specific disease burden is a key input to the Global Fund’s formula for allocation of funds for HIV/AIDS, tuberculosis, and malaria (The Global Fund, 2023).
  2. According to IHME (2023, p. 49) estimates, malaria receives comparatively high levels of development assistance for health (DAH) relative to burden. For example, DAH for malaria was more than twice the amount allocated to non-communicable diseases (NCDs) in 2023, even though malaria accounted for roughly one quarter of the disability-adjusted-life-year (DALY) burden attributable to NCDs in SSA (IHME/GHDx, 2023).
  3. IHME/GHCx, 2023.
  4. Most notably, in early 2025 the US government significantly reduced USAID’s budget (Walker & Lai, 2025), leading to the termination of the Demographic and Health Surveys (DHS) Program in February 2025 (Grown, 2025). The DHS Program was a primary source of population-level health data in many low- and middle-income countries (LMICs).
  5. For simplicity, we use the term “condition” to refer broadly to anything listed as a cause-of-death in GBD: this includes categories like drowning and road injuries, which may not be typically considered health conditions.
  6. This approach adapts the importance-neglectedness-tractability (INT) framework, widely used within the effective altruism community to prioritize causes based on their scale, existing attention, and tractability (Oehlsen, 2024; Coefficient Giving, 2025). Here, “importance” corresponds to disease burden, “neglectedness” to the strength of empirical evidence underlying burden estimates (rather than funding levels or research attention), and “tractability” to the capacity to improve data availability and quality (rather than ease of programmatic interventions).
  7. A DALY represents a year of healthy life lost due to ill-health, disability, or premature death. DALYs combine years of life lost due to premature death (YLLs) and years lived with disability (YLDs) into a single measure of disease burden, capturing both mortality and morbidity (McKee et al., 2019).
  8. The choice of 30 conditions was driven by practicality. Including more conditions would have increased data extraction time, as GBD input data must be accessed separately for each condition, while adding little additional burden coverage.
  9. Note that larger countries may also require greater resources to improve data coverage, as more extensive data collection infrastructure is needed to achieve representative national estimates. This analysis does not account for the cost or effort of data collection, which may offset some of the advantages of prioritizing high-burden, large-population contexts.
  10. See calculations.
  11. The six excluded conditions (anxiety disorders, depressive disorders, headache disorders, low back pain, dietary iron deficiency, and age-related and other hearing loss) have no mortality component and their burden consists entirely of YLDs, as shown in Appendix B.
  12. IHME’s cause-of-death star ratings summarize the proportion of deaths that are “well-certified,” combining the completeness of vital registration with the share of deaths assigned to specific, well-defined conditions rather than ill-defined or “garbage” codes. Ratings range from 0 to 5 stars, with higher values indicating more complete and reliable mortality data systems. See GBD (2023, p. 35) of GBD (2025) for more detail. These star ratings are assigned to countries as a whole and do not include the number or recency of data sources; the components of our Evidence Strength metric therefore capture different additional aspects of data availability.
  13. Source count thresholds follow a roughly logarithmic scale to reflect diminishing marginal value. We expect that the difference between 0 and 5 sources matters more for burden estimation than the difference between 20 and 25 sources. Cutoffs were chosen based on the observed distribution in SSA, where most country-condition pairs have fewer than 10 sources and only a few countries like South Africa and Mauritius exceed 25 sources for certain conditions.
  14. Recency thresholds reflect that older data becomes progressively less relevant for current burden estimation. We use finer time intervals for recent data and broader intervals for older data because we expect the difference between data from 2020 versus 2022 to be more impactful than the difference between 1998 versus 2000.
  15. We did this spot-check for interpersonal violence, diabetes mellitus, and congenital birth defects.
  16. Details on the indicators, and an explanation for their selection, are provided in Table 2.
  17. Our subjective assessment of ease of diagnosis is inherently simplistic and likely under- or over-states the ease of diagnosing certain conditions. This is particularly true for conditions that encompass multiple different sub-causes of death (e.g., neonatal disorders, maternal disorders, or lower respiratory infections), some of which may have substantially different diagnostic requirements. Because of this limitation, we treat our indicator as a heuristic rather than a robust measure, and interpret it with low confidence accordingly.
  18. In particular, diagnosis could be done accurately with a survey or verbal autopsy.
  19. Suthar et al. (2019) list operational considerations affecting the effectiveness of vital registration systems, several of which concern relationships with government officials and the effectiveness of those officials.
  20. Suthar et al. (2019) highlight large geographic areas that are sparsely populated as especially challenging for vital registration systems decentralized to community health workers.
  21. Suthar et al. (2019) found that establishing vital registration systems was cheaper and more efficient where there were existing health facilities into which it could be integrated: better organized and more widespread health facilities are likely to facilitate collecting evidence on health burden.
  22. The coding of ease of diagnosis for each condition, and our justification for this coding, is provided in Appendix D.
  23. We also tested thresholds based on DALY rates per 100,000 population; see Appendix F.
  24. For example, a funder interested in improving data for a given condition can immediately see how many contexts share the same data gap and could benefit from the same intervention.
  25. This regression has no fixed effects and does not account for country- or condition-level characteristics.
  26. The fixed effects capture systematic differences in evidence strength across countries and conditions—for instance, the fact that some countries tend to have better data across the board, or that some diseases are better studied everywhere.
  27. A negative residual implies that the data is worse than expected. We take the absolute value of the residual and multiply by the burden to create a relative priority score.
  28. For the fixed effects model, the methodology is identical except that Evidence Strength is regressed on the full set of country and condition characteristics.
  29. iNTS may have the lowest score across conditions due to (i) expensive diagnostics requiring blood cultures and lab equipment that are often unavailable in rural, low-resource settings where the disease is most common, and (ii) substantial syndromic overlap with other conditions. Because it presents as a fever without diarrhea, it may be misdiagnosed as malaria or pneumonia and frequently occurs as a co-infection with those conditions (Park et al., 2016). Patients may also be treated with anti-malarials and later receive an inaccurate cause-of-death recording.
  30. Individual component and summed country-condition pair scores are available by total DALY burden in this sheet and by DALY rate per 100,000 in this sheet. We also created a country- and condition-level dashboard for further analysis of the contribution of sub-components of the Evidence Strength and Feasibility indicators.
  31. Results for the same analysis using Burden as the DALY rate per 100,000 (such that smaller countries are comparable with larger countries) are presented in Appendix F.
  32. Note that neonatal disorders in GBD (2023) encompasses five sub-causes (each of which may also have further sub-division of cause-of-death) with varying diagnostic complexity that are not fully captured in our ease of diagnosis criterion. Our Feasibility score may therefore somewhat overstate the ease of data collection for this condition.
  33. As our threshold-based approach does not apply any ranking to pairs that meet the criteria, countries and conditions appearing less frequently are not necessarily considered lower priorities.
  34. Major infectious diseases including HIV/AIDS, tuberculosis, and malaria benefit from dedicated surveillance systems supported by initiatives such as the Global Fund to Fight AIDS, Tuberculosis and Malaria.
  35. A script to run both regression models is available here. Complete outputs for both models can be found in this spreadsheet and equations can be found in Appendix G.
  36. We chose to apply our minimum thresholds across Burden, Evidence Strength, and Feasibility, rather than present the regression results directly, because many of the highest relative priority scores indicated country-condition pairs where data is already sufficiently good (Evidence Strength ≥10) or Feasibility is low (<2). We deprioritized these data collection investments for these reasons, despite their higher overall relative priority scores. However, results without applying our minimum thresholds are available in Table G.2.
  37. The largest negative residual gap in the analysis, after applying the priority filter, is -6.67.
  38. Calculations here.
  39. Defined as (upper bound – lower bound)/mean estimate.
  40. Calculations here.
  41. We also checked the correlation when country-condition pairs with no available data were excluded and the correlation remained the same (calculations here).
  42. For example, when vital registration data for neonatal disorders is unavailable, GBD’s CODEm modeling framework generates predictions using covariates like maternal care and immunization coverage, antenatal care visits, and facility delivery rates. The model produces point estimates and uncertainty intervals based on these statistical relationships even in the absence of direct mortality observations (GBD, 2025; GBD, 2023, pp. 605-606).
  43. Scenarios B and C in this sheet.
  44. We chose not to allow Feasibility to drop to a threshold below 2 out of 4 because that implicitly included countries where gathering stronger evidence would be practically unrealistic, rather than merely challenging.
  45. Note that the relative benchmark for SSA for Evidence Strength and Feasibility in our analysis is Mauritius, a small, urbanized island with a well-functioning vital registration system that scores comparably to high-income countries using the World Bank’s Statistical Performance Indicators.
  46. Suthar et al. (2019) list operational considerations affecting the effectiveness of vital registration systems, and several link to relationships with and the effectiveness of government officials.
  47. Suthar et al. (2019) found that establishing vital registration systems was cheaper and more efficient where there were existing health facilities into which it could be integrated. Stronger health systems are also likely to reflect greater numbers of trained medical personnel and established logistics for health service delivery, which can also facilitate the collection of evidence on health burden.
  48. Suthar et al. (2019) identify both as key barriers to effective civil registration and vital statistics systems.
  49. Suthar et al. (2019) highlight large geographic areas that are sparsely populated as especially challenging for vital registration systems decentralized to community health workers.
  50. Mwanyangala et al. (2011) identify higher education levels in participants as helpful for reducing numbers of undetermined deaths gathered by verbal autopsy surveys.
  51. We found that the Government Effectiveness and Political Stability indicators from the World Bank’s World Governance Indicators had a correlation coefficient of 0.68, and decided to therefore prioritize one; we selected Government Effectiveness, as this seemed more integral to the feasibility of improving evidence on disease burden.
  52. Pairwise correlations: Government Effectiveness & Healthcare Access and Quality (HAQ) indices (0.60); Government Effectiveness & Rural access indices (0.58); Rural access index & HAQ indices (0.38).
  53. GBD (2023) includes five causes of death under neonatal disorders: neonatal preterm birth, neonatal encephalopathy due to birth asphyxia and trauma, neonatal sepsis and other neonatal infections, hemolytic disease and other neonatal jaundice, and other neonatal disorders. We followed the International Classification of Diseases’ definition (Blencowe et al., 2024) of neonatal deaths as death within 0–27 days of life and thus consider it to be primarily an event-based diagnosis. However, as there are multiple causes of varying diagnostic complexity, our scoring may not fully capture the difficulty of diagnosing all sub-conditions.
  54. WHO (2018) says that rapid diagnostic tests “can assist in making a rapid, accurate diagnosis of malaria in circumstances where microscopy-based diagnosis may be not available or unreliable.”
  55. An analysis of the GBD 2023 included 26 pathogens (GBD, 2026) as contributing to the total burden of lower respiratory infections, with Streptococcus pneumoniae accounting for the greatest number of global deaths, followed by Staphylococcus aureus and Klebsiella pneumoniae. Rapid tests are available for Streptococcus pneumoniae (Cepheid, 2016) and Staphylococcus aureus (Izumikawa et al., 2009), although not Klebsiella pneumoniae.
  56. WHO (2024) notes that “HIV can be diagnosed through rapid diagnostic tests that provide same-day results.”
  57. Diarrheal diseases comprise multiple viral, bacteria, and parasitic etiologies (e.g., rotavirus, E. coli, Giardia), some of which may require laboratory confirmation, though clinical diagnosis without specific testing is common (Kotloff, 2017).
  58. WHO (2025) recommends low-complexity automated nucleic acid amplification tests (NAATs), as the initial test for tuberculosis diagnosis (p. 3).
  59. Many structural anomalies can be identified by physical examination.
  60. In a 2018 meta-analysis of acute stroke presentation and outcomes in LMICs, Khatib et al. (2018) found that either magnetic resonance imaging (MRI) or computed tomography (CT) scans were used to diagnose the majority of patients.
  61. WHO (2025) guidance states that a “positive test result from any antigen detection test should be confirmed with culture or molecular testing to establish a definitive diagnosis.”
  62. For surveillance, WHO (2022) recommends laboratory confirmation.
  63. GBD (2023) includes 10 causes of death under maternal disorders (e.g., maternal hemorrhage, maternal sepsis and other infections, maternal obstructed labor and uterine rupture). Some of these conditions (e.g., infections) may require confirmatory diagnostic testing, while others are epidemiologically diagnosable without further testing; because this category encompasses multiple diverse causes of death, our score may not fully capture the diagnostic complexity of all sub-conditions.
  64. WHO (2020) recommends that electrocardiograms (ECGs) be available at the primary health care level, indicating that a hospital setting is not required.
  65. Andriankaja et al. (2020) note that “[p]oint-of-care blood glucose measurement may be useful to screen people with diabetes, and to assess glucose among individuals with diabetes where blood can be drawn, but laboratory tests are unavailable or untimely.”
  66. WHO (2023) identifies anthropometric tests as the key diagnostic test.
  67. The Kidney Disease: Improving Global Outcomes (KDIGO) 2024 guideline (Madero et al., 2025) recommends POC testing for creatinine and urine albumin measurement where access to a laboratory is limited, although the certainty of evidence for this recommendation was rated as low to very low.
  68. WHO (2020) states that the “presence of cirrhosis can be identified by a combination of clinical findings (oedema, ascites, variceal bleed, hepatic encephalopathy), haemogram (which may show pancytopenia; the platelets are the first to show a reduction in number), liver function tests (low serum albumin and composite scores such as APRI [Aspartate Aminotransferase-to-Platelet Ratio Index], FIB-4 [Fibrosis-4 Index]), ultrasound abdomen (nodular and shrunken liver, dilated portal vein, splenomegaly, etc.), and endoscopy (for oesophageal and gastric varices).”
  69. A laboratory-based complete blood count is recommended in WHO (2022)’s Model List of Essential In Vitro Diagnostics for hemolytic anemias, while a POC test is available for sickle cell disease.
  70. WHO (2018)’s Surveillance Standards for Pertussis recommends laboratory testing for children under 5 years of age, which is the age group at which severe effects are high.
  71. Crump et al. (2015) recommend microbiologic culture of blood or bone marrow for diagnosis, with no POC test available.
  72. WHO (2019)’s MGHAP Intervention Guide provides guidance for helping non-specialists to diagnose epilepsy without requiring EEG or imaging.
  73. WHO (2022)’s Model List of Essential In Vitro Diagnostics lists non-laboratory lateral flow rapid diagnostic tests as an acceptable diagnostic for COVID-19.
  74. By construction, a well-fitting model should produce points clustered around the best-fit line, where the actual value equals the predicted value.