Why machine learning is central to public policy

In 1854, London was struck by a devastating cholera outbreak. Public authorities responded using the best tools and scientific understanding available at the time.

Play Video

Disease was widely believed to spread through ‘bad air,’ and policy responses focused on sanitation measures consistent with that view. Yet despite these efforts, the outbreak persisted.

What eventually changed the course of events was not a failure of institutions or expertise, but the use of more detailed information. Physician John Snow mapped cholera deaths across neighbourhoods and observed a striking pattern: cases clustered around a single public water pump on Broad Street.

When access to the pump was restricted, infections declined rapidly. The episode is now remembered as an early demonstration of how additional data can sharpen public decision-making, even when institutions are acting in good faith.

Snow did not have computers, algorithms, or modern data infrastructure. What he had was a richer view of the problem-one that allowed patterns to emerge that were previously invisible.

Today, advances in data science and machine learning offer policymakers the same advantage, but at far greater scale, speed, and scope.

For decades, public policy has relied primarily on periodic national surveys, broad averages, and delayed indicators. These tools remain valuable and indispensable.

However, they are increasingly complemented by machine-learning techniques that can draw insights from administrative and transactional data generated every day. This shift quietly changes what is possible. Decision-making no longer has to depend solely on infrequent snapshots of reality; it can be informed by near-real-time signals of behaviour, not just outcomes.

At the core of this transformation is a simple idea. Machine learning allows policymakers to combine many pieces of information-each imperfect on its own-into a clearer overall picture.

Income proxies, consumption patterns, administrative records, and transaction data can be analysed together to produce individual-level or firm-level assessments that dramatically reduce information asymmetry. Policy moves from designing for the ‘average citizen’ toward targeted, evidence-based intervention.

This capability has practical implications across multiple policy areas.

In higher education funding, for example, means testing has traditionally relied on self-reported income, household surveys, and appeals processes. These methods are costly, slow, and often contested.

Machine learning offers an opportunity to complement them by combining indicators such as parental employment history, utility usage, property characteristics, and school background.

The result is not surveillance, but fairer and more defensible allocation of limited funding, with reduced gaming of the system and faster decision-making. In practical terms, means testing shifts from being declaration-based to evidence-based.

A similar logic applies to insurance pricing and social protection. Flat premiums, while simple, often penalise low-risk households and under-price high-risk behaviour.

By analysing behavioural and claims data, machine-learning models can help design fairer premium bands, expand coverage, and maintain sustainability of insurance pools. This approach naturally extends to discussions around national health insurance and universal health coverage, where balancing affordability, inclusion, and financial sustainability is critical.

Universal Health Coverage presents an even clearer case. One of the biggest challenges in targeting subsidies is identifying who is truly vulnerable, especially in economies with large informal sectors where income is difficult to observe directly.

Machine learning allows vulnerability to be inferred from patterns in health utilisation, payment behaviour, and geographic and demographic indicators. Better targeting means reduced leakage, more efficient use of public funds, and ultimately, more people covered with the same budget.

Tax policy provides another important example. Informal economic activity is often described as ‘invisible,’ leading to blunt enforcement approaches that are costly and sometimes counterproductive. Predictive analytics allows tax systems to estimate economic activity using signals such as mobile money flows, utility consumption, licensing data, and transport patterns.

This enables a policy shift-from enforcement to graduation, and from penalties to progressive inclusion. Machine learning, in this sense, allows tax systems to understand before they enforce.

Revenue forecasting at both county and national levels also stands to benefit. Traditional projections often miss turning points, detecting shocks only after revenues have already deviated from targets.

By incorporating real-time economic indicators, administrative collection data, and sector-level signals, machine-learning models can act as early warning systems, supporting more credible budgeting and fiscal planning.

At a more advanced level, machine learning enables firm-level micro-simulation. Policies rarely affect all firms in the same way. By simulating tax changes, incentives, or shocks across heterogeneous firms, policymakers can test the likely effects of interventions before implementation, reducing unintended consequences and improving policy design.

The unifying advantage across these applications is reduced information asymmetry. With sufficient high-quality data, machine learning can approximate a near-complete picture of economic behaviour-not perfectly, and not invasively, but far more accurately than traditional methods alone. Guesswork is replaced with evidence; broad assumptions give way to targeted insight.

This does not imply perfect surveillance, nor does it suggest abandoning established safeguards. Data protection and privacy are non-negotiable, and institutions are right to be cautious. In practice, however, data is often fragmented across agencies, locked in silos, or accessible only in highly constrained ways.

These limitations are understandable, but they also constrain the public value that data can generate.

The real policy challenge, therefore, is not whether to protect data, but how to unlock its value responsibly. Secure data environments, anonymization and aggregation, controlled access for policy modelling, and clear governance frameworks make it possible to balance confidentiality with public benefit. Data protection and data use are not opposites; they are complements.

Kenya already possesses much of the data needed to make this shift. The analytical tools are mature and increasingly accessible. What remains is a deliberate move toward policy-driven data governance-one that encourages responsible use of data to improve fairness, efficiency, and trust in public decisions.

Just as better data once helped resolve a public health crisis, machine learning now presents an opportunity to strengthen decision-making across education, health, taxation, and public finance. The future of public policy will belong to governments that can learn from their data-securely, responsibly, and intelligently.

Leave a Reply

Your email address will not be published. Required fields are marked *