Stronger Evidence for a Stronger DC
SW-pink.jpg

Predictive Modeling

Predictive Modeling: Helping fill information gaps and find patterns in heaps of data

 

What is predictive modeling?
Local governments are awash in data about the services they provide, the residents that make up each neighborhood, the buildings and roads within city limits. But vast amounts of data can be difficult for humans to connect and parse.

Predictive modeling can help us sift through large quantities of data to uncover trends and estimate information that would otherwise be difficult to measure. Predictive models find statistical patterns within data that we have, and then apply those patterns in ways that allow us to make educated guesses about data that we do not have. Given information about populations, places, or phenomena, we can use different types of models to estimate how likely, how much, or how similar things may be.

Importantly, compared to other methods of analysis, predictive models don’t necessarily provide insight into the specific ways that different factors influence one another. As a result, predictive models are most useful when we’re interested in approximating an outcome rather than understanding its underlying causes.

Here's an example:
Rats are a persistent health risk,1 and rat colonies grow and spread if they are not treated. The Rodent Control team at DC Health decide where to inspect for rats based on calls they receive from residents via 311. But people don’t report every rat they see, and some rat infestations may not be discovered. So how might Rodent Control identify other, unreported places that they should look?

First, they think about how rats behave. Rodent Control knows rats like old buildings and alleys that have places to hide or soft soil to burrow in. Rats also like areas with lots of food waste from restaurants, as well as garbage from homes and apartment buildings. Unfortunately, there are a lot of places with these characteristics, and there aren’t enough rodent inspectors to regularly inspect all of them. So the question becomes: how might Rodent Control decide which places to prioritize?

This is where a predictive model might be useful. We can take all the data that DC government has on locations of old buildings, alleys, restaurants, and more, and use a predictive model to make an educated guess about where rats are most likely to be located.

How do we make a predictive model?
Predictive models often require a lot of complex math, statistics, and computer programs to create, but the logic behind how we make them is actually pretty simple. Humans make predictions all the time, and predictive models aren’t that different. As an example, think of your morning commute. Say you have to be at work by 8:30 AM and you usually take the bus. On any given day, there are different factors that could make you early or late: traffic, the weather, bus delays, whether you have enough money on your Smart Trip card or not. If you’ve been taking the same route for a while, you’ve probably learned how to account for these different situations so that you leave for work on time to ensure you aren’t late.

Predictive models work in a similar way. They take a specific event, like whether you’ll arrive at work on time, and combine all the things—the so-called “features”—that make the event more or less likely to happen. Lighter than usual traffic? You’ll get there sooner. On-time bus? You’ll get there sooner. Snowy road conditions? You’ll get there later. No money on your fare card? Later.

Some of these features are more important than others. A blizzard will no doubt make you later to work than if the bus is running a couple of minutes behind schedule. A predictive model assigns different values to these features in order to provide us with useful information. So if it’s raining and the bus is running six minutes late, a predictive model will estimate exactly when you should leave the house.

As we “train” a model to calibrate the way it weighs different features, we may try slicing the data to find highly specific sets of criteria that improve the accuracy of our predictions. We can select features that align with legislation or existing research, or we can lean on statistics and computing power to try many different combinations. More complex models may have a higher accuracy rate, but their results are often harder to explain. We weigh these considerations when planning our modeling approach.

How do we know if a model does a good job of making predictions
The short answer is: we test it. Our models make educated guesses but cannot be certain about what will happen. Rather, they are telling us how likely something is to happen. In our commuting example, our model might tell us that if you leave at 7:50 AM, there’s an 80% chance of getting to work on time. But that means that there’s still a 20% chance of being late. So, when a model predicts something will happen 80% of the time, we need to make sure it actually happens about 80% of the time. We call this "validating" our model.

Validation involves comparing our model’s predictions about what will happen against what actually happens. We regularly validate our models using data that were left out of the training process. In our rodent model, we might predict where Rodent Control will find rats this month using data that we collected from the past. We can then compare where Rodent Control actually found rats this month against the model’s predictions to see how often the model was right or wrong. This tells us how well our model performs.

We can also validate our models by testing them in the real world: what we call "field validations." During field validations, we give predictions from our model to a DC government agency to test. We then compare what they find with the model’s predictions. For example, we chose 100 locations for the Rodent Control team to inspect where our model predicted they were likely to find a rat burrow. We then looked at how many times the model was right that a rat burrow was there. While the model did a good job predicting whether rats would be found in test data, it was not as accurate when we sent Rodent Control on inspections. That’s why field validations are so important.

How do we ensure predictive models are used fairly and ethically?
Predictive models can provide us with information to identify people or places that may need extra help. But predictive models can also be biased. For example, in 2016, ProPublica reported racial bias in a model that was used to decide bail amounts.2 Because the predictive model was biased, that meant that people of some races were assigned higher bail than people of other races with the same criminal history and who were accused of the same crime. The Lab is deeply committed to guarding against realities like this, so we take several steps to make sure that our predictive models reduce inequities rather than exacerbate them.

First, we think carefully about what decisions we want a predictive model to inform. Predictive models may not be appropriate for some decisions, and it is important that we understand how our models will be used before we build them. We work carefully with our agency partners to consider predictions alongside other sources of information. The Lab’s models focus on cases where the prediction can elevate the greatest needs for help, and we avoid their use for punitive practices.

Second, we need to make sure that our predictive models do not contain unintended, unexplored, and/or undeclared biases. For example, if we are developing a predictive model for finding rodents, we may consider including the number of 3-1-1 calls for rodents for each city block. That seems logical because it would make sense to send inspectors to where people are telling us there are rodents. But we also know that residents may be less likely to call 3-1-1 based on the racial and ethnic composition of their neighborhood.3 If we didn’t investigate this potential bias beforehand and adjust the model accordingly, we might have a model that only sends rodent inspectors to more privileged areas of the city where the most 3-1-1 calls come from. We develop our models carefully to try to avoid these unintended biases. This requires attention to the way the data was created, how our models work, and how factors like history and location can affect people.

Finally, predictive models can cause the greatest harm when they are opaque to the public. That is why we publish reports on how our models are built, which approaches and data were considered along the way, and how representative the model is of the city overall.4 The Lab is committed to transparency, and we aim to build understanding and trust in any predictive model we create before it goes into use.

How does The Lab use predictive modeling?
Now that you’ve read about predictive modeling, check out a few examples of how we’ve used the technique to answer these questions for Washingtonians: