Why Emergency Management’s AI Adoption Needs Equity Auditing Requirements

Consider a county EMS agency that deploys an AI-optimized resource positioning system. The system is trained on five years of computer-aided dispatch data, and its objective is to minimize average response time across the service area. By that measure, it works, and average response times improve after deployment. The agency reports the improvement to the county commissioners, and nobody flags the accompanying finding that response times in the low-income eastern corridor, which were already above the county average before deployment, have gotten even longer.

Nobody flagged it because nobody was required to look. The aggregate metric improved, so, by the standard it was given, AI succeeded. The populations for whom emergency management have historically worked the least effectively are now served by a system which has only improved service for the populations it already served the best.

That outcome is not a malfunction or an outlier. It is what happens when you deploy machine learning on historically inequitable data without equity constraints, toward an optimization target which does not require equitable distribution of improved outcomes. It will continue to happen as long as emergency management treats AI adoption as a technical and operational question rather than an equity-based one.

The Mechanism Is Not Complicated

Algorithmic bias does not require bad intentions, but a system that learns from data will learn what the data contains, including the inequities the data reflects.

Emergency management data contains a lot of these inequities. CAD response time records reflect decades of station placement decisions and staffing models which produced longer response times in low-income and minority communities. FEMA individual assistance records reflect a program that has consistently produced worse outcomes for renters, non-English speaking applicants, and manufactured housing residents. Damage assessment records reflect processes that have historically been less thorough in communities with lower political and economic power. Research by Howell and Elliott using decades of FEMA data found disaster assistance in the U.S. has consistently widened rather than narrowed wealth gaps, with white homeowners recovering more fully than Black and Latino households even after controlling for damage severity.

Training an AI on any of these datasets without first auditing them for equity is not a neutral technical decision. The system will learn to operate in a world where vulnerable populations receive less, because that is the world the data describes. That is not effective calibration. That is replicating structural inequities at machine speed.

The optimization target problem is subtler and, in some ways, worse. Inequity is produced and perpetuated precisely because the system is working correctly. In 2019, Obermeyer and colleagues documented in “Science” magazine a widely used healthcare algorithm systematically underestimated the health needs of Black patients because it used healthcare costs as a proxy for health needs. The algorithm was designed to be efficient, not discriminatory, but healthcare costs are inequitably distributed across racial groups, and the result was a system which directed fewer resources to sicker Black patients than to comparably healthy white patients.

Emergency management faces the same structural problem. Average response time, aggregate damage assessment accuracy, and overall resource utilization efficiency are all proxies for the outcomes that actually matter, and none of them require equitable distribution of improved outcomes across population groups to show improvement. A system which reduces the average response time by two minutes while increasing response times in the lowest-income zip codes by four minutes has succeeded by its own metric, but it has also failed the people who needed it most.

The third mechanism is feedback loop amplification, where operational AI systems generate new data that feeds subsequent training cycles. If a resource positioning AI concentrates units in high-income areas in cycle one, the response time data from cycle one reflects that concentration. When the system is retrained, it learns high-income areas are where rapid response is operationally normal. The disparity is encoded as a baseline assumption rather than an artifact of a specific deployment decision. Each retraining cycle potentially reinforces the pattern from the previous one, not through dramatic failures but through the steady accumulation of small optimizations that make the aggregate metric look better while the underlying equity problem gets worse.

Photos courtesy of FEMA

What Makes Emergency Management Especially Vulnerable

Every field that has grappled seriously with algorithmic bias has eventually confronted the same question: what does our historical data actually reflect? Criminal justice found risk assessment tools encoded the effects of racially disparate policing rather than actual recidivism risk. Healthcare found cost data encoded decades of differential access to care.

Emergency management has the same problem, in addition to several characteristics that make it harder to address.

The stakes here are not inconvenience. When a streaming service algorithm recommends the wrong movie, the cost is a bad evening. When a resource positioning AI produces longer response times for a patient in respiratory distress, the cost is measured differently. Emergency management AI operates in a domain where disparate outcomes are measured in preventable deaths.

The communities most likely to experience worsened outcomes from AI-driven emergency management failures are also the communities with the least capacity to identify, document, and challenge those outcomes. Low-income communities, communities of color, non-English speaking residents, people with disabilities, and people experiencing homelessness are populations that carry the most disaster risk and have the least institutional leverage to push back when a system is failing them.

That overlap is not incidental, and the harm is invisible to the people experiencing it. When a credit scoring AI denies someone a loan, they receive a denial letter. When a resource positioning AI results in a longer response time to a specific address, no notification arrives to explain an algorithm contributed to the outcome. The causal chain from system optimization to individual harm is invisible to the person harmed, which means the pressure to correct it rarely comes from below.

The data gap compounds all of this. The same pattern can be found repeatedly in that registries are outdated, single source, not geographically sortable, and inaccessible to dispatch in real time. The populations most invisible to formal emergency planning infrastructure are exactly the populations whose outcomes most need to improve. An AI trained on a registry that excludes homebound elderly patients, equipment-dependent patients, and non-English speaking residents will simply not know those populations exist.

The Accountability Vacuum

FEMA has begun developing guidance on AI adoption in emergency management. It addresses accuracy, interoperability, data governance, and workforce readiness. What it does not do is establish equity auditing requirements, mandate disaggregated outcome reporting by demographic subgroup, or create a mechanism for affected communities to identify or challenge AI-driven inequities. The guidance treats equity as a value rather than a requirement, and values without enforcement mechanisms are aspirations.

Criminal justice did not develop equity requirements for risk assessment tools proactively. ProPublica’s 2016 investigation of the COMPAS recidivism tool found the algorithm was nearly twice as likely to falsely flag Black defendants as high risk compared to white defendants. By the time this finding was published, the tool had influenced sentencing decisions for thousands of people across hundreds of jurisdictions. Unwinding it required confronting the institutional investments accumulated around it. Algorithmic impact assessment requirements eventually followed in several jurisdictions, but they came years after documented harm.

Emergency management is earlier in the adoption cycle, making this the window. Equity auditing requirements established now can become standard practice rather than retrofit, and waiting until harm is documented is a significantly more expensive way to arrive at the same conclusion.

Three Minimum Standards

A comprehensive equity auditing framework for emergency management AI is a research and policy project larger than a single article can accomplish. The minimum standards any such framework must include can and should be described now.

Disaggregated outcome measurement. Any AI system used in emergency management resource allocation, damage assessment, or risk modeling should be required to report outcomes broken down by geography, income level, race, and primary language. Aggregate performance improvement is not a sufficient measure of equity. Emergency management agencies already report response time data but requiring the same data to be reported by geographic and demographic subgroups is not technically demanding. It is politically demanding because it makes inequitable outcomes visible, and visibility is the precondition for addressing them.

Pre-deployment equity impact assessment. Before an AI system goes into emergency management operations, the training data should be audited for known equity indicators. What are the average response times by income quartile in the training dataset? What are the FEMA IA approval rates by applicant race? What populations are absent from the training data entirely, and what does their absence mean for how the system will perform in serving them? This is analogous to environmental impact assessments required before major infrastructure projects. It does not prevent deployment, but it creates documentation of known risks and accountability for monitoring the outcomes the assessment flagged as concerns.

Community accountability mechanisms. The communities most likely to be harmed by AI-driven equity failures are entitled to know when and how AI systems are influencing emergency management decisions that affect them, and to have a formal channel for raising concerns when outcomes are inequitable. Public reporting requirements, algorithmic impact statements, and community advisory processes exist in other sectors, but none of them currently exist in emergency management. Building these structures before AI adoption is normalized is substantially easier than building them after, when the systems are embedded and the institutional investments are accumulated.

Conclusion

The populations at risk from AI-driven equity failures in emergency management are not new populations with new vulnerabilities. They are the same populations emergency management has consistently failed for decades, including low-income communities, communities of color, elderly residents aging in place, people with disabilities, non-English speaking residents, and people experiencing homelessness. AI deployed without equity constraints will fail them more efficiently than systems which came before it.

The field is early enough in this adoption cycle to ask the right questions before the answers are locked in. Criminal justice shows clearly what it costs to ask them afterward.

The algorithm has already seen this call. The question is whether anyone is required to look at what it learned.

ABOUT THE AUTHOR

Emily James

Emily James is a licensed paramedic and EMS educator with more than a decade of experience in pre-hospital emergency care, clinical instruction, and healthcare operations. James holds a public health certificate and dual bachelor's degrees in sociology and psychology. She is currently pursuing a master's degree in emergency management at Columbia Southern University.

DRJ HOT ITEMS
How Artificial Intelligence Can Aid Disaster Response and Recovery
How Artificial Intelligence Can Aid Disaster Response and Recovery
Natural disasters in the U.S. are becoming more frequent and severe, putting immense strain on disaster response resources. However, the...
READ MORE >
Is Your Inclusive Messaging Backed Up By Inclusive Practices?
Subscribe to the Business Resilience DECODED podcast – from DRJ and Asfalis Advisors – on your favorite podcast app. New...
READ MORE >
Employee Well-Being During a Crisis is Your Greatest Asset
If the COVID-19 pandemic has demonstrated anything, it is we are all vulnerable to stress. Yet with every crisis, we...
READ MORE >
Bridging the Response-Recovery Divide in Disaster Management
Bridging the Response-Recovery Divide: A Unified Disaster Management Strategy
When disaster strikes, we traditionally treat response and recovery as distinct phases. This separation is rooted in practical realities. In...
READ MORE >