AI-assisted planning let US forces operate at a pace legacy methods could not match during Operation Epic Fury, helping coordinate more than 13,000 targets in 38 days, with peak demand on the Chief Digital and Artificial Intelligence Office reaching some 20 billion tokens a day.
That speed is a real advantage. But machine perception today still trails expert humans on hard recognition tasks: in 2024 US testing, reported accuracy for automated object recognition ran well below that of trained analysts.
When automated tools surface thousands of candidate items a day and a human must clear each one, the human-in-the-loop can quietly thin into a formality. As Col. Ryan Bell of the 101st Airborne recently warned, large language models do not truly understand three-dimensional space, making them ill-suited for developing courses of action.
The problem is not that the Department of Defense is buying AI, but what it is buying and the terms on which it accepts delivery.
We are procuring systems that are very good at detection, at finding a departure from a baseline, and then handing the operator a bare flag. A flag with no rationale, no supporting evidence, and no stated confidence is difficult to act on and easy to accept without question.
The failure is not in the field but in the requirement, written long before the system reaches the warfighter, and that is precisely where it can be corrected.
The Automation Bias Trap We Are Buying
Confronted with raw, opaque alerts, operators fall into two well-documented traps.
One is disuse, where distrust leads them to ignore the system. The other, more dangerous, is misuse fueled by automation bias, where operators defer to the machine because questioning it is slower than accepting it.
Observers have already warned that over-reliance on AI-assisted targeting can lead to tragedies of the kind seen in mistaken strikes.
When an operator cannot see why a model produced an alert, they cannot weigh it against other evidence, and the alert becomes a claim accepted on faith rather than a judgment subjected to scrutiny. The human decision loop is still drawn on the wiring diagram. It is no longer functioning as a check.
What makes this an acquisition problem rather than a training problem is that no amount of operator discipline can recover a rationale the system was never required to produce.
If the contract did not ask the vendor to expose the model’s reasoning, its evidence, and its uncertainty, then the operator is being asked to scrutinize a black box, and scrutiny of a black box always collapses back into faith. The defect is in the specification, and so is the remedy.

From Opaque Detection to Reviewable Evidence
The Department must stop buying detection and start buying understanding.
A system that only detects tells the operator that something departed from a baseline. A system built for understanding tells the operator what departed, where, against which baseline, with what supporting evidence, and with what stated confidence, and it states what it does not know.
It presents alternative explanations and flags missing evidence, functioning as an investigative support tool that organizes the search for a cause while leaving the determination of that cause to the human, who holds the operational context the model lacks.
None of this is a research aspiration; these are properties a contract can demand and a test can verify. A model that cannot explain its rationale, characterize its own strengths and weaknesses, and convey its uncertainty is not ready for the warfighter, and the acquisition system is the only place with the leverage to say so before the system is fielded rather than after it fails.
The Call to Action for Acquisition
This transformation will not happen on its own. Vendors optimize for what the contract rewards, and today the contract rewards detection accuracy and little else. The CDAO and the Program Executive Offices must change what they are willing to buy.
First, any AI system moving from test to production should be required to implement a standardized, explainable alert format as a condition of that transition.
The gate between a promising prototype and a fielded capability is the single highest-leverage control point in the lifecycle, and it is currently permitting opaque systems through. A model that cannot convey its rationale and its uncertainty should not pass that gate.
Second, operational testing communities must evaluate the human-machine interface with the same rigor they apply to the underlying algorithm. It is not enough to measure the false-positive rate of the code.

Testing must measure whether the system’s explanation actually improves operator decision time and reduces the rate of blind acceptance, because an interface that produces confident-looking output the operator cannot interrogate is worse than no automation at all.
Third, in alignment with Directive 3000.09 and the Government Accountability Office’s AI Accountability Framework, program managers must write continuous monitoring into the lifecycle rather than treating a system as finished at delivery.
Baselines age, the relationship between inputs and expected behavior shifts, a phenomenon known as concept drift, and detection models themselves can be targeted by adversaries who learn to stay just inside the baseline.
A fielded system must be required to signal its own reduced reliability when its inputs degrade or drift, and that requirement, too, belongs in the contract.
What the Contract Can’t Skip
Artificial intelligence should not replace human judgment. It should give human judgment something it can actually examine.
As the tempo of warfare accelerates toward machine speed, our safety and our strategic stability depend on operators who can understand what their tools are telling them, and operators can only understand what the acquisition system required the tools to explain.
If the Department keeps procuring AI that detects without explaining, it is not empowering the warfighter at all; it is buying a faster, deadlier rubber stamp, one contract at a time.

Burak Oktenli holds an MBA and is pursuing a Master of Professional Studies in Applied Intelligence at Georgetown University, where his research focuses on the governance of autonomous and AI-enabled military systems.
Author’s note: A generative AI assistant was used for language editing. All analysis, source selection, and conclusions are the author’s own.
The views and opinions expressed here are those of the author and do not necessarily reflect the editorial position of Military AI.
Have a perspective to add? See our Write for Us page.