Military members man computer workstations
U.S. Space Force Guardians monitor workstations in the Combined Space Operations Center (CSpOC) at Vandenberg Space Force Base, California. Photo: David Dozoretz/US Space Force

AI-assisted planning let US forces operate at a pace legacy methods could not match during Operation Epic Fury, helping coordinate more than 13,000 targets in 38 days, with peak demand on the Chief Digital and Artificial Intelligence Office reaching some 20 billion tokens a day. 

That speed is a real advantage. But machine perception today still trails expert humans on hard recognition tasks: in 2024 US testing, reported accuracy for automated object recognition ran well below that of trained analysts. 

When automated tools surface thousands of candidate items a day and a human must clear each one, the human-in-the-loop can quietly thin into a formality. As Col. Ryan Bell of the 101st Airborne recently warned, large language models do not truly understand three-dimensional space, making them ill-suited for developing courses of action.

The problem is not that the Department of Defense is buying AI, but what it is buying and the terms on which it accepts delivery. 

We are procuring systems that are very good at detection, at finding a departure from a baseline, and then handing the operator a bare flag. A flag with no rationale, no supporting evidence, and no stated confidence is difficult to act on and easy to accept without question. 

The failure is not in the field but in the requirement, written long before the system reaches the warfighter, and that is precisely where it can be corrected.

The Automation Bias Trap We Are Buying

Confronted with raw, opaque alerts, operators fall into two well-documented traps. 

One is disuse, where distrust leads them to ignore the system. The other, more dangerous, is misuse fueled by automation bias, where operators defer to the machine because questioning it is slower than accepting it. 

Observers have already warned that over-reliance on AI-assisted targeting can lead to tragedies of the kind seen in mistaken strikes

When an operator cannot see why a model produced an alert, they cannot weigh it against other evidence, and the alert becomes a claim accepted on faith rather than a judgment subjected to scrutiny. The human decision loop is still drawn on the wiring diagram. It is no longer functioning as a check.

What makes this an acquisition problem rather than a training problem is that no amount of operator discipline can recover a rationale the system was never required to produce. 

If the contract did not ask the vendor to expose the model’s reasoning, its evidence, and its uncertainty, then the operator is being asked to scrutinize a black box, and scrutiny of a black box always collapses back into faith. The defect is in the specification, and so is the remedy.

US Air Force air battle manager evaluates AI battle management tools during the MASH human-machine teaming experiment in Las Vegas. Photo: Debora Henley/DVIDS
US Air Force air battle manager evaluates AI battle management tools during the MASH human-machine teaming experiment in Las Vegas. Photo: Debora Henley/DVIDS

From Opaque Detection to Reviewable Evidence

The Department must stop buying detection and start buying understanding. 

A system that only detects tells the operator that something departed from a baseline. A system built for understanding tells the operator what departed, where, against which baseline, with what supporting evidence, and with what stated confidence, and it states what it does not know. 

It presents alternative explanations and flags missing evidence, functioning as an investigative support tool that organizes the search for a cause while leaving the determination of that cause to the human, who holds the operational context the model lacks.

None of this is a research aspiration; these are properties a contract can demand and a test can verify. A model that cannot explain its rationale, characterize its own strengths and weaknesses, and convey its uncertainty is not ready for the warfighter, and the acquisition system is the only place with the leverage to say so before the system is fielded rather than after it fails.

The Call to Action for Acquisition

This transformation will not happen on its own. Vendors optimize for what the contract rewards, and today the contract rewards detection accuracy and little else. The CDAO and the Program Executive Offices must change what they are willing to buy.

First, any AI system moving from test to production should be required to implement a standardized, explainable alert format as a condition of that transition. 

The gate between a promising prototype and a fielded capability is the single highest-leverage control point in the lifecycle, and it is currently permitting opaque systems through. A model that cannot convey its rationale and its uncertainty should not pass that gate.

Second, operational testing communities must evaluate the human-machine interface with the same rigor they apply to the underlying algorithm. It is not enough to measure the false-positive rate of the code. 

U.S. and Chilean air force personnel discuss intelligence mission planning during Exercise Salitre 2026 at the Combined Air Operations Center in Chile. Photo: Master Sgt. Ceaira Tinsley/DVIDS
US and Chilean air force personnel discuss intelligence mission planning during Exercise Salitre 2026 at the Combined Air Operations Center in Chile. Photo: Master Sgt. Ceaira Tinsley/DVIDS

Testing must measure whether the system’s explanation actually improves operator decision time and reduces the rate of blind acceptance, because an interface that produces confident-looking output the operator cannot interrogate is worse than no automation at all.

Third, in alignment with Directive 3000.09 and the Government Accountability Office’s AI Accountability Framework, program managers must write continuous monitoring into the lifecycle rather than treating a system as finished at delivery. 

Baselines age, the relationship between inputs and expected behavior shifts, a phenomenon known as concept drift, and detection models themselves can be targeted by adversaries who learn to stay just inside the baseline. 

A fielded system must be required to signal its own reduced reliability when its inputs degrade or drift, and that requirement, too, belongs in the contract.

What the Contract Can’t Skip 

Artificial intelligence should not replace human judgment. It should give human judgment something it can actually examine. 

As the tempo of warfare accelerates toward machine speed, our safety and our strategic stability depend on operators who can understand what their tools are telling them, and operators can only understand what the acquisition system required the tools to explain. 

If the Department keeps procuring AI that detects without explaining, it is not empowering the warfighter at all; it is buying a faster, deadlier rubber stamp, one contract at a time.


Headshot Burak Oktenli

Burak Oktenli holds an MBA and is pursuing a Master of Professional Studies in Applied Intelligence at Georgetown University, where his research focuses on the governance of autonomous and AI-enabled military systems.

Author’s note: A generative AI assistant was used for language editing. All analysis, source selection, and conclusions are the author’s own.


The views and opinions expressed here are those of the author and do not necessarily reflect the editorial position of Military AI.

Have a perspective to add? See our Write for Us page.

You May Also Like

The Pentagon Is Ready for AI’s Next Phase — If It Takes These Two Steps First

AI is ready to transform the battlefield, but the DoD must build trust in AI decision-making and standardize governance and security before it can safely and effectively scale its use.

Why Can’t the Pentagon Commit to Anthropic’s Red Lines?

The demand for AI without usage limits exposes a fundamental constitutional and ethical question: should the US military be allowed to deploy AI for mass surveillance and autonomous weapons without meaningful safeguards?

America Must Win the AI Race in the Gulf

Whoever anchors the Gulf’s AI infrastructure will shape the global balance of power — and America must ensure it’s not China.

America’s AI Advantage Relies on Leverage

In the AI race, leverage beats isolation and strategy outperforms broad restrictions.