Skip to content
EDN NEWS
security

How military exercises quietly test the assumptions written inside defense plans

A defense plan is a stack of numbered beliefs, and the exercise is the instrument built to find out which of them break first when an adversary is allowed to win.

How military exercises quietly test the assumptions written inside defense plans
AI-generated photorealistic reconstruction — not a documentary photograph.

An exercise is where a defense plan's assumptions — response times, magazine depth, command latency, alliance cohesion — must survive contact with an adversary allowed to win. The stakes were public in 2002: Millennium Challenge 2002 cost about $250 million and involved some 13,500 participants, per Defense Department statements, and its results were disputed for years.

What follows is the mechanism: how assumptions get written down, how exercises measure them, and what a passing or failing grade actually changes.

What is a plan's assumption, and how do you even write one down?

Every defense plan is a stack of numbered beliefs. A cruise missile defense plan assumes a raid arrives in salvos of a certain size, that cueing arrives within a stated number of minutes, that a stated share of interceptors fires successfully, that runway repair crews turn a cratered strip around in a stated number of hours. None of these numbers is a law of nature. Each is a claim that the plan bets lives on, and the exercise is the instrument that tests the bet.

The discipline is to write the assumptions down before the exercise, in measurable form. Air defense works is not testable. The sector can process a two-wave raid of 40 cruise missiles with the allocated sensors and retain a 60 percent interceptor magazine is — and per standard exercise-design practice in Western militaries, each such claim becomes a measurable line in the exercise's event list.

Soft assumptions get written down too. A NATO air defense plan assumes a partner's radar picture crosses the border in a stated number of minutes and that liaison procedures survive the first confusing hour; alliance exercises exist partly to time exactly those handshakes. If an assumption cannot be expressed in minutes, percentages, or counts, the design problem is not the exercise — it is the assumption.

What is the difference between a scripted exercise and free play?

A scripted exercise guarantees its training objectives: the raid arrives on schedule, the simulated missile fires, the crew gets its repetitions. Free play guarantees nothing. An opposing force — a standing unit, or a red team built for the event — is given objectives, resources, and permission to compete within the rules of the world, and the plan under test must survive whatever they invent.

The most argued case on record is Millennium Challenge 2002, a US Joint Forces Command experiment held in the summer of 2002. Its red force commander, retired Marine Corps Lieutenant General Paul Van Riper, used low-technology workarounds — reported to include motorcycle messengers once his communications were being monitored — and sank a substantial part of the blue fleet in simulation. The experiment was then extended and parts of the scenario reset, and accounts of what that reset means for the results remain disputed between participants and the command to this day. That dispute is itself the lesson: an exercise whose assumptions are tested too hard invites an argument about the test.

The caveat matters here: most public accounts of the 2002 experiment come from participants and journalists rather than its classified findings, and exercise information presented secondhand is contested and hard to verify.

How is a live exercise different from a simulated one?

Western exercise design sorts forces into three tiers, usually abbreviated LVC: live — real aircraft, ships, batteries on real ranges; virtual — real crews operating simulators; and constructive — computer models that move thousands of entities no budget could field. Each tier measures something the others cannot. A live missile firing proves a crew and a round performed; a virtual run proves a crew's decisions hold under load at tolerable cost; a constructive model proves a plan survives scale, which is the assumption live exercises are too small to touch.

The seams between the tiers are where exercises get hard, and where they are most informative. A constructive raid arriving slightly out of sync with the live players tests exactly what a plan assumes about timing — per exercise-design doctrine, the integration of live, virtual, and constructive forces is itself a skill that commands rehearse, because a wartime headquarters would face the same mixed picture of real tracks and correlated estimates.

Related stories: What force posture means and how a foreign base quietly becomes a commitment · What an airspace control order actually coordinates when jets, drones and missiles share one sky.

Who decides whether the plan passed?

Not the senior officer watching the floor. Assessment is a separate profession, with its own instruments. The exercise is built against a master event list; range instrumentation and simulation systems log every detection, engagement, and message; a white force — exercise control — adjudicates what happens where live players and simulation meet; and operations research analysts score the plan's assumptions one by one. Per US joint exercise-design practice, the outputs are an immediate debrief for the players and a formal after-action report for the staff.

InstrumentWhat it measuresWho owns it
Master event listWhether each planned event and assumption was actually exercisedExercise designers
Range and simulation logsDetection ranges, engagement outcomes, timingRange and simulation providers
White force adjudicationWhat happened at the seams between live, virtual, and constructive forcesExercise control
After-action reportWhether the plan's assumptions held, with dataAssessment and operations research cell

The rule that keeps the process honest is separation: the people who ran the force are not the people who grade it.

How does an exercise result actually change a plan?

Through a loop that serious militaries run deliberately:

  1. The after-action report documents each failed assumption with the data behind it.
  2. Staffs direct changes — doctrine, tactics, force mix, or procurement priorities.
  3. The changes get trained into the force.
  4. The next exercise cycle re-tests the specific assumptions that failed.

The loop only works if someone owns it, which is why capable militaries stand up permanent assessment cells rather than borrowing analysts per event.

The results can reach the top. NATO's 30-30-30-30 readiness goal — 30 battalions, 30 air squadrons, 30 combat ships ready to move within 30 days, announced by the alliance's military leadership in 2023 — followed an exercise-and-readiness review in which the alliance's own statements identified current response times as the binding constraint. A plan assumption about how fast forces can move was measured, found short, and converted into a public target.

Institutional examples go deeper. The Air Force's Red Flag exercise at Nellis Air Force Base, run since 1975 per the service's own account, exists because Vietnam-era loss-rate analysis showed crew survival climbing sharply after roughly the first ten combat sorties; the exercise was built to buy those first ten sorties in peacetime, against a free-play adversary, before they could be bought in war. That is assumption-testing expressed as a permanent institution.

What can an exercise never prove?

Three things, permanently. First, fidelity: the adversary is simulated or surrogate, and the real adversary's best systems never appear in full — publicly available sources do not establish how closely any exercise red force replicates a specific nation's current capabilities. Second, sample size: one exercise is one experiment, not a statistic, and an assumption that held once can fail under different weather, tempo, or attrition. Third, the observer effect: everyone knows the exercise is coming, a courtesy war never extends.

The honest formulation: an exercise validates a plan's assumptions only under the assumptions of the exercise. Printing that sentence at the top of every after-action report would make those documents shorter and more useful.

Why is finding failure the whole point?

Because the alternative is finding it in an engagement. An exercise that always ends with the plan vindicated is not an exercise; it is a ceremony with a budget line. The professional standard, visible in everything from red force rules to after-action conventions, holds that a documented failure in summer is cheaper than an undocumented one later — and the mechanism that makes the trade possible is writing assumptions down precisely enough that they can lose.

That is what the money buys. Not the exercise — the refutation.

Frequently Asked Questions

What was Millennium Challenge 2002?
A large US Joint Forces Command experiment held in the summer of 2002, costing about $250 million with some 13,500 participants, per Defense Department statements at the time. Its red force used asymmetric tactics to sink much of the blue fleet in simulation; the subsequent scenario extension and reset remain disputed between participants and the command.
What is a red force in a military exercise?
The designated adversary. A red force, opposing force, or red team is given objectives and resources approximating a real enemy and allowed to act on its own initiative within exercise rules. Its purpose is not to win but to force the plan under test to demonstrate its assumptions against resistance that refuses to cooperate with the script.
What is a master event list?
The exercise's script and test sheet in one: a sequenced list of events and scenarios built so that every assumption in the plan under test is exercised at least once and its outcome can be measured. Exercise designers write it before the event; analysts use it afterward to connect results to the specific assumptions each event was designed to check.
Do exercises ever change weapons procurement?
Yes, indirectly but measurably. When an after-action report documents that a sensor's assumed detection range or an interceptor's assumed availability did not hold, the finding enters the staff process that sets capability requirements, and those requirements then compete for budget. Exercise findings rarely buy a weapon alone, but they supply the documented shortfall a program needs.
Why do militaries rehearse against their own allies?
Because interoperability assumptions are plans too. An air defense plan that assumes a partner nation's radar picture arrives within minutes is testing communications, procedures, and trust as much as hardware. Defending alongside allies exposes exactly those seams in peacetime, at a cost measured in schedule friction rather than operational consequences.