An exercise is where a defense plan's assumptions — response times, magazine depth, command latency, alliance cohesion — must survive contact with an adversary allowed to win. The stakes were public in 2002: Millennium Challenge 2002 cost about $250 million and involved some 13,500 participants, per Defense Department statements, and its results were disputed for years.
What follows is the mechanism: how assumptions get written down, how exercises measure them, and what a passing or failing grade actually changes.
What is a plan's assumption, and how do you even write one down?
Every defense plan is a stack of numbered beliefs. A cruise missile defense plan assumes a raid arrives in salvos of a certain size, that cueing arrives within a stated number of minutes, that a stated share of interceptors fires successfully, that runway repair crews turn a cratered strip around in a stated number of hours. None of these numbers is a law of nature. Each is a claim that the plan bets lives on, and the exercise is the instrument that tests the bet.
The discipline is to write the assumptions down before the exercise, in measurable form. Air defense works is not testable. The sector can process a two-wave raid of 40 cruise missiles with the allocated sensors and retain a 60 percent interceptor magazine is — and per standard exercise-design practice in Western militaries, each such claim becomes a measurable line in the exercise's event list.
Soft assumptions get written down too. A NATO air defense plan assumes a partner's radar picture crosses the border in a stated number of minutes and that liaison procedures survive the first confusing hour; alliance exercises exist partly to time exactly those handshakes. If an assumption cannot be expressed in minutes, percentages, or counts, the design problem is not the exercise — it is the assumption.
What is the difference between a scripted exercise and free play?
A scripted exercise guarantees its training objectives: the raid arrives on schedule, the simulated missile fires, the crew gets its repetitions. Free play guarantees nothing. An opposing force — a standing unit, or a red team built for the event — is given objectives, resources, and permission to compete within the rules of the world, and the plan under test must survive whatever they invent.
The most argued case on record is Millennium Challenge 2002, a US Joint Forces Command experiment held in the summer of 2002. Its red force commander, retired Marine Corps Lieutenant General Paul Van Riper, used low-technology workarounds — reported to include motorcycle messengers once his communications were being monitored — and sank a substantial part of the blue fleet in simulation. The experiment was then extended and parts of the scenario reset, and accounts of what that reset means for the results remain disputed between participants and the command to this day. That dispute is itself the lesson: an exercise whose assumptions are tested too hard invites an argument about the test.
The caveat matters here: most public accounts of the 2002 experiment come from participants and journalists rather than its classified findings, and exercise information presented secondhand is contested and hard to verify.
How is a live exercise different from a simulated one?
Western exercise design sorts forces into three tiers, usually abbreviated LVC: live — real aircraft, ships, batteries on real ranges; virtual — real crews operating simulators; and constructive — computer models that move thousands of entities no budget could field. Each tier measures something the others cannot. A live missile firing proves a crew and a round performed; a virtual run proves a crew's decisions hold under load at tolerable cost; a constructive model proves a plan survives scale, which is the assumption live exercises are too small to touch.
The seams between the tiers are where exercises get hard, and where they are most informative. A constructive raid arriving slightly out of sync with the live players tests exactly what a plan assumes about timing — per exercise-design doctrine, the integration of live, virtual, and constructive forces is itself a skill that commands rehearse, because a wartime headquarters would face the same mixed picture of real tracks and correlated estimates.
Related stories: What force posture means and how a foreign base quietly becomes a commitment · What an airspace control order actually coordinates when jets, drones and missiles share one sky.
Who decides whether the plan passed?
Not the senior officer watching the floor. Assessment is a separate profession, with its own instruments. The exercise is built against a master event list; range instrumentation and simulation systems log every detection, engagement, and message; a white force — exercise control — adjudicates what happens where live players and simulation meet; and operations research analysts score the plan's assumptions one by one. Per US joint exercise-design practice, the outputs are an immediate debrief for the players and a formal after-action report for the staff.
| Instrument | What it measures | Who owns it |
|---|---|---|
| Master event list | Whether each planned event and assumption was actually exercised | Exercise designers |
| Range and simulation logs | Detection ranges, engagement outcomes, timing | Range and simulation providers |
| White force adjudication | What happened at the seams between live, virtual, and constructive forces | Exercise control |
| After-action report | Whether the plan's assumptions held, with data | Assessment and operations research cell |
The rule that keeps the process honest is separation: the people who ran the force are not the people who grade it.
How does an exercise result actually change a plan?
Through a loop that serious militaries run deliberately:
- The after-action report documents each failed assumption with the data behind it.
- Staffs direct changes — doctrine, tactics, force mix, or procurement priorities.
- The changes get trained into the force.
- The next exercise cycle re-tests the specific assumptions that failed.
The loop only works if someone owns it, which is why capable militaries stand up permanent assessment cells rather than borrowing analysts per event.
The results can reach the top. NATO's 30-30-30-30 readiness goal — 30 battalions, 30 air squadrons, 30 combat ships ready to move within 30 days, announced by the alliance's military leadership in 2023 — followed an exercise-and-readiness review in which the alliance's own statements identified current response times as the binding constraint. A plan assumption about how fast forces can move was measured, found short, and converted into a public target.
Institutional examples go deeper. The Air Force's Red Flag exercise at Nellis Air Force Base, run since 1975 per the service's own account, exists because Vietnam-era loss-rate analysis showed crew survival climbing sharply after roughly the first ten combat sorties; the exercise was built to buy those first ten sorties in peacetime, against a free-play adversary, before they could be bought in war. That is assumption-testing expressed as a permanent institution.
What can an exercise never prove?
Three things, permanently. First, fidelity: the adversary is simulated or surrogate, and the real adversary's best systems never appear in full — publicly available sources do not establish how closely any exercise red force replicates a specific nation's current capabilities. Second, sample size: one exercise is one experiment, not a statistic, and an assumption that held once can fail under different weather, tempo, or attrition. Third, the observer effect: everyone knows the exercise is coming, a courtesy war never extends.
The honest formulation: an exercise validates a plan's assumptions only under the assumptions of the exercise. Printing that sentence at the top of every after-action report would make those documents shorter and more useful.
Why is finding failure the whole point?
Because the alternative is finding it in an engagement. An exercise that always ends with the plan vindicated is not an exercise; it is a ceremony with a budget line. The professional standard, visible in everything from red force rules to after-action conventions, holds that a documented failure in summer is cheaper than an undocumented one later — and the mechanism that makes the trade possible is writing assumptions down precisely enough that they can lose.
That is what the money buys. Not the exercise — the refutation.
