MissionSAFE / Design study

Interpreting assurance scores: define the mission boundary

An Air Force logistics specialist and two civilian colleagues review a laptop beside stacked cargo boxes.

The operational question

In this illustrative missionsafe scenario, a evaluation lead uses AI to compare model and policy results. The workflow draws on versioned scenarios and labeled failure evidence. Its central risk is an aggregate score concealing a consequential miss. The design question is how to publish an assurance assessment for the designated review group while preserving the limits of the evaluated scenario population. This is a proposed evaluation scenario, not a report of an Archetypal customer deployment or a demonstrated operational outcome.

Define the mission boundary

The first governance decision is where the system’s authority begins and ends. Write the intended task as a bounded activity with an identifiable owner. Specify the information the workflow may use, the audience it serves, and the decisions it may support. Then identify the actions that remain outside that permission. A description of what a model can do is not an authorization to do it. A good boundary makes change visible. If the audience, information class, tool access, or intended use changes, the workflow should be able to recognize that it is operating under a new set of conditions. The owner can then decide whether an existing approval still applies or whether another review is required.

Put the control in the workflow

Place this review immediately before the team can publish an assurance assessment. The evaluation lead should see the proposed result beside the relevant parts of versioned scenarios and labeled failure evidence. Identify which statement is supported by a source, which is an interpretation, and which remains unresolved. Carry the limits of the evaluated scenario population into the decision record rather than relying on a reviewer to remember it from another screen. If the evidence does not establish the condition required for release, route the case to its owner with a concrete question. The interface should make the missing fact discoverable and the next action clear.

A test that can change the design

A task instruction expands to a new audience after the initial approval. The expected result is a visible boundary check, not a silent extension of permission. Run the case using a fixed version of the scenario and the policy under review. Ask an independent reviewer to identify the decisive fact before seeing the system’s disposition. Compare that interpretation with the result. Where they disagree, preserve both explanations and inspect whether the difference comes from the rule, the available evidence, or the interface. For interpreting assurance scores, include model scope, scoring criteria, and unresolved failures in the review packet. Repeat the test after a correction and retain the original failure as part of the evidence.

Evidence to retain

The minimum useful record connects the purpose of the task, model scope, scoring criteria, and unresolved failures, the applicable policy version, and the final disposition. Add the identity or role of the responsible reviewer, the conditions attached to approval, and the unresolved questions. If the team proceeds, distinguish the approval from an observed completion. If it stops, explain what evidence or authorization would allow another review. Keep source permissions attached to the record when it moves to the designated review group. Do not assume that permission to read the initial source includes permission to reproduce it in every downstream system.

What a result would establish

A successful run would show that this configuration recognizes the tested boundary for interpreting assurance scores and gives the evaluation lead an interpretable next step. It would not establish complete coverage of other audiences, source conditions, applications, or mission environments. Report the scope with the finding. Review any decision to publish an assurance assessment under changed conditions as a new applicability question. The strongest next experiment is usually the smallest change that could make the current conclusion false.

Review before wider use

Ask the workflow owner whether the proposed control is understandable at the point of use. Ask the policy owner whether it preserves the source requirement. Ask the evaluator whether the test can distinguish a real improvement from a change in presentation. Finally, ask the deployment owner what happens when versioned scenarios and labeled failure evidence are unavailable or the integration no longer observes the required event. Agreement among these roles should be documented as a set of decisions and remaining conditions, not compressed into an unsupported statement that the system is universally ready.

Purpose in this scenario

State the immediate task and the downstream use separately. The same output may be acceptable for an internal draft but unsuitable for a consequential decision or a broader audience. In interpreting assurance scores, the evaluation lead should apply this check to the proposed decision to publish an assurance assessment. Use model scope, scoring criteria, and unresolved failures to make the review concrete. Explain how the result changes if the condition is absent, disputed, or no longer current. Record the expected disposition before running the scenario so that the evaluator cannot quietly redefine success after seeing the output. The receiving audience is the designated review group; preserve the limitations they need to interpret the result.

Authority in this scenario

Identify the role empowered to make the decision. Record the source of that authority and any conditions attached to a delegation. In interpreting assurance scores, the evaluation lead should apply this check to the proposed decision to publish an assurance assessment. Use model scope, scoring criteria, and unresolved failures to make the review concrete. Explain how the result changes if the condition is absent, disputed, or no longer current. Record the expected disposition before running the scenario so that the evaluator cannot quietly redefine success after seeing the output. The receiving audience is the designated review group; preserve the limitations they need to interpret the result.

Audience in this scenario

Name the intended receiving group. Reassess the decision if the result is forwarded, summarized for another group, or included in a new workflow. In interpreting assurance scores, the evaluation lead should apply this check to the proposed decision to publish an assurance assessment. Use model scope, scoring criteria, and unresolved failures to make the review concrete. Explain how the result changes if the condition is absent, disputed, or no longer current. Record the expected disposition before running the scenario so that the evaluator cannot quietly redefine success after seeing the output.

Authority.

Archetypal film

Documentary footage · No dialogue · Source credits