Research / Design study

Frontier-model evaluation: build a useful decision record

A civilian software engineer discusses a project with an Air Force product manager at a shared computer desk.

The operational question

In this illustrative research scenario, a evaluation lead uses AI to compare behavior under a defined test. The workflow draws on scenario envelopes and recorded model responses. Its central risk is an aggregate score concealing a mission-critical error class. The design question is how to publish an evaluation finding for the authorized review audience while preserving the conditions covered by the evaluation. This is a proposed evaluation scenario, not a report of an Archetypal customer deployment or a demonstrated operational outcome.

Build a useful decision record

An audit trail earns its value by explaining a consequential decision after the original context has changed. Record the request’s governed purpose, the decisive facts, the applicable rule version, the disposition, and the responsible reviewer or system. Include conditions, unresolved questions, and the observed outcome when available. Keep the difference between a proposed action and a completed action explicit. Retain enough information to reconstruct the decision without treating unlimited capture as the default. Consider who may inspect the record, which source material can be linked rather than copied, and how the record should respond when a source is corrected or a policy is retired.

Put the control in the workflow

Place this review immediately before the team can publish an evaluation finding. The evaluation lead should see the proposed result beside the relevant parts of scenario envelopes and recorded model responses. Identify which statement is supported by a source, which is an interpretation, and which remains unresolved. Carry the conditions covered by the evaluation into the decision record rather than relying on a reviewer to remember it from another screen. If the evidence does not establish the condition required for release, route the case to its owner with a concrete question. The interface should make the missing fact discoverable and the next action clear.

A test that can change the design

A reviewer examines the record after the policy has changed. The expected result is a reconstructable account of the version and conditions that applied at decision time. Run the case using a fixed version of the scenario and the policy under review. Ask an independent reviewer to identify the decisive fact before seeing the system’s disposition. Compare that interpretation with the result. Where they disagree, preserve both explanations and inspect whether the difference comes from the rule, the available evidence, or the interface. For frontier-model evaluation, include test version, model configuration, and labeled failure evidence in the review packet. Repeat the test after a correction and retain the original failure as part of the evidence.

Evidence to retain

The minimum useful record connects the purpose of the task, test version, model configuration, and labeled failure evidence, the applicable policy version, and the final disposition. Add the identity or role of the responsible reviewer, the conditions attached to approval, and the unresolved questions. If the team proceeds, distinguish the approval from an observed completion. If it stops, explain what evidence or authorization would allow another review. Keep source permissions attached to the record when it moves to the authorized review audience. Do not assume that permission to read the initial source includes permission to reproduce it in every downstream system.

What a result would establish

A successful run would show that this configuration recognizes the tested boundary for frontier-model evaluation and gives the evaluation lead an interpretable next step. It would not establish complete coverage of other audiences, source conditions, applications, or mission environments. Report the scope with the finding. Review any decision to publish an evaluation finding under changed conditions as a new applicability question. The strongest next experiment is usually the smallest change that could make the current conclusion false.

Review before wider use

Ask the workflow owner whether the proposed control is understandable at the point of use. Ask the policy owner whether it preserves the source requirement.

Archetypal film

Documentary footage · No dialogue · Source credits