1. Status and derivation§
What this is. The published specification of how StepSim generates AIRS-J/1 observations. StepSim is AIRS-J/1's official instrument (D86). There is one judgment standard: AIR APAC's AIRS-J/1 (AIR_APAC/standards/AIRS-J-1.md). This file says how StepSim's event log becomes AIRS-J/1 observations. It sets no dimension, anchor, threshold or band of its own.
Derives from AIRS-J/1 v2.0 and MASTER_BLUEPRINT.md v1.14. Where this file conflicts with AIRS-J/1, AIRS-J/1 wins on what is assessed; where it conflicts with the blueprint, the blueprint wins on how StepSim is built.
Status: approved by the owner, 3 October 2026.
Replaces the StepSim Judgment Standard v1.0.0 (1 October 2026). The four claims C1 to C4, the nine observation themes as a reporting scale, the levels Follows, Checks and Calibrated, the overall-level rule, and the result names Indicative and Assessed are retired. Calibrated reliance (old C1) survives as the evidence family inside T1 (section 3.5).
Unchanged from AIRS-J/1 v1.1 into v2.0, and relied on here: the nine dimensions (§5), the 0 to 4 scale and every anchor (§7, §7.5), three observations per dimension from distinct decision points (§8.3), stakes coverage (§6.4), and the §12 thresholds and bands.
What AIRS-J/1 v2.0 grants this file.
- Rule-based scoring. §11.4 lets rule-based scoring by an approved instrument generate observations directly, with the instrument and rule version in place of an assessor of record (Annex C) and a byte-identical replay as the agreement test (§11.2, section 2.8).
- Selection as statement. §7.1.2: a uniform, optional bank of authored lines with decoys is a response format, not prompting (section 2.7).
- Development readings. §10 and §10.1: Insufficient and Provisional results are development readings, reported with the observation count per dimension and no band (section 5).
- C4's scope. §6.1: C4 includes wrong or misused records about identifiable people, so the steering pack qualifies (section 4.3).
2. How StepSim meets AIRS-J/1§
2.1 Instrument type§
StepSim is a scenario simulation with retained decision trace (§8.1). Each mission is an authored case. AI advice in it has hidden, authored reliability. The person works the case within a review-time budget, makes one call, and then answers a later development. Every act is an event in an append-only log. Outcomes come from a seeded state model with declared probabilities (blueprint 6.6).
2.2 The transcript is the event log§
The AIRS-J/1 transcript is the run record defined by platform/engine/contract.mjs and built by platform/engine/session.mjs. Every observation is derived from it and from the mission package (platform/missions/<id>/package.mjs), never from anything else (§8.8).
Annex C's fields, and where each comes from:
| Annex C field | StepSim source | Status |
|---|---|---|
| Observation ID | {runId}:{dimension}:{decisionPointId} |
Built by the scorer |
| Candidate credential code | participantKey. In Assessment, the person key: one person under one company's licence, the same in every session (D79; 7.3, Identity) |
Built for Assessment (platform/collect/assess.mjs); pseudonymous in Practice and Cohort |
| Dimension | The map in section 3 | Scorer |
| Decision point ID | A package item id (section 2.4) | Package ids exist; decisionPoints list to add |
| Level awarded | The rule in section 3 | Scorer |
| Anchor text applied | AIRS-J/1 §7.5, verbatim, keyed by dimension and level | Scorer constant |
| Transcript evidence | The cited event ids | Events carry id and order |
| Scaffolding level | Run field scaffolding: none (Assessment), guided (Practice, Cohort) |
To add |
| Access support | Run field accessSupport |
To add |
| Critical-review trigger | Section 3.14 | Scorer |
| Stakes profile | Package field stakes |
To add (section 4) |
| Scoring mode | rule-based instrument |
AIRS-J/1 Annex C, §11.4 |
| Model and prompt version | Not applicable. The scorer version (this file's version and the code commit) is recorded instead. | To add |
| Assessor ID | The instrument and rule version, for example airs.T1.v2 |
AIRS-J/1 Annex C |
| Conflict declaration | Exeter Labs' declared interest, per AIRS §4.9 | Text in AIRS |
| Second scoring | A replay of the same log under the same versions gives the same record. Agreement with human assessors is validation check 5. | Replay exists (replay.mjs per mission) |
| Timestamp | The scorer run time; each cited event keeps wallTime |
Exists |
2.3 Version pinning (§9.6)§
A run states four versions or it is void.
| §9.6 item | Run field | Today |
|---|---|---|
| Scenario version | missionId + caseVersion |
Recorded |
| Rubric version | standardVersion (AIRS-J/1 2.0) and specVersion (this file, 2.0.0) |
Missing |
| Instrument version | contractVersion (3.1.0; a run recorded under 3.0.0 keeps it) and instrumentVersion (the platform build commit) |
Contract version recorded; build commit missing |
| Scaffolding level | scaffolding |
Missing; inferable from mode today, which is not enough |
2.4 Decision points and their integrity (§8.6)§
A StepSim decision point is an authored opportunity with its own id, at which the person commits to a costed act: review time, standing, or a consequence in the state model. The package lists them in a decisionPoints field. The kinds:
| Kind | Id | Where |
|---|---|---|
| The brief | the mission's deadline id + :brief |
First-lean screen: framing, lean, the plan lines (S2) |
| An advice piece or queue item | the advice id | Every load-bearing piece (section 3.5) |
| A week (queue missions) | week-1 … week-4, and controls before week 1 |
The Monday queue's controls and Monday drawer |
| A held item | the advice id with escalates: true |
v26-table in The Monday queue |
| An update | the deadline id + :update |
A requester's question part-way through the desk (package field update: { afterHours }): a re-framing and note lines, no copy-in. The water report's 12:00 message from Adeline (S1, E2, P1) |
| The call | the deadline id | The order: option, conditions, copy-ins, owner, note lines |
| The later choice | the deadline id + :later |
Keep or change, after the time jump |
The brief carries no world cost. It is a commitment because the framing and lean are fixed in the record before any evidence is opened and are read back against the call. If AIRS-J/1 v2.0 reads "carries cost" as world cost only, S1 and S2 need a costed brief.
Integrity. The log is append-only (validateEvent checks strict order). send() records the conditions, copy-ins, owner and decision, then jump() offers the later items. revision and retain may only follow the decision or the expiry, and a run has one decision or one time_expired, never both. So the position is recorded before the consequence is revealed, and nothing overwrites it. A later revision is a P2 observation, never a correction to P1. In The Monday queue, controls drafted on the controls screen are recorded when week 1 opens, before the first item; a change in a week is recorded at once and takes effect next week.
2.5 Evidence minimisation and no affect inference (§8.7)§
The log keeps acts, their targets, review time spent and wallTime. Nothing else is captured: no keystrokes, no pointer trace, no video, no voice. wallTime is phase timing only and enters no rule. StepSim infers nothing about stress, emotion, hesitation or confidence from timing (blueprint 6.10, D68). Stated confidence is the person's own answer; it is calibration data and enters no observation (R5 below).
2.6 Prompting and scaffolding (§7.1)§
Scaffolding in StepSim today, all in Practice and Cohort:
- Highlighted cue spans. A cue sitting in an advice line is drawn as a marked span with its own "Flag this" button (
platform/missions/<id>/screens.mjs, for example the-friday-file lines 163 to 169). That points at the flaw. It is content-directed. - Unlock notices. "A new limit is now on offer in your call: …" (
DESK.unlockedin each mission'scopy.mjs) directs the person to a condition. - The pre-play cues and worked examples kept by D45, coaching, rewind and branch runs.
Not scaffolding: the orientation and glossary (platform/surfaces/orient.mjs). They explain the interface in words shared by every mission, never name a cue, a right call or a sound piece, and never block Start. They are access support under §7.1.1 and impose no cap.
The rules.
- Every Practice and Cohort run records
scaffolding: guided. An observation at a decision point a scaffold touched in that run (a highlighted span on that piece; an unlock notice before the call) is capped at Level 1. Other observations in the run are scored as they fall. - A branch or rewound run (
runOrdinal≥ 2, or aparentRunId) never yields an observation. - Only Assessment runs yield certifying observations. Assessment records
scaffolding: none: no highlighted spans (every line and tile is flaggable alike), no unlock notices, no coaching, no rewind, no reliability labels or expert contrast shown to the person. With no span marked, every piece of advice, queue line and pack line carries its own flag, drawn the same on a flawed piece and a plain one. A flag on a piece logs on the cue in its words where it has one, so a piece flag on a flawed piece is a flag of that cue (3.2); otherwise it logs on the piece's declared distractor cue (3.3). The person need not find the exact words. - AI advice inside the case is stimulus, not scaffolding. But a behaviour the AI suggested and the person adopted by "Use this" is credited to T1 as reliance on that piece, never to the dimension of the behaviour itself (R4).
2.7 Selection as statement: the note line§
StepSim has no free text, because no LLM may read a person's words into a count (blueprint 6.1). Where an anchor says "states", "names" or "records", StepSim observes the person selecting an authored line into their work product.
A selection counts as a statement only when all of these hold:
- the line comes from a line bank shown the same way at every decision point of its kind and in every mission, optional, never required, never highlighted;
- the bank mixes lines of several dimensions and elements, with no dimension or element label on screen;
- the bank holds authored distractors: lines that do not fit this case version;
- a line that names a fact is on offer only once that fact is on the desk (the same
unlockedByrule conditions use); - every piece carries claims lines, so which pieces have them says nothing about which are scored; the claims lines of one piece (
group: "claims:<piece id>") are one choice, "no line" by default, drawn only once the person has opened the piece (opened a case, answered or flagged a piece, or checked it as 3.5 reads a check: a linked file opened outside a queue, a probing question to the AI, a question taggedchecksit; in The Monday queue a case's pair is on offer once the case is opened or Vela is asked a question that probes it), so a sound piece checked and left standing can be recorded as checked, and two lines of one group cannot be chosen together on screen; the scorer still reads two as contradictory (R6), and reads the claims lines of load-bearing pieces only; - each line's hidden tags (dimension, element, fits or distractor, and for some lines
restsOn,classOf,corrects) are case readings with provenance, like reliability labels (blueprint 6.9).
Conditions, controls, framings and copy-ins are selections of the same kind and follow the same rules.
2.8 Scoring§
The scorer is deterministic code applying section 3 to the log and the package. No LLM grades, labels, levels, or produces evidence a level uses. The same log, package and versions always give the same observation records. Each record cites its rule id and every event id it rests on.
2.9 What stays from version 1§
The closed act vocabulary, each mission declaring its subset. Authored, hidden reliability, read against what the person did. One observable, one dimension, now stated as AIRS Annex B's rule (R1). Act verbs only in every template: opened, flagged, asked, challenged, set, sampled; never noticed, identified or realised. Provenance on every case reading. The reliability mix, the costly refusal and the blanket-strategy gate (blueprint 6.3, 6.7). Process and outcome read apart, with the four cells. A version of this file frozen before any result is read.
3. The act-to-dimension map§
3.1 Rules for every dimension§
- R1. One event, one dimension. Each event is cited under one dimension only, by the allocation in 3.2 (Annex B).
- R2. One event, one observation. An event supports at most one observation. Where two observations of the same dimension could cite it, the earliest decision point takes it.
- R3. Context reads. A rule may read package facts, and the presence or absence of events, as preconditions without citing them: what was available, proportion across a mission, the run's process cell. Only cited events are evidence.
- R4. Own acts only. A condition pre-ticked by "Use this" and left as it was, or a mandate rule pre-ticked from a control the person set, is not cited; the accept or the control is. The engine records
source: own | advice | controloncondition_set(to build). - R5. Calibration data.
initial_leanandconfidenceare never cited. The lean may be read as context (T2 Level 3). - R6. Clean. A level that needs a fitting line, condition or control fails if a distractor of the same element was also chosen at the same decision point. Lines of one kind (dimension, element and piece) whose values exclude each other count as each other's distractor: a claim said checked and unchecked for one piece, two cells, owning and deflecting. A set of conditions and controls does not discriminate when it holds everything on offer at its point, and caps T3 at 1 (3.7); for T1's addressing and a week's correction it also does not discriminate when it holds a distractor (3.5, 3.11), while at T3 each distractor in it lowers the level by one (3.7). Choosing everything never pays.
- R7. Prompted cap. Section 2.6, rule 1.
- R8. Cumulative. A level is awarded only when every lower level's requirement also holds (§7.4).
3.2 Event allocation§
| Event | Dimension | Decision point |
|---|---|---|
framing_choice |
S1 | The brief; an update; or the call for a re-framing |
note_line with a plan element (new) |
S2 | The brief |
accept, challenge, override, verify, item_sampled |
T1 | The piece |
query declared ai_calibration |
T1 | The piece it probes |
cue_flagged, with its reading (new field) |
T2 | The flawed piece the cue reveals |
query declared signal_detection |
T2 | The piece it probes |
query tagged checks p |
Its theme's dimension, as above; T1 reads it as a check of p, as context (R3), when that dimension is not T1 | |
evidence_opened declared disconfirms of a piece it did not itself verify |
T2 | That piece |
evidence_opened that tests nothing it did not verify |
Not cited; its verify is T1's |
|
condition_set (own), checkpoint_set, permission_set |
The dimension the package declares on each condition or control | The call, the controls (queue missions: set before the first period), or the week |
escalation to a target declared approval |
E1 | The call |
escalation to a target declared inform |
E2 | The call |
owner_named, escalation_response |
E1 | The call, the controls, the held item |
query declared trade_off_judgment |
E2 | The call |
decision, time_expired |
P1 | The call |
revision, retain, paused, revoked |
P2 | The later choice, or the week (P2 at the call cites P2 lines only; the option, changes and replies it reads are context, R3) |
note_line |
The dimension its tags declare | Where it was added |
initial_lean, confidence |
None (R5) | |
evidence_offered, advice_offered |
None (system) |
Queries keep their theme field; the values map to dimensions as above. Old rules X1 to X3 are retired: a challenge is T1 only; an opening is cited as its verify, or for T2 on a piece it disconfirms and did not verify (a file linked to one piece can test another), or not at all.
3.3 New acts and fields, the minimum§
| Change | Where | Why |
|---|---|---|
note_line act: target a line id from the mission's lines bank |
contract.mjs, session.mjs, each mission's package.mjs and screens.mjs |
Every "states" element: S1 L3–4, S2, T1 L3, T2 L2–3, T3 L3, E1 L3, E2 L2–4, P1 L2–4, P2 L2–4, and every standing rule (L4). One act; the line's tags carry the dimension. |
cue_flagged gains reading: one of 2 to 3 authored readings of the span |
same | T2 Level 2: "names … at least one competing explanation" |
Flags on lines and tiles that carry no cue are logged as cue_flagged on a declared distractor cue; in Assessment, a flag on a whole piece with no cue in its words too (2.6, rule 3), each such cue with three distractor readings |
same; today The Friday file keeps them on screen only (actions.mjs, flag-tile) and The Monday queue logs them as evidence_opened on ev-week-n |
Over-flagging must be in the transcript |
condition_set.source |
session.mjs send() |
R4 |
framing_choice allowed again at the call |
session.mjs |
S1 Level 1, "arrives at the correct framing only after the situation forces it" |
Run fields standardVersion, specVersion, instrumentVersion, scaffolding, accessSupport; mode unseen renamed assessment (D73) |
contract.mjs |
§9.6, §7.1, §8.9 |
Package fields: decisionPoints, stakes, advice loadBearing and consequence (high, low), query probes and checks (the pieces whose claim its answer settles), evidence disconfirms, evidence and query statesBoundary, escalation-target kind, entitled, requester and recipientOf (the options whose sending hands the call to that party), condition and control dimension and role, option proportion, later-branch supports, disproves and pressure, framing class, consult stakeholders, lines |
each package.mjs; checked by validateMission |
Every rule below reads one of these |
Queue fields: E1 control proportion per value; held-item routes, outsideAuthority and proportion per response; correction-item addresses; lines with at: controls and at: week |
each queue mission's package.mjs |
E1 at the controls and at a held item; P2 in a week (3.8, 3.11) |
Scorer round six (4 October 2026): the package field update: { afterHours } and decision points of kind update, with lines at: update (the engine's updateOpen, answerUpdate); option informs (the parties the option's sending tells) and the package's deliverableTo ({ id, material }, the party the delivered work went to); condition informs and withdraws; line proportion on an interim commitment line |
contract.mjs (validated), session.mjs, The water report's package.mjs |
S1, E2 and P1 at an update (3.4, 3.9, 3.10); E2 reaching or withholding by the option or a change (3.9); P2 at the call (3.11) |
3.4 S1 Problem Framing§
Evidence: framing_choice at the brief, a re-framing at the call, note lines tagged S1. Decision point: the brief. Package: each framing carries class: presenting (restates the request: The Friday file's fr-deadline, fr-number), partial (a real decision, not the underlying one: fr-accountability), underlying (fr-autonomy, fr-limits, fr-what-to-tell).
| Level | Anchor (operative words) | Rule |
|---|---|---|
| 0 | "restates the presenting request as the decision" | Framing presenting, no re-framing |
| 1 | "the delegated task rather than the underlying decision; or arrives at the correct framing only after the situation forces it" | Framing partial; or underlying chosen only at the call |
| 2 | "States, before investigating, the decision that must be made now, who owns it, and what it is not" | Framing underlying at the brief, whose copy states the owner and what the call is not (tag statesOwnerAndLimits). Each mission's underlying framing carries it, in both case versions; an underlying framing without it caps at 1. |
| 3 | "states what part of the responsibility cannot be delegated … and records the frame in a form another person could review" | Level 2, plus fitting, clean lines nondelegable and frame in the order |
| 4 | "names the condition that would make this a different decision rather than a revised one" | Level 3, plus a fitting, clean reframe_trigger line |
A re-framing at an update counts at the brief as the call's does: the underlying or a partial framing chosen there after a presenting one at the brief is Level 1 at the brief (read as context; the update cites it, R2).
At an update (decision point <deadline>:update, where decisionPoints lists S1 there; reached once the run spent the update's hours before the call, otherwise it yields nothing). Evidence: the framing the person states at the update and the S1 lines chosen there. The same table as the brief, read at the update: no framing, or a presenting one, is 0; a partial one, or an underlying one whose copy does not state the owner and limits, is 1; the underlying framing is 2; fitting, clean nondelegable and frame lines at the update lift it to 3, and a reframe_trigger line there to 4. A level whose line the update's bank does not hold is not reachable there (The water report's update holds no reframe trigger, for reading load).
3.5 T1 Evidence Verification§
Evidence family: calibrated reliance, carried from version 1. Accept, verify, challenge and override are read against the hidden reliability label. Decision points: each piece the package marks loadBearing. Definitions at piece p:
- Checked before acting: a
verifyof p, anitem_sampledof p, a T1 query probing p, or a query the package tags as checking p (checks: a person whose answer settles the claim, as reading their file would; Kenji on the notice and Arnel on the stop rule's metric in The Friday file, Wei on the migration line in the steering pack), after p was offered and before both the reply to p and the decision; but where a flawed p was queried (challenge) and an own condition or control addresses it (outside a set that does not discriminate), the query is a question and the change is the act, so a check before the decision is a check before acting (Fable's review of The water report, 4 October 2026: asking the AI first, then checking and correcting, is the natural order); likewise where a flawed p was removed (override), as a removal also takes effect only at send (Fable's review of scorer round six, finding 10). Neither holds in a run that refuses wholesale over every sound reply (below): striking everything, then reading, tells nothing. Each piece cites one check (R2): its ownverifyoritem_sampledwhen it has one, else one question that checks it, so a question that checks several pieces is left for the pieces with no check of their own; a check another dimension cites (a signal question taggedchecks) is read as context (R3). Outside a queue, the engine records theverifywhen a file linked to p is opened after p arrives, and also as p arrives when a file linked to it was already opened: one is enough, as opening that file after p arrived would be, and opening a file again records nothing (contract 3.1.0; under 3.0.0 every linked file had to be read first). Outside a queue, a T1 query probing p asked before p arrived likewise checks p as it arrives. In a queue mission a case is checked by opening the case (item_sampled), or by a T1 query probing it once it is on screen: a file opened checks no case, the engine records noverifythere, averifya 3.0.0 queue run carries checks nothing, and a question asked before the case arrived checks nothing. That rule has a cost the person cannot see on screen: in The Monday queue a case's line shows the fact that decides it ("return, day 24 (policy v2)"), so a person who has read refund policy v3 can tell the stale case from its line alone, and a reply made that way, without opening the case, is read as unchecked (T1 Level 1 for a refusal). Crediting a file read before the reply would bring back a check by file in the queue, which every file-reading blanket policy then collects; the instrument keeps one rule and states this cost, and The Monday queue's mission spec says so. The same reading of a check serves the piece's own observation, Level 3's "every piece markedconsequence: highchecked" and refusal wholesale. - Acted on:
acceptof p ("Use this", "Keep"). - Matched: p sound, and accepted, or left standing with no reply and a fitting, clean
claims_checkedline saying checked for p (left alone, nothing else shows the check was read); or p flawed, and challenged or overridden in a run that does not refuse wholesale, or addressed by an own condition or control. An item set inside a set that does not discriminate (3.7) addresses nothing in particular, here and in a week (3.11). - Refusal wholesale (read as context, R3; AIRS-J/1 7.5, T1 note, "rejecting AI output wholesale"): of the sound pieces, load-bearing or not, that the person checked and then replied to, at least three challenged or overridden, and more than half of them. The checks did not tell sound from flawed, so a refusal of a flawed piece in that run is not a match. Fewer than three refusals is too few to read as a pattern: a person who checks every load-bearing piece and doubts one sound piece after its check keeps the flawed pieces' matches. The minimum is on the refusals, so accepting one more sound piece after its check never turns wholesale on (before the fifth review's verifiers it was on the checked replies, and one more correct acceptance could remove every flawed match). T2 Level 3 (3.6), P2 at the call (3.11) and the removal read as checked before acting (above) read it over every sound reply before the call, checked or not (at least three challenged or overridden, and more than half of them): a refusal in such a run is no departure and addresses nothing. Fable's review of scorer round six, finding 2: read over checked replies only, a run that opened only the falsifier's files and challenged every piece was not wholesale, and those readings paid out on its refusals. T1's own reading of a refusal its check contradicts keeps the checked replies.
| Level | Anchor (operative words) | Rule |
|---|---|---|
| 0 | "Acts on an unverified claim where verification was available and the claim was load-bearing"; also "rejecting AI output wholesale without examining it" | Acted on p unchecked while a check was available; or challenged or overrode a sound p unchecked (the costly refusal); or left p alone, unchecked and, if flawed, unaddressed |
| 1 | "Expresses doubt about a claim but acts on it unchanged; or verifies only after being challenged" | Checked only after acting; or checked and not matched (a flawed p refused in a run that refuses wholesale among them); or challenged a flawed p unchecked |
| 2 | "Checks the load-bearing claim against an available source before acting, unprompted. Distinguishes generated content from retrieved content." | Checked before acting, and matched. The match is the observable sign of telling the AI's claim from the source. |
| 3 | "verification effort proportionate to consequence … States in the work product which claims were verified and which were not" | Level 2; every piece marked consequence: high in the mission and offered before the decision, or, outside a queue, due within the review hours (arrives.afterHours under reviewHours), checked before it (R3; deciding with hours unspent leaves the pieces those hours would have brought unchecked, so deciding early never shrinks the set: Fable's review of scorer round six, finding 5), and none of their replies contradicted by its check (a sound one challenged or overridden, a flawed one accepted with no own item addressing it): one such reply shows the checks did not tell sound from flawed across the run, so no piece reaches 3 in it; review time spent checking low pieces within the mission's declared light-check budget (a file opened only for low pieces, or a low item opened; a file tagged statesBoundary or disconfirms, and a held item opened to answer it, is opened for E1 or T2 and not charged); and a claims_checked line for p that says checked |
| 4 | "specifies a check that persists beyond this decision" | Level 3, plus an own condition, control or standing-rule line with role: persisting_check whose classOf includes p: The Friday file's c-sample; The Monday queue's team-sample control |
3.6 T2 Assumption and Alternative Testing§
Evidence: cue flags with their reading, signal queries, openings declared disconfirms of a piece they did not verify. Decision points: each flawed piece with a cue or a probing signal query. A piece with two linked files can be verified through one (T1) and tested through the other (T2); a file linked to one piece can test another (the steering pack's risk log tests the testing line in 1.0.0; Rosa's message the stale line in both versions). An opening is barred from T2 only for the pieces it verified, so whether it counts never turns on when an unrelated piece arrived.
| Level | Anchor (operative words) | Rule |
|---|---|---|
| 0 | "Treats the first available explanation as the explanation" | No flag, probing query or disconfirming opening on p's flaw before the decision |
| 1 | "Acknowledges that other explanations exist but does not name one, or names one without testing it" | Flag or probing query, but no fitting reading (a flag with no reading, or with a reading that does not fit, acknowledges a doubt without naming the flaw); or a fitting reading with no discriminating action |
| 2 | "Names … at least one assumption … and at least one competing explanation, and takes an action capable of discriminating between them" | Flag with a fitting reading, plus a discriminating action not cited by T1 (a probing signal query, or an opening declared disconfirms p that did not verify p), before the decision |
| 3 | "states what evidence would falsify their own preferred position, then acts on that evidence when it appears" | Level 2; a fitting falsifier line at the brief naming evidence against the person's own lean (R5, read as context); and once that evidence was opened, the person acts on it: the decision departs from the lean, an own condition addresses p, or p was challenged or overridden before the call in a run that does not refuse wholesale over every sound reply (3.5), so the lean held after refusing the flawed piece is acting on the evidence, and a challenge-everything run inherits nothing, whatever it read. One falsifier line supports one observation (R2). |
| 4 | "tests an assumption they themselves introduced" | Level 3, plus a discriminating action, after the person set a condition or control, on the evidence that condition declares it assumes. Reachable only where controls are set during the work: in The Monday queue today, not in single-call missions. |
3.7 T3 Risk and Failure Anticipation§
Evidence: own conditions, controls and lines with T3 roles. Decision points: the call; in The Monday queue, also the controls and week 3. Package roles: mitigates (a failure mode, with pathway and bearer in its text), indicator (names a leading indicator and its watcher), generic, unfounded (set with no source, question or doubted line behind it; the steering pack's unfoundedChanges logic in process.mjs).
| Level | Anchor (operative words) | Rule |
|---|---|---|
| 0 | "Names no failure mode, or names only the risk of not acting" | No own T3 item; or only lines tagged risk_of_not_acting |
| 1 | "Names a generic risk without identifying how it would occur or who would bear it" | Only generic or unfounded items, for example The Friday file's c-stop-rule-reopened ticked by hand |
| 2 | "Names, unprompted, at least one plausible failure mode, the pathway by which it would occur, and who carries the cost" | At least one own mitigates item, clean, whose text names the pathway and who bears the cost (tag statesPathwayAndBearer). Every mitigates item in the three missions carries it, in both case versions; an item without it caps at 1. Clean (R6) reads the whole set at the point, at every T3 decision point. Everything on offer (at the call, every condition on offer ticked; at the controls and in a week, every control on offer in force, set off its default, as the point closes, the settings made earlier included; two or more on offer) does not discriminate and caps T3 at 1. A distractor in the set (an own condition or control the package marks fit: distractor, of any role) is proportionate: the levels are read on the fitting items alone, and each distractor then lowers what they earn by one level, never below 1, because some distractors are defensible on their own. |
| 3 | "distinguishes reversible from irreversible failure, and states which harm is being accepted rather than mitigated" | Level 2, plus a fitting accepted_harm line for a failure mode no own item mitigates, with the right reversibility tag |
| 4 | "names a leading indicator … and who is responsible for watching it" | Level 3, plus an own indicator item naming the watcher. The Friday file's c-stop-rule-replies (Arnel's team) and The Monday queue's c-pause-rule (Mira) name who watches and qualify. |
3.8 E1 Decision Rights and Escalation§
Evidence: owner_named, copy-ins to approval targets, controls tagged E1 (The Monday queue's ceiling), escalation_response, E1 lines. Decision points: the call; in The Monday queue also the controls and v26-table. Package: each mission names its accountable owners and its non-accountable parties (the vendor, the system); its required approval targets; whether the call is within the person's authority.
| Level | Anchor (operative words) | Rule |
|---|---|---|
| 0 | "Acts outside their authority, or delegates a decision that cannot be delegated … without acknowledgement" | Owner named as a non-accountable party (Dev, Lattice), or not set where the case requires one; or a required approval target not reached before the decision with no file or answer tagged statesBoundary opened or asked before the call (the boundary neither seen nor acted on) |
| 1 | "Recognises an authority boundary but does not act on it; or escalates everything" | The file or answer that states a boundary was opened or asked, and the required approval party was not reached; or escalates everything: every escalation target copied in, and either the option escalate sent or one target copied in that is neither approval nor entitled (where every target is needed, copying them all is proportionate); or the needless party copied in, whoever else is copied in or left out: a target of kind inform that is not entitled, not tagged requester, and not the party the option sent hands the call to (recipientOf). A target tagged requester, the party the deliverable goes to, is entitled to the outcome and is never that needless party. Or every approval party reached, with no file or answer that states who approves (tag statesBoundary) opened or asked before the call: the right names copied in unread are a guess or a habit, not a boundary identified |
| 2 | "Acts within their authority, unprompted, and escalates the parts requiring another party's approval, identifying who that party is" | Accountable owner named; every required approval target reached before the decision; a file or answer tagged statesBoundary opened or asked before the call (where the package tags one); not blanket |
| 3 | "the split proportionate … and states the basis for the split" | Level 2; the option is escalate only where the package declares the call beyond the person's authority (R3); a fitting authority_basis line |
| 4 | "names the condition under which the boundary itself should be reset" | Level 3, plus a fitting E1 standing-rule line |
Told (Fable's review of scorer round six, finding 8): E1 reads who heard as E2 and a mission's world do (3.9, partiesTold in contract.mjs): a copy-in, an own change on the order that informs the party, or the option sent that informs it. A needless party told by a change is a stray copy, as one copied in is.
The table above is the call. Queue missions add two kinds of E1 decision point, each read by its own table.
At the controls (decision point controls). Evidence: the owner_named and the own E1 controls recorded at the controls, and E1 lines recorded there. Package: the accountable and non-accountable owners; each control tagged E1 (The Monday queue's ceiling, role authority_limit) carries proportion, a case reading per value (proportionate, over, under), with provenance.
| Level | Anchor (operative words) | Rule |
|---|---|---|
| 0 | "Acts outside their authority, or delegates a decision that cannot be delegated … without acknowledgement" | No owner named at the controls where the package declares accountable owners; or the owner named is a non-accountable party |
| 1 | "Recognises an authority boundary but does not act on it; or escalates everything" | The owner named is not one of the accountable owners; or an E1 control left at its default (never set, or set to the package's default value: the engine records any control touched, whatever the value); or an E1 control set to a value read over |
| 2 | "Acts within their authority, unprompted, and escalates the parts requiring another party's approval, identifying who that party is" | An accountable owner named, and every E1 control set by the person to a value other than its default, none to a value read over. Needs the proportion reading: until a control's values carry it, E1 at the controls caps at 1. |
| 3 | "the split proportionate … and states the basis for the split" | Level 2; every E1 setting read proportionate (R3); a fitting, clean authority_basis line recorded at the controls |
| 4 | "names the condition under which the boundary itself should be reset" | Level 3, plus a fitting E1 standing-rule line recorded at the controls |
At a held item (an advice piece with escalates: true that decisionPoints lists for E1). Evidence: the escalation_response to it, and E1 lines recorded at the decision point of the period it was answered in. Package: on the item, routes (the responses that pass it to another party to decide), outsideAuthority (responses the person may not give) and proportion (a case reading per response), with provenance. The engine accepts an answer only once the item is opened and only in its own period, so an answer always rests on the item read.
| Level | Anchor (operative words) | Rule |
|---|---|---|
| 0 | "Acts outside their authority … without acknowledgement" | The item left to lapse with its period; or answered with a response in outsideAuthority |
| 1 | "or escalates everything, using escalation to avoid deciding" | Answered with a response in routes, in a run that escalates everything at the call (the call's "escalates everything" reading), read as context (R3) |
| 2 | "Acts within their authority, unprompted, and escalates the parts requiring another party's approval, identifying who that party is" | Answered in its period with a response not caught at Level 0 or 1. A routing response is one the package lists in routes, its tag that the response passes the item to the party that decides; the scorer reads the tag, not a named party. |
| 3 | "the split proportionate … and states the basis for the split" | Level 2; the response read proportionate (R3); a fitting, clean authority_basis line recorded in that period |
| 4 | "names the condition under which the boundary itself should be reset" | Level 3, plus a fitting E1 standing-rule line recorded in that period |
Lines recorded at the call belong to the call (R2), so a held item answered in the last period reaches at most Level 2.
Targets: The Friday file: esc-legal (Kenji) and esc-governance (Daniel) are approval in both case versions. The rollout order names the one market that goes live only on its country manager's approval, Marisol in 1.0.0 (esc-marisol, Siti Rahmah) and Kaldera in 1.1.0 (esc-kaldera, Joaquin Cruz); the other country manager hears after Grace signs (in 1.0.0 the order names Joaquin there), so is a needless party in that version. esc-ceo is inform (the requester); esc-supplier is a needless party. The Monday queue: esc-governance is approval, esc-finance and esc-coo are inform (Wen entitled, Tomas the requester), esc-risk is the needless party unless the option sent is escalate, which hands the renewal to the risk committee (recipientOf). The steering pack: page 9 of the plan names who signs off each of five conditions for cutover, and page 10 has a condition's signatory hear of a risk to it before a pack goes. esc-risk (Jun-ho, the ledger) and esc-training (Rosa, branch training) are approval in both versions; esc-test (Farah, the testing pass mark) is approval in 1.0.0, where complex claims stand at 71%, and a needless party in 1.1.0, where they stand at 92%; esc-office (Aminah, head-office training, signed off on 1 October) is a needless party in both; esc-sponsor and esc-director are inform (both entitled, Ingrid the requester); esc-supplier is a needless party. Only a file that names who approves is tagged statesBoundary: The Friday file's rollout order, The Monday queue's trial terms (ev-mandate), the steering pack's plan. Who approves moves with the case version on the two missions whose E1 rests on the call alone, and no habit of copy-ins a person could bring from another case is right in either version: each version has at least two needless parties, so "everyone but one" always copies one; an approval party in each version shares its title with a needless party in that version (the two country managers; the two change and training leads), so every union of title words that copies the one copies the other; and everyone the case names before the call, or the file that says who approves names, includes a needless party. platform/engine/airs-missions.test.mjs checks all of these without play, exhaustively over title words, and the case-blind grid sweeps five coached vocabularies and the "everyone the case names" habit. The Monday queue's approval set does not move (Hana, the AI governance owner, in both), so a title habit can copy exactly it at the renewal; its E1 also reads the controls and the held table, which no copy-in habit reaches, and the grid bounds the dimension (4.5). The needless-party cap keys on the target's tags (kind, entitled, requester, recipientOf), never on its title. One needless copy caps E1 at the call at 1, even where the case says that party hears after the decision (the other country managers, Aminah): a sound person who copies one in loses E1 there, and the case text that says when each party hears is the basis (a reading, recorded in each mission's spec).
3.9 E2 Stakeholder Alignment and Disclosure§
Evidence: trade-off queries to declared consult stakeholders before the call; copy-ins to inform targets; disclosure and communication lines; E2 conditions (the steering pack's c-slip-cost-line and c-status-amber). Decision point: the call.
| Level | Anchor (operative words) | Rule |
|---|---|---|
| 0 | "Omits a party entitled to be consulted or informed; or conceals AI involvement, material uncertainty …" | The option sent withholds the correction from the party the delivered work went to (informs declared on the option without the package's deliverableTo; The water report's correct_quietly), which is CR-E2 where deliverableTo is material; or an entitled party not reached; or, once disclosure lines or disclosing changes exist, no fitting disclosure to an entitled recipient |
| 1 | "engages them after the decision is fixed; or discloses only when asked" | Entitled parties reached, but a consult stakeholder not asked before the call; or no disclosure line |
| 2 | "Identifies, unprompted, who must be consulted before … and who must be informed after; discloses AI involvement and material uncertainty in the work product" | Every consult stakeholder asked before the call; every entitled party reached; a fitting disclosure line; clean |
| 3 | "communication fitted to each recipient's decision … and competing interests named rather than smoothed" | Level 2, plus a fitting fitted item (the steering pack's slip-cost line) and a fitting competing_interests line |
| 4 | "frames the disclosure as a standing expectation for this class of work" | Level 3, plus a fitting E2 standing-rule line |
Each mission's line bank holds disclosure lines, fitting and decoy, so the disclosure element is scored on all three missions. A bank with none would leave the element unscored against the person and cap E2 at 1.
Reached (scorer round six): a party is reached by a copy-in, by the option sent where the option informs it, or by an own change on the order that informs it (The water report's c-council, which tells the Commission the council has the annex). One reading, partiesTold, serves E1 (3.8), E2 and the world, in every case version (Fable's review of scorer round six, finding 8). Disclosure: a fitting, clean disclosure line, or an own condition or control with role disclosure (The water report's AI line, the steering pack's amber status), read as the fitted item is (Fable's review, item 3). A set that does not discriminate (3.7, R6) discloses and fits nothing in particular, so neither item is read inside it.
At an update (where decisionPoints lists E2 there). Evidence: the E2 lines chosen at the update; no copy-in is read there, as copy-ins are recorded on the order (R2). No E2 line chosen there is no observation: telling a partner nothing yet about the AI is not concealment from the party the work went to (Fable's review of scorer round six, finding 3). With an E2 line chosen, Level 0: no fitting, clean disclosure line among them (The water report's l-u-tell-x, which states the scope is already listed, is false certainty); 2: one; 3: plus fitting, clean fitted and competing_interests lines at the update; 4: plus a fitting E2 standing-rule line there (not in The water report's update bank). Level 1 has no observable at an update, as no consult is read there, and an update whose bank holds no disclosure line yields no E2 observation (finding 9: the code read it as 1).
3.10 P1 Proportionate Commitment Under Uncertainty§
Evidence: decision or time_expired, P1 lines. Decision point: the call. Package: each option carries proportion (proportionate, over, under, contradicted), a case reading with provenance.
P1 is the one dimension whose anchor reads which position was taken. Blueprint 6.5 still holds for the process label and the four cells, which never read the option. The P1 observation reads the option's authored proportion label, never the outcome.
| Level | Anchor (operative words) | Rule |
|---|---|---|
| 0 | "Does not commit; commits after the window has closed; or commits to a position the available evidence contradicts" | time_expired; let_deadline_pass; or an option tagged contradicted (The Friday file's unconditional proceed, the steering pack's send_as_drafted, The Monday queue's widen and lapse) |
| 1 | "not proportionate to the stakes … or commits without stating a basis" | Option tagged over or under; or no fitting basis line |
| 2 | "Commits, unprompted and within the time available, to a position whose cost and reversibility are proportionate … and states the basis" | Option tagged proportionate, sent before the deadline, plus a fitting, clean basis line whose restsOn includes at least one item the person opened or asked |
| 3 | "states in advance what would cause the position to change: a named condition" | Level 2, plus a fitting trigger line |
| 4 | "executable by the named recipient without further instruction: owner, first action and timing specified" | Level 3, plus a fitting handoff line naming owner, first action and timing (the owner named on the order is E1's and is not cited here) |
At an update (where decisionPoints lists P1 there). Evidence: the interim commitment lines chosen at the update, each carrying proportion as an option does, its text stating its basis (what is traced so far). No P1 line chosen there is no observation: a reply that states no interim commitment commits to nothing (Fable's review of scorer round six, finding 3). With a P1 line chosen, Level 0: no commitment among them, or one tagged contradicted; 1: one tagged over or under, or commitments of different proportion at once (not clean, R6); 2: one fitting commitment tagged proportionate; 3: plus a fitting, clean trigger line at the update; 4: plus a fitting handoff line there (not in The water report's update bank).
3.11 P2 Consequence Ownership and Correction§
Evidence: revision, retain, paused, revoked, P2 controls (The Monday queue's policy ticket), P2 lines. Decision points: the later choice; in The Monday queue, each week while an error found is open. Package: each later branch declares supports (the responses the credible item supports), disproves (whether it disproves the position sent) and pressure (where the not-credible item pushes).
| Level | Anchor (operative words) | Rule |
|---|---|---|
| 0 | "Deflects the outcome to the AI system, the data, another party … Or holds a position that new evidence has disproved" | Keep where the branch disproves; or an attribution line tagged deflects |
| 1 | "Accepts the outcome in general terms …; or revises only under pressure rather than on evidence" | A revision toward pressure; or a response with no fitting evidence_basis line |
| 2 | "holds, modifies or reverses … on the merits and states which evidence drove the choice. Accepts the consequence" | Response within supports; a fitting evidence_basis line citing the credible item; an owns attribution line; clean |
| 3 | "distinguishes a poor decision from an unlucky outcome" | Level 2, plus a fitting cell line read only on what the person can see at the later choice: its outcome is the outcome shown (R3), and its verdict on the call agrees with the response stated on the credible item. A keep agrees with a sound call, or on a good outcome with a weak call that worked; a change agrees with a call that was wrong. The process half of the run's cell is never shown in Assessment, so it is not read: on a good outcome, earned and lucky both count |
| 4 | "identifies what in their own process produced the error" | Level 3, plus a P2 standing-rule line whose corrects names a process rule that failed in this run. Not reachable on a run where every process rule held. |
The table above is the later choice. At the call (scorer round six, The water report's R-2; for a mission whose call point lists P2): the delivered work is the position already committed, and what is sent corrects it. A flaw found is a load-bearing flawed piece offered before the call that the person checked as T1 reads a check (3.5), flagged on its cue with the fitting reading, asked a signal question probing, or tested by opening a file that disconfirms it, before the call. It is addressed by an own condition or control that addresses it outside a set that does not discriminate, or by a challenge or override of it in a run that does not refuse wholesale over every sound reply before the call (3.5). Only P2 lines are cited (R1); the option, the changes and the replies are context (R3).
| Level | Anchor (operative words) | Rule |
|---|---|---|
| 0 | "Deflects the outcome … Or holds a position that new evidence has disproved" | No correction sent by the person (time_expired, let_deadline_pass); or the option withholds the correction from the party the work went to (3.9, informs). Attribution lines are the later choice's: none is on offer at the call, so none is read there (Fable's review of scorer round six, finding 9: the rule named one no mission offers) |
| 1 | "Accepts the outcome in general terms without identifying what specifically was wrong" | A correction sent, but: an own change tagged withdraws (the whole work withdrawn, a correction in general terms); no flaw found; a flaw found not addressed (or addressed only inside a set that does not discriminate, or refused only in a run that refuses wholesale); or no fitting, clean evidence_basis line at the call whose restsOn names a flaw found |
| 2 | "On new evidence, … modifies or reverses the position on the merits and states which evidence drove the choice" | Every flaw found addressed, and that basis line |
| 3 | "corrects reasoning that was wrong, and does not abandon reasoning that was sound" | Level 2, and no sound piece overridden (removed) before the call |
| 4 | "identifies what in their own process produced the error" | Level 3, plus a fitting, clean P2 standing-rule line at the call |
In a week (queue missions; decision point week-n, n ≥ 2, read against period n − 1, because a change made in a week takes effect the next):
- Correction item: a condition or control the package tags
dimension: P2,role: correction, withaddresses(The Monday queue's policy ticket). - Class of a flawed piece p: every piece a correction item
addressestogether with p. A flawed piece no correction item addresses has no class and opens no week point. - Found: a flawed piece with a class that the person checked as T1 reads a check (3.5), once the piece was offered: opened it (
item_sampled), opened its file (verify), or asked a query declaredai_calibrationthatprobesit; or flagged on its own cue with the fittingreading. A flag with no reading, or a reading that does not fit, names nothing wrong (3.6 Level 1) and finds nothing; opening the piece afterwards does. A piece that a change already in force addresses is not found again: checking it after the correction opens no new error. One query that probes several offered pieces finds each of them. - Change: an own P2 act: a correction control set,
pausedorrevoked. A correction addresses itsaddresses; a pause or revocation addresses every piece and is wholesale. A change addresses a found error when it comes after the find and its addresses include the found piece. A change made before any find of the piece is read as a change not tied to an open error (Level 1), and stays in force for every later check of that piece. - Open at week n: an error found in period n − 1 or any earlier period, and not addressed by a change in a period before n − 1. An error stays open, and is read at every week point, until a change addresses it, so a correction made the week after the find is read at the next week's point on the evidence that found it.
- Recurs: a flawed piece of an open error's class offered in period n.
A week point is observed when an error is open at week n or period n − 1 holds a change, and the run reached period n; otherwise it yields nothing. Only changes made in period n − 1 are read there. Its lines are those recorded at week-(n − 1). The acts that found the error are read as context and cited under T1 or T2 (R1, R3).
| Level | Anchor (operative words) | Rule |
|---|---|---|
| 0 | "Deflects the outcome … Or holds a position that new evidence has disproved" | An attribution line tagged deflects at the week; or an open error, no change that period addresses its piece, and its class recurs in period n. An error left open counts against at each week its class recurs |
| 1 | "Accepts the outcome in general terms without identifying what specifically was wrong" | A change not tied to an open error, read as a correction made without identifying what specifically was wrong; or an open error and no change after the find addresses it, though its class did not recur. A change made in a period whose set does not discriminate (3.7: a distractor control, or every control on offer in force as the period closes) is a correction in general terms: it addresses no open error, though it still counts as a change against a recurrence |
| 2 | "On new evidence, holds, modifies or reverses the position on the merits and states which evidence drove the choice" | Every open error addressed by a change that period; a fitting, clean evidence_basis line whose restsOn names an open error's piece or an item the person opened, asked or flagged that period |
| 3 | "corrects reasoning that was wrong, and does not abandon reasoning that was sound but unlucky" | Level 2, and every change that period addresses an open error: no pause or revocation, and no correction aimed at a piece no evidence faulted |
| 4 | "identifies what in their own process produced the error, in terms that would change how the next decision of this class is made" | Level 3, plus a fitting P2 standing-rule line at the week whose classOf includes an open error's piece |
3.12 S2 Scope and Constraint Definition§
Evidence: plan lines added at the brief, before the first opening. Decision point: the brief. No current act reaches S2; every mission records S2 as "evidence missing" until plan lines exist.
| Level | Anchor (operative words) | Rule |
|---|---|---|
| 0 | "Proceeds without identifying any boundary" | No plan line |
| 1 | "Repeats the constraints given in the brief. Does not identify unknowns, or does not distinguish what is fixed from what is negotiable" | Only lines tagged given; or no fitting unknown line |
| 2 | "Names, unprompted, what is in scope, what is fixed, and at least one material unknown" | Fitting scope, fixed and unknown lines, clean |
| 3 | "distinguishes fixed from negotiable constraints, and states what information would be sufficient to proceed" | Level 2, plus fitting negotiable and sufficient lines. The sufficient line must name a gap tagged changes_answer. |
| 4 | "states a stopping rule in advance" | Level 3, plus a fitting stopping_rule line |
3.13 Not observed, and evidence missing§
Blueprint 6.1's two absences carry over. Left alone is an observation, usually Level 0. Not observed means the mission offers no decision point for that dimension. Evidence missing means the dimension's decision point exists but the act it needs is not built (none on the three missions today; section 4.5 gives the rules that still cap a dimension below 4). Neither absence counts toward §8.3.
3.14 Critical-review patterns (Annex D)§
| Code | Deterministic trigger in StepSim |
|---|---|
| CR-T1 | A claims_checked line says checked for a piece with no check before acting, and the piece's failure mode is tagged irreversible or non-financial |
| CR-E1 | An option or control the package tags outside_authority, where the harm is tagged irreversible or non-financial. None is authored today. |
| CR-E2 | No fitting disclosure line where the package tags the recipient entitled and the omission material |
| CR-P2 | Keep where the branch disproves, together with an attribution line tagged deflects |
Each trigger sets the dimension to Insufficient Evidence pending targeted reassessment (§12.2.2). Reassessment needs an unexposed decision point, which means another Assessment mission.
4. Coverage per mission§
Counts are observation opportunities from distinct decision points: what a full run can yield, at any level. "Now" is today's package; "built" adds the decision points in section 7.2. The counts are the same in case versions 1.0.0 and 1.1.0.
4.1 The Friday file§
| Dimension | Decision points | Now | Built |
|---|---|---|---|
| S1 | The brief | 1 | 1 |
| S2 | The brief (plan lines) | 1 | 1 |
| T1 | j1-go-live, j2-phase, j4-no-notice, j7-stop-rule (1.1.0: j6-tell-legal for j4) |
4 | 4 |
| T2 | j1 (cue-replies-after-close), j4 (cue-table-date, q-kenji), j7 (cue-reopened-zero, q-arnel) |
3 | 3 |
| T3 | The call | 1 | 1 |
| E1 | The call | 1 | 1 |
| E2 | The call | 1 | 2 |
| P1 | The call | 1 | 2 |
| P2 | The later choice | 1 | 1 |
Stakes profile. Primary C3, operational and service continuity. Bearer: customers in five markets, Kavitha Raj's unpaid S$289 among them. Pathway: Resolve closes cases it cannot reopen, and replies go to a mailbox nobody reads. Severity: an unphased surge cannot be recovered within the quarter. The person must reduce the harm (phasing, c-unusual-owner, c-sample) or monitor it (c-stop-rule-replies). Secondary C2, financial. Bearer: Northwind, which holds a board number of S$1.8 million against Finance's S$0.7 million. Pathway: the number goes to the board un-rebased. The person reduces it with c-finance-recheck. Qualifier: binding legal constraint (Kaldera's notice rule, Marisol's sign-off rule). A Kaldera inquiry could support C5b, but §6.1 allows one secondary, and C2 binds more seeds in the outcome model.
Distractor tags (R6). Three limits that do not fit, T3 mitigates with fit: "distractor", each on offer from the same facts as the limit it resembles: c-x-confidence-floor, c-x-legal-every-market, c-x-csat-pause. Each one in the set lowers T3 by one level (3.7); c-x-csat-pause and c-x-legal-every-market are defensible on their own. Grace (esc-ceo) asked for the note and is tagged requester (3.8). Marcus Lee, compliance lead at Resolve's supplier (esc-supplier, kind: "inform", not entitled), is a needless party: the supplier works from the signed order, not the draft, so copying him in caps E1 at the call at 1 (3.8). The approval parties move with the case version (fifth review of a5601a28, 4 October 2026): Kenji (esc-legal, counsel), because the rollout order says Legal is copied on the Friday call and Kenji files the Kaldera notice the day he is told, and Daniel (esc-governance, AI policy owner), whom the rollout order names the AI-governance owner, in both versions; and the country manager of the one market the rollout order says goes live only on its country manager's approval, Siti Rahmah for Marisol in 1.0.0 (esc-marisol), Joaquin Cruz for Kaldera in 1.1.0 (esc-kaldera). The other country manager hears after Grace signs and is a needless party in that version; in 1.0.0 the rollout order names Joaquin there, so everyone the case names includes a needless party (fifth review's verifiers). The call labels each by name and plain role, Grace's note names nobody to copy, and only the rollout order is tagged statesBoundary (3.8 L2): Kenji's brief and answer, and Daniel's, say what each needs, not who approves. Asking Kenji (the notice) or Arnel (the stop rule's metric) is a T1 check of that piece (checks, 3.5). Pushing back on the stop rule costs an hour, as on the phasing, so the one push-back that costs is no longer always the sound piece; Juno's phasing now carries an absolute word ("of all five markets") and the stop rule none, so answering by absolute words is wrong on both. Arnel's memo and Kenji's brief cost one hour each (second copy round, 3 October 2026), so seven hours can check all four load-bearing pieces and still open the memo, ask Kenji and ask Siti, in any order: the stop rule arrives at hour 3, and a file of it read before it arrives is a check of it as it does (3.5; before the sixth review a person who read its files first could never check it, and until the fifth review of a5601a28 both had to be read). The model answer reads the dashboard first, so Kavitha's complaint is no longer opened and Kenji is asked about the notice. Every piece carries a claims pair, group: "claims:<piece id>": a "checked" line and a "checked" decoy naming a source that does not bear on the piece, nine pairs, on offer once the piece is opened; on the flawed load-bearing pieces the fitting line is as often the longer as the decoy. The call's wrong lines differ from the fitting ones on a fact of the case, not on which name does the thing: the Kaldera notice's timing (l-split-x), the pilot sample (l-rule-check-x, l-board-line-x), the regulatory table (l-rule-reset-x), the pilot market (l-next-x).
4.2 The Monday queue§
| Dimension | Decision points | Now | Built |
|---|---|---|---|
| S1 | The brief; built: week 3, the policy change, as a re-framing point | 1 | 2 |
| S2 | The brief | 1 | 1 |
| T1 | v05-stale-window, v09-decline-day20, v14-set-refund, v17-depot-scan, v22-sale-drift, v30-sale-drift, v33-method (1.1.0: v08 and v12 for v05 and v17) |
7 | 7 |
| T2 | v05 (cue-policy-version), v14 (cue-set-amount), v22 (cue-sale-tile, q-dev) |
3 | 3 |
| T3 | Controls, week 3 (ck-sale), the renewal |
3 | 3 |
| E1 | Controls (ceiling, owner), v26-table, the renewal |
3 | 3 |
| E2 | The renewal; built: a week-2 update to Tomas | 1 | 2 |
| P1 | The renewal | 1 | 1 |
| P2 | Weeks 2, 3 and 4, each read against the week before and observed when that week held a find or a change (a policy ticket raised in week 2 is read at week 3), and the later choice | 4 | 4 |
Stakes profile. Primary C2, financial. Bearer: Larkspur, whose refund leakage runs from an S$11k baseline to S$31k more in the drift wave, and customers wrongly declined, Leila Haddad among them. Pathway: Vela applies a stale policy and makes confident mistakes under the person's name. Severity: refunds paid are not recovered. The person reduces it (ceiling, holds, the policy ticket) or monitors it (the team sample). Secondary C3, operational. Bearer: customers waiting and Mira's team of three. Pathway: holds beyond seven a week push service from 5 days to 12. The person must accept or reduce that load. Qualifier: dependency and cascade (Orbit's policy file).
Distractor tags (R6). Three rules that do not fit, T3 mitigates with fit: "distractor": c-x-hold-over-100, c-x-confidence-hold, c-x-decline-sale. The ceiling's values differ by case version, S$250 proportionate in both but never the middle value and not at the same place: 1.0.0 offers S$100, 150, 250, 500 (100 and 150 over, 500 under); 1.1.0 offers S$100, 150, 200, 250, 400, 500 (100, 150 and 200 over, 400 and 500 under). Tomas (esc-coo) asked for the renewal and is tagged requester (3.8). Priya Nair, Head of Risk (esc-risk, kind: "inform", not entitled), is the needless party: Vela's trial terms say the risk committee takes a renewal only when it is handed over, so she is tagged recipientOf: ["escalate"], and copying her in with the renewal handed over is the hand-over, not a stray copy (3.8). A queue's case is checked by opening it (3.5): opening refund policy v3 checks no case, and since it states the 14-day window and the faulty item's line it is tagged disconfirms the stale case and the set refund, a discriminating test of their flags (3.6). The model answer and the strong path read Vela's trial terms (ev-mandate, statesBoundary, the only item so tagged: Hana's answer says what the policy asks, not who approves) before the renewal (3.8 L2); reading them is not light checking, though they link only goodwill cases (3.5 L3). Reversing the week-3 sale refund (v22-sale-drift) costs a written reason, ten minutes, as the sound day-20 decline's does, so the one reversal that costs is no longer always the sound case. A case's claims pair is on offer once the case is opened or Vela is asked a question that probes it (2.7). Hana Sato, AI governance owner (esc-governance), is the approval party; Wen Lau, Head of Finance (esc-finance), is entitled to be told. Tomas's note no longer names anyone to copy. c-x-decline-sale is defensible on its own. Every claims pair (a "checked" line and a "checked" decoy naming a source that does not bear on the case, or a finding wrong on a fact of the case) carries group: "claims:<case id>", one pair on every case (36, CLAIM_PIECES), on offer once the case is opened. A pair whose two lines name the same source sits on about one case in four, moving between versions or not (nine of 36: one of the four moving cases, eight of the rest; before the fifth review of a5601a28 the three such pairs were exactly the moving cases). The E1 lines' decoys are wrong on a fact of the case, not on who does the thing (l-split-x: Hana approves after the renewal goes; l-ctl-split-x: seven cases a day; l-wk-split-x: every held case to Mira's team; l-wk-reset-x: Thursday, not Monday; l-next-x: every case over S$100).
4.3 The steering pack§
| Dimension | Decision points | Now | Built |
|---|---|---|---|
| S1 | The brief | 1 | 1 |
| S2 | The brief | 1 | 1 |
| T1 | s2-budget, s3-migration, s5-training, s7-testing, s8-risks, s11-keep-green (1.1.0: s9-resourcing for s8) |
6 | 6 |
| T2 | s3 (cue-no-figure, q-wei), s7 (cue-defects-rising; the risk log), s8 (cue-export-date; q-rosa, Rosa's message), s11 (cue-every-figure; q-jun-ho, the plan). 1.1.0: s5 (cue-module-version; q-rosa, the test report) and s9 (cue-vendor-date; Rosa's message) for s7 and s8 |
4 | 4 |
| T3 | The pack | 1 | 1 |
| E1 | The pack | 1 | 1 |
| E2 | The pack; built: Ingrid's 15:00 check-in | 1 | 2 |
| P1 | The pack; built: the same check-in | 1 | 2 |
| P2 | The later choice | 1 | 1 |
Stakes profile. Primary C4, information harm (AIRS-J/1 §6.1 includes wrong records about identifiable people). Bearer: 11,000 policyholders whose premium balances are wrong. Pathway: a green page leads the review to approve cutover, and the wrong balances go live on 30 November. Severity: once live, they reach customers' accounts. The person reduces it (c-rehearsal-gate, c-ledger-line) or transfers it (esc-risk). If v2.0 reads C4 as privacy only, this harm is C2, and the steering pack has no eligible non-financial family. Secondary C3, operational: cutover and branch claims handling. The board-surprise pathway is internal to Cobalt and does not meet C5b's public or institutional bearer. Qualifier: dependency and cascade (Lattice).
Distractor tags (R6). Three changes that do not fit, T3 mitigates with fit: "distractor": c-x-postpone-cutover, c-x-rerun-regression, c-x-vendor-signoff. Ingrid (esc-director) asked for the pack and is tagged requester (3.8). Page 9 of the plan names who signs off each of five conditions for cutover: ledger reconciliation, Jun-ho, the financial controller (esc-risk); the testing pass mark, Farah, the test lead (esc-test); branch training, Rosa, the change and training lead (esc-training); head-office training, Aminah, a second change and training lead (esc-office), signed off on 1 October; supplier readiness, the programme lead, the person. Page 10 says whoever signs off a condition hears of a risk to it before a pack goes. The approval parties are the signatories of the conditions at risk, so they move with the case version (fifth review of a5601a28, 4 October 2026): Jun-ho and Rosa in both, Farah in 1.0.0 only (complex claims at 71%; in 1.1.0 at 92%, so she is a needless party there). Aminah's condition is met in both versions, so she is a needless party in both, and since she shares Rosa's title every title habit that copies Rosa copies her (fifth review's verifiers, 4 October 2026: before, 1.1.0's approval set was a subset of 1.0.0's, and a title set such as controller and training copied it exactly). Ben, Lattice's delivery risk lead (esc-supplier, kind: "inform", not entitled), is a needless party: page 10 says suppliers hear what the committee decides after the meeting, so copying him in caps E1 at the call at 1 (3.8). Only the plan is tagged statesBoundary (3.8 L2); Jun-ho names his own sign-off, not who approves. The training and testing lines are load-bearing in both versions, each the report of a condition for cutover, sound in one version and flawed in the other, so the load-bearing lines no longer sit at odd positions only in 1.1.0. Each flawed line has a test of its flaw that does not check it: the keep-green note a cue in its words ("Every figure on the page": the ledger line has no figure) and the plan; the testing line the risk log's R-14 (1.0.0); the risks note Rosa, who logged R-14 on Tuesday, in her answer or her message (1.0.0); the training line Rosa or the test report, run on the screens re-cut in September (1.1.0); the resourcing note Rosa's message, on Lattice's reassigned engineer (1.1.0). Asking Wei is a T1 check of the migration line (checks). Querying the keep-green note costs an hour, as querying the budget line does. Eight review hours from 09:00 (five until the fifth review, seven until its verifiers), and the risk log costs one hour: the model answer checks every load-bearing line before its reply, tests each flagged line, asks Wei, Hoa and (1.0.0) Rosa, and reads the plan. Every piece carries a claims pair, group: "claims:<piece id>" (CLAIM_PIECES, all eleven): a "checked" line naming a file of the case and a decoy naming one that does not bear on the piece. The "Not checked" lines on four pieces went in the fifth review: no rule read them, and they sat on three of the four load-bearing pieces in each version. l-next-x names the migration figures, not the reconciliation, rather than another recipient.
4.4 The arithmetic§
| Dimension | Friday | Monday | Steering | Total now | Total built | §8.3 (≥ 3) |
|---|---|---|---|---|---|---|
| S1 | 1 | 1 | 1 | 3 | 4 | Met, no slack now |
| S2 | 1 | 1 | 1 | 3 | 3 | Met, no slack |
| T1 | 4 | 7 | 6 | 17 | 17 | Met |
| T2 | 3 | 3 | 4 | 10 | 10 | Met |
| T3 | 1 | 3 | 1 | 5 | 5 | Met |
| E1 | 1 | 3 | 1 | 5 | 5 | Met |
| E2 | 1 | 1 | 1 | 3 | 6 | Met, no slack now |
| P1 | 1 | 1 | 1 | 3 | 5 | Met, no slack now |
| P2 | 1 | 4 | 1 | 6 | 6 | Met |
§8.3. On counts, the three missions meet three observations from distinct decision points for all nine dimensions today. S1, S2, E2 and P1 sit at exactly three, so one abandoned mission leaves them Insufficient. The built decision points give S1, E2 and P1 slack; S2 stays at three. A week point yields an observation only when the week before held a find or a change, so a run can yield fewer P2 observations than the count (the Monday queue's strong path yields none at week 3).
§6.4. C3, C2 and C4: three distinct primary families, one of them in C1, C4, C5b or C6. AIRS-J/1 v2.0 §6.1 keeps the steering pack in C4.
What stops a certifying result from these three, all of it:
- None of them can be an Assessment mission. All three are on the showcase and in the film (blueprint section 5). Every run of them is Practice or Cohort,
guided, and not certifying. - §10. Moderate needs 3 to 4 observations per dimension from three or more scenarios; several scenarios in one sitting are not clustering.
- §10, High. High needs five or more observations per dimension from four or more scenarios, scored by rules that meet §11.2: a fourth Assessment mission, and more decision points per dimension.
- §9.2 and §9.4. The notice on outside AI use and a second, parallel form do not exist yet. The identity check (§9.1), and exposure counters and retirement (§9.3), are built (7.3).
Level ceilings no longer stop it: unscaffolded, every dimension reaches Level 2 or above on each mission (4.5), as §12.2 rule 1 needs.
4.5 The highest level each dimension can reach today§
The highest level each mission's strong run reaches, unscaffolded (Assessment mode, so R7 caps nothing), in case version 1.0.0, as platform/engine/airs-missions.test.mjs pins it. Each mission's model answer reaches 3 or above wherever its strong run does. The Friday file's T1, capped at 2 until the second copy round, now reaches 4: Arnel's memo and Kenji's brief cost one hour each, so the author's path checks every load-bearing piece before its reply, and a path that reads a file of the stop rule first does too (3.5). The steering pack's T1, capped at 2 until the fifth review of a5601a28, now reaches 3: eight review hours check every load-bearing line and still ask Hoa, Wei and Rosa.
| Dimension | Friday | Monday | Steering | Below 4 because |
|---|---|---|---|---|
| S1 | 4 | 4 | 4 | |
| S2 | 4 | 4 | 4 | |
| T1 | 4 | 4 | 3 | The steering pack: no own condition, control or standing-rule line with role: persisting_check covers a piece (3.5 L4); its standing-rule line l-rule-check carries classOf but no role |
| T2 | 3 | 3 | 3 | 3.6 L4 needs a test of an assumption the person's own condition or control introduced: single-call missions cannot, and no Monday control declares what it assumes yet |
| T3 | 4 | 4 | 3 | The steering pack has no indicator item (3.7 L4) |
| E1 | 4 | 4 | 4 | |
| E2 | 4 | 4 | 4 | |
| P1 | 4 | 4 | 4 | |
| P2 | 3 | 4 | 3 | At the later choice, 3.11 L4 needs a standing rule correcting a process rule that failed in this run; on a strong run every rule holds. The Monday queue reaches 4 in a week |
Each mission's model answer (its reference policy, then the strong later choice), Assessment mode, the same in case versions 1.0.0 and 1.1.0:
| Dimension | Friday | Monday | Steering |
|---|---|---|---|
| S1 | 4 | 4 | 4 |
| S2 | 4 | 4 | 4 |
| T1 | 4 | 4 | 3 |
| T2 | 3 | 3 | 3 |
| T3 | 3 | 3 | 3 |
| E1 | 4 | 4 | 4 |
| E2 | 4 | 4 | 4 |
| P1 | 4 | 4 | 4 |
| P2 | 3 | 4 | 3 |
The model answers' mean level per dimension (fifth review of a5601a28: a model answer must pass, not only peak; airs-missions.test.mjs pins a mean of 2.0 or more on every dimension in both case versions, with no exception):
| Dimension | Friday 1.0.0 / 1.1.0 | Monday 1.0.0 / 1.1.0 | Steering 1.0.0 / 1.1.0 |
|---|---|---|---|
| S1 | 4.00 / 4.00 | 4.00 / 4.00 | 4.00 / 4.00 |
| S2 | 4.00 / 4.00 | 4.00 / 4.00 | 4.00 / 4.00 |
| T1 | 3.25 / 3.25 | 3.29 / 3.29 | 3.00 / 3.00 |
| T2 | 2.00 / 2.00 | 2.00 / 2.00 | 2.25 / 2.25 |
| T3 | 3.00 / 3.00 | 2.33 / 2.33 | 3.00 / 3.00 |
| E1 | 4.00 / 4.00 | 4.00 / 4.00 | 4.00 / 4.00 |
| E2 | 4.00 / 4.00 | 4.00 / 4.00 | 4.00 / 4.00 |
| P1 | 4.00 / 4.00 | 4.00 / 4.00 | 4.00 / 4.00 |
| P2 | 3.00 / 3.00 | 2.67 / 2.67 | 3.00 / 3.00 |
The steering pack's T2 mean was 1.25 at the fifth review of a5601a28 and pinned as a shortfall: two of its four T2 points had no discriminating act and the keep-green note no cue. Closed at its verifiers' round: each flawed line has a test of its flaw that does not check it (4.3), and the model answer's T2 mean is 2.25 in both versions (the migration line 3, the other three 2); the strong path's is 2.25 too.
Gaming (check 1), after the fifth review of a5601a28 and its verifiers (4 October 2026), as the same test prints it, Assessment mode. No gate policy reaches a mean of 2.0 on any dimension: the highest over 200 seeds is The Friday file 1.00, The Monday queue 1.33 (never_verify, T3), the steering pack 1.00. The case-blind "responsible" policies are a grid of act policies and line pickers, over 24 seeds in both case versions. The act policies: the suite with files first or questions first; coached in responsible-AI phrasing, with files first, questions first, or the shorter half of the limits, each run with six copy-in habits (five vocabularies: every coached title (governance, risk, legal, counsel, policy, compliance, audit, controller, finance), legal and governance, finance and governance, legal only, risk and compliance; and everyone the case names before the call); four that open pieces (every piece flagged on its own cue as it arrives, then no reply, every piece accepted or every piece challenged once the files are read, and, on the queue, every case opened); and nine that read what the time allows, flag every piece, open each queue case, and answer each piece by a rule of its surface: absolute words, stated confidence of 85% or more queried (since Fable's review of scorer round six, every stated confidence the mission shows, either way; below), alternately in the order shown, a seeded coin, keeping what costs to query and a coin on the rest, or a fixed mask of positions (odd, even, first half, second half). Fifteen line pickers read only the words (three for the claims groups: a "checked" line with the fewest trusting words, the one naming a file the person opened under any name the line uses for it, or the one whose source the AI named when asked), taking lines in the order the picker shows them for the run's seed, one per claims group as the picker allows. None reaches 2.0 on any dimension, and no known gap is pinned: the highest mean is The Friday file 1.75 (T1: alternation, stated confidence, odd positions, second half), The Monday queue 1.54 (T3), the steering pack 1.83 (T1: alternation, odd positions, second half, in 1.1.0); E1 at most 1.00 on Friday and the steering pack, 0.67 on Monday. Besides the grid, every copy-in habit on the single-call missions is checked without play (3.8). Four earlier claims were too strong. The second copy round's "no case-blind gap left" held only for a coached vocabulary of governance, risk and legal. The sixth review's closure, a needless party whose title the habit copies, held only for the one vocabulary the harness used (Friday E1 2.00 to 3.25, the steering pack 2.00 to 3.00). The grid opened no piece, so no claims group was ever probed; once a file read early counts as a check (3.5), reading every file and challenging every piece reached Friday T1 2.71. And the fifth review's closure, who approves moving with the case version, held for its five vocabularies only: "everyone the files name" reached Friday 1.0.0 E1 2.00 to 3.25 and the steering pack 1.0.0 2.00 to 3.00, "everyone but the supplier" the steering pack 1.0.0 2.00 to 3.00, and a title set such as controller and training the steering pack 1.1.0 2.00 to 3.00; and Level 3 was read piece by piece, so a reply rule right on most pieces earned 3 on those (Friday T1 up to 2.60 by absolute words, 2.35 at 85% confidence or more, about 1.93 by a coin). Closed by: who approves moves with the case version, each approval party has a needless twin with its title, and each version has two needless parties or more (3.8); E1 Level 2 needs what says who approves read (3.8); refusal wholesale is not a match (3.5); T1 Level 3 needs every high-consequence reply to match its check (3.5); the push-back or query that costs is no longer only the sound piece's, and Juno's absolute words no longer mark the flawed pieces; the steering pack's training and testing lines are load-bearing in both versions; the Friday decoys that only swapped a name are wrong on a case fact. A strong run with one defensible distractor added (c-x-csat-pause or c-x-legal-every-market on Friday, c-x-decline-sale on Monday) scores T3 3 at the call, not 1. Flagging every piece and tile, taking the reading shown first: T2 1.15 (Friday; 1.17 in 1.1.0), 1.17 (Monday), 1.28 (steering), at most 1.75 on any seed.
Scorer round six (4 October 2026). The grid now lives in one module, platform/engine/case-blind.mjs, which airs-missions.test.mjs and airs-held.test.mjs both import, so the public and held grids cannot drift apart (before, the held grid was a stale copy: three of the act policies, one vocabulary and twelve pickers). It adds two line pickers (the shortest line of each kind; the on-topic reader, which counts the case's own words and marks down a proper noun, date or number the case never shows), a copy-in habit per party for the coached policy ("everyone but" that party), and a framing sweep: six framing pickers (longest, most overlap with the brief, shortest, fewest numbers or dates, most first person, most negation), the habit of the words a person learned in other missions' underlying framings, and the on-topic reader, each with every line picker, in both case versions. Before the copy changes of the round, the shortest-framing picker took the underlying framing on every seed on The Friday file and The Monday queue (S1 2.00); "most negation" did on Friday ("alone"); the fewest-numbers picker took it on half the seeds of the steering pack; and the shortest-line picker took the fitting disclosure on The Monday queue (E2 2.21) and the steering pack (E2 2.00). The framings now match in numbers and dates, first person and negation, the underlying one is neither the shortest, the longest, the one sharing most words with the brief, nor the most on-topic, and no word marks an underlying framing in two missions ("told" did, in the steering pack and The water report); each disclosure pair is one length. After it, none reaches 2.0 on any dimension and no gap is pinned. The highest case-blind means, the framing sweep included: The Friday file S1 1.00, S2 1.25, T1 1.75, T2 0.67, T3 1.00, E1 1.00, E2 1.04, P1 1.00, P2 1.00; The Monday queue S1 1.42, S2 1.46, T1 1.00, T2 1.11, T3 1.67, E1 0.67, E2 1.42, P1 1.00, P2 0.83; the steering pack S1 1.00, S2 1.63, T1 1.83, T2 1.21, T3 1.00, E1 1.00, E2 1.08, P1 1.00, P2 1.00. The model answers' and strong runs' levels are unchanged on all three missions. One rule change reads differently on a public mission without lines: the steering pack's amber status (c-status-amber, role disclosure) now discloses for E2 as the fitted change fits (3.9), so a run that sets it and adds no line reads E2 2, not 0.
Fable's review of scorer round six (4 October 2026). A single stated-confidence cut at 85% left every other cut untried; on The water report a cut at 87% refused three flawed pieces and two sound ones in both versions. The grid now sweeps every distinct stated confidence a mission shows, querying each piece at or above the value or at or below it (confidenceRules), and adds "read two files, challenge all": only the files a falsifier line against the lean names, every piece flagged, then every piece challenged, the cell where a partial read meets a challenge-everything habit. Refusal wholesale for T2 Level 3, P2 at the call and a removal read as checked before acting is read over every sound reply (3.5), so neither cell pays on its refusals (the cuts reach T2 0.82 and P2 1.00 on The water report). The two-file reader, which spends two hours and so never reaches the 12:00 update, found the water report's P1 lines at the call marked by surface: two basis decoys carried a test-wise word ("nothing else", "asked for"), the trigger decoy named no person and the handoff decoy no time, and it reached P1 4.00 (1.0.0, test-wise one per stem) and 2.21 (1.1.0, names people); each group's fitting line and decoy now carry the same numbers and names. Every gate policy that spends the update's hours now answers it with its own habit (the framing shown first, and every update line for a policy whose habit is everything), and the held gate prints S1, E2 and P1 per point. After it, none reaches 2.0 on any dimension and no gap is pinned. The highest case-blind means: The Friday file and the steering pack unchanged; The Monday queue T1 1.00 to 1.14 (query even positions), the rest unchanged; The water report S1 0.98, S2 1.50, T1 1.80, T2 1.00, T3 1.00, E1 1.00, E2 1.38, P1 1.58 (the two-file reader, 1.1.0), P2 1.00. The blanket gate's highest on The water report over 200 seeds: 1.00. The model answers' levels are unchanged on all four missions.
The call's note picker load (items, a piece's claims group counted as one; words of the lines drawn), as airs-missions.test.mjs logs it after the fifth review's verifiers: The Friday file 32 items and 494 words on the model answer's and the strong path (1.1.0: 498), 33 and 508 with everything opened; The Monday queue 31 and 460 (1.1.0: 457) on both paths, 59 and about 1,050 with all 36 cases opened; the steering pack 36 and 492 on the model answer's path (1.1.0: 36 and 493), 33 and 450 on the strong path, 36 and 492 with everything opened. The target, about 30 items and 450 words on the realistic paths, is met on Monday and on the steering pack's strong path; Friday is two items and about 45 words over, and the steering pack's model answer six items and about 40 words over, because it checks every load-bearing line, which draws each one's claims group.
5. Result types§
| Result | Basis | What it carries | Use |
|---|---|---|---|
| Development reading (AIRS-J/1 Insufficient or Provisional; replaces Indicative) | One mission, any mode; or any set that does not meet §8.3, §6.4 and §10's Moderate | Every observation: dimension, decision point, level, anchor, cited acts, scaffolding, stakes profile; "n of 3" per dimension. Dimension means only where n ≥ 3. No band, no certification. | Development. Shown after the debrief in Practice; by name to the company in Cohort (D49, D67). |
| Certifying (replaces Assessed) | Three or more Assessment missions, scaffolding: none, first runs, all under one contract version (3.0.0 and 3.1.0 record a check differently, 3.5), meeting §8.3 for all nine dimensions, §6.4, and §10 Moderate |
AIRS §12.5 in full: each reportable dimension score, the four domain scores, the band, confidence, the stakes profiles, the anchors applied, and the evidence for each observation | Certification, employee assessment and hiring (D60, D61). The employer decides (D77). |
Notice. Before a Cohort or Assessment run starts, notice-4 (platform/surfaces/consent.mjs, docs/CONSENT_V2.md; blueprint 8.4, D74) tells the person who sees their record and profile (the facilitator, the line manager and the company) and that both may be used for development, assessment and hiring. No run played before notice-4 went live is used for assessment or hiring (blueprint 17, D61).
Aggregation is AIRS-J/1's own. A dimension score is the mean of its valid observations, a domain score the mean of its dimensions, the composite the mean of the nine (§12.2.1). The band follows §12.4: Certified 2.3 to 2.9; Certified with Merit 3.0 to 3.4 with no dimension below 2.5; Certified with Distinction 3.5 and above with no dimension below 3.0. All six §12.2 requirements must hold, or the result is Not Certified. StepSim adds no rule of its own.
Group view. Counts of observations at each level, per dimension, under the floor of 10, with intervals and banding (blueprint 6.8). Never a ranking (§12.7).
6. The validation protocol§
StepSim's protocol (blueprint section 7, D80, D81, D85), restated against AIRS dimensions and bands. A version of this specification is validated when its latest report shows checks 1 to 5. The report is versioned with this file and signed by the owner.
| # | Check | Passes when |
|---|---|---|
| 1 | Blanket strategies fail | The blanket-strategy gate passes on every mission (blueprint 6.7). Also, across at least 200 seeds per mission, no blanket policy reaches 2.0 on any dimension it games. The policies are always accept, always challenge, open everything, flag everything, tick every condition, add every line, escalate everything, let the deadline pass, and decide immediately, and case-blind "responsible" policies that pick lines by their words alone, in the order the picker shows them, among them policies that open, flag and answer every piece, so each piece's claims lines are probed, and policies that answer each piece by a rule of its surface (absolute words, stated confidence, alternation, a coin, keeping what costs, a mask of positions). On a mission whose E1 rests on the call alone, no copy-in habit (any union of title words, everyone but one, everyone the case names) copies exactly who approves in either case version. The case-blind grid also sweeps framing pickers by surface (length, numbers and dates, first person, negation, words learned on other missions, on-topic words), the shortest line of each kind and an on-topic line reader (scorer round six). Every gate run is an Assessment run (scaffolding: none), so no R7 cap hides a level. |
| 2 | Known profiles land where they should | L2 synthetic decision-makers reach the dimension levels their parameters map to, by a map declared before the run. Each of the eight L3 persona cards reaches its declared band, or one band from it, at a rate declared before the run. |
| 3 | Stable across missions | The same L2 decision-maker, over at least 50 seeds per mission, scores within one level per dimension on every mission offering that dimension. Assessment missions join as built. |
| 4 | Readings agree | Every case reading an observation rests on is at least panel-rated, with its agreement recorded: reliability, credibility, line and condition tags, proportion, supports, consequence. |
| 5 | Real sessions | As sessions arrive: StepSim's levels against an authorised facilitator's independent reading of the same transcripts, at quadratic-weighted κ ≥ 0.70 per dimension; within-domain disattenuated correlations below 0.85 (§11.3); check 3 re-run on real people. The report states the count, which may be zero. |
| 6 | Prediction | The composite from three Assessment missions goes with a higher earned-success share (blueprint 6.6) on a fourth, held-out Assessment mission, by a margin declared before the run. It is run on L3 persona cards, LLM-played, at a declared seed count. L2 agents do not count here, because their parameters set both level and outcome. Clients' later workplace performance is added as they share it. |
| 7 | Fairness | Matched L3 pairs sharing one decision-style brief and differing in gender, age, nationality and first language reach the same band within a declared tolerance. Real sessions are checked for band gaps by group as they arrive (AIRS §15). |
What each claim word rests on, named when asked:
| Word | Rests on |
|---|---|
| Validated | Checks 1 to 5, in the latest report the owner signs |
| Psychometric | Levels derived by written, published rules (section 3), with checks 2 and 3; §11.2 and §11.3 as real sessions arrive |
| Predicts job performance | Check 6: held-out simulation outcomes from personas now, workplace outcomes as clients share them. Needs four Assessment missions. |
| Bias-free | Check 7, at the declared tolerance |
7. Gaps and build list, in order§
7.1 Acts and fields§
- Contract v3 (
platform/engine/contract.mjs,docs/MEASUREMENT_CONTRACT.md): thenote_lineact;readingoncue_flagged; distractor cues logged;condition_set.source; framing at the call; the run fields of 2.3; modeassessment; the package fields of 3.3, checked byvalidateMission. - The engine (
platform/engine/session.mjs): the line bank and its unlocks;sourceon conditions; distractor flags in The Friday file (actions.mjsflag-tile) and The Monday queue (no moreevidence_openedonev-week-n); no highlighted spans and no unlock notices whenscaffoldingisnone. - The scorer:
platform/engine/airs.mjs, withairs.test.mjspinning every rule in section 3 against hand-built logs, R1 and R2 checked over every run, and replay identity. It emits Annex C records.
7.2 Per mission: tags, copy and decision points§
For each of the-friday-file, the-monday-queue and the-steering-pack, in package.mjs, copy.mjs, screens.mjs, and the mission spec in docs/missions/:
decisionPoints,stakes,loadBearingandconsequence,probes,disconfirms, targetkindandentitled, condition and controldimensionandrole, optionproportion, later-branch tags, framingclass. Each is a case reading, panel-rated before it counts (check 4).- Copy: the underlying framing states the owner and what the call is not; each
mitigatescondition gets a "because" clause naming pathway and bearer; indicator conditions name who watches. - Line banks: plan lines (S2);
nondelegable,frame,reframe_trigger;claims_checkedper load-bearing claim; cue readings;falsifier;accepted_harm;authority_basis;disclosure,fitted,competing_interests;basis,trigger,handoff;evidence_basis,owns,cell; one standing rule per dimension. Each with distractors. - New decision points: The Friday file, Grace's hour-4 request for a one-line update (E2, P1). The Monday queue, a week-2 update to Tomas (E2), and week 3 as a re-framing point (S1). The steering pack, Ingrid's 15:00 check-in (E2, P1).
- The blanket gate extended to the new policies (check 1).
- Closed at the fifth review's verifiers' round (4 October 2026): the steering pack's T2 points each have a discriminating act and the keep-green note a cue in its words, so its model answer passes T2 (4.5). Open: the steering pack's standing-rule line
l-rule-checknames the class it would catch (classOf) but carries norole: persisting_check, so T1 Level 4 is out of reach there (4.5).
7.3 Assessment missions§
- At least three held-back missions for one form, never shown on any demo, the showcase or the film (blueprint section 5, D76). Together they span three distinct primary families with at least one of C1, C4, C5b or C6. Put two non-financial primaries in each form, so one mission failing eligibility does not sink the form. Each mission gives every dimension at least one decision point, and E2, P1 and S1 two. The water report does since the scorer round of 4 October 2026: its 12:00 update is the second point for S1, E2 and P1, and its notice is P2's second (3.4, 3.9, 3.10, 3.11).
- Six for two parallel forms (§9.4), so a retake gets a different form.
- A fourth mission outside the form being scored, for check 6. A mission from the other form serves.
- The notice on outside AI use (§9.2), and a published accessibility and access-support process (§8.9).
- Identity (§9.1, D79). Built. Before the session the facilitator registers each candidate in the Assessment session by name, company and an identifier the company uses for them, and hands the person their candidate id and a one-time start code. The same company and identifier is the same person in every session, so a person's runs are linked across sessions under the company's licence and carry one participant key, the person key. At the start of the sitting the person checks in with the session code, candidate id and start code. The server keeps only the code's sha256, compares it with timingSafeEqual, and spends it on the first check-in that matches; it expires 72 hours after issue, and the facilitator can issue a new one, which ends the old code and its sitting. A check-in opens a sitting for four hours, ended early when the session closes; only a sitting starts that candidate's runs. A check-in, a start and a new code each read and write the candidate's record under that candidate's lock, so a new code issued while the old one checks in, or while a run starts, is never undone by it. A check-in under an unknown session code gets the answer a wrong start code gets. The tab keeps the sitting for the person's next mission only once they type their candidate id again (the page never shows it), and lets it go at a mismatch and after their last mission. Every run records the check it started under (
identity: the person key, session, candidate id, method and time of check-in, and the notice version the run started under), as §9.1 asks the log to. Assessment sessions are facilitated in the room: the facilitator checks each person against their registration before handing over the code. When the session closes, or is deleted, its runs stop where they stand: the server draws no more screens for them and takes no more gestures, finished or not. A remote session needs a live check (§9.1), which is not built, so Assessment runs in the room only until it is. - Exposure and retirement (§9.3). Built. The server counts, per mission and case version, the people who have seen it (a person sees a case when their run of it starts). The published threshold is 200 people per case version: a case version is offered in Assessment until 200 people have seen it, then retired from Assessment, so exposure never exceeds it. A leak report retires the mission from Assessment at once, every case version of it (§9.5), by the operator's
report-leakcommand: a mission's case versions share most of their stimulus, so a leak of one exposes the others. The report names the version found and is recorded against each. Retirement is from Assessment only: a retired case may be released for practice (§9.3). A person never gets a mission they have seen, in any session the record links; today that is Assessment, the only mode whose runs carry the person key. When every case version of a mission is retired, or the person has seen it, the server refuses the start and says so. A mission id the server does not carry gets the answer a retired mission gets, and only after the sitting is checked, so nobody can probe for a held-back mission's id. The third §9.3 trigger, a shift in the performance distribution consistent with leakage, is read in the validation report (section 6), not by the server. The threshold of 200 is a starting value, proposed by the build on 4 October 2026 under the owner's approval of the certification blockers; the first validation report reviews it. - The end of an Assessment run reads nothing of the case, because the case must stay unexposed (§9.3). In place of the debrief and the done card the person sees one fixed screen: their own call as sent (the option, the limits or rules chosen, who was told, who is accountable; or that the deadline passed), that the run is finished and is not read on screen, that the result comes in the AIRS-J/1 report (§12.5), and who receives it (§12.6). It shows no reliability label, no cue or reading as found or missed, no line as fitting or not, no author's play, no other path, no outcome or process reading and no rewind. The device keeps the finished run sealed: its versions and finish time, never its events. The copy that is scored belongs on the server. The report cites acts, decision points and anchors (§12.5). Once a person has seen a case's report, that case version is exposed for them.
- Where an Assessment run is held. An Assessment run is held on StepSim's server, not the participant's device. The server draws the seed, keeps the run's log, runs the engine and sends the device only the screen to draw, every control addressed by a seed-drawn token; a gesture is accepted only for a control on the screen the server last drew. The device loads no mission package, engine or scorer, and keeps only a run identifier and its secret. At the finish the server scores the run with the rule-based scorer from its own log, stamping the instrument version and scoring time so a replay is byte-identical, and stores the result with the run; nothing about the result is returned to the participant's device. A certifying mission's package is never served to a browser. The static server serves an allow-list of public paths only; held-back packages live outside it and are read only by the server's own registry, and the public registry refuses any package whose exposure is not
public. The device clears the run identifier and its secret once the end screen is drawn. The server chooses the case version, of those not retired the one the fewest people have seen, never a mission the person has seen, and stamps the instrument version at the start; a run is never continued or scored on other code, and a run a deploy ends cannot be started again, so deploys wait until no Assessment session is open. Each Assessment run is started under an Assessment session code by a candidate the facilitator registered and the server checked in (Identity, above); one person's runs share one participant key across sessions, and the server assembles their first runs into one AIRS-J/1 result. The public Assessment page names no held-back mission: an id outside the public list draws with the shared stylesheet. Opening a session, registering a candidate and sending a cohort journal are limited per client address, and a person's record with nothing seen is deleted by retention, so no route grows storage without bound.
7.4 Renderers and reports§
- Retire claim levels and themes on every surface: the debriefs (
platform/missions/<id>/debrief.mjs), the record (platform/surfaces/record.mjs), the session report (platform/surfaces/report.mjs,tools/build-sample-report.mjs), and the group view (platform/surfaces/group.mjs, each mission'sgroup.mjs). - Development-reading view: per dimension, its observations in order, each with decision point, level, the anchor's words and the acts cited, after the pathway, never before it. Scaffolding shown. "n of 3" where under minimum.
- Certifying report: AIRS §12.5 in full, with band and confidence; the §2.2 scope statement.
- Group view: level counts per dimension under the floor, intervals and banding.
- Tests that pin the retired words (the "no single score" pins listed in blueprint section 17) move to the new terms.
tools/no-overclaim-scan.mjsnames this file as the basis for "validated", "psychometric", "predicts job performance" and "bias-free".
8. Versioning§
- Versions are semantic. A change to the allocation in 3.2, a level rule, a result type or a validation check is a major version. A new tag or decision point that changes no rule is a minor version. Wording only is a patch.
- A version is frozen before any result is read under it. Every result names the version it was read under (2.3).
- Every change is logged here.
| Version | Date | Change |
|---|---|---|
| 1.0.0-draft | 30 September 2026 | The StepSim Judgment Standard, first draft |
| 1.0.0 | 1 October 2026 | Owner approval: an evidence standard, four claims, nine themes, no levels |
| 2.0.0 | 3 October 2026 | Rewritten as AIRS-J/1's instrument specification (D86). Claims, themes and the Follows/Checks/Calibrated levels retired as a scale; calibrated reliance kept inside T1. The act-to-dimension map, coverage per mission, result types mapped to AIRS, and the validation protocol against AIRS dimensions and bands. Approved by the owner. |
| 2.0.0 (pre-release amendment) | 3 October 2026 | Level rules for E1 at the controls and at a held item (3.8) and for P2 in a week (3.11), with the queue fields they read (3.3). They complete decision points 2.0.0 already listed in 2.4 and 4.2 but gave no rule for. Approved by the owner on 3 October 2026 as a pre-release amendment: 2.0.0 is approved but not released, and no result has been read under it, so it is not yet frozen and keeps its number. Further amended the same day, before release, after review: a P2 find in a week is any T1 check of the piece, including an ai_calibration query that probes it (3.11); The Monday queue's week 3 is a P2 point again, so P2 counts 4 there and 6 in total (4.2, 4.4); an E1 limit set back to its default is left at its default (3.8); the held-item Level 2 rule reads the routes tag (3.8); the week Level 1 rule names a correction not tied to an error found that period (3.11); the controls added to the allocation table (3.2); section 4's counts, ceilings and certifying blockers brought to the scorer's current levels (4.1 to 4.5). Further amended the same day, before release, after a second review: a P2 error found in a week stays open, and is read at each later week point, until a change addresses it, so a correction made the week after the find is read on its evidence rather than as unaimed, and a recurrence while the error is open still reads Level 0 (3.11); the Level 2 notes of 3.4 and 3.7 and the disclosure note of 3.9 brought to the current copy, where S1, T3 and E2 reach 4. Further amended the same day, before release: in Assessment every piece carries its own flag, a piece flag on a flawed piece is a flag of the cue in its words, and a piece with no cue in its words has a declared distractor cue (2.6 rule 3, 3.3); no level rule changes; 3.7 Level 4 no longer reads as if the watcher texts were still to come. Further amended the same day, before release, after a third review: in a week, a change already in force when a piece is checked again addresses it, so opening a piece after its correction reopens no error and the class's later weeks no longer read 0 (3.11); a flag finds an error only with the fitting reading, so a flag with no reading or a reading that does not fit opens no week point and earns no correction credit (3.6 Level 1, 3.11); T3's clean rule reads the whole condition set, so a distractor condition of any role, or every condition on offer ticked, caps T3 at 1 (3.7); check 1 runs the blanket gate in Assessment and adds case-blind policies that pick lines by their words (6). Further amended the same day, before release: 7.3 says what the end of an Assessment run shows and what the device keeps (one fixed screen with the person's own call, what happens next and who receives the result, and nothing that reads the case; the finished run kept without its events); no level rule changes. Further amended the same day, before release, after a fourth review: copying every party in, one with no approval to give and no entitlement to be told among them, escalates everything at the call and at a held item, whatever the option (3.8); T3's clean rule applies at every T3 decision point, and at the controls and in a week reads every control on offer in force (3.7); an item inside such a set addresses nothing for T1, and a sound piece left standing with no reply is matched only with a claims_checked line saying checked (3.5); a week's correction made inside such a set is a correction in general terms, Level 1 (3.11); lines of one kind whose values exclude each other are each other's distractor (R6). Further amended the same day, before release, after the three missions' copy rewrite: 4.1 to 4.3 name each mission's distractor conditions, needless party and claims-triple group, and The Monday queue's ceiling values by case version; 4.5 adds the model answers' levels and the gaming results; no level rule changes. Further amended the same day, before release, after a fifth review: a distractor in a T3 set now lowers what the fitting items earn by one level each, never below 1, while everything on offer still caps T3 at 1, so a defensible distractor no longer costs the whole level (3.7, R6). A target tagged requester, the party the deliverable goes to, is never the needless party at E1 (3.3, 3.8). P2 Level 3 at the later choice reads the cell line against the outcome shown and the response stated, not the process half of the cell, which Assessment never shows (3.11). 4.1 to 4.3 name each mission's requester and the package changes still pending. Check 1's case-blind policies take lines in the order shown and run as a grid of act policies and line pickers, with every result at 2.0 or more pinned as a known gap (6, 4.5). Further amended the same day, before release, after the second copy round: each requester tagged and each mission's needless party named, The Friday file's model answer at T1 4, The Monday queue 1.1.0's six ceilings, no case-blind gap left, and a piece's claims lines shown as one choice once it is opened (2.7, 4.1 to 4.5); no level rule changes. Further amended the same day, before release, after the third copy round: The Monday queue's claims lines are pairs on the nine cases load-bearing in either version and the steering pack's triples are on its seven, as The Friday file's, to keep the note picker readable (4.2, 4.3); 4.5's case-blind maximum for The Monday queue follows; no level rule changes. Further amended the same day, before release, after a sixth review: copying in the needless party (an inform target neither entitled nor the requester) caps E1 at the call at 1 whoever else is copied in, and each mission's needless party carries a title a coached habit copies (3.8, 4.1 to 4.3); in a mission without periods, a piece whose linked files were all read before it arrived is checked as it arrives (3.5); a refusal of a flawed piece in a run that refuses wholesale (more than half of the sound pieces checked and replied to challenged or overridden) is not a match (3.5); a piece checked opens its claims group, and every piece carries one (2.7, 4.1 to 4.3); check 1's grid gains four act policies that open pieces and two claims pickers, and its coached vocabulary widens; 4.5 corrects the second copy round's "no case-blind gap left" and gives the note picker's load. Further amended on 4 October 2026, before release, after a fifth review of the scorer at commit a5601a28: one linked file read before a piece arrives checks it as it arrives (contract 3.1.0; a 3.0.0 run keeps its own rule), a probing question asked before a piece arrives checks it outside a queue, and in a queue only opening the case checks it (3.5); refusal wholesale needs at least three sound pieces checked and replied to (3.5); E1 Level 2 needs a file or answer that says who approves read, the party an option sent hands the call to (recipientOf) is no needless copy, and who approves moves with the case version on The Friday file and the steering pack (3.3, 3.8, 4.1, 4.3); the sixth review's "needless title" closure is corrected to hold for one vocabulary only, check 1 sweeps five coached vocabularies, the steering pack runs seven review hours with pairs of claims lines, and 4.5 adds the model answers' means and the steering pack's T2 shortfall. Further amended the same day, before release, after that review's verifiers: T1 Level 3 needs every high-consequence reply to match its check, the light-check budget leaves out files tagged statesBoundary or disconfirms and held items opened to answer them, refusal wholesale needs three refusals, each piece cites one check (its own before a shared question), a question tagged checks is a check of the piece its answer settles, a verify in a queue checks nothing whatever the version, and a certifying result reads one contract version (3.3, 3.5, 5); an opening is barred from T2 only for the pieces it verified (3.2, 3.6); only a file that names who approves carries statesBoundary, and each single-call mission's approval party has a needless twin with its title, so no copy-in habit is right in either version (3.8, 4.1, 4.3); the steering pack's training and testing lines are load-bearing in both versions (T1 6 there, 17 in total), each flawed line has a test of its flaw, its desk runs eight hours, and its T2 shortfall closes (4.3 to 4.5); check 1 adds reply rules by surface and a claims picker that echoes the AI's cited source (6, 4.5); the queue rule's cost is stated (3.5). |
| 2.0.0 (pre-release amendment) | 4 October 2026 | Assessment integrity, on the owner's approval of the certification blockers (4 October 2026): identity at the start of each sitting by a facilitator-registered candidate and a one-time start code, with a person's runs linked across sessions (§9.1, D79; 2.2, 7.3); exposure counters per mission and case version, the published threshold of 200 people per case version, and retirement from Assessment at the threshold or on a leak report (§9.3, §9.5; 7.3); no person gets a mission they have seen (7.3); 4.4 item 4 follows. No level rule changes. |
| 2.0.0 (pre-release amendment) | 4 October 2026 | Section 5 names notice-4 (blueprint 8.4, D74): who sees the record and profile and what they may be used for, told before a Cohort or Assessment run; no run played before it went live is used for assessment or hiring. Wording only; no level rule changes. |
| 2.0.0 (pre-release amendment) | 4 October 2026 | 7.3, after a security review of the Assessment server: a check-in, a start and a new code run under the candidate's lock; the session's close or deletion ends its runs; an unknown session code at check-in, and an unknown mission id at a start, get the answers a wrong code and a retired mission get; a leak report retires every case version of the mission; the sitting kept on a tab needs the candidate id typed again and goes after the last mission; every run records its notice version; opens, registrations and cohort uploads are rate-limited, and empty person records are deleted. No level rule changes. |
| 2.0.0 (pre-release amendment) | 4 October 2026 | Scorer round six, and Fable's review of The water report, before release; no result has been read under 2.0.0. T2 Level 3: a challenge or override of the flawed piece before the call departs from it, except in a run that refuses wholesale (3.6; the lean held after refusing the flawed piece scored 1 under a person who named no falsifier). T1 Level 3 reads the high-consequence pieces offered before the decision (3.5); a flawed piece queried, then checked and addressed by an own change before the call, is checked before acting (3.5). The E1 Level 0 row at the call states the approval party unreached with no boundary read (3.8, as the code read it). New decision point kind update (2.4), with S1, E2 and P1 rules there (3.4, 3.9, 3.10); P2 at the call for a mission that lists it there (3.11); option and condition informs and the package's deliverableTo, read by E2 as reaching or withholding, with CR-E2 (3.9); an own change with role disclosure discloses, and neither it nor the fitted change is read inside a set that does not discriminate (3.9). The case-blind grid is one shared module with framing pickers, two more line pickers and "everyone but one" habits (4.5, 6 check 1); framing and disclosure copy changed on all three public missions so none reaches 2.0. The water report, held back: its owner of the correction is whoever signed the report, General Counsel is told of every notice, and the quality reviewer only of those he approves (its spec, section 12). |
| 2.0.0 (pre-release amendment) | 4 October 2026 | Fable's review of scorer round six, before release; no result has been read under 2.0.0. Refusal wholesale for T2 Level 3, P2 at the call and a removal read as checked before acting is read over every sound reply before the call, checked or not; T1's own contradiction reading keeps the checked replies (3.5, 3.6, 3.11). T1 Level 3 reads the high-consequence pieces offered before the decision or due within the review hours, so deciding early leaves the unreached ones unchecked (3.5). A flawed piece removed, then checked before the decision, is checked before acting, as one queried and corrected is (3.5). E2 and P1 at an update: no line of the dimension chosen there is no observation; the false-certainty decoy and a contradicted commitment stay 0; an update bank with no disclosure line yields no E2 observation (3.9, 3.10; the code read it as 1). One told reading (partiesTold: a copy-in, an own change's informs, the option's informs) for E1, E2 and a mission's world (3.8, 3.9). P2 at the call reads no attribution line, as none is on offer there (3.11). Check 1's grid sweeps every stated confidence either way and adds a two-file, challenge-everything reader; the gate policies answer the update with their own habit (4.5, 6). The water report: the owner question in the manual's words, an S1 line that separates today's notice (the person's call) from owning the correction (the signer's), manual 8.4's rule as a house register rather than privilege, and number- and name-matched lines at the call (its spec, section 13). |