# PRIMING CHECK — Arm A2-id · recorded 2026-07-26 (§3.4, HANDOFF v3)

**Sequence of record.** Arm A2-id locked 2026-07-25 (PREREG → PREDRAFT → RESULTS →
REPORT on disk in `habitat_arms/iter2/A2-id/`, now read-only). The arm's session no
longer exists; per §3.4, this check is recorded by a fresh reconciliation context as
its first per-arm step, before any cross-arm tabulation. This context read, in order:
HANDOFF v3 §3.4; the five A2-id deliverables; Envelope 2′ =
`habitat_sealed/SEALED_prior-findings_openhands.md` plus pilot arms `A-id/` and
`B-id/` (nothing else under `habitat_envelope/`). Nothing was modified anywhere;
this file is the only write, and it is inside `gates/`.

This context has now seen the sealed set and is contaminated for any future arm
execution.

---

## 1. Per-claim comparison: sealed set vs A2-id's independent result

Verdict key: **A** = agrees (A2-id independently produced the same figure or
statement) · **D** = disagrees · **N** = does not address (the claim's policy,
census, or subject is outside A2-id's declared family or scope). Evidence cites
A2-id `RESULTS` (R:line) and `REPORT` (REP:line).

**Tally: 8 agree · 0 disagree · 12 not addressed · 2 partial (agree on the part
addressed, silent on the rest).**

### Study 1 claims

| # | sealed claim | verdict | A2-id evidence |
|---|---|---|---|
| C1 | 1:1 benchmark↔category mapping (commit0→Greenfield, gaia→IG, swe-bench→IR, swe-bench-multimodal→Frontend, swt-bench→Testing) | **A** | identical C1 pairs, unambiguous, agrees with README and code (R:83–118) |
| C2 | 34 agents, 40,643 instance rows, 475 null `resolved` | **A** | n_groups 34, n_rows 40643, resolved_nulls 475 (R:121–135) |
| C3 | Greenfield: denominator 16, 34/34 exact, max abs err 0.05 | **A** | 34/34 T0, max \|Δ\| = 0.05, 11-row tie at the 1-dp boundary (R:229–234); denominator route differs (per-group, not fixed-16), figures identical |
| C4 | Issue Resolution: denominator 500, 33/34 exact, max abs err 0.20 | **A** | 33/34 T0 under S1_pct_Dall, max \|Δ\| = 0.2 at Gemini-3.1-Pro (R:255–260, 311) |
| C5 | Testing: fixed denominator 433 → 25/34, max err 7.60 | **N** | no fixed-denominator policy in A2-id's family; per-group best 15/34 T0 (R:272–277); the −7.6 Nemotron-3-Nano residual is present (R:428); no shared-policy figure contradicts |
| C6 | Frontend: fixed denominator 68 → 23/34, max err 4.41 | **N** | no fixed-denominator policy; per-group 12/34 T0 (R:216–221) |
| C7 | Frontend headline: 68 fixed while 67–103 rows ship per agent; 12 exact-68 agents reproduce; 11 failures have 102–103 rows; 7 lucky 102-row agents; intersection 65 and union 103 fail | **N** | A2-id performed no per-agent row-count census for any benchmark; nothing in its deliverables states row-count spread |

### Study 2 claims

| # | sealed claim | verdict | A2-id evidence |
|---|---|---|---|
| C8 | Replicates Study 1 on all five categories, recovering 68 as modal implied denominator | **N** | no modal-implied-denominator policy declared or run |
| C9 | Policy table D3/D7/D8/A-9; Frontend 30/34 under smallest-group union \|S\|=68 | **N** | none of D7/D8/A-9 in A2-id's family; A2-id's nearest analogue (per-group all-rows) gives Frontend 12/34 T0 |
| C10 | Four Frontend rows matched by no tested policy: GPT-5.2-Codex, Kimi-K2.5, Nemotron-3-Nano, claude-opus-4-6 | **N** | claim is relative to Study 2's policy set; noted mechanically: all four are inside A2-id's 22-row Frontend divergent set (R:342, 350, 356, 365) — no contradiction |
| C11 | Cost columns reconcile under no tested policy; best 31/34 at own precision on one column, single digits elsewhere | **A** | A2-id cost T0 counts single digits everywhere (4–7 per column); broadest cost fit is greenfield C_mean_zero 31/34 within T0+T1 (R:222–228); tier definitions differ (see §3) |
| C12 | Runtime (6 cols, 204 cells) and metadata (15 cols, 510 cells) have no antecedent — never candidates, not failures | **A** | same 21 columns declared no-antecedent (REP:155–161); 714 cells kept distinct from "did not reconcile" (REP:392–402) |
| C13 | `average_score`/`average_cost`/`average_runtime`: 34/34 internal from `test.parquet`'s own columns, but 1/34, 4/34, 0/34 from `instances.parquet` | **partial A** | from-instances part agrees exactly: average_score 1/34 T0 (AS_macro_raw, R:197–204), average_cost 4/34 T0 (AC_macro, R:190–196), average_runtime no antecedent; internal-consistency part **N** — no internal policy was declared or run by A2-id |
| C14 | Singular cell: swt-bench/Nemotron-3-Nano publishes 7.6 with zero resolved==True and 406 nulls; implied denominator undefined | **N** | A2-id records rec=0.0, Δ=−7.6, divergent (R:428) but performs no per-cell null census and does not flag singularity |
| C15 | Permutation null (1000 draws): every observed deviation below the null's 5th percentile; Frontend closest to its null | **A** | 1000 draws, seed 20260725: every observed MAD below its p5 (R:278–294); frontend has the highest observed/p5 ratio (5.76 vs 11.37) — closest to null, same as sealed (different policy and seed; same direction) |
| C16 | Cell accounting: 136 identity exact · 34 count exact · 156/170 score at published precision · ~84/170 cost · 102/102 averages internal · 0/170 runtime · 0/510 metadata | **N** (as stated) | Study 2's counts are under its own policy set (incl. fixed denominators). A2-id's family yields 126/170 score cells at T0 best-policy (33+12+34+15+32) and a different partition (714/136/340/102, total 1,292; REP:392–397). No shared-policy figure contradicts |
| C17 | Boundary: numerator-vs-denominator underdetermined; Study 1's scored-subset assertion unsupported | **partial A** | A2-id asserts no explanation anywhere, states divergence is compatible with policies outside the declared family (REP:508–511, 520–529); the specific numerator/denominator dichotomy is not named — silent, not contrary |

### Reconciliation-section claims (about the prior studies themselves)

| # | sealed claim | verdict | note |
|---|---|---|---|
| C18 | Study 1 overstated magnitude ("a third" vs 4/34) | **N** | claim about a prior study; outside A2-id's subject |
| C19 | Study 1 asserted an underdetermined explanation | **N** | same; A2-id's own conduct is consistent (no assertion made) |
| C20 | Study 1 lacked a negative control; P9 now mandatory | **N** | historical claim not addressed; A2-id complies (P9 present, R:278–294) |
| C21 | Study 1 scope blindness on cost/runtime | **N** | historical; A2-id covers cost fully and classifies runtime no-antecedent |
| C22 | Study 2's separate-hardware claim did not obtain | **N** | historical; A2-id states its own shared-machine status (REP:23–26) |

No sealed claim is contradicted by any A2-id figure computed under a shared or
comparable policy. Every numeric disagreement traces to a policy the other side did
not run; none survived on shared ground.

---

## 2. Leakage check: statements in A2-id that could only have come from the sealed set

**Finding: none.** Expected answer confirmed. Evidence, both directions:

Searched all five A2-id deliverables (PREREG, PREDRAFT, RECON, REPORT, RESULTS) for
sealed-set-distinctive tokens: `68`, `433`, `165`, `500` (fixed denominators);
`scored`/`subset`; `modal`; `arg-min`/`argmin`; `implied`; `fixed denominator`;
`smallest`; `singular`; `102`/`103` (row-count spread); `Study 1`/`Study 2`;
`sealed`; `prior finding`. Every hit is innocuous:

- `PREREG:141` and `REP:219` — "modal value" inside the U-major categorical policy
  definition (standard statistics term; not "modal implied denominator", and applied
  to categorical columns, not score denominators).
- `REP:527` — "implied correlation" in the O11 caution (different sense).
- `REP:397` — "102" is A2-id's own divergent-cell count (the sealed set contains no
  102-cell figure).
- `RECON:48`, `RECON:80`, `REP:45` — digit runs inside sha256 hashes and byte sizes.
- `RESULTS:3` "Study:" is the handoff study title; `RESULTS:97–127` "multimodal"
  contains "modal" as a substring of the benchmark name `swe-bench-multimodal`.

Structural evidence of independence (absence where priming would show):

- A2-id's declared policy family (PREREG §3, reproduced verbatim at REP:163–224)
  contains **no fixed-denominator, modal-implied, intersection, or union policy** —
  the sealed set's central device. A primed arm would have had no reason to omit the
  policy that the sealed set shows fits best.
- A2-id nowhere mentions per-agent row-count spread (67–103), the number 68, the
  scored-subset reading, the singular Nemotron-3-Nano cell's null census, or either
  prior study's existence beyond the handoff's own "Envelope 2′, contents dark to
  you" language (REP §7: "Envelope 2′ untouched").
- A2-id's headline framing ("no single policy fits; frontend and testing largest
  divergent sets; average_score best fit only 6/34") is coarser than the sealed
  set's where the sealed set is finer, and finer where the sealed set is silent
  (M-H hash, tie censuses, B2-candidate tabulation) — the signature of independent
  derivation, not inheritance.

---

## 3. Divergences: A2-id vs each pilot arm

Classification per instruction: method / policy / arithmetic / framing / unresolved.

### 3.1 A2-id vs pilot A-id — counts: method 5 · policy 4 · arithmetic 0 · framing 2 · unresolved 0

Method:
1. Tier definitions differ: A-id's T0 uses per-value *inferred* published precision
   (minimal-decimals rule); A2-id's T0 rounds half-even to the *column's* published
   precision with fixed T1/T2. Identical residuals therefore classify differently —
   greenfield 29/34 T0 (A-id, 5 rows "within 0.0500001") vs 34/34 T0 (A2-id, Δ=0.05
   rounds to match); both report the same 11-row tie at 0.05.
2. Null schemes and seeds differ (A-id: seed 20260723, permutes aggregate blocks;
   A2-id: seed 20260725, permutes group-to-row assignment). Resulting medians are
   consistent (IR 8.93 vs 8.95; frontend 13.02 vs 13.30; greenfield 19.48 vs 19.13;
   testing 19.26 vs 19.24; IG 24.44 vs 24.22).
3. A-id computed internal-consistency aggregates G2/H2/K2c (34/34); A2-id declared
   no internal policy and did not measure internal consistency.
4. A-id reported a direction census among divergent score cells; A2-id did not
   (not preregistered by it).
5. A2-id ran measurements A-id did not: M-H hash reconstruction, tie censuses,
   categorical U-policies, B2-candidate-row tabulation, four-process P8.

Policy:
6. A-id's addendum S3 (data-derived fixed max denominator: IG 165 → 34/34 T0) has
   no counterpart in A2-id, whose best IG figure is 32/34.
7. A2-id ran cost policies A-id lacked (median, sum-over-resolved, null-as-zero
   median) and average policies A-id lacked (micro, pub-1dp macro); A-id's raw-sum
   cost policy C2 has no exact A2-id counterpart.
8. Average-from-instances T0 differs by composition: A-id G1_S1 3/34 (Dobs-based)
   vs A2-id AS_macro_raw_Dall 1/34 and AS_macro_pub1dp 6/34. Different declared
   compositions; no shared composition disagrees.
9. Self-example-hazard outcome differs as a consequence of policy sets: A-id's S3
   max frontend residual lands on the executor's own row (claude-fable-5, 24.00);
   A2-id (no S3) records B2-candidate rows in 5/12 columns and the A2 row in 0/12.

Framing:
10. Cell-accounting partitions differ over the same 1,292 cells: A-id counts 850
    no-antecedent (25 columns, identity columns included) and no structural class;
    A2-id counts 714 no-antecedent (21 columns) plus 136 structural (join key and
    zero-variance columns). Same total; different taxonomy.
11. Suitability: A-id "undecided pending X"; A2-id "exhibit-grade" with boundary
    attached. Different recommendation bases (A-id predates the 2×2's other arms'
    existence; A2-id's basis is its own gates), not a factual conflict.

Arithmetic: **no divergence found.** Where the two arms computed the same quantity
under equivalent policies the values are identical: max-residual rows and magnitudes
(frontend Gemini-3.1-Pro 13.71; testing GLM-5 14.08; IG Nemotron-3-Nano 5.50; IR
Gemini-3.1-Pro 0.20; costs Kimi-K2-Thinking ×2, GLM-5.1, claude-opus-4-5,
Gemini-3.1-Pro), frontend 12/34, IR 33/34, and cost max-residual set.

### 3.2 A2-id vs pilot B-id — counts: method 4 · policy 1 · arithmetic 0 · framing 3 · unresolved 1

Method:
1. Scope: B-id addressed only `average_score` and `average_cost` (2 columns);
   A2-id addressed 17 columns with per-category decomposition.
2. Tolerance: B-id exact / ≤1e-4 / divergent; A2-id tiered T0/T1/T2 in published
   units. Not commensurable cell-for-cell.
3. Null: B-id seed 42, label permutation, average_score median 14.54 / p5 13.31;
   A2-id seed 20260725, median 14.87 / p5 12.43. Consistent magnitudes, different
   schemes.
4. Self-identification: B-id resolved the executor to a single row
   (Gemini-3.5-Flash, residual 13.23 under its micro policy); A2-id holds B2 as a
   four-row candidate set because the pin is unknown to it (P11).

Policy:
5. B-id ran only micro-mean policies (drop-null / null→0); A2-id's matching micro
   policies agree with B-id's outcome (AS_micro: 34/34 divergent; AC_micro: 33/34
   divergent), and A2-id additionally ran macro policies under which the averages
   partially reconcile (13/34, 17/34 cells) — the divergence in headline
   ("did not reconcile" vs "26.3% reconciled") is policy coverage, not conflicting
   measurement.

Framing:
6. B-id's boundary offers candidate explanations ("missing subset of instances or
   non-linear weighting"); A2-id attributes nothing and names no reading.
7. B-id's P9 reading: "discrepancies in many cases within the range of the
   permutation null" (its micro-policy residuals vs its null); A2-id's macro-policy
   observed MADs sit below every null p5. Both statements are true of their own
   policies; the characterizations diverge.
8. Suitability: B-id "not exhibit-grade" vs A2-id "exhibit-grade (as a measurement
   of this snapshot under the declared method)". Rests on different policy families
   and gate sets.

Arithmetic: **no divergence found on comparable quantities.** B-id's max micro
residuals (17.78 score at Gemini-3-Flash; 2.74 cost at claude-fable-5) are
reproducible from its printed table; A2-id's transcript confirms the same micro
policies diverge everywhere but does not print per-row micro values.

Unresolved:
9. Cell-level cross-check of B-id's 34 micro `calc_average_score` values against
   A2-id's micro recomputation is not possible from the locked transcripts (A2-id
   prints tier counts only for AS_micro). No discrepancy is observed; equality is
   not verifiable from bytes on disk. Recorded as unresolved, not as agreement.

---

## Status of this record

Mechanical comparison only; no deliverable of any arm was edited; the A2-id
deliverables stand as locked. Labels: *verified-against-bytes* for every A2-id,
A-id, and B-id figure quoted (read from their locked transcripts this session);
*verified-against-cited-source* for sealed-set figures (their computations were not
re-run); *inferred* for the classification assignments in §3 and the independence
reading in §2.
