# INTEGRITY CHECK — Arm A2-id — reading-required items
**Date:** 2026-07-25 · **Checker:** post-lock integrity context (Claude, separate session; not the arm's executor session) · **Scope:** mechanics only, per instruction. Mechanical items (file presence, layout, both-impl presence, diff presence, nulls-in-RESULTS presence) were verified by the custodian's script and are **not redone here**; its three FAILs were previously diagnosed as script artifacts (work/ layout; draft files matched as reports). Nothing in the arm directory was modified; A2-id was read only.

**Files read:** `HANDOFF_v3_openhands-index_recon.md`; A2-id: `PREREG_A2-id_2026-07-25.md`, `RECON_A2-id_2026-07-25.md`, `PREDRAFT_A2-id_both-directions.md`, `RESULTS_A2-id_2026-07-25.txt` (all 822 lines), `REPORT_A2-id_2026-07-25.md`, `work/impl_A.py`, `work/impl_B.py` (work/ inspected read-only).

---

## Verdicts

| item | verdict | evidence (one line) |
|---|---|---|
| **P12** completeness floor | **PASS** | Preregistration §1 designates `test.parquet` (config `default`) as the target: 34 rows × 38 columns = 1,292 cells; RESULTS CELL TOTALS classify 340 + 102 + 714 + 136 = **1,292 of 1,292** (shortfall 0), and the per-column counts independently re-sum to the same four totals. |
| **P7** code independence | **PASS** (qualified) | `impl_A.py` (pandas/groupby) and `impl_B.py` (pyarrow read + pure-dict aggregation, no pandas import) share no module, import, or file and use disjoint aggregation substrates, with agreement checked by an external comparator (0 mismatches); qualification: both were written by the same executor in the same session, and a few short passages (join-ranking sort key, cell-classification loop) are near-verbatim parallel — independence is architectural, not authorial, which is the strongest form available within one arm ("written independently" as a process fact is CANNOT-VERIFY from bytes; the prereg discloses B was written after A from the prereg text only). |
| **P9** per-column null, coverage | **PASS** (one gap noted) | RESULTS `== NULLS ==` carries per-column entries for 16 of the 17 addressed columns — 13 computed nulls (seed 20260725, 1000 draws, median/5th/95th, observed alongside) plus 3 structural-floor degenerate notes (`agent_name`, `agent_type`, `categories_completed`, ratios stated undefined, not infinite); the one absence is `id`, which the prereg pre-declares as the join key with no policy family (classified structural), so no measured deviation exists that lacks its null — but literal "every addressed column" coverage is 16/17. |
| **L4** localization | **PASS** | Tie census reported for published and recomputed `average_score` (0 tied pairs of 561; sort strictly decidable) in both RESULTS and REPORT §4.5; rows named individually appear only under the preregistered mechanical rules (§3.5 max-\|Δ\| ties included, §3.1 executor rows, join failures — none), with no characterization of what produced any value and an explicit "no localization or nomination is offered." |
| **L5** artifact-before-magnitude | **PASS** (one borderline) | Encoding artifacts (`resolved`/`cost` nulls, zero-variance `categories_completed`, single-valued `agent_name`/`agent_type`, schema-forced `id`) are named as artifacts in REPORT §2 before any downstream magnitude in §4; borderline: §4.4's greenfield entry puts the number ("11-row tie at 0.05") and its artifact label ("exactly the 1-dp rounding boundary — an artifact of published precision") in the same sentence with the number first, rather than in a preceding sentence. |
| **L10** per-claim coverage | **PASS** (marginal) | RECON and REPORT §§1–2 label per claim; REPORT §4 assigns labels by an explicit blanket declaration ("Everything in this section is verified-against-bytes… where a row cites publisher text it is verified-against-cited-source") rather than sentence-by-sentence, and §§5–10 restate already-labeled claims without fresh labels — coverage is achieved, but partly by quantified assignment rather than literal per-claim tags. |
| **O3** null-before-magnitude | **PASS** (one exception) | §4.3 places each column's null floor (median, 5th/95th, MAD median) to the left of/before every observed quantity in the same table, and §4.4/§4.6 magnitudes follow those in-section floors; exception: the L6 "failure first" opening paragraph and the §4.2 cell-accounting shares carry match-count magnitudes ("6 of 34," "26.3%") before any null is stated — an L6-vs-O3 ordering tension resolved in favor of L6 (the omission there understates, not overstates, the observed signal, since the stated null floors are ≈0). |
| **O5** phrasing | **PASS** | No causal, evaluative, or attributive language toward the publisher found in REPORT/RESULTS/RECON; "fail/fit" is applied only to the arm's own declared policies, divergence vocabulary is the declared cell class, and §6 explicitly disclaims wrongness, process failure, and intent. |
| **O6** phrasing | **PASS** | Correspondence is derived by search and stated as tested, not presupposed (§2: publisher "nowhere states" the cross-artifact derivation; §4.1 join derived with runner-up); §10's six open questions are phrased without presupposing any state of the data (e.g., "What rule generates…" / "was not tested" / "not answerable from this corpus"). |
| **O8** underdetermination | **PASS** | Underdetermination is stated and left standing rather than resolved: §5 "Policies outside that family were not run and their outcome is unknown (O8)"; B2's pin left as a four-row candidate set per P11 instead of being resolved by choosing; the R2 tag/branch observation left explicitly unexamined; §6 states the corpus cannot adjudicate `scores.json` vs sidecar derivations. |
| **O11** adequacy | **PASS** | §6 names all four required minima: what cannot be adjudicated (cause/fault; upstream-archive agreement), what was not run (cross-references §7's unexecuted legs; §9-level cross-arm quantities), what a reader might wrongly infer (the frontend/testing "contain errors" inference, pre-rebutted with the null floor pointer), and correlation-vs-exact-reconstruction explicitly ("closeness… must not be mistaken for exact reconstruction"). |

**P12 shortfall count: 0.**

---

## Cell census as recorded in the arm's deliverables (numbers only)

Grand totals (RESULTS `== CELL TOTALS ==`, both implementations, both process runs identical):
**reconciled 340** (T0 199 · T1 71 · T2 70) · **divergent 102** · **no antecedent 714** · **structural 136** · **total 1,292**.

Per column family (summed from RESULTS `== PER-COLUMN ==`; no-antecedent families per prereg §2.3):

| column family | cols × rows | reconciled | divergent | no antecedent | structural |
|---|---|---|---|---|---|
| score (5 per-category + average) | 6 × 34 = 204 | 147 | 57 | 0 | 0 |
| cost (5 per-category + average) | 6 × 34 = 204 | 159 | 45 | 0 | 0 |
| categorical (`agent_name`, `agent_type`, `language_model`) | 3 × 34 = 102 | 34 | 0 | 0 | 68 |
| count (`categories_completed`) | 1 × 34 = 34 | 0 | 0 | 0 | 34 |
| join key (`id`) | 1 × 34 = 34 | 0 | 0 | 0 | 34 |
| runtime (`average_runtime` + 5 per-category) | 6 × 34 = 204 | 0 | 0 | 204 | 0 |
| URL (5 logs + 5 visualization) | 10 × 34 = 340 | 0 | 0 | 340 | 0 |
| metadata (`sdk_version`, `openness`, `country`, `supports_vision`, `release_date`) | 5 × 34 = 170 | 0 | 0 | 170 | 0 |
| **sum** | **38 × 34 = 1,292** | **340** | **102** | **714** | **136** |

Per-column reconciled/divergent within the addressed numeric families (RESULTS): issue_resolution_score 34/0 · frontend_score 12/22 · greenfield_score 34/0 · testing_score 22/12 · information_gathering_score 32/2 · average_score 13/21 · issue_resolution_cost 32/2 · frontend_cost 16/18 · greenfield_cost 34/0 · testing_cost 31/3 · information_gathering_cost 29/5 · average_cost 17/17 · language_model 34/0.

**REPORT vs RESULTS census:** REPORT §4.2 states 714 / 136 / 340 (T0 199 · T1 71 · T2 70) / 102 of 1,292; RESULTS states 714 / 136 / 340 (199/71/70) / 102 of 1,292 — **match**, and REPORT §4.3's per-column rec/div column matches the RESULTS per-column counts line-for-line for all 13 listed columns.

---

*This check reports coverage and consistency only; nothing here is a statement about any published value (L2). No file under `habitat_arms/iter2/A2-id/` was written or modified.*
