Vertical-sign detection — pass 6
AI3D-339 · 2026-08-03 · detect-at-predicted-site, pool verification, field-stake trap, first stop-condition panel
TL;DR — Pass 6 built a raw-point post-signature probe (compact saturated bright column + stem, searched inside small windows) and used it two ways: at lattice-predicted empty sites, and at every rejected "pool" cluster. It found and admitted 19 judge-validated devices (A1 65→81, A4_5 129→132), including 5 posts that never formed clusters at all. It also walked into and out of the best trap so far: a 5.8 m-spaced "delineator row" that is actually agricultural stakes in a ploughed field — leading to a road-context feature, three impeachments of prior "judged-real" labels, and hard evidence that regular spacing means nothing without a carriageway next to it. The first stop-condition panel convened: CONTINUE, 3–1 (Fable dissenting for STOP-2), with a convergent, bounded pass-7 queue and a unanimous recommendation to run a ~20-case human calibration pass in parallel.
1 · Two-stage raw-evidence verification
The pass-6 workhorse: for any candidate position, search an r<1.8 m window of the raw multi-record point cloud for a retroreflector signature — ≥15 points at I≥47k, from ≥2 records, xy-RMS ≤0.15 m, column reaching z≥0.9 m, with a stem beneath. Two earlier probe versions failed honestly (component clustering drowns in verge grass; a 38k "bright" threshold is meaningless in log-quantised intensity). Then a road-context stage: count saturated returns (I≥55k — lane markings, reflectors) within 15 m.
Null discipline held: phase-shifted controls 0/13, offset controls 1/144 (A1) and 5/138 (A4_5). The funnel's last stage is the point: the signature generates candidates; adversarial judges still make the admission call (precision 12–40% before judging).
2 · Detect-at-predicted-site: 5 posts that never clustered
Lattice chains regrown on the frozen pass-5 output predicted 30 unoccupied sites (plus 18 controls). Probing them recovered all three pass-5 raw-occupancy-confirmed posts and new finds — including proof that chain c2's spacing was double-counted (a "null control" at half-phase contained a textbook post).
A judged-real site detection (c1_x23_1): slender post with compact bright head, standing at a chain-extrapolated position where the detector never formed a cluster. Right panels: side views, ground at bottom.
On the anchor-rich calibrated stretches (recall's best case): base ~46% → lattice ~77% → site detection ~92%. Corridor-wide numbers remain unknown — that caveat stands in every report since pass 5.
3 · The field-stake trap (and what it broke)
A 14-candidate "curve delineator row" in A1 segs 72–73 — 5.8 m regular spacing, clean vertical strokes, bright 48–53k tops — was judged 10/10 real by Grok off the regular spacing. Opus refuted it by measurement: heights 0.6–1.6 m non-uniform, and 0.00–0.02% saturated pixels in every 25 m context — no carriageway. Fable's tiebreak view: unmistakable ploughed-field furrows. It is an agricultural stake row.
The 50 m context of "delineator" s072c19: furrow stripes, no road. Regular spacing is agricultural as often as regulatory.
- 3 prior "judged-real misses" impeached (nm02, nm05, nm42) — they sit in field rows with zero road context. The pass-3 truth stack decays when crops lack context.
- Road-context feature born: saturated-return density within 15 m separates field from roadside perfectly (stakes: exactly 0; roadside: hundreds–thousands). It also killed both signature-passing known FPs (nm61, ag14).
- Its limit is equally sharp: it cannot separate roadside vegetation from posts — that residual discrimination is visual.
- A4_5's version of the trap: young-tree plantation rows with support stakes at 4–5 m (7 of 24 sampled candidates).
4 · Refutations (kept, with controls)
| Hypothesis | Result |
| Guarded lattice fixpoint reaches the anchor-poor stretches (FP-purged anchors, teacher quorum, per-round nulls) | 1 admission, 0/12 misses reached, nulls clean. Chains cannot be willed into anchor-poor corridor. |
| RF on scalar probe + context features as auto-verifier | Roadside-only LOO AUC 0.73 — identical ceiling to pass-4's cluster-feature RF. Overall 0.877 is illusory (field items are trivially separable). |
| Saturation rule (imax ≥ 54k) as universal mechanical admitter (~85% precision on A1) | Fails corridor transfer: A4_5 runs hotter — not-real items reach 62k, real ones sit at 51.6k. |
5 · Admissions shipped (judge-validated overlay)
tools/apply_judged_admissions.py + out_eval/judged_admissions_pass6.csv (full per-row provenance) applied to the frozen pass-5 baselines → p6_final_a1 / p6_final_a45 on battlebox. Idempotent; every row is Opus+Grok adversarially judged with Fable tiebreaks.
| Corridor | Before | After | What was added |
| A1 | 65 | 81 | 5 site detections (never clustered, incl. e15 relocalized), 9 pool readmissions (2 ramp posts, 2 gore-apron posts, 1 rail-mounted device, 1 Aufsatz half-post…), 2 confirmed pass-3 misses (nm56, nm29) |
| A4_5 | 129 | 132 | 2 free-standing posts + 1 rail-mounted Aufsatzleitpfosten (from a 24-crop judged sample of 52 candidates) |
A4_5 s115c42: rail-mounted Aufsatzleitpfosten — W-beam cross-section with a slim bright-tipped element rising ~0.5 m above the rail. A device class the freestanding shape gates can never catch (pass-6 research memo, HLB §5.41).
6 · First stop-condition panel: CONTINUE, 3–1
Grok, Sol, Opus: CONTINUE. Convergent reasoning: validated-but-unshipped moves remain (field-stake mask with a 0/144 null, A4_5 corridor-wide site probing, snap-to-device relocalization, Chao2 recall reporting), and 2 of the 4 goal classes (gates, trees) have zero quality evidence — "you cannot evaluate a stop condition over a goal half of which is unmeasured" (Opus).
Fable dissent: STOP-2 — the truth stack underpinning every number is LLM-judged with zero human calibration and measured decay (3 impeachments); the next decisive step is human.
Unanimous secondary: run the ~20-case human calibration pass (web point-cloud viewer) in parallel, non-blocking. Opus set the reconvene criterion: judge-independent moves landed + Chao2 reported for both corridors + gates/trees audited or de-scoped + the 34-segment holdout burned once.
7 · Pass-7 queue (panel-bounded)
- Ship the field-stake / road-context mask into the pipeline (0/144 controls; also emits
field_stake_row as a separate experimental class — Miro's ask).
- A4_5 corridor-wide site probing (chains exist; never probed — its never-clustered count is unknown).
- Snap-to-device relocalization: ≥3 measured cases where the real device sits 1–1.5 m from a rejected candidate (weed in front of a rail-mounted post). Mechanical, judge-free.
- Chao2 capture–recapture recall from 7–16 per-record detections, both corridors (dependence-aware; never Lincoln–Petersen).
- Gates & trees: audit or de-scope — 3 gantry_or_gate on A1 never audited; trees experimental.
- Parallel, non-blocking: human ~20-case calibration — viewer + curated case list (impeached misses, disputed roadside candidates, half-posts, fp_ag14).
- Then: burn the holdout once, reconvene the panel.
Baselines: p6_final_a1 (81) / p6_final_a45 (132) on battlebox · holdout intact · all instruments in out_eval/pass6/ · reports: pass 5