Vertical-sign detection — pass 7

AI3D-339 · 2026-08-05 · the honest-measurement pass: human truth wave, A4_5 site probing, Chao2 recall, first holdout burn
TL;DR — Pass 7 stopped adding machinery and started measuring. The stop-condition panel returned CONTINUE, 4–0 unanimous (pass-6's Fable dissent withdrawn: human sessions turned out cheap and in-loop). Baselines are fully reconciled for the first time — A1 75 devices, A4_5 141 — and the old "81 / 132" bookkeeping was wrong in both directions for reasons inherited from pass 6, not from pass-7 judgement. Two human sessions rewrote the truth stack: e6 admitted, 8 items pulled (including all three A1 gantries, which are embankment canopy), and a failed calibration control overturned into a retroactive 4/4 for the human. A4_5 chain probing found 9 new devices at 155 sites with 0 of 53 controls firing. And the 34-segment holdout was burned once, under a protocol written before the burn: BEFORE 22 → AFTER 33 devices, 9 judged-real gains against 4 vegetation FP leaks. That is 1 FP per 2.25 gains against a pre-registered bar of 1 per 3 — PRIMARY failed, by a single verdict; SECONDARY passed (the road-context gate removed nothing real). Chao2 puts honest A1 corridor recall at an optimistic ceiling of 0.31–0.56, far below the 0.92 measured on anchor-rich stretches.
1 Counts reconciled · 2 The holdout burn · 3 Human truth wave · 4 A4_5 site probing · 5 Honest recall · 6 Road-context mask · 7 Veg-rescreen exhausted · 8 A regression · 9 Panel & pass-8 queue

1 · Counts reconciled — A1 75, A4_5 141

Every prior report quoted a device count that nobody could reproduce from the trees. Pass 7 fixed the counting rule first: a device is an accepted row in segment_XXX/clusters.csv, plus any verticalsigns.json detection with no accepted cluster within 0.5 m XY. Neither half alone works — gantry_or_gate clusters emit no JSON detection at all (5 across the corridors), and site admissions never form a cluster, so they live only in the JSON.

corridorCSV-accepted by typeCSVJSON-onlytotal
A1 p7_final_a1delineator 55, sign 12, pole_other 269delineator 675
A4_5 p7_final_a45delineator 118, pole_other 7, sign 4, gantry_or_gate 2131delineator 7, rail_half_post 2, gantry_or_gate 1141
count reconciliation

Row by row, nothing in the pass-7 CSVs behaved unexpectedly: 3 A1 admissions, 9 A4_5 site admissions, 8 A1 pulls, 1 A4_5 pull — all applied, none skipped, none colliding. The entire A1 −1 / A4_5 +1 drift against the ledger's projection sits in the inherited pass-6 base: a site admission (c2_i17_5x, b0 s036) that is 0.043 m from an already-accepted cluster — the same physical post admitted twice — and a JSON-only gantry detection in A4_5 segment_066 with no cluster behind it.

The build is a four-step idempotent chain over the pristine mask trees (p6 admissions → p7 admissions → s080c5 retype → removals), and the composite is byte-stable: re-running the whole chain leaves both trees md5-identical. Removals must run last — re-applying the p6 admissions after them re-admits s051c47, which pass 7 pulled. The steps are order-independent only as a closed chain.

Two bookkeeping traps worth carrying forward. judged_removals_pass7.csv row 3 (fp_nm61) declares type_was=delineator but the cluster is really a sign; the pull is keyed on cluster_id so it is correct, but the A1 histogram moves −1 sign. Combined with the s080c5 retype (+1 sign) the sign count coincidentally stays at 12 — do not read that stability as "nothing happened". Second: gantry_or_gate is now zero on A1, all three pulled as canopy.

2 · The holdout burn — first honest test in seven passes

34 A1 segments (b0 000–029 + b1 000/001/009/010) have been frozen since pass 1. Pass 7 burned them once, against a protocol committed before any holdout segment was read: exactly two configs, no re-runs after seeing results, sparse crops (<20k points in r=20 m) flagged UNJUDGEABLE rather than voted on, and success criteria declared in advance.

holdout scorecard

Scored on 32 of 34 segments (2 lost to the OOM regression in §8; BEFORE found 0 devices in both, so nothing was lost from the comparison). 20 devices agreed, 2 BEFORE-only, 13 AFTER-only — every AFTER-only item a delineator, both BEFORE-only items a sign. No item tripped the sparsity rule: the sparsest crop held 1.36 M points against a 20 000 floor.

Judging was the usual adversarial fleet — Opus and Grok independent on all 16 items, Fable tiebreaking 4 disputes by actually viewing the crops. All 4 tiebreaks went to vegetation: Opus's width/saturation calibration was upheld and Grok's roadside-position credulity bit again, exactly as in pass 6.

real delineator found only by AFTER
b0s025c90 — a real delineator the BEFORE detector missed. Compact saturated head at 0.7–0.9 m on a 0.10–0.16 m column, standing on the verge just past the edge line. 7 of the 9 real gains share this identical signature.
vegetation false positive leaked by AFTER
b0s007c14 — one of the 4 leaked FPs. Same slender roadside vertical, same position class, but no saturated head: the elevation panel shows a wispy stem, with a genuine banded post visible further right for contrast.
moving vehicle rejected by AFTER

Both BEFORE-only "signs" were non-devices, so BEFORE's precision on its unique finds is 0/2. This one, b0s011c88, is a moving vehicle on the carriageway — smeared single-record returns down the lane, an angular mass floating above the surface. The pass-7 rejections were correct.

PRIMARY: FAILED. AFTER adds 9 judged-real devices at a cost of 4 judged FPs = 1 FP per 2.25 gains, against a pre-registered bar of 1 per 3. Judged precision 0.909 → 0.879: AFTER does not strictly dominate. No post-hoc reinterpretation was allowed, and none was applied.

SECONDARY: PASSED. The road-context gate fired once on the holdout and removed a vegetation clump; zero real devices removed. The presence-gate transfer claim holds.

The reading the panel got: the pass-2..7 recall machinery generalizes — 9 real posts that BEFORE never saw, a +45% device count on unseen ground — and the FP direction is now the binding constraint. All four leaks are the same family: slender roadside vertical without a saturated head, the family the judged-admission bar catches on eval slices but the raw detector does not. A candidate rule follows immediately (not shipped, not tuned on the holdout): require a saturated/compact bright head for solo non-rail delineators. It would have killed all 4 FPs at the cost of one real find (s012c188, whose head never saturates) — net 8 gains, 0 FPs, which would have passed PRIMARY. It has to survive label-free eval-slice validation and null controls before it earns a reserve burn.

Also on record: the b1 ramp contributed 0 devices on both sides, so its distinct intensity regime (saturation counts ~100× lower than mainline) never came into play. The holdout is now consumed — no re-reads for scoring. The 46-segment reserve (b0 091–136) is untouched and consumable by pre-registration only.

3 · The human truth wave

Two sessions, both cheap. The first (2026-08-03, 24 cases, intensity-only palette, r=20 m crops) doubled as a blind calibration; the second (2026-08-05) was a re-look built after Miro's feedback — camera RGB + 60 m road context, one-click CloudCompare, ~15 min for 18 cases.

e6 with camera RGB and 55 m road context

e6 in the re-look package: 55 m of road context in camera RGB on the left, 7 m close-up on the right. Both LLM judges had downgraded it to vegetation at high confidence in pass 6. The human called it a real delineator, HIGH, twice — first at r=20 m, again with RGB and context. It is now admitted. RGB is the reason it is judgeable at all: "very hard to tell with no RGB, just intensity."

The control that overturned

Four hidden controls were seeded into the calibration set. The human scored 3/4 — the "failure" being ctrlB (s051c47), called vegetation at medium confidence against an assumed-true pass-6 pool admission. The re-look with RGB and 60 m context settled it at HIGH: it is vegetation. The LLM admission label was wrong, not the human. Controls are retroactively 4/4, the item was pulled from the baseline, and human authority moved up accordingly.

Sol's caveat, on record: a retroactive 4/4 obtained by re-adjudicating the one failure is not independent calibration. It is evidence, not a clean score.

What the humans changed

directionitems
Admitted (+3 A1)e6 (real HIGH ×2, overturns the pass-6 LLM downgrade), i11 and s053c57 (both real HIGH — both overturning a Fable vegetation/"gore tuft" vote; Grok was right on i11)
Pulled (−8 A1, −1 A4_5)s051c47 (the ctrlB overturn), nm61, ag16 (fence), ag26, ag27, all three A1 gantries (s086c546, s087c113, s087c191), s078c59 (A4_5)
Kept, retypeds080c5: human says "positively a device, but not a half-post — a sign" → admission survives its pull-unless-rescued clause, type changed delineator → sign
Kept acceptedag11 (human "real sign", LOW — no pull authority, watched), ag14 (human unsure twice — stays documented as an accepted FP)
Miss ledgernm28 impeached (veg HIGH); nm02/05/42 field-stake impeachments now FINAL; s091c42 demoted from mislocated-miss to admission candidate (sign, MED — insufficient alone); s059c30 stays listed, snap-to-device stays parked

The gantry pulls matter beyond the arithmetic. All three had been flagged by a report render two days earlier — 4.7–5.7 m tall, 8–9 m long masses on an embankment 20–28 m lateral and 6.5–8.7 m below the lane markings, with no overhead span anywhere. The human confirmed all three as canopy. The gantry_or_gate class on A1 is now zero-confirmed and de-scoped; the three A4_5 gantry items remain unaudited — Sol flagged that gap as the reason the panel dossier's "all reconvene conditions met" was too strong.

Two standing rules came out of this

The measured bias the humans exposed: the LLM stack over-calls vegetation on ambiguous roadside verticals — four overturns now (e6, i11, s053c57, and ctrlB in the other direction). That is precisely why Opus flagged the holdout's 4 "vegetation FP" labels as stack-generated and put human adjudication of them at the top of the pass-8 queue.

4 · A4_5 site probing — 155 sites, 0/53 controls

A1 was depth-exhausted after pass 6, but A4_5's lattice chains had never been probed. Regrowing them corridor-wide produced 155 sites: 52 interior gaps, 50 extrapolations, and 53 phase-shifted controls — controls interleaved so the probe cannot know which is which. The pass-6 raw-point signature probe was run at every one, with the road-context prefilter in front.

The probe fired at 10 sites and none of the 53 controls. One pair was a duplicate of the same physical device found from two chains, admitted once → 9 admissions, every one judged by Opus + Grok (one Fable tiebreak on c2_i134_1, where the type is ambiguous — striped panel plus banded post pair — but the presence of a device is not). A4_5 has no camera RGB, so the admission bar there stays higher than on A1 by construction.
rail-mounted half post on A4_5

c1_x126_1 — one of two new rail_half_post devices. The elevation panel shows the W-beam guardrail cross-section with a 0.44 m bright-tipped stub rising above the rail top, saturating at 65535. This is the Aufsatzleitpfosten class: ~91% of A4_5 is railed, freestanding shape gates can never catch these, and there is no published MLS detector for them. A rail-relative search — extract the rail, then look for short vertical attachments in rail coordinates — is the last big recall lever, researched in pass 7 and still unbuilt.

5 · Honest recall — and the caveat stack that comes with it

Recall has been the unmeasured defect since pass 1. Pass 7 finally measured it, and the honest answer is uncomfortable. Design: an occasion is one drive (one run3 record), never a spatial block. For every segment the detector was re-run once per record with only that record feeding candidate accumulation, while the ground DEM and road corridor were still built from the full record list — so an occasion measures a single pass's detection ability, not its context availability. 1220 single-drive detector runs across 166 jobs, all exit 0. Segments are used only as bootstrap resampling units.

Chao2 recall ceilings
Chao2 lower-bounds N, therefore Sobs/N̂ is an OPTIMISTIC recall figure — never a lower bound. Every estimator that is less conservative about heterogeneity puts N higher and recall lower (A1 spans N = 48–124 across Chao2-bc, Chao2-classic, jackknife-1/2 and M0). Worse, detectability decays with distance from the road, so the estimator is structurally blind to the worst cases: a device with p≈0 in all T drives contributes nothing to Q1 or Q2, and no closed-population estimator can learn a class it never samples. Direct evidence of that class exists here — 51 of 78 A1 baseline devices were never re-found by any single drive.
corridorTS_obscoverageN̂ (Chao2-bc)recall ceilingbootstrap 95% CI
A1 (eval slice, 65 segs)8270.2087.40.31[0.11, 0.43]
A4_5 (70 of 141 segs)12370.8641.00.90[0.13, 1.00]

The caveat stack is recorded in full, because the number is only as honest as its caveats: the single-drive protocol is a degraded survey, not a replicate (min_points=300 is tuned for the fused cloud, so clusters die on point count — per-drive capture probability is ≈0.05/0.09); novel single-drive detections are mostly artifacts and are excluded from the primary population (34 of 41 A1 novels sit on clusters the full run explicitly rejected, and their type mix is wrong — 20 "signs" among 41, against 12 in the entire A1 baseline); the single_record_transient veto is structurally armed in every occasion (it keys on n_records_present ≤ 1, true for every one-drive cluster — it biases against capture, so it cannot have inflated these figures); occasion labels are per-segment, so no time-effect model is identifiable; effort is unequal (T spans 5–10 and 1–14); and positive dependence between drives remains, which inflates overlap and makes recall optimistic on top of Chao2's own lower-bound property.

A methodological discovery worth more than the numbers: the earlier capture_recapture/ and cr2/ instruments were split-half Chapman estimators, and they are now superseded. Capture and recapture there are strongly dependent (same sensor, same rules, same geometry) — the exact case the research memo rules out. Worse, what they were actually measuring was spatial coverage overlap, not detection capture: a 0.93 coverage figure would have manufactured a "99% recall" claim. Coverage is not capture.

6 · The road-context mask, shipped — and honestly a no-op

Pass 6's field-stake trap produced the road-context feature; pass 7 shipped it (commits 989c0e0 + b4ee4cc). Three design decisions, all deliberate:

Validation: 209/209 accepted devices across both corridors survive the gate. The mask re-run reproduced the pass-6 accepted-cluster set exactly on both corridors — which is the honest finding: on current detector output the mask is a no-op. Its value is insurance on the admission paths, where the field-stake family actually enters, not a scoring win. Claiming otherwise would be inventing credit.

road-context gate rejection on the holdout

The gate's single holdout firing — b0 s026 c43, and the SECONDARY criterion's whole evidence base. roadctx_n_sat = 0 within 15 m; the candidate (circled) sits ~20 m off the carriageway among vegetation stripes, road at upper left. Critically the domain guard behaved correctly rather than being bypassed: the wider neighbourhood holds 52 812 saturated returns, far above the 1000 floor, so the gate was legitimately armed and the local zero is a genuine reading, not coverage noise — exactly the distinction b4ee4cc was written to draw. Both judges and the geometry agree: vegetation.

7 · Veg-rescreen: the pool is exhausted

Since the humans proved the LLM stack over-calls vegetation, the obvious move was to re-screen pass-6's judged-vegetation rejections for hidden e6-class reals — specifically those that nonetheless passed the raw signature gate. 14 candidates qualified; the road-context prefilter dropped 5 immediately (zero saturated returns inside saturation-rich segments, including a b1 s008 trio that is a field row); a fresh refute-framed Opus+Grok fleet took the remaining 9.

8 of 9 confirmed vegetation on geometry — measured 8–14 m beyond the pavement edge. The single "rescue", pool_s037c72, turned out to have been already admitted in pass 6: the selection query keyed on any vegetation vote rather than the final stack verdict. A no-op, though the fresh fleet re-confirmed it unanimously at HIGH and read it as a rail-mounted stub (a rail_half_post retype candidate). Conclusion: e6-class hidden reals in already-judged rejections are exhausted. A1 stays at 75. This direction is closed and should not be re-proposed.

8 · A regression the burn found

crown_circles raises a MemoryError (radius.py via features.py:899) that kills b0 segments 015 and 016 — ~6% of A1 — even run sequentially at 85 GB. It surfaced only because the holdout burn ran the current config over segments nobody had touched since pass 1. It is unfixed. Both dead segments had 0 BEFORE devices so the burn's scoring was unaffected, but a detector that crashes on 6% of a corridor cannot be declared finished — Sol made it a mandatory hygiene item, and it is judge-independent work.

9 · Panel: CONTINUE, 4–0 — and the pass-8 queue

Fable's seat was sealed first (commit 43d7505, before any other seat had read the dossier) to keep the panel from anchoring on the strongest voice. It came back CONTINUE, dissolving the pass-6 STOP-2 dissent on its own terms: the objection was that the truth stack had no human calibration, and human sessions turned out cheap, repeatable and in-loop.

seatverdictconfcore argument
Fable (sealed first)CONTINUEmedstop-2 dissolved by proven cheap human sessions; the holdout exposed a concrete testable defect with clean reserve data left to test it on
GrokCONTINUEmed-highstop-1 premature — label-free moves remain; the saturated-head rule must pass eval-slice validation + null controls before it touches the reserve
OpusCONTINUEhighthe PRIMARY failure is one verdict wide and its 4 FP labels come from the measured veg-biased stack; 46 reserve segments unconsumed; the detector still crashes on ~6% of A1
SolCONTINUEmedPRIMARY genuinely failed; the saturated-head rule is hypothesis material, not shipping evidence; the crown_circles regression is a judge-independent next step. Flags dossier overreach: A4_5 gantries unaudited, and retro-4/4 is not independent calibration

Pass-8 queue (panel-directed)

Then reconvene. Directions now formally closed and not to be re-proposed: Chapman capture–recapture, the imax≥54k rule (transfer-fails), zbin_count_cv, the scalar RF verifier (AUC ceiling 0.73), the guarded lattice fixpoint, spacing-as-evidence, saturation counts as radiometric anchors, veg-rescreen of judged rejections, and A1 chain-site probing.

Baselines: p7_final_a1 (75) / p7_final_a45 (141) on battlebox · holdout consumed, reserve b0 091–136 intact · instruments in out_eval/pass7/ · zero human training labels · reports: pass 6 · pass 5