Vertical signs — pass 2: the first shipped recall fix
What shipped (7 commits, all AI3D-339)

bffa014+f27a33a+7ffc887— the recall fix:seed_bright_percentile=95(new config key) plus a seed-referenced twin of the brightness fraction (hi_intensity_fraction_seed, CSV column 51) that only the delineator gates read. The p98 fraction feeding the locked RF verifier is byte-identical; setting the key tonullrestores the pre-split detector exactly. 241 tests pass.27d20d1+95b57ce— A1 unblocked: segments without run4 road-surface files run with an empty corridor (distance gates disarm instead of crashing), and the resultinginfis clamped before RF scoring.e906559—judged_gains.csv: machine-judged verdicts for all 56 gains — the second auto-label seed batch (first: 24 → now 33 cross-half consensus pairs).5dcf762— pass-2 Codex research memo (see pass-3 queue).- Fresh re-baseline replaces the unreproducible
out_final: current master honestly scores 103 detections (out_final's 127 included verifier-drift accepts).
The judging: every gain presumed FP until it survived refutation
All 56 census gains of the p95 variant got marked zoom crops (16 m + 50 m, baseline/variant/candidate markers). Two independent refute-framed auditors (Opus, Grok) judged the 33 free-side gains — they agreed on 32/33, unanimous on all 3 refutations; I tiebroke the one disagreement with the crop evidence. A third audit (Opus, with the rail-mounted-reflector hypothesis explicitly allowed) judged the 23 railed/unknown-side gains: only 7 survived.

Mechanism: it was never the seed pass
A seed-only p95 run reproduced just 3/37 judged-real gains — the recall lever is the classification brightness floor, not seeding. The p98 all-points reference is self-referential (one bright guide panel starves every weakly-sampled post head in the segment), but it feeds the frozen verifier, so the shipped form splits the reference: delineator gates read the p95 twin, sign/panel/plate gates and the RF keep p98. The final variant reproduced 29 judged-real gains with 3 judged-FP (sub-noise specks) and zero unjudged surprises.
Instruments re-run under the shipped config
| pass 1 (p98) | pass 2 (shipped) | ||
|---|---|---|---|
| Delineator agreement 2m/(n₁+n₂) | 0.404 | 0.432 | halves find it too — not a full-run artifact |
| Delineator recall ceiling | 0.693 (N̂ 126) | 0.700 (N̂ 159) | more detections against a larger estimated population |
| Sign agreement | 0.000 | 0.000 | untouched by design; top pass-3 defect |
| All-type agreement | 0.375 | 0.410 | |
| Cross-half consensus pairs (auto-label seed) | 24 | 33 |

Chapman caveats from pass 1 still apply: same-detector halves are positively dependent, so these are relative comparisons, not absolute recall.
Abschnitt 1: the win does not transfer
Before/after on the agreed evaluation slice (branch_000 segs 030–090 minus 65–67, branch_001 segs
002–008; run4-less data handled by the new corridor fallback):
before 8 + 2 detections, after 8 + 2. The A4_5 fix moved nothing on A1.
Open hypotheses, in order: (1) genuinely sparse roadside furniture on this stretch, (2) different intensity
calibration shifting both percentiles together (the memo's range/incidence-normalization point), (3) posts
buried in the hedge line where clustering merges them into vegetation. Needs the same crop-level diagnosis
A4_5 got — queued as the first A1-specific pass-3 item. The before-baseline is frozen at
/mnt/d/ai3d-aml/vsigns_eval/a1_before_20260801 for the final report.
What didn't work / was deliberately not done
- Sign-side p95 rejected: judged 3 real / 9 FP — extending the relaxation to the sign path would have shipped mostly vegetation texture. Sign gates stay frozen at p98.
- Seed-only relaxation rejected: 3/37 reproduction killed the "it's the seed percentile" hypothesis from the pass-1 report — the sweeps' attribution was wrong, only the census diff plus judging caught it.
- 16 of 23 railed-side gains refuted — including both largest clusters (24k-point merged rail section, 5.8k-point rail glare). Raw census gains remain worthless without side-joins and judging.
- Both big battlebox runs each needed one unblocking fix (missing run4 crash, RF inf crash) — A1 had never actually been runnable with the current detector.
Pass-3 queue (ranked, research-memo-informed)
- Sign density-fragility: replace the segment-p98 reference for the sign path with range/incidence-normalized intensity residuals + "k independent high-residual hits" (memo §3; no controlled RA1/RA2/RA3 MLS separability exists — measure, don't assume).
- Static-FP cues for the judged FP families: multi-radius plate regularity vs foliage (6 vegetation FPs), rail-corridor profile (2 glare FPs), per-pass splitting for the 7 sub-noise specks (memo §1: no literature resolves their physical origin — instrument them).
- A1 transfer diagnosis: crop-level audit of A1 middle segments; check intensity distribution vs A4_5 before touching thresholds.
- Rail-reflector topology (memo §4): rail-coordinate peak detection as a separate
rail_reflectorclass; the seg 095–097 chain is the seed evidence. - Weak-supervision verifier: 33 consensus pairs + 56 judged verdicts as the seed; LFs with abstention, PU objective, segment-group isolation.
Judge panel: not convened
None of the four stop conditions is plausibly true: improvements without human annotation demonstrably continue (this pass shipped one), no human judgement is currently blocking (the adversarial audits resolved everything they were asked), more data isn't the binding constraint, and the results are visibly not perfect (sign agreement 0.000, A1 non-transfer, one 2×50 m lattice hole). The panel gates stopping, not continuing.