Vertical signs — pass 2: the first shipped recall fix

AI3D-339 · A4_5 corpus (143 segments) + Abschnitt 1 · 2026-08-01 · follows pass 1 (instruments)

TL;DR. Pass 1's instruments paid off: the brightness-percentile relaxation they flagged was adversarially judged (37 real / 19 FP over all 56 census gains), its mechanism isolated, and shipped in a form that keeps the RF-verifier contract intact. A4_5 delineators: 79 → 111, with the 4 lost signs each visually confirmed as false positives. Split-half delineator agreement rose 0.404 → 0.432. Two honest negatives: sign fragility is untouched (by design — sign-side gains judged 3 real / 9 FP), and the win does not transfer to Abschnitt 1 (8 → 8 detections; diagnosis queued).

What shipped (7 commits, all AI3D-339)

Detections by run

The judging: every gain presumed FP until it survived refutation

All 56 census gains of the p95 variant got marked zoom crops (16 m + 50 m, baseline/variant/candidate markers). Two independent refute-framed auditors (Opus, Grok) judged the 33 free-side gains — they agreed on 32/33, unanimous on all 3 refutations; I tiebroke the one disagreement with the crop evidence. A third audit (Opus, with the rail-mounted-reflector hypothesis explicitly allowed) judged the 23 railed/unknown-side gains: only 7 survived.

Judged verdicts
Typical real post
Sustained (seg 062): compact bright speck at the paved edge with an occlusion notch — the signature all 26 lattice posts share. Cluster h_max ≈ 1.08 m.
Gore beacon tiebreak
Tiebreak (seg 134): inside a junction gore. The long occlusion shadow means a real vertical object; the baseline already accepts its twin at the island tip (cyan).
50m lattice
The label-free corroboration: chain 062–072's recovered posts sit on the StVO 50 m Leitpfosten lattice — the 99.7 m gap is exactly one still-missed slot. The baseline had found one post in this half-kilometre. Chain 126–140 interleaves the same way with existing detections.

Mechanism: it was never the seed pass

A seed-only p95 run reproduced just 3/37 judged-real gains — the recall lever is the classification brightness floor, not seeding. The p98 all-points reference is self-referential (one bright guide panel starves every weakly-sampled post head in the segment), but it feeds the frozen verifier, so the shipped form splits the reference: delineator gates read the p95 twin, sign/panel/plate gates and the RF keep p98. The final variant reproduced 29 judged-real gains with 3 judged-FP (sub-noise specks) and zero unjudged surprises.

Lost sign is an FP
All four baseline signs "lost" by the change were inspected: this one (seg 114) sits mid-lane — a motion smear. The others: a carriageway ghost (seg 025, a known QC-FP segment), vegetation texture (seg 038), a gore paint fleck (seg 133). Losing them is a correction, not a cost.

Instruments re-run under the shipped config

pass 1 (p98)pass 2 (shipped)
Delineator agreement 2m/(n₁+n₂)0.4040.432 halves find it too — not a full-run artifact
Delineator recall ceiling0.693 (N̂ 126)0.700 (N̂ 159) more detections against a larger estimated population
Sign agreement0.0000.000 untouched by design; top pass-3 defect
All-type agreement0.3750.410
Cross-half consensus pairs (auto-label seed)2433
Agreement comparison

Chapman caveats from pass 1 still apply: same-detector halves are positively dependent, so these are relative comparisons, not absolute recall.

Abschnitt 1: the win does not transfer

A1 segment 050
A1 branch_000 seg 050: vegetation walls both roadsides — a different world from A4_5's open verges. Zero detections here, before and after.

Before/after on the agreed evaluation slice (branch_000 segs 030–090 minus 65–67, branch_001 segs 002–008; run4-less data handled by the new corridor fallback): before 8 + 2 detections, after 8 + 2. The A4_5 fix moved nothing on A1. Open hypotheses, in order: (1) genuinely sparse roadside furniture on this stretch, (2) different intensity calibration shifting both percentiles together (the memo's range/incidence-normalization point), (3) posts buried in the hedge line where clustering merges them into vegetation. Needs the same crop-level diagnosis A4_5 got — queued as the first A1-specific pass-3 item. The before-baseline is frozen at /mnt/d/ai3d-aml/vsigns_eval/a1_before_20260801 for the final report.

What didn't work / was deliberately not done

Pass-3 queue (ranked, research-memo-informed)

  1. Sign density-fragility: replace the segment-p98 reference for the sign path with range/incidence-normalized intensity residuals + "k independent high-residual hits" (memo §3; no controlled RA1/RA2/RA3 MLS separability exists — measure, don't assume).
  2. Static-FP cues for the judged FP families: multi-radius plate regularity vs foliage (6 vegetation FPs), rail-corridor profile (2 glare FPs), per-pass splitting for the 7 sub-noise specks (memo §1: no literature resolves their physical origin — instrument them).
  3. A1 transfer diagnosis: crop-level audit of A1 middle segments; check intensity distribution vs A4_5 before touching thresholds.
  4. Rail-reflector topology (memo §4): rail-coordinate peak detection as a separate rail_reflector class; the seg 095–097 chain is the seed evidence.
  5. Weak-supervision verifier: 33 consensus pairs + 56 judged verdicts as the seed; LFs with abstention, PU objective, segment-group isolation.

Judge panel: not convened

None of the four stop conditions is plausibly true: improvements without human annotation demonstrably continue (this pass shipped one), no human judgement is currently blocking (the adversarial audits resolved everything they were asked), more data isn't the binding constraint, and the results are visibly not perfect (sign agreement 0.000, A1 non-transfer, one 2×50 m lattice hole). The panel gates stopping, not continuing.

Artifacts: out_eval/recall_sweeps/judged_gains.csv · out_eval/cr2/ · out_eval/research_pass2_memo.md · battlebox /mnt/d/ai3d-aml/vsigns_eval/{seedgate95,cr2_*,a1_*_20260801} · re-baseline out_baseline_20260801. Generated by the AI3D-339 improvement loop, Fable 5 orchestrating.