Vertical-sign detection, pass 1: measure before you tune
record_coverage_ratio) that flags 5/6 transient FPs at an 8.7% base rate;
a census showing 3 thresholds block 433 near-miss clusters (the pass-2 triage
pool); and two clean discards — the plate-orientation cue (measured dead) and 3 of 4 seed-gate
relaxations. One infrastructure finding: out_final is no longer reproducible from
master (the shipped RF verifier post-dates it) and must be re-baselined before the next census.
No detection quality was changed this pass — deliberately. Every default stays;
pass 2 spends the new instruments on actual improvements.What went well
- Instruments over knobs. All four deliverables are measurement tools; every one produced a finding that redirects pass 2 away from something we would otherwise have wasted effort on.
- The additive-only discipline held. The feature work ships with a byte-identical regression gate on all 44 pre-existing CSV columns (verified twice, including after deletion cleanup); 237 unit tests pass.
- Skeptical-by-default worked. Two ideas died by measurement, not opinion: the plate-normal orientation cue and the seed-span/min-points relaxations. Cheap deaths — each cost hours, not a shipped regression.
What didn't
- Handoff item 4's premise was wrong. The plate-normal-vs-road-tangent scalar assumed verge strips / road paint / guardrail tops are extended planes. Measured: they are metre-scale blobs (planarity ratio 1.2–1.8); the scalar's distribution is identical whether the normal is well-defined or noise (medians 0.181/0.182/0.180 across planarity strata). deleted
- The recall number is softer than hoped. Split-half capture–recapture with one detector is a capture-overlap completeness index, not unbiased recall (positive dependence pushes it optimistic; half-data thinning pushes the other way). Literature check (Chao et al.) confirms. Reported as ceilings, useful mainly for relative comparisons.
- Baseline drift found late.
out_finalpredates the shipped RF verifier; master vetoes 3 of its accepts (all QC-confirmed FPs — the verifier is right), and a fresh default sweep drifts further (79/16 vs 88/30 delineators/signs). Every diff-vs-out_final this pass carries that caveat; re-baseline queued before pass 2's censuses. - Delegate wrangling cost real time. Two runs died by premature backgrounding and one timed out with zero output before escalation to a stronger model succeeded.
Threshold-margin census (item 3)
min_accept_h_max_m=0.9is the workhorse gate: binding constraint for 55 of 127 reconstructible accepts.- Three thresholds block huge near-miss pools while binding almost no accepts:
panel_min_hi(227),gantry_len_major_m(106),plate_hi_intensity_fraction_weak(100). These 433 one-conjunct rejects are the pass-2 auto-annotation triage pool — render crops, judge with vision agents, find whether any real signs are single-threshold-blocked. - 18 config thresholds are fully inert on this corpus; ~9 conjunct families (row rejection, gantry pairing, ML veto…) are uncoverable from CSVs and stay unmeasured.
Record-multiplicity + planarity features (items 2, 4)
- shipped Six appended
clusters.csvcolumns (additive-only, uncommitted → now committed): 4 multiplicity fields +plate_planarity_ratio.record_coverage_ratio(passes-that-saw-it / passes-that-could-have) flags 5/6 judged transient FPs at a 8.7% base rate, neutralizes the one-record-segment confound, and correctly ignores static clutter. - not a rule yet 1 of 4 judged reals is also flagged (a confirmed sign seen by 1 of 12 passes). Next measurement: coverage-ratio distribution over the full 143-segment accepted set, transient-family candidates only.
- Taxonomy result: multiplicity is a transient detector; the static FP family (paint/verge/rail-top) is misclassification and needs shape/context cues — the two families need different weapons.
- deleted
plate_normal_costan+ the corridor tangent estimator — measured dead (see "What didn't").
Label-free recall instrument (item 1)
- Overall ceiling 0.735 [0.665, 0.797]; delineators 0.693; signs 0.491.
- The sharp finding: sign agreement = 0.000 — halves found 7 and 6 signs, none matching (27 in the full run). Sign detection (plate evidence) collapses under point thinning. This is a robustness defect measurable and fixable without labels.
- 24 cross-half consensus detections dumped — the first auto-label seed. 6 of 123 full-run detections are single-record (fragile tail, first QC targets).
- Instrument upgrade queued (research-informed): per-pass independent redetection with capture histories instead of half-splitting; report as completeness index with dependence warning.
Seed-gate sweeps, FP-first framing (round-4 suspects)
| Variant | Gains (free/railed side) | Verdict |
|---|---|---|
| brightness p98→p95 | 33 / 19 | judge in pass 2 — the load-bearing knob; free-side gains are well-sampled (median 995 pts) |
| span 0.80→0.65 | 3 / 0 | discard — negligible |
| min_points 300→150 | 2 / 2 | discard — admits exactly the under-300 reject class |
| combo | 39 / 26 | discard — brightness + min-points bleed-in |
Framing per the retracted recall gap: census gains are FP candidates until an overlay says otherwise; gains on guardrailed sides are near-certain FPs. 50 intensity overlays fetched for the judging pass (agent-first, then Miro only if ambiguous).
Research step (Codex, web)
Full memo: out_eval/research_pass1_memo.md (5 questions, ~30
sources). Highest-value takeaways: (1) per-pass independent redetection + capture-history
vector is the literature-standard form of our coverage idea (You et al. NeurIPS'22 persistence
prior, Barnes/Maddern self-supervised distraction removal); (2) the static-FP family wants
height-sliced footprint topology, road-station continuity, and surface-conformality cues — all
label-free; (3) our Chapman use needed the dependence caveat it now carries; (4) the auto-label
path is abstaining labeling-functions → probabilistic label model → PU-learned verifier with
road-section isolation; (5) German RPS/VwV priors are quantifiable: 50 m Leitpfosten lattice
0.5 m off the paved edge at ~1 m height, sign lower edge ~2 m at ~1.5 m lateral, gantry
clearance ~4.5 m — usable as soft likelihoods, not hard gates.
Pass 2 queue (ranked)
- Re-baseline
out_finalwith current master (verifier on) — prerequisite for every future census. Then freeze the A1 "before" baseline (branches 000+001 middle segments, excl. 65–67) for Miro's end-of-goal before/after report. - Vision-judge the pass-1 candidate pools (agent-first): 33 free-side brightness gains + top near-miss families (227 panel_min_hi / 100 plate_hi_weak). Output: auto-labels + a keep/kill verdict on the p95 brightness relaxation.
- Static-FP cues: height-sliced footprint topology + road-station continuity + surface conformality (additive features, then rules or LFs).
- Sign density-robustness fix: diagnose why plate evidence dies at half density; likely make plate gates density-aware (ratio-based, not count-based).
- Weak-supervision verifier (Miro-approved ML without human labels): rules-as-labeling-functions with abstention + consensus/coverage seeds → PU verifier; strict segment-group isolation.
- Instrument v2: per-pass redetection capture histories replacing half-split.
Judge-panel status: not convened this pass — no stop condition is plausibly true: improvements clearly remain that need neither human annotations (weak supervision path is open), nor new data (A1/A4_5 unexhausted), and the results are far from perfect (sign fragility, unjudged FP pools). The panel gates stopping, not continuing; it will be convened when a pass ends with a plausible stop claim.
Artifacts
- Reports:
out_eval/{threshold_census,features_pass1,capture_recapture,recall_sweeps}/report.md, memoout_eval/research_pass1_memo.md, stateout_eval/SESSION_STATE.md - Code: 6 appended CSV columns in
src/.../detect.py+features.py(+49 tests),tools/threshold_margin_census.py,tools/recall_capture_recapture.py - Battlebox: sweeps under
/mnt/d/ai3d-aml/vsigns_eval/; baseline~/dev/3dai.iolabs.pointcloud.verticalsigns/out_final(drift caveat) - Sibling page: veg-overhang zoom walkthrough (workstream A, same day)