Vertical signs — pass 4: the vegetation-texture veto
The rule that shipped
Veto a delineator or sign verdict when plate_thickness_m ≥ 0.05 and
hi_intensity_fraction_seed ≥ 0.668. Physically: a bush or a plastic shelter tube is a
thick blob that is uniformly bright at the seed threshold, while every judged-real marker
is either a thin plate (protected by the thickness conjunct no matter how bright — a plate is
supposed to be all-bright) or a bright-head-dark-post delineator whose seed fraction stays
moderate. Both cuts sit mid-gap on the 118 judged pass-3 clusters; the binding real (ag12,
diamond+rectangle plates on a mast) sits at fraction 0.650 vs the 0.668 cut.
The transfer failure the validation loop caught
The pass began with a new instrumented feature: zbin_count_cv, the coefficient of
variation of per-height-bin point counts (a marker concentrates mass at its plate; foliage fills
every bin evenly). On the offline instrument it was the best vegetation cue ever measured here
(AUC 0.854). Implemented in the pipeline and validated by a full re-run, it
collapsed to AUC 0.542 — and the first-cut rule would have cost 3 judged reals (ag02, ag12,
nm47).
Two lessons now in the working rules: thresholds are only ever derived on pipeline-computed
values, and every candidate rule gets a full-rerun diff before it counts as shipped. The feature
stays in clusters.csv as a diagnostic; the context-flavoured version of the idea
(ring occupancy) is queued properly as a neighbourhood feature.
All seven removals, each with a verdict
The one removal that had never been judged (an accepted "sign" in A1 segment 049) got the full adversarial treatment: crops rendered, Opus and Grok independently asked to refute the sign hypothesis. Both called it vegetation — a ~2.5 m uniformly bright cylinder with no plate flare and no head band, at the ragged edge of a canopy bank: a retroreflective tree-growth shelter tube (Grok high confidence, Opus medium).
Detection state after pass 4: A4_5 127 accepted (114 delineators, 4 signs — all four judged real, sign precision 33% → 67% → 100% of judged across passes 2→3→4), A1 57 (delineators ~40, signs 13). Remaining known FPs: the two dim scrub bands (too dark for this veto), two thin objects (yard cone, plant stem — the Codex memo confirms thin bare stems are a genuinely unresolved single-view ambiguity), and two other-manmade disputeds.
Lattice instrument: supporting prior only
Spacing-to-nearest-accepted-delineator separates in median (judged real 30 m vs judged veg 77 m) but overlaps badly — scrub bands sit on lattice positions too, and after fixing a self-match bug the collinearity cue mostly evaporated. Verdict: the lattice is never a standalone admit rule; it becomes useful as a phase-locked prior combined with texture gates — which is exactly what the research memo's HLB table enables (below).
Codex research memo (Q1–Q4)
- Thin veg (Q1): usable only as weak veto; bare stems vs posts is an unresolved single-view ambiguity in the MLS literature — stop trying to solve it with cluster geometry, use lattice/RF/cross-pass evidence.
- Weak-supervision RF refresh (Q2): full recipe — confidence-binned soft labels with Laplace shrinkage, grouped cross-fitting, Platt (not isotonic) calibration, kNN source-support gate, conformal veto-only abstention, and the selective-labels caveat: our labels estimate P(real | detector fired), never recall.
- Lattice prior (Q3): StVO Zeichen 620 = 50 m mainline; the HLB/BMV-1992 radius-dependent curve table (3 m at R=20 up to 50 m at R≥600) with transition spacings; no sourced ramp/gore lattice — the prior must switch off at topology breaks; rail-mounted reflectors are a separate modelling class.
- Tree shelters (Q4): plane-vs-cylinder RANSAC residual score on the central 80% height + 5 cm slice features (line-vs-circle residuals, radius slope, azimuth span); Tubex shelters are 73–105 mm tubes at 1.2 m — same size regime as posts, geometry of curvature is the discriminator.
Full memo: out_eval/research_pass4_memo.md (committed).
Fresh-eyes Fable strategist (new per Miro)
A context-free Fable subagent read only the session state + memo and judged the loop itself. Its verdict, condensed:
- Chase recall, not precision. 6 residual FPs at ~97% judged precision is diminishing returns; ~12 judged-real misses on A1 plus unmeasured recall is the real defect.
- Recall has never been measured against anything external — every metric so far conditions on the detector firing. Proposed fix: a blind lattice-site occupancy audit — predict post sites from confident lattice stretches, probe raw-point occupancy at predicted-but-undetected sites. Label-free recall measurement + free hard examples for the RF.
- The A4_5 count contradiction: rail-corrected expectation was 26–51 freestanding posts; we accept 114 delineators. Either the correction was wrong or a chunk of the 114 are rail-mounted reflectors conflated into the class — potentially a hidden FP pool bigger than the known 6. Needs a mass audit by roadside/rail-presence.
- LLM judges are correlated (same 2D crops, same blind spots). Use physics (lattice phase, cross-pass occupancy) as arbiter of record; test verdict stability under different renders for disputeds.
- Freeze a blind holdout now, before pass-5 tuning — four passes have tuned on the same slice.
- Stop conditions: none true yet. Closest is #2 (human judgement) for the judge-split sparse family — but only after lattice + cross-pass instrumentation still leaves it ambiguous, in which case Miro gets a point-cloud viewer scoped to ~15 disputed objects.
Pass-5 queue (merged: strategist × memo × pass-4 evidence)
- Freeze a blind holdout (untouched A1 + A4_5 segments) before touching any threshold — methodology fix, zero cost.
- Lattice-prior recovery of the ~12 judged-real sparse misses: fit phase + spacing per carriageway side (50 m straight / HLB curve table), admit near-misses that phase-lock within learned residuals, off at gores; validate with adversarial crops + spacing-residual distributions.
- Blind lattice-site occupancy audit — first external recall number, feeds the stop-condition evaluation directly.
- Cross-pass occupancy arbiter for every disputed/dim object (raw-point occupancy per driving pass at the object location) — physics tiebreak for the correlated-judges problem; also probes the dim scrub bands ag26/ag27.
- A4_5 count-contradiction audit (114 vs 26–51): delineator-vs-rail-reflector mass classification by rail presence and mounting height.
- RF refresh per memo Q2 recipe — last, once the above have generated more and harder labels.