Pass 5 — the lattice pass
Vertical-sign detection improvement loop · AI3D-339 · 2026-08-03 ·
report 5 of N · previous:
pass 4
TL;DR. Pass 5 turned the German delineator-spacing regulation into a shipped detector stage.
- Shipped: a corridor-level delineator_lattice post-pass — accepted delineators seed chain growth; rejected clusters that phase-lock into an anchor-rich chain and look like an HLB post are admitted. Full-rerun validated: A1 57 → 65 (+8, all judged real), A4_5 127 → 129 (+1 judged real, +1 documented FP), 0 removals.
- First external recall number: on the two calibrated 25 m stretches the base detector had found only ~50% of true posts; lattice admission lifts those stretches to ~83%. Three missed posts never even clustered — confirmed by raw-point occupancy probes at lattice-predicted sites.
- A4_5 count contradiction resolved: the "114 vs 26–51 expected" excess is an interchange, not an FP pool — ramp curves legally carry 7–8 m spacing (HLB), and 12/12 sampled detections are visually real.
- Honest negatives: the isolation-score instrument was refuted; the 12 previously-judged-real misses remain unrecovered (their stretches have too few anchors to seed chains).
- Not stopped. Fresh-eyes strategist: no stop condition holds; pass-6 queue below. One ask for Miro: install the web point-cloud viewer now — a ~20-case human calibration pass would bound LLM-judge error for the whole loop.
1 · The A4_5 count contradiction is an interchange, not an FP pool
Pass 4 left a live suspicion: 114 accepted A4_5 delineators vs a rail-corrected expectation of
26–51. Half of the "excess" sits in segments 131/132/134 — which turn out to be a full interchange.
The nearest-neighbour spacings there (7.1–7.7 m, plus 1.2–2 m gore-nose groups) match the HLB
curve table for R≈70–80 m ramps, not mainline. All 12 sampled detections are unambiguous posts,
bollards or gore devices in close-up; 4 spot checks were re-confirmed by the adversarial fleet
(4/4 real). Side lesson: on_road_fraction=1.0 /
dist_to_road=0 is unreliable at interchanges — never use it as an FP cue there.
2 · Fitting the lattice through the posts themselves
Road-centerline models failed on this corridor (junctions + dual carriageway), so rows are fitted
directly through the candidates: every accepted delineator seeds chain growth over a gated pool of
rejected clusters; each step predicts the next site at the tracked base spacing, snaps within 3 m,
and tolerates up to 6 missing sites. A null test (same growth, pool positions randomized
within each segment) determines which chain classes carry real signal:
High-skip and 1–2-anchor chains appear as often in randomized pools — chance. The
signal lives in anchor-rich, low-skip chains.
Chain 0 (segs 32–36): a textbook 25 m row along the carriageway edge. Orange rings =
chain members; red rings = predicted-but-empty sites, later probed for raw occupancy.
3 · What the adversarial fleet said
36 crops (+2 late A4_5 crops) judged by Opus and Grok independently, refute-framed, 28/36
agreement, Fable tiebreak on the rest. Outcomes: all 8 A1 admissions from the two 6-anchor chains
are real; all 6 candidates from 1–2-anchor chains are vegetation (exactly as the null test
predicted); 3 of the occupied predicted-empty sites hold real posts the detector never clustered.
Typical admission (s034 c207): slender ~1 m post with brighter head at the verge —
sparse scanline sampling of a 12 cm HLB post. Rejected by the base detector as "unclassified";
phase-locked at 25 m in a 6-anchor chain.
4 · The first external recall number
A clean post with bright cap at a lattice-predicted site (~1 m prediction offset) —
no cluster, no detection. Three of these are confirmed; they can only be recovered by
detect-at-predicted-site (pass-6 queue #1).
Caveat (strategist): these stretches were selectable because they
had 6 detected anchors — the measurement is recall's best case, an upper bound. Corridor-wide
recall is still unknown; measuring it via the A4_5 road model + rail registry is pass-6 queue #3.
5 · The shipped rule and its validation
| Gate | Value | Why |
| chain anchors | ≥ 4 accepted delineators | null test: 1–2-anchor chains are chance (their candidates judged 6/6 vegetation) |
| spacing residual | < 15% of base | lattice regularity (HLB: 50 m mainline, 3–50 m by curvature) |
| pool height | 0.8–1.4 m | HLB post = 1.00 m above pavement edge |
| pool plate | ≤ 0.05 m | post is a plate, not a blob |
| pool bright seed fraction | ≥ 0.15 | a delineator has a retroreflector; the one judged-veg candidate inside a 6-anchor chain had 0.000 |
Implementation: lattice.py post-pass invoked at the end of
detect.main() over the whole invocation (chains cross segment borders);
idempotent; lattice_admission=False restores byte-identical outputs.
New append-only CSV column z_ground lets the pass rebuild full detection
dicts without reloading NPZ. 263 tests pass (11 new).
Full-rerun diff (p5_final vs p4_final): A1 +8 / −0 (exactly the judged-real
set), A4_5 +2 / −0. No other change.
The one bad admission, on the record: s078 c59 (A4_5) — both judges say
vegetation (a shrub in a row that happens to phase-lock at 49.5 m against 11 anchors). Its bright
fraction (0.182) is a near-tie with the dimmest judged-real admission (0.177): no honest geometric
threshold separates them, so it ships as a documented FP instead of a p-hacked cut. Judged
admission precision: 9/10.
6 · Negative results (kept on purpose)
- Isolation score refuted. The hypothesis that pass-4's transfer failure hid a usable
"posts stand in empty space" signal is wrong on this corridor: real posts stand near rails and
embankment structure, and sparse vegetation is often the isolated thing (AUC 0.26–0.73, wrong
direction). Discarded.
- The 12 judged-real misses stay unrecovered. They neither phase-lock (anchor-poor
stretches) nor separate on any cluster-internal feature. They are the concrete case for pass-6
detect-at-predicted-site and for the human calibration pass.
7 · Research memo (Codex) highlights
- Q1 lateral placement: HLB §5.2 — posts ≥ 0.50 m outside the paved edge; RQ 31 derivation
puts the outer row ~6.1–6.4 m from the right-lane centre. Derived prior, not a norm.
- Q2 rail furniture: shortened Aufsatzleitpfosten (≈0.55 m product examples)
continue the 50 m sequence on rails; small direct-rail reflectors (7–10 cm, 2–5 cm proud) have
no sourced universal spacing. Different classes — don't merge.
- Q3 recall benchmarks: no published recall curve exists for ~1 m delineators in MLS
— the lattice-site occupancy audit is genuinely novel measurement.
- Q4 gore furniture: Leitplatte (Z 626) = broad bright slab, lower edge ~0.6 m;
supports the interchange audit's gore-assembly reading.
8 · Pass-6 queue and stop status
Stop-condition verdict (fresh-eyes strategist): not stopped. Recall ~50%
measured on best-case stretches kills condition 4; raw occupancy proves the data still contains
unexploited signal (kills 3); rule-level moves remain (kills 1); condition 2 is close but the
machine path isn't exhausted.
- Detect-at-predicted-site: lower clustering thresholds only inside windows around
lattice-predicted empty sites, then apply the existing HLB shape + retroreflectivity gates. The 3
confirmed never-clustered posts are the guaranteed yield.
- FP purge + guarded lattice fixpoint: remove the ~6 known accepted FPs from anchoring,
then iterate admission with an original-anchor quorum and a per-iteration null test.
- Corridor-wide recall on A4_5 via road model + rail registry — turns the anecdote into a
denominator. Holdout stays sealed until goal.
- RF refresh as soft verifier over #1's candidates, fed lattice/context features
(cluster-internal features are capped at AUC ~0.73).
- Gates and trees: measure or explicitly de-scope — the goal names four classes and two
are currently unreported.
Ask for Miro: install the web point-cloud viewer now (parallel, cheap
lead time). Strongest use: one human calibration pass over ~20 hard judged cases (the 12 misses +
8 admissions) to bound LLM-judge error — the whole loop's precision/recall numbers currently rest
on machine judges with no human calibration.
Artifacts
- commits: c664f6b (lattice admission + z_ground), 798a74c (blind holdout), 42741f9→798a74c amended
- battlebox: /mnt/d/ai3d-aml/vsigns_eval/{p5_final_a1, p5_final_a45} = new frozen baselines (65 / 129 accepted)
- holdout: out_eval/HOLDOUT.md — A1 b0 000–029 + b1 000/001/009/010 sealed; b0 091–136 reserve
- research memo: /tmp/research_pass5_memo.md (to be committed as out_eval/research_pass5_memo.md)
- verdicts: /tmp/p5_verdicts_{opus,grok}.csv, /tmp/p5_verdicts2_{opus,grok}.csv, /tmp/p5_admission_verdicts.csv