Roadside device inventory — pass 9

AI3D-339 · 2026-08-06 · the pass that found the devices sitting on top of the guardrails — and showed the filters had been throwing real ones away
In one paragraph. Pass 9 found about 110 additional roadside devices — marker posts, delineators and signs — that the detection pipeline had either never looked for or had actively discarded. These are draft finds: every one has been confirmed by independent automated judges, but they are pending final human sign-off before they enter the official inventory. If they all hold, the A1 corridor goes from 75 to 128 devices (+71%) and the A4/5 corridor from 184 to 218. Three findings drove it: A1 turned out to carry more guardrail than A4/5, and searching the space just above the guardrail crest found 55 confirmed marker posts nobody had counted; a deliberate control experiment on the candidates the system rated "weak" found that 43% of them were real, meaning the filters had been far too conservative; and re-examining 47,000 previously discarded detections recovered 23 more real devices. A repaired completeness estimate says the pipeline, running unaided, finds at most about 63% of the devices actually present — and that the reviewed inventory has now overtaken what the pipeline can find on its own. The four-member expert review panel voted unanimously, 4–0, for exactly one more engineering pass — build these discoveries into the pipeline, prove them on held-back data — and then conclude the project.
1 The numbers · 2 Devices on top of the guardrail · 3 The filters were too strict · 4 Mining the discard pile · 5 How much is still missing · 6 Two fixable defects · 7 Two scope decisions · 8 The panel's decision · 9 What happens next

1 · The numbers

corridorinventory before pass 9draft additionsdraft inventorychange
A175+53128+71%
A4/5184+34218+18%
both corridors259+87346+34%
plus: confirmed devices standing in road stretches the current inventory files do not span+23kept with their evidence

The pass reviewed 291 candidate locations. 117 were judged real devices; after removing 7 that turned out to be the same physical object found twice by two different searches, 110 draft admissions remain. 87 of those sit in the road stretches the two corridor inventories actually cover; the other 23 are real devices in stretches those files do not span, so they are held with their evidence rather than counted.

Read these as draft, not as delivered. Every number above comes from automated judging. Under the process used since pass 8, nothing enters the official inventory until a human has personally reviewed the imagery for each find. That review session is the first and blocking item of the next pass. One number in particular needs it: roughly half of the A1 guardrail finds were decided by a third judge breaking a disagreement between the first two, so the review session doubles as the first direct measurement of how reliable that tie-breaking judge actually is.

2 · Devices on top of the guardrail

Some roadside markers are not free-standing posts at all — they are short stubs bolted onto the crest of the guardrail beam. A search that looks for a post-shaped object rising out of the ground will miss every single one of them. Pass 8 built a search that works in guardrail coordinates instead — find the beam, measure its true crest height, look at what sticks up above it — and that search added 43 devices to A4/5.

Pass 9 pointed the same search at A1, on the assumption that A1 is lightly railed and the exercise would be cheap. The assumption was wrong in the most useful possible way. A1 carries more modelled guardrail than A4/5 — roughly 17.5 km across the searched stretch against A4/5's 16.3 km, over fewer road segments. Stated carefully: that is a comparison of guardrail length, not of the fraction of each corridor that is railed.

stagecount
candidate locations above the guardrail crest652
candidates in the high-confidence group (bright reflective head, seen on more than one survey pass)73
judged real marker posts55
of those, inside the stretch the inventory covers32 of 41 candidates = 78%
A1 marker post standing on the guardrail crest, top-down and side views
The family archetype on A1. Left: the road from above, candidate circled, sitting on the guardrail line between the carriageway and the tree band. Right: two side views along the road — the cyan line is the fitted guardrail crest, and a short bright stub stands clear above it. Both judges: real, high confidence.
A1 marker post on the guardrail with vegetation nearby
The same signature with vegetation in frame. The bushy mass to the left stays below the crest line; the marker is the one compact object standing above it. Measuring relative to the guardrail rather than to the ground is what makes that distinction possible.

3 · The filters were too strict — a control experiment says so

The guardrail search sorts its candidates into a strong group — bright reflective head, and seen on more than one survey pass — and a weak tail. Pass 8 judged only the strong group and expected the tail to be essentially empty. Pass 9 judged the A4/5 tail as a deliberate control experiment, with that near-zero expectation written down in advance.

24 of 56 weak candidates were real devices — 43%. The strong-candidate test is therefore not a boundary between real and not-real. It is a ranking aid worth roughly a factor of two in precision, and nearly half the real devices in this corridor's tail sit below it. Any production rule built on that test alone silently forfeits them.

The immediate consequence: the equivalent A1 tail — 579 candidate locations nobody has looked at — can no longer be assumed empty. Sampling it is on the queue for the final pass.

weak candidate confirmed as a real guardrail marker post
Judged real. One of the 24. The side views resolve the guardrail beam and its bare support posts, with a slender marker standing on the crest. It failed the strong-candidate test only because its reflective head did not saturate the sensor and it was captured on a single survey pass.
weak candidate rejected as vegetation
Judged not a device. The other side of the same experiment: above the beam sits a small forked whisker that leans rather than standing plumb, rooted in the vegetation bank behind the rail. 32 of the 56 were rejected like this — the control measures precision honestly in both directions.

4 · Mining the discard pile

Every time the detection pipeline runs it records not only what it accepted but every cluster of points it rejected, with the measurements behind the decision and the reason. Nobody had ever looked at that pile: 47,159 rejected detections across the two corridors, none of them reviewed by anyone.

Pass 9 screened them against nine physical criteria — shape, height, brightness, position relative to the road — all derived either from the pipeline's own accepted devices or from the handful of cases where a human had previously overruled the machine. No human labelling was needed and no re-run of the pipeline: the data was already on disk. The screen cut 47,159 down to roughly 200, of which the top 60 were rendered and judged.

23 of 60 were real devices — 38%, three times the prior expectation. 14 on A1, 9 on A4/5. Real delineators and posts that the pipeline had found, measured, and then thrown away.
rescued delineator standing at the verge in front of a tree line
Rescued from the discard pile on A1: a clean bright post at the verge line, guardrail and vegetation behind it. Unambiguous once anyone looks at it — and previously discarded without review.
rescued delineator on an embankment above the carriageway
A second rescue, on an embankment above the carriageway. This is the geometry that trips up ground-referenced shape tests: the object is perfectly normal, the ground beneath it is not flat.
rescued marker post above a guardrail beam on A4/5

A third rescue, on A4/5 — and the clearest single picture in the pass of what these searches look at. The side views resolve the guardrail beam, its regularly spaced bare support posts, and one bright marker head standing above the crest. The pipeline had this cluster, measured it, and rejected it.

5 · How much is still missing

Counting what you found is easy. Estimating what you missed is the hard part, and the project's previous attempt at it was too weak to quote — it only ever saw about a fifth of the population it was trying to measure.

Pass 9 rebuilt the estimate. The survey data contains multiple independent passes over the same road; drawing eight independent half-samples of those passes and comparing what each finds yields a statistical estimate of the total population, including the devices no sample found. The rebuilt instrument sees 94% of the population it is estimating, against 20% before, and it was validated against synthetic populations of known size before its output was believed.

quantityvalue
estimated devices present on the A1 evaluation stretch98 (range 98–147)
found by the detection pipeline running unaided62
pipeline completenessat most ~63% (range 42–63%)
the inventory as shipped after pass 8, including human-reviewed additions75
the draft inventory after pass 9128
The reviewed inventory has overtaken the pipeline. The pipeline alone finds 62 devices on that stretch; the reviewed inventory already stands at 75 and, if the pass-9 finds are ratified, at 128. The discoveries are real, but they currently live in review records rather than in the software — which is precisely why the panel directed the final pass to build them into the pipeline instead of continuing to search.

Stated honestly, as the panel insisted: the 63% figure is a diagnostic bracket, not a measurement of the population. The method is optimistic by construction — a cross-check where the true answer was known shows it under-reports hidden devices by more than tenfold — so real completeness is lower, not higher. It is quotable as "the pipeline alone is far from complete", not as a precise percentage.

6 · Two fixable defects in the pipeline

Both were found this pass, both are concrete, and neither needs a single item of human-labelled training data to fix.

Defect 1 — combining the survey passes discards real devices

The completeness experiment produced a by-product: 22 locations the half-samples of the data agree on, but that the full combined run does not report. Judged blind, 10 of the 22 were real devices — and all ten were unanimous, both independent judges agreeing without a tie-break. This is the least disputable evidence in the pass.

Merging all survey passes together makes the pipeline see less, not more. Real devices visible in thinner slices of the data are lost when everything is combined — most likely because merging fuses a marker into a neighbouring structure until it no longer looks like a marker. The mechanism is not yet proven, and root-causing it is the first engineering item of the final pass: a fix whose mechanism is unknown cannot be trusted as a rule.
real delineator that the combined pipeline run does not report

One of the ten. A bright marker standing above the guardrail at the verge, plainly visible from above and in both side views — and absent from the combined run's output. Both judges: real delineator, high confidence, no tie-break needed.

Defect 2 — a placeholder value that triggers a rejection

One of the internal shape measurements — "is this cluster upright?" — does not always measure anything. When a cluster is wider than a set limit, the routine returns a placeholder value of zero instead of reporting that it cannot tell. That placeholder then reads as "definitely not upright" to a downstream rule, which rejects the cluster as a passing vehicle or as noise.

The scale of it: the placeholder fires on 65% of everything A1 discards — and on none of the devices A1 accepts. Two of the three known cases where a human previously overruled the machine sit inside that blind spot. It follows that the 38% hit rate from the discard pile in section 4 is a floor on what that pile holds, not an estimate of it: the screen could only see the narrow-cluster part. The fix is small and is queued for the final pass.

7 · Two scope decisions

Trees are out of scope, 4–0. Tree detection was carried in the goal from the start, but there has never been usable training data for it: the vegetation model that scores trees was never trained on human-verified examples, so its apparently strong accuracy figures measure agreement with machine-generated guesses rather than with reality. Continuing to tune it would be measuring nothing. The class is formally removed, re-openable only on the two things that would make it meaningful — real human-annotated tree examples and a proper completeness audit. This is a scoping decision, not a failure: it stops the project spending effort where no evidence can be obtained.

Valley-mounted reflectors: searched, none present. Some countries fit reflectors inside the trough of the guardrail beam rather than on top of it, and that class was structurally invisible to every search built so far. Pass 9 built a dedicated search for it and judged 80 candidates: zero in-trough reflectors. Five turned out to be discrete markers at or above the crest, four of them the same objects the guardrail search had already found. Honest limit: the search was seeded on bright returns, so the defensible claim is "no bright trough reflector exists on this corridor", and dim ones remain untested. That matches the German standard, which treats this mounting as exceptional use.

8 · The panel's decision — unanimous, and terminal

Each pass ends with a four-member review panel of independent automated reviewers. Each writes its verdict without seeing the others', and each must state in advance what evidence would prove it wrong. Pass 8 ended 3–1, with one member arguing to stop.

Pass 9: CONTINUE, 4–0. All four converged, independently and unprompted, on the same shape: exactly one more bounded pass — build the discoveries into the pipeline, prove them on held-back data — and then stop. Confidence ranged from 68% to 88%. The member who voted to stop last pass changed position: his objection was that completeness was unobservable, and the repaired estimate substantially answers it.
the panel's reasoning, in short
No stop condition is currently true. The pipeline is demonstrably incomplete, and the two defects holding it back are cheap, concrete and fixable without any human labelling. Stopping with known free repairs on the table is not defensible.
But continuing to search is no longer the valuable activity. The searches have delivered; the findings now need to be turned into software and tested against data the project has never read.
Therefore the next pass is pre-committed as the last one. Each member wrote down in advance the specific outcome that would make them vote to stop at its end — and each expects to.

The panel also recorded five points where the pass-9 write-up overstated its own evidence, and the sections above reflect the corrected versions — most importantly that the completeness figure is a diagnostic bracket rather than a measurement.

9 · What happens next

  1. Human ratification session — first, and blocking. Every one of the 110 draft finds is reviewed against its imagery by a human. Nothing enters the official inventory before that. Several parked decisions ride along, including what to do with the 23 finds that sit outside the current inventory files. The session also produces the first direct measurement of how often the tie-breaking judge is right.
  2. Find the root cause of the defect that makes the combined run discard real devices — before writing any rule that claims to fix it.
  3. Build the discoveries into the pipeline: the guardrail-relative search, a rescue stage for the discard pile, and both defect fixes. The acceptance test is strict — a clean rebuild must reproduce the ratified inventory automatically, with no regressions, and the pipeline's unaided count on the A1 evaluation stretch must rise from 62 to at least 67.
  4. Run the held-back fresh-data test, last. A stretch of road never used during development, read once, against criteria fixed in advance. It has been deferred twice for good reasons; this time it runs.
  5. Then conclude — a before/after report on the A1 corridor, comparing the inventory at the start of the project with the ratified final one.

Baselines before pass 9: A1 75 devices, A4/5 184. Draft after pass 9, pending sign-off: A1 128, A4/5 218. Zero human training labels used anywhere in the project to date. Previous report: pass 8 · full project story: the journey so far