Roadside asset inventory from mobile LiDAR — final report

Corridor A1, before and after · 7 August 2026 · for colleagues and management

A survey vehicle drives the motorway with a laser scanner and cameras and returns tens of millions of 3-D measurement points per kilometre. The task of this project was to turn that raw material into a trustworthy inventory of the roadside furniture — delineator posts, traffic signs, guardrail-mounted reflectors, overhead sign structures — automatically, and to be honest about how much of it we are actually catching. This report closes the project and states the final numbers.

75 → 125
confirmed devices on corridor A1 (+67%)
141 → 216
confirmed devices on corridor A4_5 (+53%)
1.2%
measured error rate of the automated judges, against human verdicts
~400
candidates individually reviewed against point-cloud evidence
0
hand-drawn training labels used to build any of it
4–0
unanimous review-panel decision to conclude the programme

The result

Two motorway corridors were inventoried. Every device added over the course of the project was image-verified from the raw scan by independent automated review panels and then ratified by a human expert. Nothing was added on the detector's own say-so.

CorridorBefore (1 August)Final (7 August)Change
A175 devices125 devices+50  (+67%)
A4_5141 devices216 devices+75  (+53%)

The A1 gain is not spread evenly — it is concentrated in classes the original system was effectively blind to. Breaking the corridor down by device type shows where the inventory actually grew:

Device type on A1BeforeFinalChange
Delineator posts (the ~1 m verge posts)6170+9
Traffic signs1219+7
Guardrail-mounted half-posts034+34
Other poles22
Total75125+50

The single largest finding of the project is the row in the middle. Where a safety barrier runs along the verge, the full-height delineator post is replaced by a short reflector stub bolted to the top of the rail. The original system reported none of these on A1 — they are half the height of a normal post and mounted on a large metal structure that dominates the scan around them. Thirty-four of them are now in the inventory, each one confirmed from the imagery.

Short reflector post standing on top of a safety barrier
A barrier-mounted half-post, confirmed. Left, seen from above: the ringed device sits on the narrow strip between the white edge line and the tree line. Right, seen side-on: a short stub standing clear above the crest of the safety barrier. The lower panel looks along the barrier and shows why these are hard — the rail and its regular support legs are a far larger object than the small reflector sitting on top of them.
A second barrier-mounted reflector post seen from two angles
A second confirmed barrier post, and the reason the reviewers were sure. Looking along the barrier, two stubs appear at regular installed spacing, and the cross-view shows a bright flat reflector plate at the rail top. Regular spacing plus a manufactured reflective surface is a pattern vegetation does not produce.

Why the numbers are trustworthy

A detector that reports more objects is not automatically better — it may simply be wrong more often. The rule enforced throughout the project was that a larger count is not a gain until the evidence has been examined, one candidate at a time.

Every candidate is judged from rendered evidence, not from a score

For each proposed device the raw scan is rendered into a standard evidence package: a wide view showing where the object stands relative to the road, and a to-scale close-up of the object itself. The decision is made from those pictures. The detection pipeline never gets to vouch for its own output.

Independent judges, briefed to argue for rejection

Each candidate is assessed by two independent automated review panels, each explicitly instructed to build the case against the candidate. Where they disagree, a third, more capable reviewer looks at the imagery and breaks the tie. Agreement between reviewers who are designed to disagree is worth far more than one confident opinion.

The judges themselves were measured against a human

The obvious question about an automated judge is how often it is wrong. We can answer it with a number. A human expert re-reviewed the full set of tie-broken decisions: 83 cases, of which he overturned 1 — an error rate of 1.2%, against an acceptance threshold of 30%. On the cases where both automated panels already agreed, he overturned none out of 27. The judging instrument is not a black box we are asking anyone to take on faith; it is calibrated, and the calibration is published.

The human's verdicts are authoritative wherever the two disagree. He struck five candidates from the final count — three confirmed as vegetation, one as part of the barrier structure rather than a device in its own right, and one left deliberately unresolved — and corrected the recorded type on 22 more. Every one of those corrections is reflected in the 125 and the 216 above.

A delineator post confirmed on previously unexamined road
One of the 51 real devices found on the sealed road. Left: the aerial view — the ringed object stands on the verge immediately outside the white edge line, with the tree line well behind it. Right: side-on, a slim post carrying a bright striped reflector head, sitting exactly in the height band a delineator should occupy.
A second confirmed delineator post on previously unexamined road
Another sealed-road find: a compact bright post rising about a metre out of the verge on a slip-road island, with a clean straight-sided outline quite unlike the feathery grass tufts beside it. Devices like this are genuinely present and genuinely missed. The difficulty was never finding them — it was telling them apart from the several false alarms that arrive alongside each one.

The held-back test — the honest centrepiece

From the very start of the project, 40 road segments were frozen and never looked at. No rule was ever tuned on them, no candidate from them was ever reviewed, nobody involved saw a single image from them. The success criteria and the judging rules were written down and locked before the test was opened. It was opened exactly once, at the end, and the result is the result.

This is the only part of the project that measures what the system does on road it has never seen — which is the only question that matters for deploying it anywhere else. Four changes went into the test:

Change under testReal devices foundFalse alarmsOutcome
Aggressive barrier-post search1985not shipped
Aggressive recovery of discarded candidates1040not shipped
Conservative bug fix — measurement error in the shape test20shipped
Conservative bug fix — more robust object sizing40shipped

The two aggressive rules found plenty of genuinely real devices — and buried them in false alarms. Roughly four out of five things they flagged were not devices (18–20% precision). On the road they had been developed against, those same rules had run at 40–96% precision. They were not shipped.

The two conservative changes were not clever new search rules at all — they were straightforward defect fixes in existing measurements. Between them they found 6 real devices and produced 0 false alarms. Those two are now live in production.

This is the system working exactly as designed

It is tempting to read "two out of four changes rejected" as a disappointing result. It is the opposite. Both rejected rules looked excellent right up until the moment they met unfamiliar road; had we shipped on development-set evidence, we would have put a pipeline into production that generates four false entries for every real one, and we would not have found out until somebody tried to use the inventory. The held-back test caught it before a single wrong record reached anyone. That is what the test is for, and it paid for itself the one time it was used.

The pattern is also informative rather than random. Bug fixes transferred perfectly; tuned rules did not. A defect in our own measurement code is wrong on every road, so fixing it helps on every road. A threshold tuned on one corridor's geometry encodes that corridor — and quietly stops working on the next one. This is the central technical lesson of the project and it is written into the handover.

What was searched corridor-wide, and closed

The programme finished by systematically sweeping every remaining place a device could be hiding, so that nothing is left merely assumed. Each of these is now closed with an answer rather than an open question.

SearchOutcome
Barrier-mounted postsSwept across both corridors. 34 added on A1, 71 now on A4_5. The largest single class of previously missed devices, and the reason A1 grew 67%.
Reflectors in road troughsSuspected as a further hidden class. Eighty candidate positions examined; the class is simply not present on these corridors. Closed as absent rather than left as a nagging doubt.
The rejected-detection archiveTens of thousands of candidates the pipeline had discarded over its lifetime were re-examined. 23 real devices recovered and returned to the inventory, and the underlying defect that discarded them was identified and fixed.
Overhead sign gantriesThe last unverified class. A corridor-wide sweep found 9 gantries on A1, confirmed by the automated judges from colour imagery. These await final human sign-off and are not yet counted in the 125.
TreesFormally removed from scope, 4–0. Tree classification had never had human-verified examples to learn from, so its apparently strong accuracy measured agreement with machine guesses, not with reality. Reinstatable only if genuine human-annotated tree data becomes available.
Overhead motorway sign gantry reconstructed from LiDAR points in camera colour
An overhead sign gantry on A1, in true camera colour. Top: the carriageway from above — the vehicles, the lane markings and the green verge are all measured 3-D points, not a photograph. Below, looking across the road: the blue sign panel hangs high above the traffic on its support. Nine of these were located on A1.
A second overhead sign gantry in camera colour
A second A1 gantry, same evidence format, the blue panel over the lanes again unmistakable in the cross-road view. Overhead structures were the last class the project had never verified, so these are reported separately and are not included in the confirmed 125 until a human has signed them off.
A device whose recorded type was corrected during the final review
Not every correction is an addition. This device was on record as a barrier-mounted post — but the review found no barrier anywhere near it; the nearest is almost 12 metres away. The along-road view on the right shows what it actually is: a free-standing post rising clear of the ground returns. The device is real and stays in the inventory; only its classification changed. Twenty-two records had their type call reviewed and corrected this way in the final review — and where the class genuinely moved, it consistently became more specific, never vaguer.

The verdict — why the programme is being concluded now

The project was governed by a standing rule: it continues until a panel of four independent reviewers agrees that a defined stopping condition has been met. That panel has now returned a unanimous 4–0 decision to stop, each seat reaching it independently, and each having committed in advance to the criteria that would trigger it.

The reasoning is the held-back test. The approach used throughout this project was deliberately label-free: no human ever drew a training example; the system worked from geometry and reflectivity, using hand-designed rules. That approach has now been driven to its limit. It has extracted what it can, and the held-back test shows precisely why it cannot extract more: hand-tuned rules that are selective enough to be useful on the road they were built for are not selective enough on any other road.

The programme produced its own successor's training data

The stopping condition the panel invoked was written at the outset and says, in effect: stop when further progress requires supervised machine learning trained on real labels. That condition is now met — and the striking thing is that the project has manufactured exactly those labels as a by-product of verifying its own numbers:

The project began with zero labelled examples. It ends with a labelled dataset, and the two rejected aggressive rules are retained not as detectors but as candidate generators for that next phase: they find nearly everything that is there, they simply cannot tell the real ones apart. Deciding which is which is exactly what a trained model does well and a hand-written threshold does badly.

What's next

Reopen conditions, agreed in advance and recorded: this programme restarts only if new corridor data with camera coverage arrives, if a trained scorer clears the fresh-road false-alarm bar, if a review ruling removes devices already counted, or if the detection code is changed at all — in which case the change requires its own regression suite and its own pre-registered fresh-road test.

Honest numbers — what these figures do and do not claim

FigureStanding
125 devices on A1 · 216 on A4_5Ratified counts. Every device backed by reviewed point-cloud evidence and confirmed by a human expert. Reproducible: the inventory rebuilds byte-for-byte from the recorded decisions.
1.2% judge error rateMeasured, not estimated — 1 overturn in 83 human-reviewed decisions.
6 real devices found by the shipped fixesMeasured on road never used for development, with zero false alarms.
Completeness of the inventoryBounded, and knowingly incomplete. The held-back test found roughly 1.3 real, previously unrecorded devices per road segment on fresh road. Real devices remain undetected on both corridors; we can say so with evidence, which is better than not knowing. Closing that gap is the job of the machine-learning phase.
9 overhead gantries on A1Awaiting final human sign-off. Deliberately excluded from the 125 until signed off. We would rather under-report than report an unverified class.
One disputed A1 candidateFlagged for a second look. The expert's verdict was "probably vegetation, but not sure", so it is excluded. If it later resolves as real, A1 becomes 126.

Scope note: everything above concerns two motorway corridors surveyed with the same vehicle and sensor configuration. The central finding of the held-back test — that rules calibrated on one road do not carry to another — applies to this work as much as to anyone else's. Any new corridor should be assumed to need its own calibration and its own verification round until shown otherwise.

Related reading