Update, 7 August 2026 — the project has since concluded. The figures in this report are the interim ones of 3 August and have been superseded: the final ratified inventory is 125 devices on A1 and 216 on A4_5. The open items listed at the bottom — the held-out area, the corridor-wide coverage figure, the unconfirmed gantries — have all been closed out. See the closing chapter at the end of this page, or the full final report.
A survey vehicle drives the motorway carrying a laser scanner and cameras. It comes back with tens of millions of 3-D points per kilometre. Our job is to turn that into a usable inventory of the roadside furniture — delineator posts, traffic signs, guardrail-mounted reflectors, overhead sign structures — automatically, and to know how much of it we are actually catching.
This report covers where the accuracy stands today, how each number is checked, and a few things we found along the way that were not in anyone's plan. It is built without a single hand-drawn training label: the system works from geometry and reflectivity, and everything it proposes is checked against the raw scan before it counts.
Four classes of roadside device, in decreasing order of how well we have them nailed down:
Signs — plate-like objects on a stem — are the third class. They are well handled now, but only after moving traffic was separated out from them; that turned out to be the dominant error on one corridor and is worth its own section below.
The system also proposes gantry-mounted signs, but this is the least-verified class: only a handful of candidates exist and none has been through expert review yet. A first look at their evidence crops did not show the overhead structure we would expect, so these are queued for the next review session before we report any number for them. We would rather leave a class blank than report it as verified.
A detector that reports more objects is not automatically better — it may simply be wrong more often. So the working rule throughout has been: a bigger count is not a gain until the evidence has been checked, one candidate at a time. Five mechanisms enforce that.
For each proposed device we render a standard package straight from the raw scan: a wide view showing where the object stands, and a to-scale close-up of the object itself. The judgement is made from those images, not from the detector's own confidence — the detector does not get to vouch for itself.
Each candidate is assessed by more than one independent automated reviewer, each explicitly asked to make the case for rejecting it. Agreement between reviewers that disagree by design is worth far more than a single confident opinion. Only candidates that survive this are admitted to the count.
An expert judged 24 cases blind, four of which were hidden controls with a known answer. The controls came back correct, which is what licenses us to treat the expert's high-confidence verdicts as the final word — and on two cases they overturned the automated reviewers. The direction of the disagreement is useful in itself: the automated reviewers are too quick to write off ambiguous roadside verticals as vegetation, so we now re-screen that specific group.
Any search method can manufacture discoveries out of noise. So we run the identical search at decoy positions where no device should exist. If it "finds" things there, the method is fooling itself and the result is thrown out. In the most recent corridor-wide search, none of the 53 control positions produced a detection, while the real predicted positions produced 9 survivors.
A section of road has been frozen from the start and has never been looked at during development. It gets used exactly once, at the end, to produce the honest final number — with the success criteria and the judging rules written down in advance, so the test cannot be quietly re-scoped after the fact.
Six changes account for essentially all of the improvement. None of them is a bigger model; all of them came from looking at what the system was getting wrong.
| Step | What changed | Effect |
|---|---|---|
| Measure first | Built the ability to see where candidates were being lost before changing any detector logic. | No count change — set the rule everything else follows. |
| Brightness fix | One bright sign panel was setting the brightness bar for its whole neighbourhood, suppressing every faint post near it. | +37 confirmed delineators on A4_5. |
| Vehicles out | Moving traffic was being counted as roadside signs. | 6 false "signs" removed. |
| Second corridor | A size limit was cutting off posts whose base merges into verge grass. | A1 went from nearly blind to a working inventory. |
| Vegetation out | Bushes and plastic tree-shelters look like bright verticals; a shape-and-texture test separates them from real markers. | A whole family of false positives closed. |
| Direct search | Instead of waiting for the detector to assemble an object, predict where the next post should be from its confirmed neighbours and ask the raw scan whether one is there. | A1 65→81, A4_5 129→132, and recall to 92% on calibrated stretches. |
Panel-shaped, the right size, standing beside the carriageway — and driving. The tell is pure physics: the corridor is driven several times, so anything genuinely bolted to the ground appears in every pass, while a vehicle appears in exactly one and is somewhere else by the next. That single check removed six false devices and closed the largest error family on the corridor.
A row of stakes, evenly spaced, correct height, running in a straight line — and hundreds of metres from any road. It fooled the automated reviewers ten times out of ten, because regular spacing feels like strong evidence and is not: crops, plantations and fence lines are all regularly spaced. The fix was to stop treating spacing as proof and start checking road context directly. Near a real road there are hundreds of bright lane-marking returns; in a ploughed field there are none. Two previously accepted "misses" were withdrawn on that basis.
Some real posts return so few laser points that the detector never assembles them into an object — they are simply absent from its output, not rejected. Predicting where a post ought to be and then interrogating the raw scan at that spot recovered five such posts that had never registered in any form. This is the single biggest source of remaining upside.
One stretch reported far more delineators than the per-kilometre norm allowed, which looked like a false-positive cluster. It was a full motorway interchange: real devices at the tighter spacing used on ramps. The lesson has repeated often enough to be a rule here — when a count disagrees with an expectation, check the expectation before blaming the detector.
| Figure | Standing |
|---|---|
| 81 devices on A1 · 132 on A4_5 | Verified counts. Every device backed by reviewed point-cloud evidence. |
| 6 vehicle false-signs removed | Verified. Confirmed individually and by the physics check. |
| +9 devices in the latest pass | Under review. They pass every automated check including the null test; the review that admits or rejects them is running. |
| 92% found on calibrated stretches | Best-case estimate. True on well-surveyed stretches with a local expert count. Not yet a corridor-wide claim. |
| Corridor-wide coverage | Not yet measured honestly. An earlier method for estimating this was found to be invalid and discarded rather than reported — it would have flattered us badly. |
Scope note: everything above concerns two motorway corridors surveyed with the same vehicle and sensor configuration. Transfer to other corridors has repeatedly failed for rules that looked solid on one — thresholds calibrated on one road do not simply carry over. Any new corridor should be assumed to need its own calibration and its own verification round until shown otherwise.
Everything above was written mid-flight. This is the ending. Four days later the three open items had all been closed, the held-out area had been opened, and a panel of four independent reviewers voted unanimously to conclude the programme. The full write-up is the final report; this is the short version.
Four improvements went into the one-shot test. The two ambitious search rules — the ones that had looked best during development — found plenty of real devices and buried them in false alarms: roughly four wrong for every one right, on road they had never seen. On their development road the same rules had run between 40% and 96% correct. They were not shipped. The two unglamorous bug fixes found 6 real devices with zero false alarms, and went live.
Two out of four rejected sounds like a poor result. It is the opposite: without that test we would have put a pipeline into production that invents four bogus records for every real one, and nobody would have noticed until somebody tried to use the inventory. The frozen area paid for itself the single time it was used.
The pattern also explains itself. Bug fixes transfer; tuned rules do not. A defect in our own measurement code is wrong on every road, so fixing it helps everywhere. A threshold tuned against one corridor's geometry has quietly memorised that corridor. This is the same lesson as the scope note above, now measured rather than suspected.
This project was built with zero hand-drawn training labels — that was the point, and the tile at the top of this page still says so. That approach has now been pushed to its limit and the held-out test shows exactly where the limit is: hand-written rules selective enough to be useful on their own road are not selective enough on anybody else's.
The way past that is a trained model, which needs labelled examples — and the project turns out to have manufactured them as a by-product of checking its own work: about 400 candidates judged individually, 156 human verdicts over them, and from the sealed test 51 confirmed devices plus 143 confirmed false alarms on road nothing was ever fitted to. The two rejected rules are kept, not as detectors but as candidate generators for that next phase — they find nearly everything that is there; they simply cannot tell which ones are real. That last judgement is precisely what a trained model does well and a hand-written threshold does badly.
So the programme ends having produced two things: an inventory that is 67% larger and individually verified, and the training data for whatever replaces it.