Sheet 06
Evidence
How accurate the reader is, how its confidence was calibrated, where it fails, how fast it runs, and whether the agent does what it should. Every figure on this page comes from files in the repository and can be regenerated with one command.
Technical report (PDF)Architecture diagram
- Accepted readings off by 2% of the span or more
- 3 of 231
- Gauges scored
- 340 incl. 40 real
- Agent scenarios passing
- 8 of 8
- OpenCV
- 5.0.0
Method
Ground truth. Photos of real gauges don't come with their true reading, so the main sets are rendered: a generator draws a gauge face (scale, ticks, numbers, needle, red zone, brand text), places it in 3D at a random angle, lights it, adds glare, blur, sensor noise and JPEG compression. The needle's value is known exactly.
Two sets the model never saw. The confidence model was fitted on 600 other rendered gauges. The normal set allows up to 35° tilt and occasional glare; the hard set allows up to 50°, glare on almost half the photos and heavy blur.
Error is |read − true| as a share of the scale's span, so 1% on a 0–16 bar gauge is 0.16 bar. A reading is accepted when no photo check blocks it and its confidence is at least 90%; otherwise the agent asks for a new photo.
Real photos. Photos of real gauges from Wikimedia Commons (CC0, CC BY and CC BY-SA), each read by eye from the photo for its true value. Where a dial prints two scales, the reading is scored against the scale the reader used. The 37 in the development set were used while building the reader (its failures on them guided fixes), so their numbers are optimistic. Blind photos were labelled before the reader ever saw them and the reader was not changed afterwards; see Blind tests. A further 46 photos the reader shouldn't accept (two needles, two gauges in one frame, the back of a gauge) test whether it declines.
Results by set
| Set | Gauges | Read at all | Within 2% (all reads) | Accepted | Accepted within 2% | Median error, accepted | Sent back |
|---|---|---|---|---|---|---|---|
| Rendered, normalRendered gauges, up to 35° tilt, some glare and blur | 200 | 200 | 99.0% | 178 (89%) | 99.4% | 0.29% | 22 |
| Rendered, hardRendered gauges, up to 50° tilt, heavy glare and blur | 100 | 74 | 81.1% | 41 (41%) | 97.6% | 0.35% | 59 |
| Real photos, developmentReal photos from Wikimedia Commons, read by eye (development set) | 37 | 29 | 65.5% | 12 (32%) | 91.7% | 0.96% | 25 |
| Real photos, blindReal photos from Wikimedia Commons, read by eye, never used for tuning (blind set) | 3 | 1 | 100.0% | 0 (0%) | – | – | 3 |
Blind tests
Why. A reader tuned on its own test photos will look better than it is. So real photos were set aside, read by eye first, and the reader was run on them once.
Blind batch 1 (15 dials, 11 photos to decline). On its one blind run the reader accepted 4 readings; 50% of them were within 2% and 1 was off by more than 5%. It wrongly accepted 1 of the 11 photos it should have declined. The batch exposed three weaknesses: print in the blank part of a dial taken for the needle, handwheels and pipe ends taken for a second gauge, and needles resting below the first printed number. The first two were fixed, and the batch joined the development set.
Blind batch 2 was drawn at random from Commons after those fixes and read by eye before the reader saw it. Most random photos turned out not to be readable gauges (19 of them, used as photos to decline), so only 3 dials could be scored: too few for a percentage, so each one is listed below. It wrongly accepted 0 of the 19 photos it should have declined.

No reading read; by eye 30 °C
Refused (blur, scale).
Palagiri, CC BY-SA 3.0

No reading read; by eye 0 atm
Refused (blur, scale).
Tim Sheerman-Chase, CC BY 4.0

15.1 read; by eye 15.1 bar
0.1% of the scale off. Sent back, 79% confidence.
Michel le tigre, CC BY-SA 3.0
Development photos, one by one
Every real photo in the set: what the reader said, what the gauge shows by eye, and whether the agent would have accepted the reading. Credits are the photographers'; derived images keep the photos' licences.

180 read; by eye 185 psi
1.3% of the scale off. Accepted.
RAF-YYC from Calgary, Canada, CC BY-SA 2.0

67.0 read; by eye 69 °F
1.3% of the scale off. Accepted.
Conant, CC BY-SA 3.0

70.6 read; by eye 70 °F
0.4% of the scale off. Accepted.
Conant, CC BY-SA 3.0

72 read; by eye 74 °F
1.0% of the scale off. Accepted.
Mike Shaw, CC BY-SA 4.0

25 read; by eye 0 psi
0.6% of the scale off. Sent back, 72% confidence.
Hey Paul from State College, PA, USA, CC BY 2.0

No reading read; by eye 0 MPa
Refused (scale).
Deyan Levski, CC BY-SA 3.0

No reading read; by eye 0 psi
Refused (scale).
Pbsouthwood, CC BY-SA 4.0

No reading read; by eye 0 bar
Refused (glare_needle, scale).
CEphoto, Uwe Aranas, CC BY-SA 3.0

No reading read; by eye 0 kg/cm²
Refused (scale).
Nord68, CC BY-SA 4.0

-38.17 read; by eye 0 bar
95.4% of the scale off. Sent back, 57% confidence.
Amfeli, CC BY 4.0

0.46 read; by eye 0.5 bar
1.0% of the scale off. Accepted.
Antti, CC BY 2.0

9 read; by eye 0 kg/cm²
0.2% of the scale off. Sent back, 84% confidence.
Peter Southwood, CC BY-SA 4.0

771 read; by eye 110 psi
0.7% of the scale off. Sent back, 82% confidence.
Ncysea, CC BY-SA 4.0

22.6 read; by eye 22 psi
1.1% of the scale off. Sent back, 61% confidence.
In Transit, CC BY-SA 4.0

594 read; by eye 580 kPa
1.4% of the scale off. Sent back, 72% confidence.
Fumikas Sagisavas, CC0

610 read; by eye 640 kPa
3.0% of the scale off. Sent back, 81% confidence.
Fumikas Sagisavas, CC0

834 read; by eye 800 psi
0.8% of the scale off. Accepted.
leapingllamas, CC BY 2.0

59 read; by eye 62 psi
1.1% of the scale off. Sent back, 77% confidence.
Clem Rutter, Rochester, Kent., CC BY 3.0

No reading read; by eye 22 °C
Refused (scale).
Cjp24, CC BY-SA 3.0

No reading read; by eye 0 bar
Refused (scale).
R. Henrik Nilsson, CC BY 4.0

56.4 read; by eye 122 psi
41.0% of the scale off. Sent back, 16% confidence.
Pete Birkinshaw from Manchester, UK, CC BY 2.0

245 read; by eye 250 psi
1.2% of the scale off. Accepted.
Joe Mabel, CC BY-SA 4.0

46 read; by eye 48 psi
0.7% of the scale off. Accepted.
Mark Cross, CC BY 4.0

40.3 read; by eye 39.5 °C
0.8% of the scale off. Sent back, 65% confidence.
Levin Holtkamp, CC BY-SA 4.0

-55.5 read; by eye 23.3 °C
117.7% of the scale off. Sent back, 4% confidence.
Pogrebnoj-Alexandroff, CC BY 3.0

10.7 read; by eye 0 psi
35.7% of the scale off. Sent back, 5% confidence.

0.4 read; by eye 0 psi
0.4% of the scale off. Accepted.
Palagiri, CC BY-SA 3.0

6.7 read; by eye 0 psi
6.1% of the scale off. Accepted.
Delusion23, CC BY-SA 4.0

10 read; by eye 0 psi
4.8% of the scale off. Sent back, 80% confidence.
Momcilo Stokanovic, CC BY-SA 4.0

70 read; by eye 82 kg/cm²
4.8% of the scale off. Sent back, 43% confidence.
Централизованный информационный портал Республики Башкортостан, CC BY 4.0

0.63 read; by eye 0 kg/cm²
0.6% of the scale off. Sent back, 51% confidence.
Cjp24, CC BY-SA 4.0

15.01 read; by eye 0 bar
8.8% of the scale off. Sent back, 19% confidence.
Shixart1985, CC BY 2.0

No reading read; by eye 590 kPa
Refused (blur, scale).
N509FZ, CC BY-SA 4.0

4.4 read; by eye 0 psi
2.0% of the scale off. Accepted.
Peter Southwood, CC BY-SA 4.0

-2.3 read; by eye -2 °F
0.3% of the scale off. Accepted.
Tony Webster from Minneapolis, Minnesota, United States, CC BY 2.0

No reading read; by eye 14 °C
Refused (scale).
Quant, CC BY-SA 3.0

97.9 read; by eye 0 psi
65.3% of the scale off. Sent back, 27% confidence.
Joe Mabel, CC BY-SA 4.0
Photos it should decline
Real photos outside what Dialed reads. The right answer is not to log a number. It accepted 1 of 46; those are listed first.

needle below the printed scale (oven thermometer at room temperature)
Accepted (wrongly).
Gmhofmann, CC BY-SA 3.0

two needles, two scales
Not accepted, 0% confidence.
Shixart1985, CC BY 2.0

pyrometer with three scales
Not accepted, 56% confidence.
Wilfredor, CC BY-SA 4.0

two needles
Refused (scale).
Ivor the driver, CC BY-SA 4.0

several needles, faded
Refused (scale).
Sanjay Acharya, CC BY-SA 4.0

gauge tiny in the frame
Refused (blur, multiple, scale).
Schickdavid3, CC BY-SA 4.0

gauge tiny in the frame
Refused (tilt, scale).
Xnatedawgx, CC BY-SA 4.0

the back of a gauge, no face
Refused (tilt, blur, scale).
CEphoto, Uwe Aranas, CC BY-SA 3.0

extreme close-up, part of a dial
Refused (tilt, small, blur, glare, scale).
Cory Denton from Saskatoon, CC BY 2.0

two gauges in one photo
Not accepted, 37% confidence.
Lupus in Saxonia, CC BY-SA 4.0

dirty glass, scale unreadable
Refused (scale).
Shixart1985, CC BY 2.0

compound vacuum and pressure scale
Not accepted, 75% confidence.
Marc Ryckaert (MJJR), CC BY 3.0

two needles, two scales
Refused (tilt, scale).
Shixart1985, CC BY 2.0

gauge tiny in the frame
Not accepted, 21% confidence.
Mds08011, CC BY 4.0

open mechanism
Refused (scale).
Frank John-Lorenz Aus Schrott-Kleinteilen von Frank John-Lorenz zusammengefügt , CC0

mechanism and a loose dial
Refused (scale).
Daderot, CC0

two gauges in one photo
Not accepted, 97% confidence.
VeRdUgO PY, CC BY-SA 4.0

a valve, no gauge
Refused (tilt, scale).
Shixart1985, CC BY 2.0

contact gauge with two set-point needles
Not accepted, 27% confidence.
S.J. de Waard, CC BY 2.5

pump with no dial
Refused (small, blur, scale).
Radler22, CC BY-SA 4.0

pump with no dial
Refused (small, blur, scale).
Radler22, CC BY-SA 4.0

fire-extinguisher gauge with no numbers
Refused (tilt, blur, scale).
User:connortk, CC BY-SA 3.0

two gauges in one frame
Refused (blur, glare, scale).
Lupus in Saxonia, CC BY-SA 4.0

gauge mechanism, no dial face
Refused (scale).
CristianChirita, CC0

back of a gauge
Refused (tilt, scale).
Amfeli, CC BY 4.0

second, red set pointer
Refused (scale).
Omnibus-Trip, CC BY-SA 3.0

two gauges in one frame
Not accepted, 7% confidence.
Pogrebnoj-Alexandroff, CC BY 3.0

several small gauges on a test board
Refused (tilt, scale).
National Institute of Standards and Technology, Public domain

small dial with an inset second dial
Refused (tilt, blur, scale).

engine-order telegraph, words not numbers
Not accepted, 0% confidence.
unknown, CC BY 4.0

gauges seen edge-on
Refused (tilt, glare, scale).
Alan Murray-Rust, CC BY-SA 2.0

the side of a gauge, no face
Refused (tilt, small, blur, scale).
National Gauge Co, CC0

the back of a gauge
Refused (tilt, blur, scale).

the back of a gauge
Refused (small, blur, scale).

the back of a gauge
Refused (small, blur, scale).

steep angle, most of the scale out of view
Refused (blur, scale).
National Gauge Co, CC0

a computer screen, not a gauge
Refused (tilt, scale).
Ayratayrat, CC BY-SA 4.0

no gauge
Refused (tilt, small, blur, scale).
Department of Energy. National Nuclear Security Administration. Sandia National Laboratories. 3/1/2000, Public domain

two gauges in one frame
Not accepted, 8% confidence.
Daderot, CC0

several gauges in one frame
Not accepted, 98% confidence.
Rémi Kaupp, CC BY-SA 3.0

sight glass, no gauge
Refused (blur, scale).
Rosser1954, CC BY-SA 4.0

two needles on one dial
Refused (scale).
RobbieMcConnel, CC BY-SA 3.0

tyre inflator with no dial
Refused (tilt, scale).
R. Henrik Nilsson, CC BY 4.0

gauge cut off by the frame, behind a bottle
Refused (small, blur, scale).
Ayratayrat, CC BY-SA 4.0

compound gauge, vacuum and pressure on one dial
Not accepted, 83% confidence.
Joe Mabel, CC BY-SA 3.0

two gauges in one frame
Not accepted, 1% confidence.
unknown, Public domain
Where errors come from
The worst cases
The largest errors in each set, whether or not they were accepted, and photos the reader refused outright. Accepted readings are marked; none of these large errors got through.

-0.43 vs 3.48 bar
24.4% off, 32° tilt, sent back (low confidence)

205.14 vs 197.58 °F
3.8% off, 30° tilt, accepted

No reading
45° tilt, sent back (tilt, blur, scale)

-63.84 vs 147.96 psi
132.4% off, 65° tilt, sent back (tilt, range)

No reading
57° tilt, sent back (tilt, blur, scale)

No reading
39° tilt, sent back (blur, scale)

No reading
45° tilt, sent back (tilt, blur, scale)

244.46 vs 198.33 °F
23.1% off, 60° tilt, sent back (tilt, blur)

-29.68 vs -7.84 °C
27.3% off, 17° tilt, sent back (low confidence)

-34.66 vs 53.46 psi
88.1% off, 32° tilt, sent back (blur)

-27.25 vs 5.73 bar
206.2% off, 29° tilt, sent back (blur)

-137.04 vs 47.22 psi
115.2% off, 30° tilt, sent back (blur, glare)
Does the agent do the right thing?
Eight scenarios, each a rendered photo pushed through a full capture against the sample plant's history, run once with the rule engine and once with the model. Pass means the outcome, and for holds the breach, is the one the guards should produce.
| Scenario | Expected | Rule engine | Model |
|---|---|---|---|
| Normal readingPI-102 | LOGGED | Pass, 3 tools | Pass, 3 tools |
| Bearing warming all weekTI-201 | HOLDdrift | Pass, 3 tools | Pass, 4 tools, 1 refused by guards |
| Filter pressure drop over limitPI-105 | HOLDdifferential | Pass, 3 tools | Pass, 4 tools, 1 refused by guards |
| Air receiver over its alarm limitPI-106 | HOLDalarm high | Pass, 3 tools | Pass, 4 tools, 1 refused by guards |
| Intake below its alarm limitPI-101 | HOLDalarm low | Pass, 3 tools | Pass, 4 tools, 1 refused by guards |
| Shaky photoPI-104 | RE-SHOOT | Pass, 2 tools | Pass, 2 tools |
| Photo turned 62° awayPI-107 | RE-SHOOT | Pass, 2 tools | Pass, 2 tools |
| Wrong gauge photographedPI-107 | CHECK TAG | Pass, 2 tools | Pass, 3 tools |
Model runs used Amazon Nova Micro on Bedrock.
How the vision result changes what happens next
Three recorded runs. In each, the next tool the model chose depended on what the OpenCV 5 reading returned: a week's drift led to a held work order, blocking photo checks led to a re-shoot with their reasons, and a printed range that disagreed with the tag led to a tag check.
read_gauge63.5 °C, 98% confidence; printed range 0–120compare_historyRising 2.7 °C a day over the past weeklog_readingloggedhold_work_order“Pump P-1 bearing temperature drift”, medium priority, waiting for a supervisor
Amazon Nova Micro on Bedrock; each tool was chosen after reading the previous result.
read_gaugeno reading; blocking checks: blur, scalerequest_reshootThe photo is blurred. Hold still or tap to focus. Can't read the scale numbers. Get closer or reduce glare.
Amazon Nova Micro on Bedrock; each tool was chosen after reading the previous result.
read_gauge109.61 bar, 99% confidence; printed range 0–160 (does not match the tag)read_gaugeno reading; blocking checks: noneflag_wrong_gaugeThis doesn't look like PI-107: the scale reads 0.0–160.0, the tag says 0.00–10.00. Check the tag and photograph the right gauge.
Amazon Nova Micro on Bedrock; each tool was chosen after reading the previous result.
Speed and the two DNN engines
OpenCV 5 ships a new DNN engine next to the classic one. Measured on the same CPU, they suit the two text models differently, so the reader loads each model with the engine that runs it faster.
The rest of the pipeline is classical OpenCV and takes tens of milliseconds; reading the printed numbers is most of the time.
On AWS Lambda (3008 MB, arm64, Sydney), measured from the function's own logs: a cold start took 3.9 s, and warm reads took 2.2 s median and 2.9 s at p90 over 5 photos. Upload and download time comes on top and depends on your network.
| Model | Classic engine | New engine | Used |
|---|---|---|---|
| PP-OCRv3 text detector, 640 × 640 | 357.4 ms | 210.1 ms | New |
| CRNN recogniser, one word | 43.8 ms | 179 ms | Classic |
Responsible use
- No people in the pictures. The reader looks for dials; there is no face or person detection anywhere, and photos are only kept on a round, as evidence for the work order they support.
- A person decides. The agent never releases a work order or touches a control system.
- Uncertainty is shown, not hidden. Every reading carries its confidence and the checks behind it; struck-through figures stay visible.
- Known gaps are listed on the read page, and refused photos say why.
- Not a safety system. Alarms and trips stay in the plant's control system.
Reproduce every number
python scripts/get_models.pydownloads the two text models and checks their hashes.python vision/synth.py data/synth200 200and the hard set renders the gauges.python vision/calibrate.pyfits the confidence model on its own 600 gauges.python scripts/evidence.pyscores both sets and rebuilds this page's data.python scripts/agent_eval.py --modelruns the eight agent scenarios.
Code and data: github.com/RohanGlitched/dialed






