Dialed

Sheet 06

Evidence

How accurate the reader is, how its confidence was calibrated, where it fails, how fast it runs, and whether the agent does what it should. Every figure on this page comes from files in the repository and can be regenerated with one command.

Accepted readings off by 2% of the span or more
3 of 231
Gauges scored
340 incl. 40 real
Agent scenarios passing
8 of 8
OpenCV
5.0.0

Method

Ground truth. Photos of real gauges don't come with their true reading, so the main sets are rendered: a generator draws a gauge face (scale, ticks, numbers, needle, red zone, brand text), places it in 3D at a random angle, lights it, adds glare, blur, sensor noise and JPEG compression. The needle's value is known exactly.

Two sets the model never saw. The confidence model was fitted on 600 other rendered gauges. The normal set allows up to 35° tilt and occasional glare; the hard set allows up to 50°, glare on almost half the photos and heavy blur.

Error is |read − true| as a share of the scale's span, so 1% on a 0–16 bar gauge is 0.16 bar. A reading is accepted when no photo check blocks it and its confidence is at least 90%; otherwise the agent asks for a new photo.

Real photos. Photos of real gauges from Wikimedia Commons (CC0, CC BY and CC BY-SA), each read by eye from the photo for its true value. Where a dial prints two scales, the reading is scored against the scale the reader used. The 37 in the development set were used while building the reader (its failures on them guided fixes), so their numbers are optimistic. Blind photos were labelled before the reader ever saw them and the reader was not changed afterwards; see Blind tests. A further 46 photos the reader shouldn't accept (two needles, two gauges in one frame, the back of a gauge) test whether it declines.

  • Rendered gauge, true value 12.4514 bar
  • Rendered gauge, true value 12.3284 bar
  • Rendered gauge, true value 3.9556 psi
  • Rendered gauge, true value 6.3559 bar
  • Rendered gauge, true value 9.1914 bar
  • Rendered gauge, true value 0.2912 psi
  • Rendered gauge, true value 0.7079 bar
  • Rendered gauge, true value 3.4784 bar
  • Rendered gauge, true value 197.5845 °F

Results by set

SetGaugesRead at allWithin 2% (all reads)AcceptedAccepted within 2%Median error, acceptedSent back
Rendered, normalRendered gauges, up to 35° tilt, some glare and blur20020099.0%178 (89%)99.4%0.29%22
Rendered, hardRendered gauges, up to 50° tilt, heavy glare and blur1007481.1%41 (41%)97.6%0.35%59
Real photos, developmentReal photos from Wikimedia Commons, read by eye (development set)372965.5%12 (32%)91.7%0.96%25
Real photos, blindReal photos from Wikimedia Commons, read by eye, never used for tuning (blind set)31100.0%0 (0%)––3

Blind tests

Why. A reader tuned on its own test photos will look better than it is. So real photos were set aside, read by eye first, and the reader was run on them once.

Blind batch 1 (15 dials, 11 photos to decline). On its one blind run the reader accepted 4 readings; 50% of them were within 2% and 1 was off by more than 5%. It wrongly accepted 1 of the 11 photos it should have declined. The batch exposed three weaknesses: print in the blank part of a dial taken for the needle, handwheels and pipe ends taken for a second gauge, and needles resting below the first printed number. The first two were fixed, and the batch joined the development set.

Blind batch 2 was drawn at random from Commons after those fixes and read by eye before the reader saw it. Most random photos turned out not to be readable gauges (19 of them, used as photos to decline), so only 3 dials could be scored: too few for a percentage, so each one is listed below. It wrongly accepted 0 of the 19 photos it should have declined.

  • File:Industrial instrumetns-dial guage stem thermometer..jpg

    No reading read; by eye 30 °C

    Refused (blur, scale).

    Palagiri, CC BY-SA 3.0

  • File:Contents Gauge on Siebe Gorman Self-Contained Compressed Air Diving Apparatus (SCCADA).jpg

    No reading read; by eye 0 atm

    Refused (blur, scale).

    Tim Sheerman-Chase, CC BY 4.0

  • File:Manometre hpz.jpg

    15.1 read; by eye 15.1 bar

    0.1% of the scale off. Sent back, 79% confidence.

    Michel le tigre, CC BY-SA 3.0

Development photos, one by one

Every real photo in the set: what the reader said, what the gauge shows by eye, and whether the agent would have accepted the reading. Credits are the photographers'; derived images keep the photos' licences.

Photos it should decline

Real photos outside what Dialed reads. The right answer is not to log a number. It accepted 1 of 46; those are listed first.

Where errors come from

0%2%5%10%0°15°30°45°60°camera tilt
Error against camera tilt. Each dot is one gauge. Filled dots were accepted; rings were sent back. Large errors cluster at steep angles, and the reader turns them away.
0%50%100%50%60%70%80%90%100%confidence thresholdacceptedwithin 2%
The threshold trade-off. As the confidence threshold rises, fewer readings are accepted (coverage) and the accepted ones get cleaner. The agent accepts at 90%.
0%0%50%50%100%100%stated confidence
Is the confidence honest? Readings grouped by stated confidence, against the share that were actually within 2%. Points near the diagonal mean the number means what it says.
Blurred49Turned too far38Numbers unreadable26Glare on the face22Low confidence16Glare on the needle2Needle off the scale2
Why photos were sent back. The checks that fired on rejected photos, from both sets.

The worst cases

The largest errors in each set, whether or not they were accepted, and photos the reader refused outright. Accepted readings are marked; none of these large errors got through.

Does the agent do the right thing?

Eight scenarios, each a rendered photo pushed through a full capture against the sample plant's history, run once with the rule engine and once with the model. Pass means the outcome, and for holds the breach, is the one the guards should produce.

ScenarioExpectedRule engineModel
Normal readingPI-102LOGGEDPass, 3 toolsPass, 3 tools
Bearing warming all weekTI-201HOLDdriftPass, 3 toolsPass, 4 tools, 1 refused by guards
Filter pressure drop over limitPI-105HOLDdifferentialPass, 3 toolsPass, 4 tools, 1 refused by guards
Air receiver over its alarm limitPI-106HOLDalarm highPass, 3 toolsPass, 4 tools, 1 refused by guards
Intake below its alarm limitPI-101HOLDalarm lowPass, 3 toolsPass, 4 tools, 1 refused by guards
Shaky photoPI-104RE-SHOOTPass, 2 toolsPass, 2 tools
Photo turned 62° awayPI-107RE-SHOOTPass, 2 toolsPass, 2 tools
Wrong gauge photographedPI-107CHECK TAGPass, 2 toolsPass, 3 tools

Model runs used Amazon Nova Micro on Bedrock.

How the vision result changes what happens next

PerceptionOpenCV 5 reads the value,confidence, photo checks,printed unit and rangeDecisionthe model picks the nexttool; guards in codeallow or refuse itActionlog, ask for a re-shoot,check the tag, or holda work orderPersona supervisor approvesor rejects every heldwork orderevery tool call, refusal and timing is stored with the photo as a trace

Three recorded runs. In each, the next tool the model chose depended on what the OpenCV 5 reading returned: a week's drift led to a held work order, blocking photo checks led to a re-shoot with their reasons, and a printed range that disagreed with the tag led to a tag check.

Bearing warming all week TI-201 HOLD
  1. read_gauge 63.5 °C, 98% confidence; printed range 0–120
  2. compare_history Rising 2.7 °C a day over the past week
  3. log_reading logged
  4. hold_work_order “Pump P-1 bearing temperature drift”, medium priority, waiting for a supervisor

Amazon Nova Micro on Bedrock; each tool was chosen after reading the previous result.

Shaky photo PI-104 RE-SHOOT
  1. read_gauge no reading; blocking checks: blur, scale
  2. request_reshoot The photo is blurred. Hold still or tap to focus. Can't read the scale numbers. Get closer or reduce glare.

Amazon Nova Micro on Bedrock; each tool was chosen after reading the previous result.

Wrong gauge photographed PI-107 CHECK TAG
  1. read_gauge 109.61 bar, 99% confidence; printed range 0–160 (does not match the tag)
  2. read_gauge no reading; blocking checks: none
  3. flag_wrong_gauge This doesn't look like PI-107: the scale reads 0.0–160.0, the tag says 0.00–10.00. Check the tag and photograph the right gauge.

Amazon Nova Micro on Bedrock; each tool was chosen after reading the previous result.

Speed and the two DNN engines

OpenCV 5 ships a new DNN engine next to the classic one. Measured on the same CPU, they suit the two text models differently, so the reader loads each model with the engine that runs it faster.

The rest of the pipeline is classical OpenCV and takes tens of milliseconds; reading the printed numbers is most of the time.

On AWS Lambda (3008 MB, arm64, Sydney), measured from the function's own logs: a cold start took 3.9 s, and warm reads took 2.2 s median and 2.9 s at p90 over 5 photos. Upload and download time comes on top and depends on your network.

ModelClassic engineNew engineUsed
PP-OCRv3 text detector, 640 × 640357.4 ms210.1 msNew
CRNN recogniser, one word43.8 ms179 msClassic

Responsible use

Reproduce every number

  1. python scripts/get_models.py downloads the two text models and checks their hashes.
  2. python vision/synth.py data/synth200 200 and the hard set renders the gauges.
  3. python vision/calibrate.py fits the confidence model on its own 600 gauges.
  4. python scripts/evidence.py scores both sets and rebuilds this page's data.
  5. python scripts/agent_eval.py --model runs the eight agent scenarios.

Code and data: github.com/RohanGlitched/dialed