TL;DR
- I added an “Enhance” button to my photo viewer, backed by RealESRGAN running in ComfyUI on the Mac Studio. Its stated job was to help me read model plates and serial numbers in blurry surplus photos.
- I benched it against a plate where I knew the right answer. At small glyph sizes it did the opposite of its job: crisp, confident, wrong text.
G1312B-60050came back as01D12B-(0)00. - It is not a bad model pick. Four different upscalers garbled the same lines the same way.
- The usual advice (pre-upscale with a classical filter first) changed nothing. The failure is the generative prior, not a starved input.
- What I do now: the original pixels stay the source of truth, there is a text-safe sharpen view that invents nothing, and the AI view composites real pixels back over anything that looks like fine detail.
- A wrong model number off an enhanced photo is a wrong comp, which is a wrong bid. This one was worth catching before it spent money.
Why I wanted it
The most valuable pixels in a surplus lot photo are the smallest ones. The difference between a generic median price and an exact-model price is often a part number on a data plate, shot from a few feet away at a resolution where a 13-point label is about twelve pixels tall. I already pull the full gallery for every lot and archive it, so I had the best pixels the seller provided. They just were not enough for some plates.
Super-resolution looked like the obvious tool. I already run image generation locally for the blog (the ComfyUI on Apple silicon setup is the same box), and an upscaling model is a few nodes in a graph. So I wired it up: a button in the photo lightbox that sends the image to ComfyUI, upscales it, stores a 2x copy, and shows it. I also added a step to strip the auction site’s watermark before upscaling, because the model happily sharpened a diagonal watermark into a feature.
I was pleased with it for about four days.
The bench
The thing that finally bothered me was that Enhanced photos looked great. Crisp text, clean edges. I had never checked whether the text was right, because I only used the button on plates I could not already read. That is the exact condition under which you cannot catch a hallucination.
So I made a plate I could check. I rendered an equipment data plate at high resolution with known strings, deliberately full of characters that blur into each other (0/O, 1/I, 5/S, 8/B, 6/G), then degraded it the way a lot photo degrades: blur, downscale, JPEG. Then I ran the exact graph the app sends, and read the result back against the truth. To be clear, the plate is synthetic. The strings are made up. What is real is the degradation, the model and the pipeline.
At about 12 pixels per glyph, which is a plate shot from a few feet away and the common case:
| Path | Large (24pt) values | Fine print (13pt) |
|---|---|---|
| Classical Lanczos 2x | correct | correct |
| RealESRGAN x4plus | correct | garbled |
Garbled looked like MADE IN GERMANY becoming MAOE IJI GEIWANY, and PART NO. becoming PART ND.. Not noise. Plausible, well-formed, sharp.
At about 5 pixels per glyph, the case where a plate is in a wide shot and you most want help:
| Truth | After RealESRGAN |
|---|---|
G1312B-60050 | 01D12B-(0)00 |
DE64O5S812 | DIGIO35012 |
5065-9931 | 3005-0001 |
REV B08 | DLV D10 |
Everything was invented. The output had no uncertainty in it anywhere. It looked like a photograph of a real plate, which it was not.
Four things that ruled out the easy fixes
- It is the architecture, not the checkpoint. I had four upscaling models installed. Three more (a sharpness-tuned one, an 8k-trained one, an illustration-tuned one) garbled the same lines the same way. A model that has learned what text looks like will paint text-like strokes where there is no information.
- Classical pre-upscaling does not help. The standard advice is to enlarge with Lanczos first so the network gets a bigger input. At 1.5x and 2x it changed nothing. The failure is the prior, not an under-fed network.
- There is no cheap pixel check. I tried a whole-image fidelity gate and a column-darkness correlation. The garbled output scored 4.5 against a gate at 12 (so it passed), and the correlation was 0.985 for an honest copy against 0.989 for a garbling one. Catching hallucinated glyphs needs a text recogniser. I looked at what the research literature does, and the recent text-aware super-resolution papers all name the general-purpose upscalers, diffusion ones included, as baselines that garble text. Swapping to a bigger, fancier upscaler would have been a regression.
- The gate I had was a gate against the wrong failure. My fidelity check was designed to catch an upscaler that wandered off from the source in broad structure. It was blind to the glyph-level lie.
What I do instead
Originals first. The archived original is the source of truth. The viewer’s derived copies are for human eyes only; the identification code works from the archived originals. This is the cheapest rule and the most important one.
A sharpen view that invents nothing. The viewer’s first button is now a plain classical path: Lanczos enlargement, a tonal stretch, an unsharp mask. No model. It runs in the web pod in well under a second (about 650 ms, after my first guess of 40 ms turned out to be fiction), needs no GPU (so it works when the Mac Studio is asleep, which was the most common reason “Enhance unavailable” showed up), and its output is soft but truthful. On the five-pixel plate, every value was still recoverable by zooming. The pattern that actually reads a serial is sharpen, then zoom, then look.
Crops, not hallucinations. When the identification step reports that it could not actually read the model number, it asks the model where the label is, crops that region out of the original photo, enlarges it with a plain classical resample, and re-reads only the identifiers. The identification code never calls the AI upscaler. Zooming into real pixels cannot invent a character.
The AI view stays, with a leash. Enhance is still useful for the thing it is good at: judging condition (a scratch, a cracked bezel, a bent pin). So I kept it but stopped it touching text. After the model pass, a detector finds regions of dense fine detail and composites a plain enlargement of the original pixels back over them. The network keeps the smooth frame it handles well and never redraws a glyph. Legible text stays legible because nothing redrew it, and illegible text stays soft instead of becoming crisply wrong. On the test plate, the three fine-print lines that had come out as gibberish came back correct.
Two design choices there are worth stealing:
- The detector is deliberately not text-specific. Knurled knobs and object outlines get protected too. Over-masking costs a little prettification. Under-masking costs a fabricated serial. I know which I prefer.
- The threshold is absolute, not a percentile. A percentile lights a fixed fraction of the image by construction, so a photo with no text got as much mask as a photo full of it. An absolute threshold is what lets a smooth panel come out with no protection at all.
Also: if the mask step errors, the code serves the unprotected result rather than nothing, and logs it loudly. I am not thrilled about that, but the state it protects against (a fabricated serial) is also the state of not having this module, so failing loud and open seemed like the lesser evil.
I also bumped the cache version, because every cached enhanced copy from before the change was pure unprotected output with possibly-fabricated text, and the cache is checked by existence.
The general lesson
Any tool that makes an image look better is optimizing for looks. For OCR-adjacent work what you want is evidence, and those objectives diverge exactly where the information runs out. The more convincing the output, the less you should trust it as a measurement.
I ran this tool for days on plates I could not otherwise read, and nothing ever told me it was wrong. That is the real finding: a quality-looking result and a correct result are different properties, and you can only tell them apart if you have a case where you know the answer. Which is the argument for keeping a small golden set of cases with known answers, and benching every change against it. I will write that method up separately.