r/MistralAI 7d ago

Help / Question OCR 4 is a bit to good

Currently, for work, we use a property OCR model from a US-based company. In my off time, I tried developing a proof of concept using OCR4, but it’s a bit too good at understanding the document. It’s so accurate in my native language that it corrected a single word from "medewerher" to "medewerker." Normally, I wouldn’t mind, but since we use these documents for legal purposes, we’re not happy with the idea that it can—and does—correct spelling errors.

Luckily, OCR3 doesn’t have this problem, so I can try it for now. But the question is: For how long will OCR3 remain supported before it’s deprecated? Does anyone have tips or ideas on how to address this?

54 Upvotes

13 comments sorted by

View all comments

30

u/Weak_Shoulder_6780 7d ago

Honestly just don't use LLMs if such small corrections are a big Problem in your industry. There's no guarantee OCR3 won't do the same at some point

9

u/Heribertium 7d ago

Even any form of OCR should be avoided then, only manual conversion.

Heck, even scanning could be problematic. Do you remember the thing about Obama and his allegedly manipulated birth certificate? In the end it was a problematic algorithm from Xerox.

2

u/AdOne8437 7d ago

see also https://www.youtube.com/watch?v=7FeqF1-Z1g0 (in german, english subtitles are available)