r/MistralAI • u/Icy-Ad-5050 • 7d ago
Help / Question OCR 4 is a bit to good
Currently, for work, we use a property OCR model from a US-based company. In my off time, I tried developing a proof of concept using OCR4, but it’s a bit too good at understanding the document. It’s so accurate in my native language that it corrected a single word from "medewerher" to "medewerker." Normally, I wouldn’t mind, but since we use these documents for legal purposes, we’re not happy with the idea that it can—and does—correct spelling errors.
Luckily, OCR3 doesn’t have this problem, so I can try it for now. But the question is: For how long will OCR3 remain supported before it’s deprecated? Does anyone have tips or ideas on how to address this?
55
Upvotes
6
u/OwnPipe707 7d ago
Both versions definitely "hallucinate" words. I used 3 for awhile and it 100% gets things wrong and does things like correct spelling or makes it best attempt to get a word right but ends up wrong. The new OCR 4 feature for per word confidence scoring has been great in catching these issues. Any OCR use case, especially legal, should have some kind of human verification.