go ahead... be a heretic | |
PerlMonks |
PDF::OCR2 results not what I was hoping forby nysus (Parson) |
on Feb 08, 2016 at 15:33 UTC ( [id://1154639]=perlquestion: print w/replies, xml ) | Need Help?? |
nysus has asked for the wisdom of the Perl Monks concerning the following question: I'm trying to OCR this document. The results are disappointing to say the least. The output was basically just random characters with lots of blank lines. None of the text appearing in the original PDF was recognized. When I tried on a "clean" document converted straight to PDF from a word processing document, the module worked fine. So apparently OCR2 just doesn't have the logic to pull text from more sophisticated documents that are scanned. I know that the PDF::OCR2 just provides an interface to tesseract/imagemagick so this probably isn't the best forum for this question but I'm hoping someone can give me some advice that points me in the right direction. I'm interested in pulling out the time and location data from the document.
$PM = "Perl Monk's";
Back to
Seekers of Perl Wisdom
|
|