DST609 has asked for the wisdom of the Perl Monks concerning the following question:
I am using pdftotext to extract text from multipage pdf's and it works great. I need to look for certain lines counting from the top of each page. I can find out how many pages are in the pdf and I see that pdftotext has first page / last page parameters but was looking for a way to do this more efficiently then say, invoking multiple file handles on the same document (running a page loop with the pdftotext filehandle within the loop).
Below is the code I am using to extract the data of the entire document
my $i=0; open (FILE, "pdftotext -layout multipage.pdf - |"); while(<FILE>) { $i++; my($line) = $_; print "\n<div class=\"line\"><div>$i</div>$line</div>"; } close FILE;
|
---|
Replies are listed 'Best First'. | |
---|---|
Re: XPDF pdftotext page loop
by Anonymous Monk on Oct 11, 2012 at 04:13 UTC | |
by Anonymous Monk on Oct 11, 2012 at 04:13 UTC | |
by DST609 (Novice) on Oct 11, 2012 at 04:37 UTC | |
by Kenosis (Priest) on Oct 11, 2012 at 06:07 UTC | |
by Anonymous Monk on Oct 11, 2012 at 12:59 UTC |
Back to
Seekers of Perl Wisdom