Beefy Boxes and Bandwidth Generously Provided by pair Networks
Problems? Is your data what you think it is?
 
PerlMonks  

Re: Extracting text from pptx

by vagabonding electron (Chaplain)
on Apr 18, 2013 at 15:15 UTC ( #1029374=note: print w/replies, xml ) Need Help??


in reply to Extracting text from pptx

I am a beginner too, but here is my attempt. It is very quick and very dirty (since it takes header and footer from each slide as text as well as the text) but it does extract the text from the pptx (at least from my pptx :-) ) and perhaps it could be useful.
#!/perl use strict; use warnings; use Archive::Zip qw( :ERROR_CODES ); use XML::Twig; my @text; my $file = "The full path to your pptx"; my $zip = Archive::Zip->new(); $zip->read( $file ) == AZ_OK or die "Unable to open Office file\n"; my @slides = $zip->membersMatching( "ppt/slides/slide.+\.xml" ); for my $i ( 1 .. scalar @slides ) # to sort them. { my $content = $zip->contents( "ppt/slides/slide${i}.xml"); my $twig= XML::Twig->new( #keep_encoding=>1, twig_handlers => { 'a:t' => \&topicref_processing, }, ); $twig->parse( $content ); } sub topicref_processing { my($twig, $ppttext) = @_; push @text, $ppttext->text(); } use FindBin qw($Bin); open my $out, ">:encoding(UTF-8)", "$Bin/test_ppt_PM.txt" or die "$!"; print {$out} "$_\n" for @text;

Log In?
Username:
Password:

What's my password?
Create A New User
Node Status?
node history
Node Type: note [id://1029374]
help
Chatterbox?
[shmem]: glasswalk3r: just because I'm a bit annoyed
[shmem]: trying to build a RPM for some perl with a bunch of modules. Trying perlbrew, trying alien.

How do I use this? | Other CB clients
Other Users?
Others avoiding work at the Monastery: (4)
As of 2017-03-25 22:15 GMT
Sections?
Information?
Find Nodes?
Leftovers?
    Voting Booth?
    Should Pluto Get Its Planethood Back?



    Results (313 votes). Check out past polls.