Beefy Boxes and Bandwidth Generously Provided by pair Networks
No such thing as a small change
 
PerlMonks  

Re: Parsing HTML/XML with Regular Expressions (HTML::TreeBuilder::XPath)

by tangent (Priest)
on Oct 18, 2017 at 02:34 UTC ( #1201545=note: print w/replies, xml ) Need Help??


in reply to Parsing HTML/XML with Regular Expressions

In my previous comment I mentioned that I could not find a way to pass the attribute empty_element_tags from HTML::TreeBuilder to HTML::Parser. Looking at the source code for HTML::TreeBuilder I found this:
our @ISA = qw(HTML::Element HTML::Parser); # This looks schizoid, I know...
So I've learnt something there! I can call empty_element_tags(1) and now it works.
use HTML::TreeBuilder::XPath; my $file = 'example.html'; my @result; my $tree = HTML::TreeBuilder::XPath->new; $tree->empty_element_tags(1); # calls this on HTML::Parser $tree->parse_file($file); $tree->eof; my @divs = $tree->findnodes('//div[@class="data"]'); for my $div (@divs) { my $text = $div->as_text || ''; $text =~ s/\W//g; push(@result, $div->attr('id') . "=$text"); } print join(', ',@result);
Output:
Zero=, One=Monday, Two=Tuesday, Three=Wednesday, Four=Thursday, Five=F +riday, Six=Saturday, Seven=Sunday

Replies are listed 'Best First'.
Re^2: Parsing HTML/XML with Regular Expressions (HTML::TreeBuilder::XPath)
by fishy (Monk) on Oct 18, 2017 at 07:08 UTC
    Great!
    Thanks.

Log In?
Username:
Password:

What's my password?
Create A New User
Node Status?
node history
Node Type: note [id://1201545]
help
Chatterbox?
and all is quiet...

How do I use this? | Other CB clients
Other Users?
Others chanting in the Monastery: (3)
As of 2017-12-16 23:00 GMT
Sections?
Information?
Find Nodes?
Leftovers?
    Voting Booth?
    What programming language do you hate the most?




















    Results (459 votes). Check out past polls.

    Notices?