Beefy Boxes and Bandwidth Generously Provided by pair Networks
We don't bite newbies here... much
 
PerlMonks  

Re: Re: Re: Scraping HTML: orthodoxy and reality

by ff (Hermit)
on Jul 09, 2003 at 05:38 UTC ( #272581=note: print w/replies, xml ) Need Help??


in reply to Re: Re: Scraping HTML: orthodoxy and reality
in thread Scraping HTML: orthodoxy and reality

For sure, this approach is expensive cpu-wise, etc., but if I need a solution that works right away then "module fetches/renders HTML into text", combined with regex processing that at least I know how to do, IS a solution. Sure, per RBFuller, "... if the solution is not beautiful, I know it is wrong" but if those cycles won't be used for anything else, who cares? This bear of little brain would have his program done.

So, assuming that efficiency doesn't matter, I'm still fishing for something like building the $html object via a LWP 'get' as above and then turning it into text that I can examine with regexen. (However, since this is turning a golden object into lead, I'll do some more digging as you suggest, like re-reading this thread's Data::Dumper/HTML::TableExtract example! :-)

  • Comment on Re: Re: Re: Scraping HTML: orthodoxy and reality

Replies are listed 'Best First'.
Re: Re: Re: Re: Scraping HTML: orthodoxy and reality
by chanio (Priest) on Jul 09, 2003 at 07:17 UTC
    I am just starting my studies with Perl, and of course, with modules I have less experience.

    But if you could print an HTML file to a plain text printer the result sent to a file would be just what you saw at the screen, right?

    Then you would treat it like text...

Log In?
Username:
Password:

What's my password?
Create A New User
Node Status?
node history
Node Type: note [id://272581]
help
Chatterbox?
[Mr. Muskrat]: (considered) Re^3: Please help with Regexp::Common needs to be reparented to Re^2: Please help with Regexp::Common
[Mr. Muskrat]: Is it just me or does that timeout issue seems to be happening more often lately?
[Corion]: Mr. Muskrat: I'm not sure if it really happens more often, but I don't exactly know either
[LanX]: yep
[LanX]: more often for some weeks now
[Corion]: I think I'll have to manually (as god) intervene with that node, as the simple reparenting didn't seem to fix the parent/child relationship of the nodes
[Corion]: I think I have an idea but I'll have to open a ticket with Pair.com on that - hopefully I get to that on the weekend
LanX imagines a burning thorn bush
[Mr. Muskrat]: Thank you!
[Mr. Muskrat]: Oh that is odd. I got the message that it was reparented but yeah, it didn't actually do it. lol

How do I use this? | Other CB clients
Other Users?
Others meditating upon the Monastery: (11)
As of 2017-01-19 16:24 GMT
Sections?
Information?
Find Nodes?
Leftovers?
    Voting Booth?
    Do you watch meteor showers?




    Results (170 votes). Check out past polls.