Beefy Boxes and Bandwidth Generously Provided by pair Networks
Don't ask to ask, just ask
 
PerlMonks  

Re: Comparison of HTML documents with Perl

by talexb (Canon)
on Feb 24, 2005 at 21:29 UTC ( #434266=note: print w/ replies, xml ) Need Help??


in reply to Comparison of HTML documents with Perl

I've had good luck with HTML::Parser, but I think what you're asking for could end up being infinitely complicated.

You want to end up doing a high level compare, not a line by line or word by word compare. If you have to handle anyone's HTML, that could be impossible. If you're trying to version your own HTML, that will probably be easier -- you can focus on just a few tags.

I think I'd start by comparing the structure of the document for changes, and see where a paragraph has been added or a table row has been deleted. From there, I'd look at the text within the parts whose structure has not changed. Boy, that's a really interesting project. Let us know how it turns out.

Alex / talexb / Toronto

"Groklaw is the open-source mentality applied to legal research" ~ Linus Torvalds


Comment on Re: Comparison of HTML documents with Perl

Log In?
Username:
Password:

What's my password?
Create A New User
Node Status?
node history
Node Type: note [id://434266]
help
Chatterbox?
and the web crawler heard nothing...

How do I use this? | Other CB clients
Other Users?
Others lurking in the Monastery: (5)
As of 2014-10-26 02:03 GMT
Sections?
Information?
Find Nodes?
Leftovers?
    Voting Booth?

    For retirement, I am banking on:










    Results (149 votes), past polls