Beefy Boxes and Bandwidth Generously Provided by pair Networks
Syntactic Confectionery Delight
 
PerlMonks  

Re^4: Module to extract text from HTML

by bliako (Monsignor)
on Feb 28, 2024 at 14:15 UTC ( [id://11157947]=note: print w/replies, xml ) Need Help??


in reply to Re^3: Module to extract text from HTML
in thread Module to extract text from HTML

If I understood correctly that you are in control of websites and the formatting of their content, perhaps you could add some tags to the content by means of html comments or, better, custom attributes for html tags <p "data-purpose"="description" "data-index"="1">blah blav</p> and then you just reconstruct the text content from html.

Replies are listed 'Best First'.
Re^5: Module to extract text from HTML
by Bod (Parson) on Mar 01, 2024 at 15:47 UTC
    you are in control of websites and the formatting of their content

    No - although I am testing it on our own websites, it will be required to read our customers' websites over which we have no control.

Log In?
Username:
Password:

What's my password?
Create A New User
Domain Nodelet?
Node Status?
node history
Node Type: note [id://11157947]
help
Chatterbox?
and the web crawler heard nothing...

How do I use this?Last hourOther CB clients
Other Users?
Others goofing around in the Monastery: (5)
As of 2024-05-20 05:11 GMT
Sections?
Information?
Find Nodes?
Leftovers?
    Voting Booth?

    No recent polls found