I like the code from roboticus
. A slight reformulation is below.
You can capture all of the names from the "match global" in one expression.
For most of my "web scraping" code, very fast performance is not that important. Neither is being super general purpose. If I can write a short regex in 5 minutes that gets me what I want, then I go with it and if the web page changes in a year, then I write another 5 minute regex.
What makes sense in your application has to do with what you are "scraping", how often the page format changes, what the impact of that will be (maybe boss calling you at midnight - or just some thing that you have to get "done this week"). Mileage varies.
There are some very fine HTML parsing modules and they can be used to make much more general solutions. However, I often write one regex to get to a hunk of html that has what I want and then write another regex like below to extract what I want from that hunk of stuff. Write as much code as you need, but don't write more than you have to. And no matter what you do, this HTML stuff is a very "fragile" interface - meaning that your code will break at the whim of the web developer.
my $data =<DATA>;
$data =~ tr/\n/ /; #turn \n's into spaces
my (@names) = $data =~ m/<a[^>]*>(.*?)<\/a>/g;
<a href="foo">Jon.Martinez</a><li>gabba, gabba,
Jones</a><p>Gazebo!</p><a href="baz">Rob Oticus</a><a>Joe Blow</a>
Posts are HTML formatted. Put <p> </p> tags around your paragraphs. Put <code> </code> tags around your code and data!
Titles consisting of a single word are discouraged, and in most cases are disallowed outright.
Read 237057 if you're not absolutely sure you're posting in the right place.
Please read these before you post! —
Posts may use any of the 29281:
You may need to use entities for some characters, as follows. (Exception: Within code tags, you can put the characters literally.)
- a, abbr, b, big, blockquote, br, caption, center, col, colgroup, dd, del, div, dl, dt, em, font, h1, h2, h3, h4, h5, h6, hr, i, ins, li, ol, p, pre, readmore, small, span, spoiler, strike, strong, sub, sup, table, tbody, td, tfoot, th, thead, tr, tt, u, ul, wbr
Link using PerlMonks shortcuts! What shortcuts can I use for linking?
See 17558 and other pages linked from there for more info.
| & || & |
| < || < |
| > || > |
| [ || [ |
| ] || ] ||