Re: Top five words by occurrence

by Anonymous Monk
on Jul 19, 2005 at 11:04 UTC ( #476035=note: print w/replies, xml ) Need Help??

in reply to Top five words by occurrence

Is this because I am using split? Is there a better way to go about this.
Yes. You're splitting on whitespace, and there's no whitespace between uncomfortable and its following comma. Instead of splitting on whitespace, you might want to extract sequences of word characters - instead of
my @words = split;
you'd write:
my @words = /\w{5,}/g;
with the added benefit of not having to test of word length anymore, you're extracting words consisting of at least 5 characters.
I am sure I will start missing words that have apostrophes too.
Indeed. Extracting word characters will miss words containing apostrophes. Or hyphens. Extracting words from a random text, where the words can contain punctuation is not a trivial thing to do.


Node Type: note [id://476035]
As of 2020-02-21 11:28 GMT
