### Comment on

 Need Help??

Here's a fairly simple method to measure similarity of strings. Given \$one and \$two, apply text compression to each and to their concatenation. The ratio of the size of them compressed together to the sum of the separately compressed sizes measures their similarity. The smaller, the closer.

```#!/usr/bin/perl

use Compress::Zlib 'compress';

# Usage:   \$arrayref = similarity( LIST)
# Returns: AoA reference to string similarity table for LIST
sub similarity {
my (%single, @ret) = map {\$_ => length compress \$_} @_;
for my \$this (@_) {
push @ret, [
map {
(length compress \$this . \$_)
/ (\$single{\$this} + \$single{\$_})
} @_
];
}
\@ret;
}

my @titles = (
q(The Last Public Hanging In Old West Virginia - Flatt and Scruggs
+),
q(Flatt_and_Scruggs__The_Last_Public_Hanging_In_Old_West_Virginia)
+,
q(Rainy Day Woman Number 12 and 35 - Flatt and Scruggs),
q(Rainy Day Woman Number Twelve and Thirty-five - Bob Dylan),
);

my \$results = similarity @titles;

for my \$this (@\$results) {
print pack('A6' x @\$this, map {sprintf '%4.3f', \$_} @\$this), \$/;
}

__END__

0.529 0.715 0.784 0.841
0.708 0.529 0.887 0.870
0.784 0.863 0.536 0.748
0.848 0.863 0.739 0.532
Note that 0.500 is the ideal minimum for that, so subtracting .5 from those would give more impressive differences.

I saw this technique described in a SciAm recently. Will update if I can find out which.

After Compline,
Zaxo

In reply to Re: similar texts !? by Zaxo
in thread similar texts !? by bugsbunny

Title:
Use:  <p> text here (a paragraph) </p>
and:  <code> code here </code>
to format your post; it's "PerlMonks-approved HTML":

• Posts are HTML formatted. Put <p> </p> tags around your paragraphs. Put <code> </code> tags around your code and data!
• Titles consisting of a single word are discouraged, and in most cases are disallowed outright.
• Read Where should I post X? if you're not absolutely sure you're posting in the right place.
• Posts may use any of the Perl Monks Approved HTML tags:
a, abbr, b, big, blockquote, br, caption, center, col, colgroup, dd, del, div, dl, dt, em, font, h1, h2, h3, h4, h5, h6, hr, i, ins, li, ol, p, pre, readmore, small, span, spoiler, strike, strong, sub, sup, table, tbody, td, tfoot, th, thead, tr, tt, u, ul, wbr
• You may need to use entities for some characters, as follows. (Exception: Within code tags, you can put the characters literally.)
 For: Use: & & < < > > [ [ ] ]
• Link using PerlMonks shortcuts! What shortcuts can I use for linking?

Create A New User
Chatterbox?
and all is quiet...

How do I use this? | Other CB clients
Other Users?