Beefy Boxes and Bandwidth Generously Provided by pair Networks
good chemistry is complicated,
and a little bit messy -LW
 
PerlMonks  

Re: Re: Re: Re: Unicode and locales

by mirod (Canon)
on Nov 12, 2002 at 11:51 UTC ( #212254=note: print w/replies, xml ) Need Help??


in reply to Re: Re: Re: Unicode and locales
in thread Unicode and locales

I think it's time for a benchmark here:

Using perl 5.8.0, on Linux (Mandrake 9.0) on a rather fast machine (Athlon dual-processor 1.8):

#!/bin/perl -w use strict; use Benchmark( 'cmpthese'); use Encode; use Text::Iconv; use Unicode::Map8; use Unicode::String qw(utf8); use utf8; my $enc= 'latin1'; my $convert_iconv = Text::Iconv->new( 'utf8', $enc); my $convert_unicode = Unicode::Map8->new ($enc); my $text= <DATA>; chomp $text; # lets just check the output! print "Encode : ", encode("iso-8859-1", $text), "\n"; print "Text::Iconv : ", $convert_iconv->convert( $text), "\n"; print "Unicode::Map8 : ", $convert_unicode->to8 (utf8($text)->ucs2), " +\n"; print "regexp : ", latin1( $text), "\n"; # now benchmark cmpthese( 500000, { 'Encode' => sub { encode("iso-8859-1", $text); + }, 'Text::Iconv' => sub { $convert_iconv->convert( $text +); }, 'Unicode::Map8' => sub { $convert_unicode->to8 (utf8($t +ext)->ucs2); }, 'regexp' => sub { latin1( $text); + }, }); sub latin1 { my $text=shift; $text=~s{([\xc0-\xc3])(.)}{ my $hi = ord($1); my $lo = ord($2); chr((($hi & 0x03) <<6) | ($lo & 0x3F)) }ge; return $text; } __DATA__ texte soupçonné d'être plein de caractÚres accentués

Results:

Encode : texte souponn d'tre plein de caractres accentus Text::Iconv : texte souponn d'tre plein de caractres accentus Unicode::Map8 : texte souponn d'tre plein de caractres accentus regexp : texte souponn d'tre plein de caractres accentus Benchmark: timing 500000 iterations of Encode, Text::Iconv, Unicode::M +ap8, regexp... Encode: 6 wallclock secs ( 4.91 usr + 0.02 sys = 4.93 CPU) @ + 101419.88/s (n=500000) Text::Iconv: 2 wallclock secs ( 2.20 usr + 0.00 sys = 2.20 CPU) @ + 227272.73/s (n=500000) Unicode::Map8: 7 wallclock secs ( 7.66 usr + 0.00 sys = 7.66 CPU) @ + 65274.15/s (n=500000) regexp: 6 wallclock secs ( 5.65 usr + 0.01 sys = 5.66 CPU) @ + 88339.22/s (n=500000) Rate Unicode::Map8 regexp Encode Tex +t::Iconv Unicode::Map8 65274/s -- -26% -36% + -71% regexp 88339/s 35% -- -13% + -61% Encode 101420/s 55% 15% -- + -55% Text::Iconv 227273/s 248% 157% 124% + --

Note: I am not an expert in using Benchmark, so please let me know if my test is flawed.

Log In?
Username:
Password:

What's my password?
Create A New User
Node Status?
node history
Node Type: note [id://212254]
help
Chatterbox?
and the web crawler heard nothing...

How do I use this? | Other CB clients
Other Users?
Others about the Monastery: (4)
As of 2018-12-19 06:09 GMT
Sections?
Information?
Find Nodes?
Leftovers?
    Voting Booth?
    How many stories does it take before you've heard them all?







    Results (85 votes). Check out past polls.

    Notices?