Beefy Boxes and Bandwidth Generously Provided by pair Networks
Problems? Is your data what you think it is?
 
PerlMonks  

Re^12: Text::CSV encoding parse()

by slugger415 (Monk)
on Aug 21, 2019 at 17:56 UTC ( #11104827=note: print w/replies, xml ) Need Help??


in reply to Re^11: Text::CSV encoding parse()
in thread Text::CSV encoding parse()

I can't really give you the whole shebang but here are a couple of URLs, including the first one which has the spanish characters.

https://www.ibm.com/support/knowledgecenter/es/search/¿Cuales son las +partes de una cadena de conexión??scope=SSGU8G_12.1.0|https://www.ibm +.com/support/knowledgecenter/es/SSGU8G_12.1.0/com.ibm.jdbc_pg.doc/ids +_jdbc_011.htm|0|1|1|0 https://www.ibm.com/support/knowledgecenter/search/onsmsync?scope=SSGU +8G_12.1.0|https://www.ibm.com/support/knowledgecenter/SSGU8G_12.1.0/c +om.ibm.sec.doc/ids_lb_002.htm|1|1|1|1

Thanks!

Replies are listed 'Best First'.
Re^13: Text::CSV encoding parse()
by Tux (Abbot) on Aug 22, 2019 at 08:48 UTC

    The problem with you pasting the data here inside the code tags, does not reflect the binary compatibility of your actual data.

    If I download this snippet, the code works fine:

    $ cat test.csv https://www.ibm.com/support/knowledgecenter/es/search/&#65533;Cuales s +on las partes de una cadena de conexi&#65533;n??scope=SSGU8G_12.1.0|h +ttps://www.ibm.com/support/knowledgecenter/es/SSGU8G_12.1.0/com.ibm.j +dbc_pg.doc/ids_jdbc_011.htm|0|1|1|0 https://www.ibm.com/support/knowledgecenter/search/onsmsync?scope=SSGU +8G_12.1.0|https://www.ibm.com/support/knowledgecenter/SSGU8G_12.1.0/c +om.ibm.sec.doc/ids_lb_002.htm|1|1|1|1 $ perl -CEO -MData::Peek -MText::CSV_XS -wE'my$c=Text::CSV_XS->new({se +p_char=>"|",auto_diag=>1,binary=>1});while(<>){$c->parse($_);DPeek fo +r$c->fields}' test.csv PV("https://www.ibm.com/support/knowledgecenter/es/search/\277Cuales s +on las partes de una cadena de conexi\363n??scope=SSGU8G_12.1"...\0) PV("https://www.ibm.com/support/knowledgecenter/es/SSGU8G_12.1.0/com.i +bm.jdbc_pg.doc/ids_jdbc_011.htm"\0) PV("0"\0) PV("1"\0) PV("1"\0) PV("0"\0) PV("https://www.ibm.com/support/knowledgecenter/search/onsmsync?scope= +SSGU8G_12.1.0"\0) PV("https://www.ibm.com/support/knowledgecenter/SSGU8G_12.1.0/com.ibm. +sec.doc/ids_lb_002.htm"\0) PV("1"\0) PV("1"\0) PV("1"\0) PV("1"\0)

    The *output* is, as you could see, iso-8859-1 (latin1) instead of your expected utf-8, because the source data is iso-8859-1 (or a variety thereof) and does not require an upgrade to utf-8.

    You can however make the data utf-8 by decoding your source data:

    $ perl -CEO -MEncode=decode -MData::Peek -MText::CSV_XS -wE'my$c=Text: +:CSV_XS->new({sep_char=>"|",auto_diag=>1,binary=>1});while(<>){$c->pa +rse(decode("utf-8",$_));DPeek for$c->fields}' test.csv PV("https://www.ibm.com/support/knowledgecenter/es/search/\357\277\275 +Cuales son las partes de una cadena de conexi\357\277\275n??s"...\0) +[UTF8 "https://www.ibm.com/support/knowledgecenter/es/search/\x{fffd} +Cuales son las partes de una cadena de conexi\x{fffd}n??scope=SSGU8G_ +12.1.0"] PV("https://www.ibm.com/support/knowledgecenter/es/SSGU8G_12.1.0/com.i +bm.jdbc_pg.doc/ids_jdbc_011.htm"\0) PV("0"\0) PV("1"\0) PV("1"\0) PV("0"\0) PV("https://www.ibm.com/support/knowledgecenter/search/onsmsync?scope= +SSGU8G_12.1.0"\0) PV("https://www.ibm.com/support/knowledgecenter/SSGU8G_12.1.0/com.ibm. +sec.doc/ids_lb_002.htm"\0) PV("1"\0) PV("1"\0) PV("1"\0) PV("1"\0)

    Note that Text::CSV_XS *only* decodes to utf-8 if it needs to or is explicitly told to: it needs to be able to deal with pure binary data.


    Enjoy, Have FUN! H.Merijn

Log In?
Username:
Password:

What's my password?
Create A New User
Node Status?
node history
Node Type: note [id://11104827]
help
Chatterbox?
and the web crawler heard nothing...

How do I use this? | Other CB clients
Other Users?
Others chanting in the Monastery: (4)
As of 2019-10-18 04:25 GMT
Sections?
Information?
Find Nodes?
Leftovers?
    Voting Booth?
    Notices?