Beefy Boxes and Bandwidth Generously Provided by pair Networks
Your skill will accomplish
what the force of many cannot
 
PerlMonks  

Comment on

( #3333=superdoc: print w/ replies, xml ) Need Help??

The first thing I’d say about this program is ... “always use real variables,” not “implied” stuff like $_ which can very easily get away from you.   Then, “start with a test case.” One example string, that you can verify gets correctly parsed by the regular-expression you intend to use.   You can even use a “Perl one-liner” like this:

perl -e 'my $str = "chr1:4777082-4777141"; \ my ($foo) = $str =~ /([0-9]+)/; print "foo is $foo\n";' foo is 1 ... oops, that's not right ... mike$ perl -e 'my $str = "chr1:4777082-4777141"; \ my ($foo) = $str =~ /[:]([0-9]+)/; print "foo is $foo\n";' foo is 4777082 ... correct.

Now, write your program, something like:

while (my $str = <INFILE>) { my ($foo) = $str =~ /[:]([0-9]+)/; # WE TESTED THIS die "Something's Wrong with $str!" unless ($foo); # NEVER ASSUME, NEVER ASSUME print OUTFILE "FILE=$foo\n"; # NOTICE DOUBLE-QUOTES }

We used the one-liners to verify the actual regular-expression parsing, to quickly get it right (and to uncover a subtle bug in the first attempt), then wrote code that is above all else, clear to do the actual work.

Within that program, we also added a die statement that will cause the program to test its assumption that every single record of the input will be handled correctly.   (We could make this test even-more aggressive if we knew that every record of the input file should contain seven digits, preceded by a ":" and followed by a "-", with: (see bold-faced parts)
/[:]([0-9]{7})[-]/)
Thus, the very fact that the program runs normally to completion, with a very stringent regular-expression that must be matched every time, is a strong indication that both the input data and the resulting output are correct.


In reply to Re: handling files using regular expression by sundialsvc4
in thread handling files using regular expression by rocketperl

Title:
Use:  <p> text here (a paragraph) </p>
and:  <code> code here </code>
to format your post; it's "PerlMonks-approved HTML":



  • Posts are HTML formatted. Put <p> </p> tags around your paragraphs. Put <code> </code> tags around your code and data!
  • Read Where should I post X? if you're not absolutely sure you're posting in the right place.
  • Please read these before you post! —
  • Posts may use any of the Perl Monks Approved HTML tags:
    a, abbr, b, big, blockquote, br, caption, center, col, colgroup, dd, del, div, dl, dt, em, font, h1, h2, h3, h4, h5, h6, hr, i, ins, li, ol, p, pre, readmore, small, span, spoiler, strike, strong, sub, sup, table, tbody, td, tfoot, th, thead, tr, tt, u, ul, wbr
  • You may need to use entities for some characters, as follows. (Exception: Within code tags, you can put the characters literally.)
            For:     Use:
    & &amp;
    < &lt;
    > &gt;
    [ &#91;
    ] &#93;
  • Link using PerlMonks shortcuts! What shortcuts can I use for linking?
  • See Writeup Formatting Tips and other pages linked from there for more info.
  • Log In?
    Username:
    Password:

    What's my password?
    Create A New User
    Chatterbox?
    and the web crawler heard nothing...

    How do I use this? | Other CB clients
    Other Users?
    Others about the Monastery: (14)
    As of 2015-07-06 13:11 GMT
    Sections?
    Information?
    Find Nodes?
    Leftovers?
      Voting Booth?

      The top three priorities of my open tasks are (in descending order of likelihood to be worked on) ...









      Results (74 votes), past polls