Beefy Boxes and Bandwidth Generously Provided by pair Networks
laziness, impatience, and hubris

Re: How to deal with Huge data

by roboticus (Chancellor)
on Jan 23, 2007 at 14:09 UTC ( #596095=note: print w/replies, xml ) Need Help??

in reply to How to deal with Huge data


Another way you might be able to do the job is with a file merge. To do so, sort both files on the key(s) of interest, then read records in order and merge them as appropriate.


#!/usr/bin/perl -w use strict; use warnings; open F1, 'sort -k3 mergefile.1|' or die "opening file 1"; open F2, 'sort -k2 mergefile.2|' or die "opening file 2"; open OUF, '>', 'mergefile.out' or die "opening output file"; my @in1; my @in2; sub getrec1 { @in1 = (); if (!eof(F1)) { (@in1) = split /\t/, <F1>; chomp $in1[2]; } } sub getrec2 { @in2 = (); if (!eof(F2)) { (@in2) = split /\t/, <F2>; chomp $in2[2]; } } sub write1 { print OUF "$in1[2]\t$in1[0]\t$in1[1]\tnull\tnull\n"; getrec1; } sub write2 { print OUF "$in2[1]\tnull\tnull\t$in2[0]\t$in2[2]\n"; getrec2; } sub writeboth { print OUF "$in1[2]\t$in1[0]\t$in1[1]\t$in2[0]\t$in2[2]\n"; getrec1; getrec2; } # Prime the pump getrec1; getrec2; while (1) { last if $#in1<0 and $#in2<0; if ($#in1<0 or $#in2<0) { # Only one file is left... write2 if $#in1<0; write1 if $#in2<0; } elsif ($in1[2] eq $in2[1]) { # Matching records, merge & write 'em writeboth; } elsif ($in1[2] lt $in2[1]) { # unmatched item in file 1, write it & get next rec write1; } else { # unmatched item in file 2, write it & get next rec write2; } }
Example output:

root@swill ~/PerlMonks $ cat mergefile.1 15 20 foo 22 30 bar 30 33 baz 14 22 fubar root@swill ~/PerlMonks $ cat mergefile.2 alpha baz 17.30 gamma foobar 22.35 gamma bar 19.01 delta fromish 33.03 sigma bear 14.56 root@swill ~/PerlMonks $ ./ root@swill ~/PerlMonks $ cat mergefile.out bar 22 30 gamma 19.01 baz 30 33 alpha 17.30 bear null null sigma 14.56 foo 15 20 null null foobar null null gamma 22.35 fromish null null delta 33.03 fubar 14 22 null null root@swill ~/PerlMonks $

Log In?

What's my password?
Create A New User
Node Status?
node history
Node Type: note [id://596095]
[Discipulus]: planetscape welcome back! (or is well comeback?)
[Discipulus]: fore every 2 good old monks that come back, we will accept 1 clinical
[planetscape]: Might be "well, come back!" Discipulus ;-)
[stonecolddevin]: that's a 2.0 k/d ratio, i'll take it
[choroba]: #cbstream still down?
[erix]: yeah, annoying
[choroba]: last ambrus: 1 week ago :-(

How do I use this? | Other CB clients
Other Users?
Others pondering the Monastery: (10)
As of 2017-06-22 21:02 GMT
Find Nodes?
    Voting Booth?
    How many monitors do you use while coding?

    Results (530 votes). Check out past polls.