Beefy Boxes and Bandwidth Generously Provided by pair Networks
more useful options
 
PerlMonks  

Comment on

( #3333=superdoc: print w/ replies, xml ) Need Help??

Magic numbers do not solve the problem completely. Old versions of MS Word use a file format that is also used by other products. For example, old Crystal Reports files are often misidentified as MS Word files.

But the OP gave a nice clue: *.docx -- that means a newer MS Word, stuffing XML into a ZIP file. So, if we talk ONLY about newer MS Word files in zipped XML format, the file can be testet easily: Try to open the file using Archive::ZIP (it is a ZIP file, after all). Look inside the archive, try to find an XML file with the name MS Word uses for the content, unpack that file from the archive. Use an XML parser to see if the file has one of the well-known type declarations for MS Word (or similar products, if you want to allow OpenOffice and friends). Should any of those steps fail, the input file is not a valid *.docx MS Word file.

If also old MS Word files (*.doc) have to be processed, you need a second test routine that can properly detect the old MS Word binary dump formats. There are several, one for each version. Testing "magic numbers" works well most of the times, but you may have false positives (see above).

Word can also read and write Rich Text Format files (*.rtf). If this format has to be processed, you need a third test. Again, "magic numbers" may work here. CPAN has some RTF readers, using them to test for a valid file should give less false positives.

Alexander

--
Today I will gladly share my knowledge and experience, for there are no sweeter words than "I told you so". ;-)

In reply to Re^2: OLE and WORD docs by afoken
in thread OLE and WORD docs by forinti

Title:
Use:  <p> text here (a paragraph) </p>
and:  <code> code here </code>
to format your post; it's "PerlMonks-approved HTML":



  • Posts are HTML formatted. Put <p> </p> tags around your paragraphs. Put <code> </code> tags around your code and data!
  • Read Where should I post X? if you're not absolutely sure you're posting in the right place.
  • Please read these before you post! —
  • Posts may use any of the Perl Monks Approved HTML tags:
    a, abbr, b, big, blockquote, br, caption, center, col, colgroup, dd, del, div, dl, dt, em, font, h1, h2, h3, h4, h5, h6, hr, i, ins, li, ol, p, pre, readmore, small, span, spoiler, strike, strong, sub, sup, table, tbody, td, tfoot, th, thead, tr, tt, u, ul, wbr
  • Outside of code tags, you may need to use entities for some characters:
            For:     Use:
    & &amp;
    < &lt;
    > &gt;
    [ &#91;
    ] &#93;
  • Link using PerlMonks shortcuts! What shortcuts can I use for linking?
  • See Writeup Formatting Tips and other pages linked from there for more info.
  • Log In?
    Username:
    Password:

    What's my password?
    Create A New User
    Chatterbox?
    and the web crawler heard nothing...

    How do I use this? | Other CB clients
    Other Users?
    Others wandering the Monastery: (6)
    As of 2014-08-02 07:33 GMT
    Sections?
    Information?
    Find Nodes?
    Leftovers?
      Voting Booth?

      Who would be the most fun to work for?















      Results (55 votes), past polls