http://www.perlmonks.org?node_id=404808


in reply to Analysing five years of blogging

Being a BioGeek, you must surely know about Markov Chains. Perhaps you could use that volume of writing to generate transition state statistics from word to word and see how well it performs at writing articles similar to Tom.

Other options include simple word frequency counts, and possibly analysis of the domains to which he links (just a simple frequency count based on the second-level domain name might be interesting).

Unfortunately I'm working on my thesis, so I don't have spare time to play with more data. Best of luck with the project. :)