comment on

Hi bibliophile,

It is a very good thought, indeed. However, you would need to think carefully and extensively on what kind of features the articles you are interested in have in common. You could use some sort of data clustering (FCM, maybe?) to help you with this task. You would then need to find a way to extract those features consistently. Finally, you could use a classifier to filter the raw data and present you only with the stuff you are interested in. When you design the classifier, try to incorporate a confidence index that tells you how reliable the results are. In this way, you could play with the outputs until you are happy with the results. Does it make sense?

Cheers,

lin0

In reply to Re^2: RFC: Presentation on Machine Learning with Perl by lin0
in thread RFC: Presentation on Machine Learning with Perl by lin0

Are you posting in the right place? Check out Where do I post X? to know for sure.
Posts may use any of the Perl Monks Approved HTML tags. Currently these include the following:
<code> <a> <b> <big> <blockquote> <br /> <dd> <dl> <dt> <em> <font> <h1> <h2> <h3> <h4> <h5> <h6> <hr /> <i> <li> <nbsp> <ol> <p> <small> <strike> <strong> <sub> <sup> <table> <td> <th> <tr> <tt> <u> <ul>
Snippets of code should be wrapped in <code> tags not <pre> tags. In fact, <pre> tags should generally be avoided. If they must be used, extreme care should be taken to ensure that their contents do not have long lines (<70 chars), in order to prevent horizontal scrolling (and possible janitor intervention).
Want more info? How to link or How to display code and escape characters are good places to start.


Don't ask to ask, just ask
	PerlMonks