Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

While your comment applies on a broader sense (we need to test both for cases in which we expect the algorithm to find similarity and cases in which we expect to find dissimilarity - lest we embrace confirmation bias), I believe your reference to machine learning is a bit off.

In this case, we are not looking at a model training, we are just evaluating the outcome of a (static) algorithm. On the other hand, in the context of ML you should try to include as much information as possible in your training dataset. This is a concept that comes by in many fields, some of them just tangentially related to machine learning (e.g. excitation signals for system identification, speaking both in terms of frequency domain and differential equations).

tl;dr: there is no "training set" to speak in the presented article - hence, concepts such as overfitting and information content (of the training set) do not quite apply here.



I wasn't making a direct comparison between machine learning and this particular task. The article just made me think of that machine learning problem, and I would like to see more examples that show a wide range of results. For instance, there are no examples of false positives -- how often would that happen in an image processor? Probably a lot, but I wouldn't know it based on this particular article.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: