The short story is that I tried for 500 images and then got 5.1% error on the test set. Looking at my mistakes, I then tried to differentiate two sources of error: a genuine problem with the dataset (e.g. many correct answers in an image, or incorrect label), and errors I felt could be eliminated by an ensemble of very committed humans who were even better than me at classifying dogs :) And that optimistic error rate is approx 3%.
The short story is that I tried for 500 images and then got 5.1% error on the test set. Looking at my mistakes, I then tried to differentiate two sources of error: a genuine problem with the dataset (e.g. many correct answers in an image, or incorrect label), and errors I felt could be eliminated by an ensemble of very committed humans who were even better than me at classifying dogs :) And that optimistic error rate is approx 3%.