I think a better formulation is that the test informs us about the likelihood that the observed number of genes (20 out of 60 in the example above) could be occurring by random chance alone considering the frequency of annotations in the database.
The test is about the enrichment, not the selection of the genes. The gene selection process may have been valid and with no error whatsoever. The selection does not factor into whether the genes are enriched for a given annotation or not.
The annotation may fail to pass for both reasons: the selection was incorrect, or the genes are not actually enriched for the annotation.
I think the latter part of the answer has the same interpretation that I give, but you do start out talking about a random selection of the genes.