Perfect, thanks!
Also, this might be worth opening as a new question, but could I get you opinion on the Ortholog conjecture itself? After asking the question above, I found a paper which is essentially saying the assumption that orthologs share function is flawed, which would be slightly worrying! When I digged into it a little there seems to be people on both sides of the fence.
Moses Stamboulian, Rafael F Guerrero, Matthew W Hahn, Predrag Radivojac, The ortholog conjecture revisited: the value of orthologs and paralogs in function prediction, Bioinformatics, Volume 36, Issue Supplement_1, July 2020, Pages i219–i226, https://doi.org/10.1093/bioinformatics/btaa468
I think when they started Pfam, the idea was to have entries for individual protein families - thus the name. This in turn means that the intent was to have only orthologs as members of each Pfam entry. However, this is really difficult to do without experiments, and HMMs are so sensitive that they inevitably pull in matches to non-orthologs. I think that's why Pfam eventually ended up being a mix of mostly protein families (orthologs) and superfamilies (a mix of ortho-, para- and homologs).
Ah very interesting, thanks for your comment! Would you consider all members of a Pfam (= a protein family, not a superfamily, unless I misunderstand) to be orthologous then? I thought that they were grouped into Pfams according to domain sharing, which doesn't necessarily imply orthology
Pfam is domain-centric, so I was taking about orthology at the domain level. In many cases, especially in prokaryotes, proteins have a single domain, so the domain orthology is the same as orthology in the general sense.
There is a large number of single-domain and single-copy proteins in prokaryotes, which by definition means that all their members are orthologs. Just about all ribosomal proteins, translation factors, single-chain polymerases, etc belong to orthologous Pfam entries.