Thank you!!
One more question:
I'm thinking I will keep only the longest transcript for each gene based on the GTF and then extract the protein sequences. I assume this preprocessing is required? I have the following formats, and the mouse faa seems to have proteins for multiple transcripts of the same gene.
For example.
Mouse
ENSMUSP00000070648.4|ENSMUST00000070533.4|ENSMUSG00000051951.5|OTTMUSG00000026353.2|OTTMUST00000065166.1|Xkr4-001|Xkr4|647
Chicken
NP_001001127.1 endothelin receptor type B precursor [Gallus gallus]