Hi, thank you for taking your time to help me out.
Within this protein they identified regions with amino acid repeats whose number vary between organisms and seems to correlate with body size. Hence why I am gathering many sequences from species of different taxa like primates, rodents, carnivora with the idea to then carry out MSA to identify these regions and count the number of repeating amino acids in each species and correlate it to body size and maybe other sensible adaptations.
About validating the sequences predicted with QuickProt, I was thinking about creating a phylogenetic tree with the QuickProt predictions and with sequences from RefSeq genomes. My idea was that, say I get 10 new primate sequences, I would then create a phylogenetic tree with a few curated sequences (Human, Rhesus Monkey, Mice, Rat and others from other taxonomic groups). My expectation would be that if the predictions are correct, they would cluster with those from the same taxonomic group. However, I am also skeptical about validating them this way. Because I am using the same sequences (human and rhesus monkey) for protein-to-genome alignment and the gene itself is quite conserved across organisms, hence it seems almost obvious that they are bound to cluster together.