Thank you for the response. You have clarified a big doubt. Yesterday I used Cd-hit for this purpose and clustered sequences at 0.85 threshold with no other criteria (i plan to check some lower and higher thresholds too). On most test sequences, i got significantly higher bit-scores with clustering compared to the non-clustered data.
my average bit scores are somewhere around 300 for whole data and 350 for clustered sequences.
Many proteins in my database have only 5 - 10 representative sequences. Will it be a good idea to make profiles for such proteins ?
Can you please also tell me if hmmcalibrate function is still available in hmmer 3, as it was in hmmer 2 but i cant find any reference to it in hmmer 3 docs ?
I will be really thankful if you can answer any of these questions. Even if you dont, I would still like to thank you a lot for the answer to the original question.