Hi Damian, thank you so much for your great explanation and clarification. Now my assumption is that, from my proteomic analysis, I should remove those genes which are either processed transcripts or pseudogenes, and only keep protein coding genes. As STRING also does not contain transcripts or pseudogenes, and therefore could not identify them. I was wondering if you could also let me know about other cases. For example, in my MS list, there are two NEDD4L proteins with two different IDs which I do not know which one I should keep or remove from the dataset, and what "(fragment)" means in K7ENS6.
protein name in my dataset
K7ENS6 (NEDD4L) ---> E3 ubiquitin-protein ligase NEDD4-like (Fragment) OS=Homo sapiens OX=9606 GN=NEDD4L PE=1 SV=1
A0A1B0GVY1 (NEDD4L) ---> E3 ubiquitin-protein ligase NEDD4-like OS=Homo sapiens OX=9606 GN=NEDD4L PE=1 SV=1
Also, for PRSS2, STRING returns PRSS3P2 which seems to be a different protein.
protein name in my dataset ---> String output
A6XMV9 (PRSS2) ---> PRSS3P2 (Q8NHM4)
ENSG00000275896 (PRSS2) ---> PRSS3P2 (Q8NHM4)
I would highly appreciate your great help. Best, Farah