I am interested in performing enrichment analysis on disease-resistance plant proteins. I also have all the protein IDs and sequences extracted from the RiceRelativesGD V4.0 database. However, I think they have changed the fasta headers with their own assigned protein IDs, due to which I am having issues in performing the following:
- In silico gene expression analysis.
- Phylogeny analysis
- Construction of protein-protein network in STRING Database
In the above points, such as in 1st, the gene expression analysis requires gene IDs, which can be automatically fetched by the Gene Expression Atlas, but due to the protein IDs and some species not present in the Atlas, I am unable to do the analysis. Similarly, in 2nd, for phylogeny analysis, I need Multiple Sequence Alignment, which I am unable to do as during P-BLAST, it is not showing any major hits. In addition, in 3rd point, the STRING database requires species names which, in my case, I am unable to find.
Kindly suggest some other tools or ways to do the enrichment analysis on such proteins.
The following are the IDs and their respective species:
scaffold41.166
OGLUM01G32470
ONIVA01G41900
OPUNC01G35650
scaffold76.301
LPERR12G17120
ORUFI01G40100
OMERI01G33870
OPUNC12G16300
OsR498G0203451400.01
scaffold29.669
scaffold15.384
Zlat_10044396
Zlat_10027152
ONIVA12G18960
scaffold25.271
Chr2_RaGOO.1323.mRNA1
Chr1_RaGOO.4384.mRNA1
OB06G28290
Zlat_10021111
*Echinochloa crus galli
Oryza glumaepatula
Oryza nivara
Oryza punctata
Echinochloa crus galli
Leersia perrieri
Oryza rufipogon
Oryza meridionalis
Oryza punctata
Oryza sativa subsp Indica
Echinochloa crus galli
Echinochloa crus galli
Zizania latifolia
Zizania latifolia
Oryza nivara
Echinochloa crus galli
Oryza sativa f spontanea indica group
Oryza sativa f spontanea indica group
Oryza branchyantha
Zizania latifolia*
gene-ontology