Hi,
I am working with annotating my metagenomes through the MG-RAST pipeline, annotating them with taxonomy and function. For this study we are interested in understanding phosphorus metabolism. To summarise, MG-RAST has a filtering option where you can filter by different levels in subsystems or refseq. When I use the filtering option to pull out phosphorus metabolism, I recover more functions identified as phosphours metabolism, and the abundnace of some functions related to phosphorus metabolism are consistantly lower when I do this, as compared to simply looking at all annotated functions without any filters. I can understand that some functions will have multiple higher level classifications depending, but I don't understand why the abundance of some functions would change when filtering. Can someone help me understand what is happening?
I submitted my raw sequencing reads to MG-RAST and joined paired reads through the upload tool. I analyzed my data using SEED's subsystem for function and refseq for taxonomy. I left the settings at default and never changed them (e.val cutoff of 5, % ident of 60, min length of 15, min abundance of 1, using representative hit). I will use my E81 sample as an example to describe the issue I am encountering. When I create a filter for phosphorus metabolism (seeds lvl 1) I get 13,901 reads identified as P. metabolism. However, when I do not filter for anything and simply download subsystems annotations at function level as a .tsv I get a different result. When you download subsystems annotations at function level, it gives all levels that a gene may fall into, for example: Phosphorus Metabolism NULL P_uptake_(cyanobacteria) Phosphate ABC transporter, periplasmic phosphate-binding protein PstS (TC 3.A.1.7.1) When I pull out all of the phosphorus metabolism lvl 1 functions from this TSV, I get 14,930 reads identified as being phosphorus metabolism. Doing this also recovers less functions than when filtering, for example, 'Sodium-dependent phosphate transporter' was simply classed as membrane transport at lvl one, but when sorting for phos. metabolism it shows up as a phosphorus metabolism function. It would make sense that this function would be classified as both phosphorus metabolism and membrane transport, because it transports phosphorus across cell membranes. What would determine though what lvl. 1 function is displayed when downloading the data? Another issue I've noticed is that the number of reads for some functions changes when sorting by seeds subsystems lvl 1. When I sort by phos.metabolism the number of reads that assigns to alkaline phosphatase (EC 3.1.3.1) stays the same (1123 reads) as when I don't sort by anything. But when I sort by phosphorus metabolism, the number of reads identified as Phosphate ABC transporter, periplasmic phosphate-binding protein PstS (TC 3.A.1.7.1) changes from 1113 reads (unsorted) to 912 reads (sorted by phos.metabolism). Why might something like this happen, an what should I do in response? How many reads should I consider to be associated with phosphorus metabolism? How many reads belong to PtsS, why are there less reads when sorting by phosphorus metabolism?
The other functions with this issue are : inorganic pyrophosphatase (EC 3.6.1.1) Phosphate ABC transporter, periplasmic phosphate-binding protein PstS (TC 3.A.1.7.1) Phosphate regulon transcriptional regulatory protein PhoB (SphR) Phosphate starvation-inducible ATPase PhoH with RNA binding motif Polyphosphate kinase (EC 2.7.4.1) Putative outer membrane TonB-dependent receptor associated with haemagglutinin family outer membrane protein
All of these have lower abundance values when sorting by phosphorus metabolism. Other phosphorus metabolism related functions do not change when sorting by phosphorus metabolism.
I noticed a similar issue with nitrogen metabolism functions, just used phosphorus metabolism as an example.
0 answers
No answers yet.
Log in to answer this question.