This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Which read counts file from miRDeep2 should be considered for differential expression analysis?

Hello everyone,

I would appreciate some help with my data. I want to do Differential Expression profiling for miRNA samples.

Using the miRDeep2 program and algorithm, I'm in trouble when dealing with results to make a proper table in order to use it as input for DE software such as EdgeR or DEseq2.

Which .csv should be used for read count?

miRNAs_expressed_all_samples_15_09_2017_t_07_26_27.csv

or

miRNA_expressed.csv

Also, there are duplicate values in the miRNA_expressed.csv. Should I consider the average or only one of them with highest value?

miRNA read_count precursor

ssc-let-7a 66271 ssc-let-7a-1

ssc-let-7a 66388 ssc-let-7a-2

ssc-let-7c 22747 ssc-let-7c

mirna mirdeep2 mirbase rna-seq

1 answer

You should use the miRNAs_expressed_all_samples_15_09_2017_t_07_26_27.csv as input for edgeR or DESeq.

miDeep2 runs bowtie with -a (report all alignments) and --best --strata, which means the same read will be mapped to multiple precursor miRNAs, so I believe you should use only one of the counts for each mature miRNA - I would use the one with higher counts.

Log in to answer this question.