How to deal with extreme values of k in RUV-seq?
I discovered that a certain lab is abusing RUVseq by using a high value of k to squash variance and obtain a large number of DEGs. Have you ever seen this, and if so, what did you do?
ruvseq
• 1,654 views
•
link
updated
by
LChart
•
written
by
telroyjatter
0 answers
No answers yet.
Log in to answer this question.
More posts like this
-
How to extract DEGs from GEO database
written by harrydo1892 •Hello everyone, I’m a newbie to bioinformatics and I’m trying to extract Differentially Expressed Genes (DEGs) from the GEO database for my research. I’ve tried …
-
GO analysis: Why did gene set with high cluster frequency from up-regulation category will disappea…
written by greymanTo further clarify, I did two GO analysis: (i) up, (ii)down, and (ii) all DEGs. The output from up regulation category showed 120 gene sets, …
-
How to interpret Site Frequency Spectrum? (SFS, allele frequency spectrum)
written by jaredmeyers •I have a Site Frequency Spectrum graph with K value on the x-axis and Number of Variants on the y-axis. I'm not sure what the …
-
high sequence duplication ddRAD
written by gubrinsHello, I'm relatively new to NGS analyses. I'm working with single-read ddRAD data of a non-model species and we just obtained the fastq files. The …
-
Effect of sample size on log fold change in gene expression analysis
written by Gene_MMP8I am doing gene expression analysis using a set of 228 patients (52 positive and 176 negatives). I found a list of 17 DEGs that …
-
Is it ok to use genotype replicates for association studies?
written by rimgubaevHello everyone! I got 90 individuals genotyped in 3 replicates (genotyping by sequencing), so as a result, my vcf file contains information on 270 samples. …
-
Normalizing DNA sequencing reads using DNA spike in [sanity check]
written by nattzy94 •My main goal is to quantify absolute abundance of a known bacterial sample. My samples have either E. coli, K. pneumoniae or both. These are …
-
RNA-SEQ: Which pairwise comparison for cells applied with low and high levels of mechanical forces?
written by Muha0216 •hi guys, i am a total noob when it comes to RNA-seq and nobody in my lab has done this technique before since we are …
-
Appropriate to apply RUVSeq to exon bin counts for DEXSeq?
written by AdamcHi, I have been working on an RNA-Seq dataset on which I applied RUVSeq (ruvr method) at the gene-count level to account for some unexpected …
-
Standard error of cross-validation estimates in software ADMIXTURE
written by manfred.mayer5 •Hello, I'm currently using the software "*ADMIXTURE*" to calculate the most probable number of genetic groups (K) in a large panel of 38 landraces (old …
Analyze the same data with a different method and compare the results? One can abuse a wrench by using it as a hammer under certain circumstances. Many genomic methods are much like a screen, filtering candidate genes by some often arbitrary threshold, with the idea to then pursue those genes further with other methods (i.e. it's a hypothesis generating method). In some cases, experiment conditions dictate this is all one can do, in other cases it might simply be sloppy, flawed thinking or design, and a colossal waste of time and resources. More information would be needed to make that call (though it appears you already have?). FWIW, you might alter the title of your post to be more objective...something like "how to deal with extreme values of k in RUV Seq", or something more directly related to your concern (Validate? Justify? Explore? something something of k-values).
Thank you, I changed the title of the post. I could re-analyze the data and demonstrate the value of k chosen is much too high, but I guess I'm wondering what others' experiences have been with this.
Thanks again.
If you think it is a problem you could make a figure that plots number of DEGs with k from 0 to what they use. If your concerns are true you might see sort of a trend that number of DEGs is a function of k, I have seen things like that before. Problem eith methods like RUV is that there is no (imo) automated yet robust way go estimate a good k.
It's not clear based on this post what you are seeking to accomplish. Are you asking how to issue a formal rebuttal? How to warn a colleague about the potential misuse of statistics?
By the way, large values of
kin RUVseq should not generate a large number of DEGs when properly applied as each factor removes a degree of freedom. Is RUVseq being misapplied (i.e., by using the corrected counts in DE instead of the original counts with factors entering as covariates)?