I'm a student learning bioinformatics from scratch, and our lab is planning to perform ATAC-seq on rare population cells in HSCs ( around 5k-20K ) for the first time ever. I need your help. The cells will be sorted before the transposase reaction, but we are having difficulties in getting a comparable or consistent number of cells across the conditions because the counting method is not reliable with a small number of cells. ( 4 biological replicates on control and sample, no technical replicates ) My boss is skeptical about making libraries out of samples with different cell numbers. I'm wondering if this is serious issues and if we can address the negative consequence of this later through processing. I'm sorry if the question was difficult to understand due to poor English.
2 answers
Your English is excellent - it's more coherent and better edited than most native speakers on the site.
As for your question, the number of cells isn't as important as DNA quantity/quality at the end of your prep. Ultimately, I wouldn't worry much about this, as ATAC has an excellent signal to noise ratio and pretty much any processing pipeline accounts for biases like sequencing depth. This is particularly true if you're only comparing whether a region is accessible or not (i.e. a binary peak presence/absence study) rather than trying to compare signal across samples (quantitative comparisons, which are a tougher nut to crack).
Seconding jared.andrews07, your english is excellent, no worries.
As for your question, ATAC-seq is pretty robust towards variable cell numbers. Much more important than the absolute number is the viability of the cells and the absence of death cells which is why I would always check by microscope using e.g. Tryphan Blue staining on a tiny aliquot. Treat your cells gently during the preparation, always keep on ice with FCS. Sorting into cold PBS + 2%FCS is recommended. Try to keep the sorting time as short as possible. c-Kit enrichment prior to the sort definitely helps with that, e.g. using magnetic CD117 beads.
I've generated ATAC-seq data from FACS-sorted rare(r) cell populations myself recently ranging from somewhat ~ 8k to 50k cells and the results were not confounded by the variable cell numbers. I recommend you use the OmniATAC protocol that was recently published and uses NP-40, Digitonin and Tween-20 as detergents which I can recommend. We generated quite many ATAC-seq datasets in the past but the best so far with this protocol in terms of low read duplication rate, large number of callable peaks, high FRiPs (fraction of reads per callable peaks per sample) ~ 40-60%, high library complexity. We typically do 2x50bp sequencing.
If you need further advice, feel free to ask. You can also join the Biostars slack which might make discussion a little easier: biostar.slack.com: Chat for the biostars community
Log in to answer this question.