Hi there,
I am working on a Xenium dataset that includes two housekeeping genes in the panel. I wonder how to make use of those genes. My ideas were to use them somehow in the normalization (as normalization factors?) and/or for doublet detection. Indeed those could provide a baseline of transcripts level useful to compare cells. This is especially important as my Xenium samples were generated from sectioning slides of 5um which will contain some cells entirely and some others are expected to be cropped.
Anyone ever used housekeeping genes for such purposes?
Thank you very much for any insights!
1 answer
Two genes is too few to normalise on. A Xenium housekeeping gene gives you maybe tens of counts per cell, so a size factor built on two of them carries huge Poisson noise - total counts across the panel is far better behaved and is what people use. Housekeeping genes in a panel are really slide-level chemistry QC, not per-cell normalisation.
They won't help with doublets either, since a doublet has roughly 2x everything and that's already in total counts. And in imaging-based spatial your artefact usually isn't a droplet doublet but segmentation error - transcripts assigned to a neighbour, or two cells merged. Catch that with cell area versus total counts and cells co-expressing mutually exclusive markers, or re-segment with Baysor or Proseg.
Normalisation can't fix the 5 um cropping either - a half cell genuinely has fewer transcripts, so filter on area and read the data as composition rather than absolute expression.
Log in to answer this question.