Thank you for your insight. I initially had 28k UMIs, but after running the clusterer from UMI-tools, only 2.8k UMI clusters remained.
Do you know if this indicate that the true number of unique molecules is much lower than expected, or that the sequencing depth was insufficient and more sequencing is needed?
What kind of data is this?
lineage tracing with cas9, we try to capture specific genomic region contain gRNA
There is an extensive answer below that may cover all use cases but it may be interesting to find out where the UMI's are coming from? (adapters for library prep or some other way).