Thank you, sir! I will try both methods and observe any differences and correlations. I will carefully read the paper you cited. Before delving into it, I have one more question: How do you determine which approach is correct and more suitable (removing or retaining duplicate reads) if you observe differences in downstream analysis (e.g., differentially expressed genes are different)? I'm not sure if there are different considerations for single-cell and bulk data. I know it's too early to discuss since I haven't obtained any results, but I still wish to seek your recommendations.
In the paper you cited, they concluded that solely removing duplicated reads based on their mapping coordinates would introduce substantial bias. Does this mean that not deduplicating reads would be better in cases where UMI is not available, such as SMART-seq2 data?