Dear Kevin, thank you for your thoughtful answer and practical suggestions. I really appreciate your time.
Important note: I should point out that sequence similarity ≠ homology (I'm sure you know it as well as I do). Just to quickly remind it to everybody reading this post: "We infer homology when two sequences or structures share more similarity than would be expected by chance. Common ancestry explains excess similarity (other explanations require similar structures to arise independently); thus excess similarity implies common ancestry". (see https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3820096/) "Phrases like “sequence (structural) homology”, “high homology”, “significant homology”, or even “35% homology” are as common, even in top scientific journals, as they are absurd, considering the above definition. In all of the above cases, the term “homology” is used basically as a glorified substitute for “sequence (or structural) similarity”" (see https://www.ncbi.nlm.nih.gov/books/NBK20255/ for more thoughts on this essential topic).
Returning to the question:
I guess, I should modify my intervals list for variant calling by excluding poorly covered regions (those originally targeted regions that are prone to ambigous mapping) and just admit that these are not analyzable using the current approach. What do you, guys, think?
Which factors one must consider while establishing read length threshold?
Trimming bases is carried out primarily at the 3'-end, right? As far as I've heard, if you are in need of trimming the 5'-end, then your data is rather lousy and trimming is flogging a dead horse and is generally not recommended. I may be wrong.
Suggestions are welcome! Have a great day!