Hi @Genomax and @Rob,
Thanks for the note and my apologies for the delay.
@Genomax, couldn't agree more in relation to the short insert library (and negative inner distance) as a central concern. We don't anticipate this going away in the future wet lab, so we are erring on the side of caution with regards to data assumptions.
@Rob, thanks very much for the detail. It is reassuring to know that due to the transcriptome alignments and paired-end reads, this is a win-win situation. We had a further think about your points, added with extra reading from the COMBINE-lab GitHub. Specifically, we found your information regarding the values --fldMean and --fldSD being used for prior parameters of a normal distribution which is then truncated on the left at 0 very helpful ( post 127 https://github.com/COMBINE-lab/salmon/issues/127 ).
We took a look at our fastq files, pre and post trimming. On average, the Median is lower than the mean (albeit by approx. 10) and a histogram shows a positive right skew (e.g. 1.032, when using "skewness" in the R "moments" package.) Are heavy tails adjusted/detected by Salmon, causing an update to the prior distribution and possibly transforming the gaussian probabilistic model? I hope that made some sense.
Would the positive skew have any sort of impact on TPM output, and would it be preferential to not use TPM for further analysis downstream?
Thanks in advance, Chris
If you have a reference available then paired-end sequencing allows one to estimate the length of the library fragment being sequenced by inferring how far apart the two reads map/align on the reference. One of the exceptions would be, if the library fragment captured a breakpoint (e.g. two ends map to two different chromosomes). In that case it is not possible to estimate insert size.
You have a "short" insert library. There is no solution for this specific issue except making a new prep/library.