Hello and thank you very much for your answers. My barcodes are random 8mers that I do not know beforehand. They are tagged to each sequence that is ordered and gets amplified (I am sorry if I am not explaining it perfectly, I am new to this).
They key thing is that nobody knows these 8mers beforehand.
Regarding Sparrow's answer, I am not sure I get it completely...So far what I have done is basically map all my reads against my reference and isolate the respective region to my NNNNNNNN region in the construct (which obviously refers to the barcodes).
And what I have is basically an Excel sheet with some thousands of barcodes and their respective frequency in my sequencing reads.
What do you mean I should do next? If, for example, I have barcode ACCTAATT that is found in 10,000 reads, take all reads and create a consensus sequence out of them? And how do I use this to decide if another barcode that has been found only in 500 of my sequencing reads is a true barcode or not?
Also, is there some paper you might have come across where they do it and maybe you can refer me to?
Thank you very much!
What are those barcodes exactly, unique molecular identifiers?