Hi I'm new to bioinformatics (i'm undergrad). I have 2 genomes, same species, probably different strains. I've used get_homologues to define the core genome, but i'm having trouble to find the singletons. I have already read several times both the manual and tutorial available and i've tried to used parse_pangenome_matrix to find which genes are present in A and absent in B, and it seems that there isn't any! (file with genes present in set A and absent in B (0) ).
How can it be possible, if the core genome is smaller than both genomes?
1 answer
If you have a core gene set and complete gene set, then you can consider subtracting core genes from complete gene set using simple shell command.
For instance,
grep -w -v -f core_gene_set.txt complete_gene_set_sampleA.txt >singleton_sampleA.txt
grep -w -v -f core_gene_set.txt complete_gene_set_sampleB.txt >singleton_sampleB.txt
Here,
- -w, means it will search for exact word provided in core gene set file from the complete_gene_set.txt file
- -v, means it will search for non-matching lines provided in core gene set file from the complete_gene_set.txt file
- -f, means it will grep PATTERN provided in core gene set file from complete_gene_set.txt file
Log in to answer this question.
If they are the same species this is not so surprising.