This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Finding Unique genes in Roary output

ello , i did pan genome analysis to 60 strain using roary

in Rtab file i found certains number but i didn't find name of genes , i found this

1215 1261 1679 1887 2167 2548 3281 3405 3671 3837 3885 4536 4437 4438 4572 4614 4623 4655 4654 4648 4631 4831 5083 5168 5209 5270 5307 5383 5349 5627 5764 6370 6393 6332 6646 6805 6973 6986 6971 7027 7266 7202 7412 7441 8197 8249 8298 8299 8309 8333 8339 9161 9319 9370 9409 9432 9466 9431 9365 9677 9641 9703 9883 9878 9913 9915 9990 10002 10108 10142 10134 10204 10089 1078 758 598 879 1476 1766 1878 2034 2105 2119 2119 2523 3220 3881 4042 4230 4295 4249 4277 4364 4425 4499 4559 4625 4688 4841 5071 5209 5206 5319 5379 5505 5733 5842 5996 6304 6332 6346 6279 6286 6346 6464 6456 6520 6842 6914 6904 6954 6935 7534 7558 7714 7717 7733 7869 8159 8148 8258 8197 8208 8297 8274 8216 8977 8977 9139 9142 9154 9913 9808 9865 10075 10089 1338 1528 2530 2524 2456 .........

where the number of colone is the number of strains i used (the informations above are just exemple )

also in gene_absence_presence_file i got somthing like this .

1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
bcp 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
fabZ 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
rplT 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
rplA 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
group_6054 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
rpsQ 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
rplE 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
pstA 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
purE 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
rpsT 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
mdh 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
ctaD 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1

with colones are the strains , (didn't apair in the exemple above and lines are genes , basically i can extract number of unique genes using this file , by counting all the times where 1 appear in one colonne ( so that means that this gene is specefic to this strain but i don't knowhow to that )

and i don't know how to extract number of unique genes for each strain using this file

i want to know is there's any file contain information about unique genes i want to know if there's any way to extract number of unique genes for each strain,

thank you

assembly alignment software error

1 answer

From what I can tell, you need to find rows in gene_absence_presence_file where the sum of values is 1. These couple of lines should do that:

awk '{$1=""; print $0}' gene_absence_presence_file | awk '{$1=""; print substr($0,2)}' | awk '{sum=0; for (i=1; i<=NF; i++) { sum+= $i } print sum}' > sums.txt

paste sums.txt gene_absence_presence_file | awk '{print $2, $1}' | sort -nk2 > genes_sorted.txt

The first part of awk command removes column #1 (gene name), while the second part removes trailing spaces (probably not necessary). The last part sums up the rows and redirects into sums.txt. Finally, you take the 1st column (gene names) and merge it with column sums, and sort the final output by the number of gene occurrences. Genes with fewest occurrences in your strains will be near the top, and hopefully some of them will have a count of 1. To me that sounds unlikely for 60 strains, but who knows.

Thank you very much for your comment

i used the two commands you wrote and i used gene_presence_absence.Rtab

but in the finla result ( gene_sorted.txt) i got juts genes , and there presence ( number of strains where the gene exist , like this ;

apaG_2 1
argB_2 1
argD_3 1
argD_4 1
argR_2 1
argS_1 1
argS_2 1
aroA' 1
artM_2 1
artP_2 1
artQ_2 1
asd_2 1
asd2_1 1
asd2_2 1
aspS_1 1
atpA_2 1
atpB_1 1
atpB_2 1
atpB_3 1
atpC_2 1
atpD_2 1
atpD_3 1
atpD_4 1
atpG_2 1
atpG_3 1
atpG_4 1
atpH_3 1
bamA_1 1
bamB_2 1
bamB_4 1

but what i want is to know the number of specefic genes in each strain , something like this

strain1 , 100 strain2 34 strain3 67 strain4 80 and so on

Thank you

Log in to answer this question.