This has worked perfectly. Thank you very much!
I have a very large text file (9.2GB) that contains data in two columns (named All_SNPs.txt). I would like to use this file to convert my chr:pos to rs in my plink files. The file looks like this:
1:116342 rs1000277323
1:173516 rs1000447106
1:168592 rs1000479828
1:102498 rs1000493007
However, plink produced an error stating that All_SNPs.txt contains duplicates in column 1. I have tried different ways of removing the duplicates; specifically, I attempted the following line commands:
awk '!x[$1]++ { print $1,$2 }' All_SNPs.txt > All_SNPs_nodup.txt
sort All_SNPs.txt | uniq -u > All_SNPs_nodup.txt
cat All_SNPs.txt | sort | uniq -u > All_SNPs_nodup.txt
However, each time I faced the same two problems: 1) either, the code will not make any difference and I will receive the same file back (this is for line commands that uses cat) or 2) I will receive an error : the procedures halted: have exceeded (this is for sort | uniq, and awk).
I will be very grateful for any ideas of how I can make this work. Thank you very much.
PS, this file is far too large to open it in R.
4 answers
Cut and uniq to find duplicates, then grep -v them away. Relatively quick on 389MB dummy file.
time sort -V All_SNPs.txt | cut -f 1 | uniq -c | perl -ane 'if($F[0] ne "1"){print "$F[1]\t$F[0]\n";'} > All_SNPs.dup.chr-pos.txt
real 0m57.428s
user 3m19.091s
sys 0m3.152s
time cut -f 1 All_SNPs.dup.chr-pos.txt | grep -wvf - All_SNPs.txt > All_SNPs.nodup.txt
real 0m2.516s
user 0m1.697s
sys 0m0.229s
Split the text file into smaller one using the split command. You could split by size, as example:
split -b 200m filename
This will produce files named ' xaa, xab, xac..' . Now use awk, but with simpler syntax
awk -F"\t" '!seen[$1]++' xa*
And after that, join files using a sample cat into the destination file.
How big is the chance that two dups end up in different files?
If you sort first I guess the chance is very small no?
You can use awk to split your file according to the chromosomes:
awk '{ split($1, a, ":"); print $1"\t"$2 >> a[1]".txt"; }' All_SNPs.txt)
That assures you that you don't miss two dups. If you like to reduce the file size a bit, you can remove the chr: :
awk '{ split($1, a, ":"); print a[2]"\t"$2 >> a[1]".txt"; }' All_SNPs.txt)
This will also allow you to keep a bit more items in the hash.
This has worked. Thank you very much for the suggestion!
Did you try:
awk '!seen[$1]++' All_SNPs.txt > All_SNPs_nodup.txt
If it doesn't work I think you need better hardware...
Plink has basic mechanism to deal with dups
--list-duplicate-vars <require-same-ref> <ids-only> <suppress-first>
https://www.cog-genomics.org/plink/1.9/data#list_duplicate_vars
Then the duplicated vars can be excluded using --exclude plink.dupvar
Log in to answer this question.
Using datamash:
datamash is available in brew, conda, apt repos.
using tsv-utils :
Also try with a hashtable, like a python dict. No idea how much memory it would require but you can try if nothing else works.
You could try to convert the input file to a valid
vcfand use thanbcftools sortandbcftools norm -N -d noneto remove the duplicates. At the end you can convert back to the input format.