Thanks a million for this! I think I would have needed several months to code something comparable.
I have tried to run it on a test VCF dataset of my own with one individual and two SNPs:
#CHROM POS ID REF ALT QUAL FILTER INFO FORMAT Test_Ind
Sh05 1738002 . T C . . AC=5;AN=12 GT:AD:DP:GC 0/0/0/0/0/0/0/1/1/1/1/1:122,89:211:168.534
Sh05 1738019 . G A . . AC=3;AN=12 GT:AD:DP:GC 0/0/0/0/0/0/0/0/0/1/1/1:159,61:220:8.92207
With the proposed method, I obtain:
#CHROM START END COUNT N-VARIANTS (POS\tALT)+
Sh05 1738002 1738019 78 2 1738002 C 1738019 G
Sh05 1738002 1738019 1 2 1738002 T 1738019 G
Sh05 1738002 1738019 1 2 1738002 T 1738019 A
Sh05 1738002 1738019 6 2 1738002 T 1738019 G
...
[several dozen lines]
...
Sh05 1738002 1738019 2 2 1738002 T 1738019 G
Sh05 1738002 1738019 2 2 1738002 T 1738019 A
There seems to be an issue in that it aggregates results only for the first category "CG". I don't know to which extent this is complicated to solve?
Otherwise the results are very comprehensive, thanks again.