Thank you so much for this one liner and to all those who helped. Combined these scripts have allowed me to achieve my goal:
Bhlh 533
Myb_Related 397
C2H2 355
Nac 352
C3H 337
Erf 297
Myb 265
Wrky 263
B3 249
Bzip 217
Far1 196
G2-Like 167
Gras 157
M-Type_Mads 134
Trihelix 134
Arf 128
Mikc_Mads 123
Hd-Zip 120
Lbd 115
Gata 91
Can you show an example of the annotations, to see delimiters and possible ids?
Sure, LibreOffice is frozen at the moment the two columns are TAB delimited and the annotations within are delimited with '^'. The identifier I need is the one before the first ^, however my plan was to take the new list of annotations and open in LibreOffice, specifying it as '^' delimited and then copying the first column (the protein ID) into a new list, and then use this script:
To parse out the gene name.
Then I was going to use this script to get gene family counts:
As soon LibreOffice stop freezing I'll post an example of the two columns. I was going to at first but it didn't look pretty because the annotations were wider than the page.
best, -j
Finally LibreOffice unfroze (wish I could force it to use all 4 cpu)!
Here is an example of the columns:
This is messy, and you're naming your list as the same as what you're opening your file as.
Try this:
Thank you, that is much cleaner.