Thanks for that. So, I tried what you suggested on my tab-separated file (20,970 entries) and the output has 19,954 entries. I still need to be able to compare the IDs in my IDs file with the tab-separated file. Your script only uses the latter. How would I do that?
I wish to know how to extract subset of data from a file given a set of IDs.
The ID file format
A0A1H3TMV2
Z8X0S9
A0A2M6TTB8
A0A2Q5TOM7
A0A1U4MIN3
X9RTW0
The tab-separated file has the following format:
A0A1H3TMV2 A0A1H3TMX3
A0A1H3TMV2 M1P8Z9
A0A1H3TMV2 A0A3M1XSM6
Z8X0S9 A0A1U9ZQ4
Z8X0S9 ZRT98WEN2
Z8X0S9 Z8X0U4
A0A2Q5TOM7 YYY8IB3R9U
Z8X0S9 HR915ERQN
Z8X0S9 GLE2WI9Z7
Z8X0S9 FFF883REW3
A0A2Q5TOM7 A0A2Q5TMX5
In the complete tab-separated file, each ID may show up 10s of times. Observe that the same Query ID can be associated with different neighbor ID and I must be able to only select the correct pair. For example, here are the correct pairs I would like to extract from this set:
A0A1H3TMV2 A0A1H3TMX3
Z8X0S9 Z8X0U4
A0A2Q5TOM7 A0A2Q5TMX5
Sometimes the query ID and its neighbor are different only by two letters near the end. At other times, they can be off by three. The differences are unlikely to exceed 4 characters near the end.
Now, I can sort the ID file without any issues. However, I can't do that for the tab-separated file because the matched pairs will no longer exist.
1 answer
How about (comparison of two strings excluding 3 characters from right end, adjust as needed)?
$ awk -F "\t" '{if (substr($1,0,length($1)-3) == substr($2,0,length($2)-3)) {print $0}}' file
A0A1H3TMV2 A0A1H3TMX3
Z8X0S9 Z8X0U4
A0A2Q5TOM7 A0A2Q5TMX5
You can use grep -f to recover just the entries you need.
Neither grep -F nor -Fwf yield much of anything. I just get 3 entries back.
You should use grep (with ID's of your interest) on the results file that you produced with awk code above.
my replies aren't showing up anymore...why?
I am unable to do what you suggested. I am only getting the last entry. Any thoughts?
File with ID you want
$ more ID
A0A1H3TMV2
A0A2Q5TOM7
Output from awk
$ more original
A0A1H3TMV2 A0A1H3TMX3
Z8X0S9 Z8X0U4
A0A2Q5TOM7 A0A2Q5TMX5
Final extraction of ID's you want
$ grep -f ID original
A0A1H3TMV2 A0A1H3TMX3
A0A2Q5TOM7 A0A2Q5TMX5
Log in to answer this question.
Have you looked at UNIX
sortwhich can sort based on columns?I can't sort the tab-separated file because line information will be lost. Each row represents a linked pair and I am trying to get the pair that is very similar at the word level (if that makes sense).
Sorting by column 2 followed by 1
Sorting by column 1
Ok, fair enough. Is this newly sorted file amenable to awk or grep to extract only the entries I want? I still don't know how to do that. For example, my IDs file has one entry for Z8X069, but the tab-separated file has 30 - 100 entries. Within the latter, only one correct match exists. That would be Z8X0U4. The good news is that the match at least has the first three same characters. Do you know of a way to get the output I want given your sort strategy? Thank you very much.