This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Retrieving matched subsets of sequence IDs from a file

I wish to know how to extract subset of data from a file given a set of IDs.

The ID file format

A0A1H3TMV2
Z8X0S9
A0A2M6TTB8
A0A2Q5TOM7
A0A1U4MIN3
X9RTW0

The tab-separated file has the following format:

A0A1H3TMV2      A0A1H3TMX3
A0A1H3TMV2      M1P8Z9
A0A1H3TMV2      A0A3M1XSM6
Z8X0S9          A0A1U9ZQ4
Z8X0S9          ZRT98WEN2
Z8X0S9          Z8X0U4
A0A2Q5TOM7      YYY8IB3R9U
Z8X0S9          HR915ERQN
Z8X0S9          GLE2WI9Z7
Z8X0S9          FFF883REW3
A0A2Q5TOM7      A0A2Q5TMX5

In the complete tab-separated file, each ID may show up 10s of times. Observe that the same Query ID can be associated with different neighbor ID and I must be able to only select the correct pair. For example, here are the correct pairs I would like to extract from this set:

A0A1H3TMV2      A0A1H3TMX3
Z8X0S9          Z8X0U4
A0A2Q5TOM7      A0A2Q5TMX5

Sometimes the query ID and its neighbor are different only by two letters near the end. At other times, they can be off by three. The differences are unlikely to exceed 4 characters near the end.

Now, I can sort the ID file without any issues. However, I can't do that for the tab-separated file because the matched pairs will no longer exist.

sequence

However, I can't do that for the tab-separated file because the matched pairs will no longer exist.

Have you looked at UNIX sort which can sort based on columns?

I can't sort the tab-separated file because line information will be lost. Each row represents a linked pair and I am trying to get the pair that is very similar at the word level (if that makes sense).

Sorting by column 2 followed by 1

$ sort -k2,2 -k1,1 file
A0A1H3TMV2      A0A1H3TMX3
Z8X0S9          A0A1U9ZQ4
A0A2Q5TOM7      A0A2Q5TMX5
A0A1H3TMV2      A0A3M1XSM6
Z8X0S9          FFF883REW3
Z8X0S9          GLE2WI9Z7
Z8X0S9          HR915ERQN
A0A1H3TMV2      M1P8Z9
A0A2Q5TOM7      YYY8IB3R9U
Z8X0S9          Z8X0U4
Z8X0S9          ZRT98WEN2

Sorting by column 1

$ sort -k1,1 file
A0A1H3TMV2      A0A1H3TMX3
A0A1H3TMV2      A0A3M1XSM6
A0A1H3TMV2      M1P8Z9
A0A2Q5TOM7      A0A2Q5TMX5
A0A2Q5TOM7      YYY8IB3R9U
Z8X0S9          A0A1U9ZQ4
Z8X0S9          FFF883REW3
Z8X0S9          GLE2WI9Z7
Z8X0S9          HR915ERQN
Z8X0S9          Z8X0U4
Z8X0S9          ZRT98WEN2

Ok, fair enough. Is this newly sorted file amenable to awk or grep to extract only the entries I want? I still don't know how to do that. For example, my IDs file has one entry for Z8X069, but the tab-separated file has 30 - 100 entries. Within the latter, only one correct match exists. That would be Z8X0U4. The good news is that the match at least has the first three same characters. Do you know of a way to get the output I want given your sort strategy? Thank you very much.

1 answer

How about (comparison of two strings excluding 3 characters from right end, adjust as needed)?

$ awk -F "\t" '{if (substr($1,0,length($1)-3) == substr($2,0,length($2)-3)) {print $0}}' file
A0A1H3TMV2      A0A1H3TMX3
Z8X0S9  Z8X0U4
A0A2Q5TOM7      A0A2Q5TMX5

Thanks for that. So, I tried what you suggested on my tab-separated file (20,970 entries) and the output has 19,954 entries. I still need to be able to compare the IDs in my IDs file with the tab-separated file. Your script only uses the latter. How would I do that?

You can use grep -f to recover just the entries you need.

Neither grep -F nor -Fwf yield much of anything. I just get 3 entries back.

You should use grep (with ID's of your interest) on the results file that you produced with awk code above.

my replies aren't showing up anymore...why?

I am unable to do what you suggested. I am only getting the last entry. Any thoughts?

File with ID you want

$ more ID
A0A1H3TMV2
A0A2Q5TOM7

Output from awk

$ more original
A0A1H3TMV2      A0A1H3TMX3
Z8X0S9  Z8X0U4
A0A2Q5TOM7      A0A2Q5TMX5

Final extraction of ID's you want

$ grep -f ID original
A0A1H3TMV2      A0A1H3TMX3
A0A2Q5TOM7      A0A2Q5TMX5

I know what you mean and it works in the small example. However, it must be a file format problem or something that it only lists 3 entries when I do this in the entire file. Go figure.

If you want to put a bigger example up at pastebin.com I can take a look.

Log in to answer this question.