This is a test version of Biostars. For the public version, visit https://www.biostars.org.
removing prob IDs that match to more than one AGI ID

Hi

I have this list of gene IDs

Affy ID         AGI      
244901_at       ATMG00640      
244902_at       ATMG00650      
244903_at       ATMG00660      
244904_at       ATMG00670      
244905_at       ATMG00680      
244906_at       ATMG00690      
244907_at       ATMG00710      
244908_at       ATMG00720      
244909_at       ATMG00740;AT2G07686      
244910_s_at     ATMG00750;AT2G07686      
244911_at       ATMG00820      
244912_at       AT2G07783;ATMG00830      
244913_at       ATMG00840;AT2G07682      
244914_at       ATMG00850;AT2G07682      
244915_s_at     ATMG00860;AT2G07682      
244916_at       ATMG00880;ATMG00870;AT2G07682      
244917_at       ATMG00880;ATMG00870;AT2G07682      
244918_at       ATMG00890      
244919_at       AT2G07768;ATMG00960

As you consider, some _at IDs match to more than one AGI ID. then how I can remove such _at IDs from my list (which are match to more than one AGI)

Thank you

gene r

Do you only want to remove the duplicate ATM id's and keep the _at ID? The solution below would remove that _at ID line.

Thank you, I want to remove the _at ID match to more than one AGI id

1 answer

Does it need to be done in R?

In Excel: =FIND(";",B2)

Python:

import csv
with open('inputfile.csv','r') as r, open('outputfile.csv','w') as w:
    for line in r:
        if ';' not in line:
            w.write(line)

Log in to answer this question.