removing prob IDs that match to more than one AGI ID
Hi
I have this list of gene IDs
Affy ID AGI
244901_at ATMG00640
244902_at ATMG00650
244903_at ATMG00660
244904_at ATMG00670
244905_at ATMG00680
244906_at ATMG00690
244907_at ATMG00710
244908_at ATMG00720
244909_at ATMG00740;AT2G07686
244910_s_at ATMG00750;AT2G07686
244911_at ATMG00820
244912_at AT2G07783;ATMG00830
244913_at ATMG00840;AT2G07682
244914_at ATMG00850;AT2G07682
244915_s_at ATMG00860;AT2G07682
244916_at ATMG00880;ATMG00870;AT2G07682
244917_at ATMG00880;ATMG00870;AT2G07682
244918_at ATMG00890
244919_at AT2G07768;ATMG00960
As you consider, some _at IDs match to more than one AGI ID. then how I can remove such _at IDs from my list (which are match to more than one AGI)
Thank you
• 1,771 views
•
link
1 answer
Does it need to be done in R?
In Excel: =FIND(";",B2)
Python:
import csv
with open('inputfile.csv','r') as r, open('outputfile.csv','w') as w:
for line in r:
if ';' not in line:
w.write(line)
• 0 views
•
link
Log in to answer this question.
Do you only want to remove the duplicate ATM id's and keep the _at ID? The solution below would remove that _at ID line.
Thank you, I want to remove the
_atID match to more than one AGI id