This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Dataset cleaning through python

I have a dataset in the tsv file which contains gene information. First, upload it in the pandas' data frame and now I want to remove all missense mutations present in data through the 'mutation somatic status' columndata

My code:

chunks=pd.read_csv("CosmicGenomeScreensMutantExport.tsv",chunksize=1000000,sep='\t')
 dfList = []

 for df in chunks:
     dfList.append(df)

 df = pd.concat(dfList,sort=False)

After removing missense mutations I want to isolate only those records that contain gene P23

Can anyone help me in this?

python pandas

1 answer

if your df is as posted images, try one of the two below:

print(df[(df['Mutation Description'] != '<missense_SO_Term>') & (df['Gene name'] == '<Gene_symbol>')])
print(df.query('(`Mutation Description` != "<missense_SO_Term>") & (`Gene name` == "<Gene_symbol>")'))

Please replace <missense_SO_Term> with appropriate text for missense mutations, and <Gene_symbol> with appropriate name for P23 in your data.

I also want to remove N.A, null and duplicated values from my dataset

This code : print(df[(df.'mutation somatic status' !="missense") & (df.'Gene name'=="TP53")]) is giving me syntax error

I have updated the code with appropriate column names. Please replace SO term and Gene name with appropriate values and also check the column names.

Log in to answer this question.