How do I keep the row-names ? or I can simple make a subset out of the original data set and cbind.data.frame().
I have a dataframe which contains value of log2fold change but it contains inf and NA values i searched all over stack-exchange tried their solution but most of them seems not working few of them works but its not giving me the desired output ,need help it should't be that difficult i suppose .
my data frame
GENE_NAME `HSC_VS_CMP` `HSC_VS_GMP``HSC_VS_Monocytes`
ACTL6A -0.20084399 0.297430 -0.350876000
ACTR8 -0.2925280 -0.158551 1.10747
AICDA inf NA inf
ANP32B -0.6549 -0.615725 -0.35858
I get error like this " default method not implemented for type 'list' "
So please suggest me how to remove all the inf and NA from the data frame
2 answers
Alternative (borrowed from this SO answer):
# test data
d <- mtcars
# add NAs and Inf.
d[1,1] <- NA
d[2,2] <- NA
d[2,1] <- Inf
d[1,2] <- Inf
head(d)
mpg cyl disp hp drat wt qsec vs am gear carb
Mazda RX4 NA Inf 160 110 3.90 2.620 16.46 0 1 4 4
Mazda RX4 Wag Inf NA 160 110 3.90 2.875 17.02 0 1 4 4
Datsun 710 22.8 4 108 93 3.85 2.320 18.61 1 1 4 1
Hornet 4 Drive 21.4 6 258 110 3.08 3.215 19.44 1 0 3 1
Hornet Sportabout 18.7 8 360 175 3.15 3.440 17.02 0 0 3 2
Valiant 18.1 6 225 105 2.76 3.460 20.22 1 0 3 1
# the magic:
d <- do.call(data.frame, lapply(d, function(x) {
replace(x, is.infinite(x) | is.na(x), 0)
})
)
# note that you lose the row names.
head(d)
mpg cyl disp hp drat wt qsec vs am gear carb
1 0.0 0 160 110 3.90 2.620 16.46 0 1 4 4
2 0.0 0 160 110 3.90 2.875 17.02 0 1 4 4
3 22.8 4 108 93 3.85 2.320 18.61 1 1 4 1
4 21.4 6 258 110 3.08 3.215 19.44 1 0 3 1
5 18.7 8 360 175 3.15 3.440 17.02 0 0 3 2
6 18.1 6 225 105 2.76 3.460 20.22 1 0 3 1
That is inconvenient, isn't it? I would just assign the modified data.frame to a different variable (in the do.call part) and then copy the row names from the original data to the modified one.
is.na() works on dataframes. For the Infs, lapply() a function:
d2 = lapply(d, function(x) {
if(any(is.infinite(x))) {
x[is.infinite(x)] = 0
}
return(x)
})
d2 = as.data.frame(d2)
I still see inf in my data frame
When I do it I don't see Inf, so show a reproducible example.
This works for me, no problem (for example, using the dataset in my answer). Note, in the example data you included infinite is specified as "inf" not Inf (R's way). So maybe it is related?
OK, even if in your original data you had "inf" instead of "Inf" it seems read.table (and probably friends) ignore case (see ?read.table). So it is read properly and shouldn't be a problem (e.g. read.table(text = "inf 0 10").
Some guesswork, I would also check for: "inf" and "NA" as character, assign NA, then use complete.cases()
Log in to answer this question.
What have you tried (with code)? Presuming you just want to remove the rows, it's a simple (A) apply(d, 1, function(x) anyis.na(x) || is.infinite(x))) and then (B) subsetting accordingly.
Alternative without using
apply:Note: either way, you may need to take special care of the GENE_NAME column. I suspect your are having trouble with that in your attempts so far (but as mentioned by Devon, please show us the code).
I can take out the GENE_NAME column and apply the same
You can't
is.infinite()dataframes, which is probably what resulted in the originally reported error.what is the way to do it ?the way you suggested ,I know there are multiple ways but since Im learning and the at the same time I have to use them in the data sets so i have to just read and see which one is working.
So can you show me how do i remove NA ,inf and 0 if any from my data frame in a concise code
You are right! (didn't do enough testing...) Find this a bit inconsistent since
is.na()works fine. Further searching points to this SO post where a viable solution is to implementis.infinite.data.framemethod, e.g.:I tried this
}
What is the desired output?
well I want to replace inf and NA with 0 .