Lets face it, SAS (aka Statistical Analysis Software) is an archaic program that is difficult to learn, cumbersome to code, and slow to implement calculations. While graphics are improving, graphs are still hard to customize and lag far behind those of other programs. In addition, yearly licensing fees were more than enough to scare me away, especially since there are a lot of open source statistical programing packages that are just as good, if not better than SAS.
So what are these alternatives to SAS? Here is my attempt to make a list of valid alternatives to SAS. Some may be more powerful or better tailored to specific niches than others, but all are capable of carrying out basic statistical tests. The list is alphabetical with major SAS contenders highlighted in bold.
Epi Info - simple CDC tool for clinicians
Excel - basic, but there are some useful plugins
gretl - open source with econometrics focus
KXEN - seems more geared towards industry
Mathematica - used it doing physics research, seemed powerful at the time, but a long time ago
MATLAB - again, maybe more for engineers and such, but has statistical functionality
Minitab - quite accessible to those with little computer knowledge
Octav - open source, similar to MATLAB
OpenEpi - online, JAVA based statistical applets
Python - more of a programming language, but has statistical packages you can load
R - open source, thousands of libraries and active community of users
S-Plus - similar to R, S based dialect, could not find a link to this
Sage- unified interface of open source packages
SPSS - point and click, with some scripting options, crashes a lot
Stata - command line or point and click, powerful but not bloated
STATISTICA - markets themselves as a user friendly SAS alternative
That's about as exhaustive of a list as I can come up with. If I missed something important please comment and let me know about it. Hopefully I have been able to convince you there are plenty of excellent alternative statistical programming packages to SAS that are worth your time looking into. In my opinion, R and Stata are the two clear front runners, particularly R since it is free and open source. Happy hunting for a SAS replacement!
A repository of programs, scripts, and tips essential to
genetic epidemiology, statistical genetics, and bioinformatics
Welcome to the Genome Toolbox! I am glad you navigated to the blog and hope you find the contents useful and insightful for your genomic needs. If you find any of the entries particularly helpful, be sure to click the +1 button on the bottom of the post and share with your colleagues. Your input is encouraged, so if you have comments or are aware of more efficient tools not included in a post, I would love to hear from you. Enjoy your time browsing through the Toolbox.
Showing posts with label analysis. Show all posts
Showing posts with label analysis. Show all posts
Wednesday, January 15, 2014
Tuesday, January 14, 2014
Remove Rows with NA Values From R Data Frame
Rows with NA values can be a pesky nuisance when trying to analyze data in R. Here is a short primer on how to remove them.
There are two primary options when getting rid of NA values in R, the na.omit/is.na commands and the complete.cases command. Both are part of the base stats package and require no additional library or package to be loaded. Below are examples of how the two work with a data frame called data and a variable called var.
The na.omit/is.na commands work as follows:
na.omit(data) - will only select rows with complete data in all columns
data[rowSums(is.na(data[,c(2,3,5)]))==0,] - will only select rows with complete data in columns 2, 3, and 5
var[!is.na(var)] - will only select values of a variable not equal to NA
The complete.cases command works as follows:
data[complete.cases(data),] - will only select rows with complete data in all columns
data[complete.cases(data[,c(2,3,5)]),] - will only select rows with complete data in columns 2, 3, and 5
var[complete.cases(var)] - will only select values of a variable not equal to NA
I use both commands at times, but ultimately prefer the complete.cases command for the cleaner syntax and generalizability. Hope this helps you remove those NA's from your data. If you have additional tips or questions please leave a comment below.
There are two primary options when getting rid of NA values in R, the na.omit/is.na commands and the complete.cases command. Both are part of the base stats package and require no additional library or package to be loaded. Below are examples of how the two work with a data frame called data and a variable called var.
The na.omit/is.na commands work as follows:
na.omit(data) - will only select rows with complete data in all columns
data[rowSums(is.na(data[,c(2,3,5)]))==0,] - will only select rows with complete data in columns 2, 3, and 5
var[!is.na(var)] - will only select values of a variable not equal to NA
The complete.cases command works as follows:
data[complete.cases(data),] - will only select rows with complete data in all columns
data[complete.cases(data[,c(2,3,5)]),] - will only select rows with complete data in columns 2, 3, and 5
var[complete.cases(var)] - will only select values of a variable not equal to NA
I use both commands at times, but ultimately prefer the complete.cases command for the cleaner syntax and generalizability. Hope this helps you remove those NA's from your data. If you have additional tips or questions please leave a comment below.
Subscribe to:
Posts (Atom)