This is a list of data science software and platforms used in data science, which includes programming languages, programming environments, machine learning frameworks, data engineering tools, statistical software, data analysis, plotting, MLOps systems, and more.
Programming languages
Development environments These interactive notebooks, IDEs, and platforms provide specialised development environments.
Apache Zeppelin Architect — Eclipse (software) CoCalc Dataiku Data Science Studio FreeMat GNU Octave Google Colab DataSpell Jupyter Notebook / JupyterLab Kaggle Notebooks MATLAB O-Matrix PyCharm RStudio SAS (software) and SAS Studio Spyder Visual Studio Code
Machine and deep learning software The Machine learning / deep learning tools support development in those fields.
Data engineering Examples of Data engineering tools.
Apache Airflow Apache Flink Apache Hadoop Apache Kafka Apache NiFi Apache Spark Dask Data build tool (dbt)
Data mining Examples of Data mining tools.
Data mining
Proprietary
Database management
List of RDBMS
Proprietary
Data warehouses Data warehouse environments include:
Amazon Redshift Snowflake Google BigQuery Microsoft Azure Synapse Teradata Vertica
Data lakes Data lake environments include:
Apache Hadoop Cloudera Databricks Delta Lake Amazon S3 Google Cloud Storage Azure Data Lake
Algorithms Apriori algorithm – frequent itemset mining and association rule learning in market basket analysis Backpropagation – algorithm for training artificial neural networks using gradient descent Decision Trees – tree-based algorithm for classification and regression Expectation–maximization algorithm – iterative procedure for maximum likelihood estimation with latent variables Gradient descent – iterative optimization algorithm for minimizing a loss function ID3 algorithm – used to generate a decision tree from a dataset K-Means – clustering algorithm based on minimizing within-cluster distances K-Nearest Neighbors (KNN) – instance-based learning and classification method Linear regression – estimation method for predicting a dependent variable based on independent variables Logistic regression – classification algorithm for predicting a binary outcome Naive Bayes – probabilistic classifier based on Bayes' theorem Ordinary least squares – estimation method for parameters in linear regression PageRank – graph-based algorithm for link analysis and search ranking Principal component analysis – technique to reduce high-dimensional data while preserving variance Q-learning – reinforcement learning algorithm for learning optimal actions Random forest – ensemble of decision trees for improved classification or regression Sequential minimal optimization – solver for training support vector machines Stochastic gradient descent – randomized variant of gradient descent for large-scale machine learning Support Vector Machines (SVM) – algorithm for finding a hyperplane to separate classes
Statistical software
Open-source
Public domain CSPro Dataplot Epi Map X-13ARIMA-SEATS
Freeware BV4.1 MINUIT WinBUGS Winpepi
Proprietary
Data processing Tools for Data processing and analysis:
Data and information visualization Software for Data visualization:
Plotting software Software for plotting data to support processing and visualise results.
Maps and geospatial visualization
ArcGIS Carto Epi Map GeoDA Google Earth Engine Leaflet Mapbox MountainsMap QGIS
Machine learning MLOps and model deployment:
BentoML Data Version Control (DVC) Kubeflow MLflow Seldon Core Streamlit TensorFlow Serving Weights & Biases
Data repositories
Kaggle – platform for data science competitions, datasets, and notebooks. OpenML – collaborative platform for sharing datasets, algorithms, and experiments. University of California, Irvine Machine Learning Repository Zenodo – open-access repository supported by CERN and the EU.
Educational data science software
Kaggle – online platform for data science education, competitions, datasets, and collaborative learning. KNIME – open-source data analytics platform used for teaching data science, machine learning, and workflow-based analysis. RapidMiner – used in academic research and education for data mining and machine learning. Statistics Online Computational Resource (SOCR) – online tools and instructional resources for statistics education. Tanagra (machine learning) – data mining software developed for research and teaching purposes. TinkerPlots – explore and analyze data through visual modeling.
See also
Business intelligence software List of data science journals List of R software and tools Lists of mathematical software and List of open-source software for mathematics List of numerical-analysis software List of numerical analysis topics List of numerical libraries List of open-source data science software Common Crawl – nonprofit that crawls the web and freely provides its archives and datasets to the public under an MIT License
References
External links 20 Tools for Data Scientists | Pragmatic Institute igorbarinov/awesome-data-engineering
