Causal inference from observational data is the goal of many data analyses in the health and social sciences. However, academic statistics has often frowned upon data analyses with a causal objective. The introduction of the term "data science" provides a historic opportunity to redefine data analysis in such a way tha…
Data Science is currently a popular field of science attracting expertise from very diverse backgrounds. Current learning practices need to acknowledge this and adapt to it. This paper summarises some experiences relating to such learning approaches from teaching a postgraduate Data Science module, and draws some learn…
Today, the prominence of data science within organizations has given rise to teams of data science workers collaborating on extracting insights from data, as opposed to individual data scientists working alone. However, we still lack a deep understanding of how data science workers collaborate in practice. In this work…
Defines data science as a natural ecosystem with challenges and missions.
problem Challenges and missions in data science due to 5D complexities and data life cycle phases.
method Systemic and data-centric view of data science as a fusion of data universe and its challenges, formalizing a general-purpose architecture.
result Essential data science as a natural ecosystem integrating specific disciplines and high-impact applications.
Foundation models alter medical data science workflow, challenging veridical data science principles.
problem Foundation models disrupt traditional data science practices in medicine.
method Critically examined the medical foundation model lifecycle and its deviation from veridical data science principles.
result Foundation models challenge veridical data science principles of predictability, computability, and stability.
ML methods improve planetary science data analysis.
problem Insufficient use of ML in planetary science.
method Ten recommendations for integrating ML in planetary science.
result Expanding planetary science insights from large datasets.
Data science enhances knot theory by analyzing invariant relations.
problem Understanding the complex relations between knot invariants.
method Topological data analysis applied to knot theory.
result New insights into long-standing conjectures about knots.
The goal of this article is to inspire data scientists to participate in the debate on the impact that their professional work has on society, and to become active in public debates on the digital world as data science professionals. How do ethical principles (e.g., fairness, justice, beneficence, and non-maleficence) …
Donoho's JCGS (in press) paper is a spirited call to action for statisticians, who he points out are losing ground in the field of data science by refusing to accept that data science is its own domain. (Or, at least, a domain that is becoming distinctly defined.) He calls on writings by John Tukey, Bill Cleveland, and…
Machine learning improves wildfire science and management, but requires expert knowledge.
problem Improving wildfire science and management through AI.
method Review of popular ML approaches and their application in six wildfire science domains.
result Opportunities exist for applying more advanced ML methods in wildfire science.
Data science models, although successful in a number of commercial domains, have had limited applicability in scientific problems involving complex physical phenomena. Theory-guided data science (TGDS) is an emerging paradigm that aims to leverage the wealth of scientific knowledge for improving the effectiveness of da…
This paper explores data science applications in economics using a taxonomy of models and hybrid models showing higher accuracy.
problem Investigating data science applications in economics.
method Systematic literature review using Prisma method.
result Hybrid models showed higher prediction accuracy than other algorithms.
Proposes a methodology to improve data science ROI by addressing key business questions.
problem Companies often fail to maximize data science value, focusing on basic analysis.
method Categorizes and answers 'The Big Three' questions using data science methods.
result Shows how to apply the methodology to real business use cases.
This study analyzes data science vocabulary changes over 13 years.
problem Understanding evolution of data science terms over time.
method Exploratory Data Analysis, Latent Semantic Analysis, Latent Dirichlet Analysis, N-grams Analysis.
result Identified new vocabulary and its incorporation into scientific literature.
Kan extensions help in data science extrapolation and learning.
problem Generalizing functions over larger sets in data science.
method Kan extensions in category theory applied to data science problems.
result Kan extensions can be used to derive classification and clustering algorithms.
New geometric methods improve optimization and data science problems.
problem Improving optimization and data science problems.
method Geometric tools for high-dimensional optimization and statistical data science.
result New algorithms and statistical guarantees for optimization and data science.
Paper relaxes optimal transport using convex functions for data science.
problem Optimal transport problem on finite spaces.
method Relaxation via strictly convex functions (Kullback-Leibler divergence, Bregman divergences). Gradient descent iterative process.
result Mathematical foundations and iterative process for the relaxed optimal transport problem.
New method shows data-driven causal studies can be misleading.
problem Misattribution of causality in data-driven earth science studies.
method Subsample-based ensemble approach for robust causality analysis.
result Transfer entropy-based causal graphs can be spurious.
Provides a compendium of data sources for various applications.
problem Lack of comprehensive data sources for data science, machine learning, and AI.
method Compilation of diverse data sources across multiple application areas.
result A comprehensive list of data sources for data scientists and machine learning experts.
Ridge regularization simplifies model complexity in data science.
problem Overfitting in statistical models.
method Adding a penalty on the magnitude of coefficients.
result Effective in reducing model complexity and improving generalization.
FinTech uses data science and AI to transform finance.
problem Transforming finance with data science and AI.
method DSAI techniques including complex system methods, quantitative methods, etc.
result DSAI enables smart FinTech for various financial sectors.
This paper uses information theory to improve risk modeling in big data.
problem Insufficient application of information theory in actuarial science.
method Explores information theory to uncover performance limits of insurance big data systems.
result Guidance for risk modeling and actuarial pricing systems.
New model estimates species population trends from citizen science data.
problem Interannual confounding in citizen science data.
method Double Machine Learning framework to estimate population change and propensity scores for confounding adjustment.
result Spatially detailed trend estimates from citizen science data with low error rates.
Conversion of raw data into insights and knowledge requires substantial amounts of effort from data scientists. Despite breathtaking advances in Machine Learning (ML) and Artificial Intelligence (AI), data scientists still spend the majority of their effort in understanding and then preparing the raw data for ML/AI. Th…
Kaggle chronicles 15 years of competitions, innovation, and data science.
problem Exploring 15 years of data science competitions and innovations.
method Longitudinal trend analysis and exploratory data analysis of millions of kernels and discussion threads.
result Kaggle is a growing platform with diverse use cases and adaptable Kagglers.
The Prescriptive Canvas improves business outcomes by directly prescribing actions based on predictions.
problem Sub-optimal performance in business projects due to a two-step approach of prediction and decision-making.
method The Prescriptive Canvas methodology for framing and communicating actions directly based on predictions.
result Improves framing and communication across stakeholders for successful business impact.
Paper develops an attention mechanism for long-term scientific impact prediction.
problem Predicting the long-term impact of scientific papers based on citation records.
method Develops an attention mechanism to predict long-term scientific impact.
result Emphasizing the limited attention can better stand on the shoulders of giants.
New method calculates discrete curvature using effective resistances.
problem Calculating discrete curvature on graphs.
method Effective resistances to calculate curvature on graph nodes and links.
result Relation to established discrete curvatures and convergence to continuous curvature.
Machine learning's data-centric philosophy conflicts with natural sciences' standards.
problem Conflict between machine learning's ontology and epistemology and natural sciences' practices.
method Identifying and analyzing contexts where ML can be beneficial or harmful in natural sciences.
result ML can enhance trustworthiness in causal inference but introduces biases in emulation and labeling.
MIM adds indicator variables to improve model performance on incomplete data.
problem Missing data in incomplete data sets.
method Missing Indicator Method (MIM) and Selective MIM (SMIM).
result MIM improves model performance for informative missing values and high-dimensional data.
This paper explores a real-world fundamental theme under a data science perspective. It specifically discusses whether fraud or manipulation can be observed in and from municipality income tax size distributions, through their aggregation from citizen fiscal reports. The study case pertains to official data obtained fr…
Citizen science projects are successful at gathering rich datasets for various applications. However, the data collected by citizen scientists are often biased --- in particular, aligned more with the citizens' preferences than with scientific objectives. We propose the Shift Compensation Network (SCN), an end-to-end l…
Study assesses 'big data' in materials science, highlighting challenges.
problem Understanding what constitutes 'big data' in materials science.
method Selected examples of machine learning models, data quality, and infrastructure requirements.
result Big data presents unique challenges in materials science.
Data science predicts user interest for midwifery content.
problem Improving midwives' learning and preventing maternal and newborn deaths.
method Forecasting methods using user-generated logs from online learning apps.
result Determining future user interest in midwifery content types.
The study explores how machine learning can enhance scientific research.
problem Improving scientific models with machine learning.
method Analysis of data-driven models versus manually added variables in regression.
result Complex models may not always improve over simpler ones in scientific contexts.
The reproducibility of scientific research has become a point of critical concern. We argue that openness and transparency are critical for reproducibility, and we outline an ecosystem for open and transparent science that has emerged within the human neuroimaging community. We discuss the range of open data sharing re…
Review of clustering methods for functional data across various fields.
problem Identify heterogeneous morphological patterns in continuous functions.
method Comprehensive review and systematic taxonomy of existing methods.
result Proposes a new taxonomy linking functional data clustering to conventional multivariate methods.
Data science principles enhance AI interpretability for better user control.
problem Risks from opaque AI models without clear impacts.
method Synthesizes principles from interpretability literature, emphasizing audience goals.
result Illustrates basic techniques and criteria for evaluating interpretability.
Building and expanding on principles of statistics, machine learning, and scientific inquiry, we propose the predictability, computability, and stability (PCS) framework for veridical data science. Our framework, comprised of both a workflow and documentation, aims to provide responsible, reliable, reproducible, and tr…
SCIENCE improves prediction intervals for individual causal effects.
problem Wide prediction intervals limit practical utility of causal inference.
method Surrogate-assisted conformal inference for efficient individual causal effects.
result SCIENCE produces more efficient prediction intervals for individual causal effects.
Pipeline for comparing trading algorithms in finance and crypto.
problem Disconnected research and applications in algorithmic trading.
method General pipeline for designing, programming, and evaluating trading strategies.
result Systematic comparison of trading algorithms in finance and crypto.
LLMs can select predictive features without seeing training data.
problem Selecting the best features for prediction tasks.
method Zero-shot prompting LLMs with feature names and task descriptions.
result LLMs consistently identify predictive features across various mechanisms.
The paper uses a graph autoencoder to learn unbiased plant-pollinator interaction embeddings.
problem Sampling bias in citizen science data affects ecological network analysis.
method Bipartite graph variational autoencoder with HSIC for fairness.
result The method mitigates sampling bias and provides unbiased embeddings.
Data scientists guide to streamflow prediction and flood forecasting.
problem Forecasting floods and predicting streamflow volume.
method Explains hydrologic concepts and machine learning applications.
result Helps data scientists understand streamflow prediction.
ML4Chem offers a user-friendly platform for developing and deploying machine learning models in chemistry.
problem Developing and deploying machine learning models in chemistry and materials science.
method User-experience design, six core building blocks: data, featurization, models, model optimization, inference, and visualization.
result Ease of use and functionality of the atomistic module for neural networks and kernel ridge regression.
Paper evaluates synthetic retail data for fidelity, utility, and privacy.
problem Ensuring accurate synthetic data in retail.
method Differentiates between continuous and discrete data, measures fidelity and utility, and uses Differential Privacy for privacy.
result Validated framework for reliable and scalable synthetic data evaluation.
The authors seek financial datasets to benchmark feature engineering methods on US market data.
problem Improving predictive models for financial data science competitions.
method Feature engineering methods applied to multivariate time-series data from the US market.
result Predictive power of models tested against Numerai-Signals targets.
The paper uses data science to predict stock trends of Amazon, Apple, Google, and Microsoft.
problem Short-term market movement prediction for major tech stocks.
method Combination of technical analysis and machine/deep learning for trend classification.
result Generated labels for data set: +1 (buy), 0 (hold), -1 (sell).