New method shows data-driven causal studies can be misleading.
problem Misattribution of causality in data-driven earth science studies.
method Subsample-based ensemble approach for robust causality analysis.
result Transfer entropy-based causal graphs can be spurious.
ERFit identifies dynamic equations from data with minimal supervision.
problem Data-driven sparse system identification in science and engineering.
method Entropic Regression method.
result ERFit package simplifies sparse system identification for various applications.
The study explores how machine learning can enhance scientific research.
problem Improving scientific models with machine learning.
method Analysis of data-driven models versus manually added variables in regression.
result Complex models may not always improve over simpler ones in scientific contexts.
Data science enhances knot theory by analyzing invariant relations.
problem Understanding the complex relations between knot invariants.
method Topological data analysis applied to knot theory.
result New insights into long-standing conjectures about knots.
DUE framework models unknown equations from data using deep learning.
problem Unknown equations in complex systems.
method Data-driven modeling using deep learning techniques.
result Framework capable of learning various types of unknown equations.
Data scientists guide to streamflow prediction and flood forecasting.
problem Forecasting floods and predicting streamflow volume.
method Explains hydrologic concepts and machine learning applications.
result Helps data scientists understand streamflow prediction.
Geotechnics adopts data-driven methods from materials informatics.
problem Soil complexity and lack of comprehensive data.
method Leveraging deep learning and transfer learning for feature extraction.
result Revolutionary potential of advanced computational tools in geotechnics.
Centuries of development in natural sciences and mathematical modeling provide valuable domain expert knowledge that has yet to be explored for the development of machine learning models. When modeling complex physical systems, both domain knowledge and data provide necessary information about the system. In this paper…
Enhances Bayesian model selection for high-dimensional problems.
problem Bayesian model selection for high-dimensional problems.
method Proximal nested sampling with data-driven priors.
result Improves model selection for log-convex likelihood models.
The optimization of composition and processing to obtain materials that exhibit desirable characteristics has historically relied on a combination of scientist intuition, trial and error, and luck. We propose a methodology that can accelerate this process by fitting data-driven models to experimental data as it is coll…
Conversion of raw data into insights and knowledge requires substantial amounts of effort from data scientists. Despite breathtaking advances in Machine Learning (ML) and Artificial Intelligence (AI), data scientists still spend the majority of their effort in understanding and then preparing the raw data for ML/AI. Th…
TNet combines DL with physics models to solve inverse problems efficiently.
problem Solving inverse problems with limited data and physics constraints.
method Model-constrained deep learning approach using TNet.
result TNet solutions are as accurate as traditional methods but faster.
New method handles complex systems with discontinuous, heavy-tailed noise.
problem Handling discontinuous, heavy-tailed Lévy noise in stochastic systems.
method Developed nonlocal Kramers-Moyal formulas for SDEs with multiplicative Lévy noise.
result Validated framework for discovering interpretable SDE models from data.
Data science principles enhance AI interpretability for better user control.
problem Risks from opaque AI models without clear impacts.
method Synthesizes principles from interpretability literature, emphasizing audience goals.
result Illustrates basic techniques and criteria for evaluating interpretability.
A new neural network model uses polynomial chaos theory to improve neural signal processing.
problem Redundant neural signal representation in DANNs.
method Employing arbitrary polynomial chaos theory to construct orthonormal representations in DANNs.
result Improves neural signal processing by reducing redundancy and enhancing orthogonality.
Continued reliance on human operators for managing data centers is a major impediment for them from ever reaching extreme dimensions. Large computer systems in general, and data centers in particular, will ultimately be managed using predictive computational and executable models obtained through data-science tools, an…
As the Industrial Internet of Things (IIoT) grows, systems are increasingly being monitored by arrays of sensors returning time-series data at ever-increasing 'volume, velocity and variety' (i.e. Industrial Big Data). An obvious use for these data is real-time systems condition monitoring and prognostic time to failure…
Study optimizes CANN for actuarial tasks using RSM.
problem Optimizing hyperparameters for neural networks in actuarial science.
method Factorial design and response surface methodology (RSM).
result Reduced hyperparameter optimization from 288 to 188, achieving near-optimal performance.
Paper presents a workflow for reliable unsupervised learning in science.
problem Lack of standardization in unsupervised learning workflows for reproducible scientific discoveries.
method Structured workflow including data preparation, modeling, validation, and communication.
result Illustrates the importance of validation in unsupervised learning.
EcoCast predicts biodiversity risks using satellite data and citizen science records.
problem Unprecedented shifts in species distributions due to climate change and habitat loss.
method Spatio-temporal model using sequence-based transformers and continual learning.
result Promising improvements in forecasting bird species distributions compared to Random Forest.
A main challenge of data-driven sciences is how to make maximal use of the progressively expanding databases of experimental datasets in order to keep research cumulative. We introduce the idea of a modeling-based dataset retrieval engine designed for relating a researcher's experimental dataset to earlier work in the …
LLMs can select predictive features without seeing training data.
problem Selecting the best features for prediction tasks.
method Zero-shot prompting LLMs with feature names and task descriptions.
result LLMs consistently identify predictive features across various mechanisms.
Interpretable ML helps discover insights from big data.
problem Validating data-driven discoveries from complex datasets.
method Statistical and machine learning techniques for interpretable models.
result Challenges in validating data-driven discoveries remain.
Data-driven decision-making often overestimates benefits due to the winner's curse.
problem Accurate policy evaluation in data-driven decision-making.
method Model-based policy evaluation using estimated models from data.
result Model-based methods can produce large, spurious reported benefits even when true effects are zero.
Recent research has helped to cultivate growing awareness that machine learning systems fueled by big data can create or exacerbate troubling disparities in society. Much of this research comes from outside of the practicing data science community, leaving its members with little concrete guidance to proactively addres…
CGD improves diffusion models' out-of-distribution generalization.
problem Reliable sampling from high-value regions beyond training data.
method Context-guided diffusion (CGD) using unlabeled data and smoothness constraints.
result Substantial performance gains across various diffusion processes.
Meta-learning, or learning to learn, is the science of systematically observing how different machine learning approaches perform on a wide range of learning tasks, and then learning from this experience, or meta-data, to learn new tasks much faster than otherwise possible. Not only does this dramatically speed up and …
DeepCausalMMM models marketing impacts using deep learning and causal inference.
problem Traditional MMM approaches struggle with non-linear dynamics and temporal patterns.
method Combines deep learning, causal inference, and marketing science. Uses GRUs for temporal patterns and DAG structure for channel dependencies.
result Captures non-linear dynamics and temporal patterns in marketing impacts.
Bayesian machine scientist uncovers accurate models from data.
problem Challenging scientific problems require interpretable mathematical models.
method Bayesian approach using Markov chain Monte Carlo to explore model space.
result Out-of-sample predictions are more accurate than existing methods.
WeldNet reduces complex dynamics to simpler, manageable segments.
problem Complex, high-dimensional time-dependent datasets from physical processes are costly to simulate.
method Windowed Encoders for Learning Dynamics, splitting time domain into windows for nonlinear dimension reduction and propagator training.
result WeldNet captures nonlinear latent structures and dynamics, outperforming existing methods.
New method improves local precipitation predictions using video diffusion.
problem Limited high-resolution local precipitation predictions due to computational costs.
method Extends video diffusion models to capture conditional distribution of high-resolution patterns.
result Method outperforms state-of-the-art baselines in CRPS, MSE, and precipitation distribution.
Proposes LVGP for multi-source data fusion in science and engineering.
problem Differences in quality and comprehensiveness of data sources.
method Latent Variable Gaussian Process (LVGP) framework.
result Improved predictions for sparse-data problems.
Labor market institutions are central for modern economies, and their polices can directly affect unemployment rates and economic growth. At the individual level, unemployment often has a detrimental impact on people's well-being and health. At the national level, high employment is one of the central goals of any econ…
The goal of data-driven algorithm design is to obtain high-performing algorithms for specific application domains using machine learning and data. Across many fields in AI, science, and engineering, practitioners will often fix a family of parameterized algorithms and then optimize those parameters to obtain good perfo…
Time-series data is being increasingly collected and stud- ied in several areas such as neuroscience, climate science, transportation, and social media. Discovery of complex patterns of relationships between individual time-series, using data-driven approaches can improve our understanding of real-world systems. While …
Many decision problems in science, engineering and economics are affected by uncertain parameters whose distribution is only indirectly observable through samples. The goal of data-driven decision-making is to learn a decision from finitely many training samples that will perform well on unseen test samples. This learn…
Hybrid Bayesian MOT uses neural networks to improve model aspects, achieving state-of-the-art performance.
problem Improving multiobject tracking performance across various scenarios.
method Hybrid approach combining neural network enhancements with Bayesian estimation and belief propagation.
result State-of-the-art performance in autonomous driving dataset evaluation.
The numerical solution of large-scale PDEs, such as those occurring in data-driven applications, unavoidably require powerful parallel computers and tailored parallel algorithms to make the best possible use of them. In fact, considerations about the parallelization and scalability of realistic problems are often criti…
Quantum dynamics reveals hidden geometric structure in data.
problem Understanding complex, high-dimensional datasets through geometric structure.
method Introducing semiclassical and microlocal analysis to data analysis.
result First tractable algorithm for approximating wave dynamics and geodesics on data manifolds.
Researchers analyze and compare nonparametric meta-learners for estimating heterogeneous treatment effects.
problem Evaluating treatment effectiveness in empirical science, especially when effects vary among individuals.
method Theoretical analysis of four meta-learning strategies, focusing on plug-in estimation and pseudo-outcome regression.
result Theoretical insights guide algorithm design and reveal relative strengths of different learners under various data-generating processes.
Joint models are a common and important tool in the intersection of machine learning and the physical sciences, particularly in contexts where real-world measurements are scarce. Recent developments in rainfall-runoff modeling, one of the prime challenges in hydrology, show the value of a joint model with shared repres…
Method extracts governing laws from non-Gaussian stochastic systems data.
problem Modeling complex dynamics with non-Gaussian Lévy noise.
method Data-driven method to extract stochastic dynamical systems from noisy data.
result Established a theoretical framework and numerical algorithm to compute Lévy jump measure, drift, and diffusion.
This review covers AI in finance, challenges, techniques, and opportunities.
problem Challenges and opportunities in AI applications in finance.
method Comprehensive categorization and overview of AI research in finance over decades.
result A dense roadmap of AI challenges, techniques, and opportunities in finance.
Estimates price sensitivity from transaction data using a novel odds ratio method.
problem Estimate price sensitivity from transaction-level data with partially observed treatment assignments.
method Recursive partitioning procedure with adversarial imputation for robust estimation.
result Validated on synthetic data and applied to three case studies, demonstrating heterogeneity in treatment effects.
Machine learning algorithms typically rely on optimization subroutines and are well-known to provide very effective outcomes for many types of problems. Here, we flip the reliance and ask the reverse question: can machine learning algorithms lead to more effective outcomes for optimization problems? Our goal is to trai…
This report aims to improve trust in AI by explaining machine learning models.
problem Understanding and trusting automated decision-making systems.
method Survey and distillation of literature on explainable machine learning.
result Survey findings help practitioners understand and apply explainable methods.
Network analysis detects insider trading by flagging coordinated trades.
problem Detecting insider trading due to limited labelled data.
method Data-driven network approach using SEC trade data.
result Algorithm identifies insider trading clusters with high accuracy.
The paper develops sampling methods for ocean phenomena based on temperature and salinity measurements.
problem Improving oceanographic sampling with limited resources.
method Design criterion based on uncertainty in excursions of vector-valued Gaussian random fields.
result Demonstrates effective exploration of ambiguous regions for data-driven sampling.