Study integrates machine learning with SAA for optimizing decisions based on uncertain parameters and covariates.
problem Optimizing decisions under uncertain parameters and covariates.
method Data-driven frameworks integrating machine learning prediction models within SAA for scenario generation.
result Consistent and asymptotically optimal solutions under certain conditions, with finite sample guarantees.
Optimal data-driven formulations are found for learning and decision-making with historical data.
problem Designing optimal learning and decision-making formulations from historical data.
method Define a yardstick for measuring formulation quality, then construct an optimal formulation that is uniformly closer to the true cost.
result Existence of three distinct out-of-sample performance regimes with corresponding optimal formulations.
Enhances Bayesian model selection for high-dimensional problems.
problem Bayesian model selection for high-dimensional problems.
method Proximal nested sampling with data-driven priors.
result Improves model selection for log-convex likelihood models.
Improved MCMC sampling for expensive, irregular likelihoods.
problem Bayesian inference challenges with irregular, expensive likelihoods.
method Adapt subset samplers, introduce data-driven proxies, adaptive controller.
result Improved HINTS algorithm achieves best sampling error in fixed budget.
Unified framework for data-driven priors in Bayesian inverse problems
problem Bayesian inverse problems
method Unified framework using score functions
result Evaluation of four data-driven priors
Stochastic cutting planes improve data-driven optimization speed.
problem Data-driven Mixed-Integer Nonlinear Optimization problems.
method Stochastic version of cutting-plane method.
result Stochastic algorithm converges to ε-optimal solution with high probability.
Framework improves data-driven ROMs for complex systems using Bayesian operator inference.
problem Improving the quality of data-driven reduced-order models for complex dynamical systems.
method Develops an active learning framework using Bayesian operator inference to identify and select training parameters.
result The proposed adaptive sampling strategy consistently yields more stable and accurate ROMs than random sampling.
Designs a robust data-driven decision-making model to handle multiple overfitting sources.
problem Overfitting in data-driven models due to statistical error, data noise, and data misspecification.
method Holistic distributionally robust optimization formulation combining Kullback-Leibler and Lévy-Prokhorov approaches.
result Guaranteed holistic protection against statistical error, data noise, and data misspecification.
Optimizes decisions without knowing the true distribution using historical data.
problem Optimizing decisions without knowing the true distribution.
method Combines sampling and bisection search algorithms to solve an optimization problem.
result Proves sufficient conditions for local out-of-sample optimality.
A new method for support vector regression using a data-driven insensitive parameter.
problem Determining an optimal insensitive parameter in support vector regression.
method A data-driven approach to approximate the insensitive parameter by minimizing a generalized loss function based on the likelihood principle.
result The proposed method outperforms traditional support vector regression methods and has lower computational costs.
Unified framework for subgraph-enhanced GNNs, improving prediction accuracy and reducing computation time.
problem Limited understanding of subgraph-enhanced GNNs and their relation to the Weisfeiler-Leman hierarchy.
method Theoretical framework, theoretical expressivity results, and data-driven subgraph sampling methods.
result Data-driven subgraph-enhanced GNNs outperform non-data-driven methods in predictive performance.
Develops robust MDPs for unknown disturbances with performance guarantees.
problem Unknown disturbance distribution in MDPs.
method Empirical distribution, sublevel set of distance function, weak convergence, concentration inequality.
result Robust optimal value function converges to true optimal value function with increasing sample sizes.
New method improves combinatorial optimization by overcoming inefficient sampling.
problem Wandering in contours: sampling similar solutions for long periods.
method Reheated Gradient-based Discrete Sampling with a novel mechanism.
result Superior performance across various combinatorial optimization problems.
Paper proposes a robust hypothesis testing method using Sinkhorn distance.
problem Hypothesis testing for small samples.
method Data-driven approach using Sinkhorn uncertainty sets.
result The method provides a more flexible detector compared to Wasserstein robust test.
Study optimizes sampling to avoid extreme tail risks in unknown heavy-tailed distributions.
problem Identify optimal alternative with minimal extreme tail risk from unknown heavy-tailed distributions.
method Data-driven sequential sampling policies to maximize likelihood of selecting the optimal alternative.
result Proposed methods outperform existing approaches in identifying the optimal alternative.
End-to-end algorithm for controlling bilinear systems with probabilistic noise.
problem Controlling bilinear systems with noisy data.
method Proposes an end-to-end algorithm using statistical learning theory and robust controller design.
result Derived finite sample identification error bounds and structurally suitable for control.
This research designs a data-driven partition to test independence between continuous variables.
problem Testing independence between continuous random variables.
method Empirical log-likelihood statistic and data-driven tree-structured partition.
result Strongly consistent test of independence over probability families.
Bayesian nonparametrics improves data-driven risk optimization under distributional uncertainty.
problem Improving out-of-sample performance in machine learning models due to distributional uncertainty.
method Combining Bayesian nonparametric theory and decision-theoretic preferences to propose a robust optimization criterion.
result The proposed robust optimization procedure provides favorable statistical guarantees and tractable approximations.
New method handles complex systems with discontinuous, heavy-tailed noise.
problem Handling discontinuous, heavy-tailed Lévy noise in stochastic systems.
method Developed nonlocal Kramers-Moyal formulas for SDEs with multiplicative Lévy noise.
result Validated framework for discovering interpretable SDE models from data.
This paper analyzes data-driven Newsvendor problems and finds a wide range of possible regrets.
problem Guessing the number drawn from an unknown distribution with asymmetric costs.
method Unified analysis using the notion of clustered distributions and new lower bounds.
result The entire spectrum of achievable regrets from 1 / n 1/\sqrt{n} 1/ n to 1 / n 1/n 1/ n is possible. Data-driven decision-making often overestimates benefits due to the winner's curse.
problem Accurate policy evaluation in data-driven decision-making.
method Model-based policy evaluation using estimated models from data.
result Model-based methods can produce large, spurious reported benefits even when true effects are zero.
Proposes robust assortment optimization from observational data.
problem Real-world scenarios often violate assumptions of stable customer preferences and correct choice models.
method Develops a robust framework that accounts for potential distributional shifts in customer choice behavior.
result Uncovered the notion of ``robust item-wise coverage'' as the minimal data requirement for sample-efficient robust assortment learning.
InVAErt networks use data-driven methods for system synthesis and identifiability analysis.
problem Model synthesis and identifiability analysis for complex systems.
method Deterministic encoder and decoder, normalizing flow, variational encoder, loss function penalty coefficients, latent space sampling.
result Validation through various system types, demonstrating effectiveness of the framework.
This paper presents a data-driven approach to model planar pushing interaction to predict both the most likely outcome of a push and its expected variability. The learned models rely on a variation of Gaussian processes with input-dependent noise called Variational Heteroscedastic Gaussian processes (VHGP) that capture…
Many decision problems in science, engineering and economics are affected by uncertain parameters whose distribution is only indirectly observable through samples. The goal of data-driven decision-making is to learn a decision from finitely many training samples that will perform well on unseen test samples. This learn…
As an efficient and scalable graph neural network, GraphSAGE has enabled an inductive capability for inferring unseen nodes or graphs by aggregating subsampled local neighborhoods and by learning in a mini-batch gradient descent fashion. The neighborhood sampling used in GraphSAGE is effective in order to improve compu…
OptCS optimizes model selection after conformal inference, controlling FDR and power loss.
problem Challenges in model selection for conformal inference, especially when limited labeled data and many model choices are available.
method OptCS framework that allows valid statistical testing after flexible data-driven model optimization, using novel multiple testing procedures.
result Valid conformal p-values constructed despite substantial data reuse, maintaining FDR control.
This study analyzes decision-making in diverse environments where past data may not predict future outcomes.
problem How to make decisions when past data is not indicative of future outcomes due to unobserved confounders.
method Developed a framework to analyze and bound the performance of data-driven policies in heterogeneous environments.
result Established a method to upper bound the asymptotic worst-case regret of policies and analyzed the performance of Sample Average Approximation (SAA).
We analyze the impact of the sampling interval on the estimation of Kramers-Moyal coefficients. We obtain the finite-time expressions of these coefficients for several standard processes. We also analyze extreme situations such as the independence and no-fluctuation limits that constitute useful references. Our results…
We analyze the (unconditional) distribution of a linear predictor that is constructed after a data-driven model selection step in a linear regression model. First, we derive the exact finite-sample cumulative distribution function (cdf) of the linear predictor, and a simple approximation to this (complicated) cdf. We t…
We propose a data-driven approach to solve multiscale elliptic PDEs with random coefficients based on the intrinsic low dimension structure of the underlying elliptic differential operators. Our method consists of offline and online stages. At the offline stage, a low dimension space and its basis are extracted from th…
A new method improves the interpretability of data-driven models in ironmaking processes.
problem Lack of transparency in machine learning models used in industrial processes.
method Combines Variational Autoencoder (VAE) with Local Interpretable Model-agnostic Explanations (LIME) for better model interpretability.
result Improved local fidelity of local interpretable linear models compared to LIME.
A neural network method estimates densities from characteristic functions.
problem Estimating fixed-horizon probability densities from empirical characteristic functions.
method Data-driven Fourier-mixture neural-network method trained in Fourier space.
result Competitive performance and clear gains on heavy-tailed targets.
We propose flow-based likelihoods to accurately capture non-Gaussian data.
problem Bypassing the Gaussian assumption in scientific analyses.
method Use optimization targets of flow-based generative models to reconstruct likelihoods.
result Flow-based likelihoods can accurately capture non-Gaussian data, improving parameter constraints.
We present a novel distribution-free approach, the data-driven threshold machine (DTM), for a fundamental problem at the core of many learning tasks: choose a threshold for a given pre-specified level that bounds the tail probability of the maximum of a (possibly dependent but stationary) random sequence. We do not ass…
Empirical mode modeling improves state-space analysis of noisy data.
problem Analyzing nonlinear systems with noisy data.
method Combining empirical mode decomposition with empirical dynamic modeling.
result Empirical mode modeling enhances state-space representations in noisy data.
New framework calibrates decision robustness using inverse conformal risk control.
problem Inadequate robustness levels in decision-making due to ad hoc choices.
method Constructs valid estimators to trace miscoverage-regret Pareto frontier.
result Provides distribution-free, finite-sample guarantees on robustness levels.
We introduce a methodology for nonlinear inverse problems using a variational Bayesian approach where the unknown quantity is a spatial field. A structured Bayesian Gaussian process latent variable model is used both to construct a low-dimensional generative model of the sample-based stochastic prior as well as a surro…
New insights into bias-variance tradeoff for data-driven optimization under local misspecification.
problem Understanding the relative performance of SAA, IEO, and ETO under local misspecification.
method Developed a local misspecification perspective using contiguity theory in statistics.
result Explicit expressions for decision bias and geometric understanding of variance.
Bayesian autoencoders discover physics from noisy data.
problem Challenges in identifying governing equations and coordinates from noisy, low-data real-world data.
method Bayesian SINDy autoencoders with hierarchical Bayesian sparsifying prior and adaptive empirical Bayesian method.
result Better physics discovery with lower data and fewer training epochs, along with valid uncertainty quantification.
Develops a machine learning approach for solving AC-OPF problems.
problem Nonlinear and computationally demanding AC chance-constrained OPF problem.
method Uses Gaussian process regression to approximate AC power flow equations.
result Demonstrates competitive and promising results compared to state-of-the-art approaches.
New algorithm improves knowledge transfer in dynamic decision-making.
problem Utilizing data from existing ventures to improve decision-making in new ventures.
method Proposes Transferred Fitted Q Q Q -Iteration algorithm for estimating optimal action-state function Q ∗ Q^* Q ∗ . result Significantly improved final learning error of Q ∗ Q^* Q ∗ function. We demonstrate the application of an algorithmic trading strategy based upon the recently developed dynamic mode decomposition (DMD) on portfolios of financial data. The method is capable of characterizing complex dynamical systems, in this case financial market dynamics, in an equation-free manner by decomposing the s…
Deep learning models predict S&P500 option hedge ratios.
problem Optimizing hedging strategies for S&P500 index options.
method Feedforward neural network with time to maturity, delta, and sentiment variables.
result Deep learning model outperforms traditional hedging methods.
Flexible multi-task learning framework using summary statistics.
problem Data-sharing constraints in healthcare settings.
method Proposes a flexible multi-task learning framework utilizing summary statistics and adaptive parameter selection.
result Systematic non-asymptotic analysis and simulations demonstrate the method's performance.
Improves robustness of high-dimensional regression with rank objective and group lasso regularization.
problem Heavy-tailed noise and outliers in high-dimensional regression.
method Non-smooth Wilcoxon score based rank objective, group lasso regularization, data-driven tuning rule, proximal augmented Lagrangian method.
result Robust estimator with finite-sample error bound and efficient computational method.
New method for fair resource allocation in AI-aware networks with unknown utility functions.
problem Fair resource allocation in AI-aware communication networks with unknown utility functions.
method Distributed, data-driven bilevel optimization approach to learn surrogate utility functions.
result The proposed algorithm learns from data to autotune surrogate utility functions for unknown utility functions.
HAL accelerates the generation of training sets for accurate interatomic potentials.
problem Generating accurate and transferable interatomic potentials is time-consuming and requires expert input.
method HAL framework using a physically motivated sampler with a biasing term to drive high uncertainty configurations.
result HAL-generated training databases for alloys and polymers predict macroscopic properties with high accuracy.