Most structure inference methods either rely on exhaustive search or are purely data-driven. Exhaustive search robustly infers the structure of arbitrarily complex data, but it is slow. Data-driven methods allow efficient inference, but do not generalize when test data have more complex structures than training data. I…
Automated synthesis of inductive invariants is an important problem in software verification. Once all the invariants have been specified, software verification reduces to checking of verification conditions. Although static analyses to infer invariants have been studied for over forty years, recent years have seen a f…
Hybridizes physical and data-driven methods for predicting physicochemical properties.
problem Predicting physicochemical properties accurately using limited data.
method Distills physical method predictions into a prior model and combines with sparse experimental data using Bayesian inference.
result Significant improvements in predicting activity coefficients at infinite dilution compared to baselines and ensemble methods.
Framework improves data-driven ROMs for complex systems using Bayesian operator inference.
problem Improving the quality of data-driven reduced-order models for complex dynamical systems.
method Develops an active learning framework using Bayesian operator inference to identify and select training parameters.
result The proposed adaptive sampling strategy consistently yields more stable and accurate ROMs than random sampling.
HI-SIGMA improves sensitivity in high-dimensional statistical inference with data-driven background models.
problem Performing high-dimensional statistical inference with complex backgrounds in high-energy physics.
method HI-SIGMA uses generative ML models to learn signal and background distributions, incorporating systematic uncertainties.
result HI-SIGMA provides improved sensitivity compared to classifier-based methods.
Most of Markov Chain Monte Carlo (MCMC) and sequential Monte Carlo (SMC) algorithms in existing probabilistic programming systems suboptimally use only model priors as proposal distributions. In this work, we describe an approach for training a discriminative model, namely a neural network, in order to approximate the …
Bayesian neural networks improve uncertainty in data-driven VFMs for oil and gas wells.
problem Uncertainty and robustness in data-driven VFMs for oil and gas wells.
method Bayesian neural networks with variational inference for uncertainty quantification.
result Variational inference provides more robust predictions on future data.
With the advent of modern data collection and storage technologies, data-driven approaches have been developed for discovering the governing partial differential equations (PDE) of physical problems. However, in the extant works the model parameters in the equations are either assumed to be known or have a linear depen…
Improved MCMC sampling for expensive, irregular likelihoods.
problem Bayesian inference challenges with irregular, expensive likelihoods.
method Adapt subset samplers, introduce data-driven proxies, adaptive controller.
result Improved HINTS algorithm achieves best sampling error in fixed budget.
Bayesian imaging uses neural networks to learn prior knowledge from data.
problem Performing Bayesian inference in imaging problems with limited prior knowledge.
method Constructs a data-driven prior on a sub-manifold of the image space using neural networks, and performs Bayesian computation on this manifold.
result Established the existence and well-posedness of the posterior distribution and moments, and demonstrated superior performance compared to existing methods.
BI-EqNO improves Bayesian inference with flexible neural operators.
problem Inaccurate estimation of marginal likelihoods in approximate Bayesian methods.
method Equivariant neural operator framework for generalized approximate Bayesian inference.
result BI-EqNO enhances both deterministic and stochastic approaches to Bayesian inference.
ADML combines debiased learning with data-driven model selection for efficient inference.
problem Debiased machine learning estimators can be unstable and biased in nonparametric models.
method Data-driven model selection techniques combined with debiased machine learning.
result ADML estimators yield superefficient inference for pathwise differentiable parameters.
Paper discovers valid IVs from data without domain knowledge.
problem Inferring causal effects from observational data with latent confounders.
method Data-driven algorithm based on partial ancestral graphs (PAGs).
result Discovering valid IVs leads to accurate causal effect estimation.
Graph neural networks detect structural perturbations from time series data.
problem Detecting structural causes of disturbances in complex systems.
method Graph neural network approach to infer structural perturbations from functional time series.
result Data-driven approach outperforms typical reconstruction methods and meets Bayesian inference accuracy.
Extends dimension reduction to data-driven settings without gradients.
problem Gradient-based dimension reduction limitations in data-driven settings.
method Score ratio matching framework, tailored parameterization, regularization, eigenvalue deflation.
result Outperforms standard score-matching for problems with low-dimensional structure.
OptCS optimizes model selection after conformal inference, controlling FDR and power loss.
problem Challenges in model selection for conformal inference, especially when limited labeled data and many model choices are available.
method OptCS framework that allows valid statistical testing after flexible data-driven model optimization, using novel multiple testing procedures.
result Valid conformal p-values constructed despite substantial data reuse, maintaining FDR control.
We introduce a framework using Generative Adversarial Networks (GANs) for likelihood--free inference (LFI) and Approximate Bayesian Computation (ABC) where we replace the black-box simulator model with an approximator network and generate a rich set of summary features in a data driven fashion. On benchmark data sets, …
As an efficient and scalable graph neural network, GraphSAGE has enabled an inductive capability for inferring unseen nodes or graphs by aggregating subsampled local neighborhoods and by learning in a mini-batch gradient descent fashion. The neighborhood sampling used in GraphSAGE is effective in order to improve compu…
MAGI-X learns unknown dynamics from data without numerical integration.
problem Difficult to propose ODEs in closed-form for complex systems.
method MAGI-X uses neural networks within a manifold-constrained Gaussian process framework.
result MAGI-X achieves competitive accuracy in fitting and forecasting with reduced computational time.
Probabilistic methods improve SHM by learning from noisy, incomplete data.
problem Noisy and incomplete SHM data, lack of prior labels.
method Probabilistic algorithms for semi-supervised, active, and multi-task learning.
result Probabilistic methods enhance SHM by incorporating new data.
VB-DeepONet uses Bayesian inference to improve DeepONet's predictions and uncertainty quantification.
problem Overfitting and lack of uncertainty quantification in DeepONet.
method Variational Bayes approach to approximate posterior distribution, reducing computational cost.
result VB-DeepONet alleviates DeepONet's limitations and provides uncertainty quantification.
Bayesian inference learns free energy landscapes from experimental data.
problem Characterize the free energy landscape of classical many-body systems from experimental data.
method Combines non-parametric Bayesian inference with physically-motivated constraints to automate the construction of approximate free energy functionals.
result Inference algorithms yield a probability distribution over free energy functionals, leading to highly accurate analytic expressions.
We introduce physics informed neural networks -- neural networks that are trained to solve supervised learning tasks while respecting any given law of physics described by general nonlinear partial differential equations. In this two part treatise, we present our developments in the context of solving two main classes …
Maximum Likelihood Estimation (MLE) is the bread and butter of system inference for stochastic systems. In some generality, MLE will converge to the correct model in the infinite data limit. In the context of physical approaches to system inference, such as Boltzmann machines, MLE requires the arduous computation of pa…
Bayesian Gaussian Process ODEs enhanced with normalizing flows for improved flexibility and accuracy.
problem Limitations of standard Gaussian Process ODEs in modeling complex scenarios.
method Introducing normalizing flows to reparameterize the ODE vector field, developing a data-driven variational learning algorithm.
result Improved accuracy and uncertainty estimates for Bayesian Gaussian Process ODEs.
We consider the problem of unveiling the implicit network structure of node interactions (such as user interactions in a social network), based only on high-frequency timestamps. Our inference is based on the minimization of the least-squares loss associated with a multivariate Hawkes model, penalized by ℓ1 and t…
The paper proposes a method to improve Bayesian inference for periodic data using data-driven priors.
problem Efficiency in approximating posterior distribution in models with periodicity.
method Construct a prior distribution from data using a Gaussian process with a periodic kernel, approximated using adaptive importance sampling.
result The proposed method improves the marginal posterior distribution of the period parameter.
We propose a method to clean covariance matrices of nonstationary systems by using time-independent eigenvalues.
problem Noise in covariance matrices of nonstationary systems with time-independent eigenvalues.
method Data-driven approach to use independent eigenvalues encoding long-term influence of future on present.
result Our method outperforms optimal stationary methods for filtering covariance matrix and its inverse.
Recent studies have shown that information disclosed on social network sites (such as Facebook) can be used to predict personal characteristics with surprisingly high accuracy. In this paper we examine a method to give online users transparency into why certain inferences are made about them by statistical models, and …
Generative model emulates climate model for 100-year forecasts.
problem Challenges in accurately simulating long-term climate data.
method Integrates DYffusion with SFNO for stable, accurate climate simulations.
result Achieves near gold-standard performance for climate model emulation.
Data-driven decision-making often overestimates benefits due to the winner's curse.
problem Accurate policy evaluation in data-driven decision-making.
method Model-based policy evaluation using estimated models from data.
result Model-based methods can produce large, spurious reported benefits even when true effects are zero.
Kernel methods accurately predict Hamiltonian systems from data.
problem Data-driven simulation of Hamiltonian systems.
method Two-step and one-step kernel-based methods for identifying and forecasting Hamiltonian systems.
result Framework achieves accurate, data-efficient predictions across various benchmark systems.
This paper integrates LLMs into SCD to improve causal inference accuracy.
problem Challenges in acquiring domain expert knowledge for causal models.
method Statistical causal prompting (SCP) for LLMs and prior knowledge augmentation for SCD.
result LLM-KBCI and SCD augmented with LLM-KBCI approach ground truths more closely.
FaIRGP model improves climate emulation with physical interpretability.
problem Lack of physical interpretability in data-driven emulators.
method Bayesian approach to a data-driven emulator of energy balance equations.
result Demonstrates skillful emulation of global and spatial surface temperatures.
Bayesian approaches have become increasingly popular in causal inference problems due to their conceptual simplicity, excellent performance and in-built uncertainty quantification ('posterior credible sets'). We investigate Bayesian inference for average treatment effects from observational data, which is a challenging…
We propose flow-based likelihoods to accurately capture non-Gaussian data.
problem Bypassing the Gaussian assumption in scientific analyses.
method Use optimization targets of flow-based generative models to reconstruct likelihoods.
result Flow-based likelihoods can accurately capture non-Gaussian data, improving parameter constraints.
Study examines inference methods after variable selection in Cox models.
problem Bias and misleading inference after variable selection in Cox models.
method Simulation study of inference procedures for Lasso and adaptive Lasso in Cox models.
result Performance of inference procedures varies, with debiased Lasso showing promise.
Robo-advisor uses ML to optimize investment performance.
problem Maximizing investment performance with historical data.
method Inverse optimization and deep reinforcement learning.
result Robo-advisor consistently outperformed S&P 500.
Interpretable ML helps discover insights from big data.
problem Validating data-driven discoveries from complex datasets.
method Statistical and machine learning techniques for interpretable models.
result Challenges in validating data-driven discoveries remain.
Many important schemes in signal processing and communications, ranging from the BCJR algorithm to the Kalman filter, are instances of factor graph methods. This family of algorithms is based on recursive message passing-based computations carried out over graphical models, representing a factorization of the underlyin…
The relationship between statistical dependency and causality lies at the heart of all statistical approaches to causal inference. Recent results in the ChaLearn cause-effect pair challenge have shown that causal directionality can be inferred with good accuracy also in Markov indistinguishable configurations thanks to…
This paper presents a fast Bayesian filtering technique for state estimation.
problem Bottleneck in Bayesian inference for state estimation from noisy sensor data.
method Processor-native uncertainty tracking for uncertainty propagation and inference.
result Deterministic approximate filtering with up to 805x speedup and competitive accuracy.
Our goal is for agents to optimize the right reward function, despite how difficult it is for us to specify what that is. Inverse Reinforcement Learning (IRL) enables us to infer reward functions from demonstrations, but it usually assumes that the expert is noisily optimal. Real people, on the other hand, often have s…
Novel framework combines tree-based discretization and ILP matching for causal inference.
problem Challenges in identifying causal relationships from observational data.
method Combines tree-based discretization and ILP matching for causal inference.
result Yields computational efficiency and less biased ATT estimates.
Proposes a physics-informed VAE for disentangling physics from confounding influences.
problem Challenges in inferring and predicting physical systems under partial knowledge.
method Physics-informed variational autoencoder with adversarial training.
result Model successfully disentangles known physics from confounding influences.
Deep learning, an area of machine learning, is set to revolutionize patient care. But it is not yet part of standard of care, especially when it comes to individual patient care. In fact, it is unclear to what extent data-driven techniques are being used to support clinical decision making (CDS). Heretofore, there has …
Proposes a method to represent high-dimensional covariates for causal inference.
problem Inefficient and unreliable causal inference with high-dimensional covariates.
method Machine-learning-assisted covariate representation approach.
result Statistical reliability and performance guarantees for proposed methods.
Paper introduces RVNP to improve SBI in misspecified models.
problem Misspecification in simulation-based inference leads to unreliable posterior estimation.
method RVNP uses variational inference and error modeling to bridge the simulation-to-reality gap.
result RVNP can recover robust posterior inference without hyperparameters or priors.