This paper tackles infinite-dimensional diffusion bridge simulation using operator learning.
problem Challenges in simulating diffusion bridges for modeling natural data due to intractable drift terms and continuous data representations.
method Merges score matching techniques with operator learning to directly learn infinite-dimensional bridges.
result Demonstrates high efficacy in simulating diffusion bridges for various applications, including real-world biological data.
Graph Neural Networks model 3D granular flow simulations.
problem Accurate modeling of complex 3D granular flow processes.
method Graph Neural Networks approach to simulate 3D granular flow using LIGGGHTS.
result Machine learning trajectories match physical granular flow processes.
The paper proposes using path signatures for better inference in time series data.
problem Simulation models with time series data often lack tractable likelihood functions.
method Approximate Bayesian Computation with path signatures to handle sequential data.
result Theoretical guarantees on the resultant posteriors for Bayesian parameter inference.
Federated learning studies separate client data and distribution gaps.
problem Understanding performance differences in federated learning across different datasets.
method Proposed a framework to disentangle out-of-sample and participation gaps.
result Dataset synthesis strategy is crucial for realistic simulations of federated learning generalization.
Ultrasound (US) is one of the most commonly used imaging modalities in both diagnosis and surgical interventions due to its low-cost, safety, and non-invasive characteristic. US image segmentation is currently a unique challenge because of the presence of speckle noise. As manual segmentation requires considerable effo…
Posterior SBC validates inference conditionally on observed data.
problem Validating inference for specific observed data.
method Simulation-based calibration checking (SBC) adapted to use posterior parameters.
result Validates inference conditionally on observed data.
Improves efficiency of simulators that fail to return.
problem Computational inefficiency in simulators that don't return for certain inputs.
method Trains a conditional normalizing flow to propose perturbations.
result Increased computational efficiency of simulators.
Adaptive calibration improves model accuracy with fewer simulations.
problem Inefficient calibration of complex models using fixed designs.
method Bayesian adaptive experimental design to optimize simulation runs.
result The method achieves better parameter estimation with fewer simulations.
Framework synthesizes programs for simulating complex models and estimating parameters.
problem Parameter estimation for complex models requires manual encoding of fixed model structures.
method Combines LLMs for program synthesis with neural simulation-based inference.
result Identifies plausible model families from open-ended prompts with high accuracy.
FinRL-Meta creates diverse market environments for DRL in finance.
problem Inaccurate financial data and diverse market environments challenge DRL in finance.
method Open-source data processing tools, hundreds of market environments, and multiprocessing.
result FinRL-Meta improves DRL accuracy and speed in financial simulations.
Paper proposes a new framework to improve policy optimization by aligning real and simulated data distributions.
problem Inaccurate model estimation leads to performance degradation in model-based reinforcement learning.
method Introduces unsupervised model adaptation to minimize the IPM between real and simulated data distributions.
result Achieves state-of-the-art performance in sample efficiency on various continuous control tasks.
Discovery of causal relations from observational data is essential for many disciplines of science and real-world applications. However, unlike other machine learning algorithms, whose development has been greatly fostered by a large amount of available benchmark datasets, causal discovery algorithms are notoriously di…
Proposes a new Alzheimer's disease simulator for causal effect estimation.
problem Lack of suitable benchmarks for evaluating causal effect estimators in real-world healthcare data.
method Developed a simulator of Alzheimer's disease using ADNI dataset, incorporating various parameters to model complexities.
result Compared estimators of average and conditional treatment effects using the new simulator.
In this work we present a modified neural network model which is capable to simulate Markov Chains. We show how to express and train such a network, how to ensure given statistical properties reflected in the training data and we demonstrate several applications where the network produces non-deterministic outcomes. On…
This paper presents a simulator-assisted training method (SimVAE) for variational autoencoders (VAE) that leads to a disentangled and interpretable latent space. Training SimVAE is a two-step process in which first a deep generator network(decoder) is trained to approximate the simulator. During this step, the simulato…
INTAGS uses interactive simulation to improve realism in multi-agent systems.
problem Challenges in developing realistic agent-based simulators for multi-agent systems.
method INTAGS introduces a novel metric to evaluate the difference between real and synthetic multi-agent systems, optimizing a stochastic policy in reinforcement learning to adapt to interactive sequential decision-making environments.
result INTAGS generates more realistic market data compared to state-of-the-art approaches.
Bayesian calibration improves ABMs for predicting travel patterns.
problem Calibrating ABMs for accurate travel pattern predictions.
method Gaussian Process emulator with deep learning dimensionality reduction for high-dimensional, non-stationary data.
result Improved accuracy in predicting travel patterns using traffic flow data.
OmniFold uses deep learning to deconvolve high-dimensional simulations.
problem Removing detector distortions and accounting for noise processes in high-dimensional simulations.
method OmniFold is a deep learning-based approach for maximum likelihood deconvolution.
result OmniFold can remove detector distortions and account for noise processes and acceptance effects.
NPE trains neural networks to approximate posterior distributions in SIR models from final outcome data.
problem Computational challenges in Bayesian inference for SIR models with final outcome data.
method Neural posterior estimation (NPE) using a logNormal posterior approximated by a neural network.
result NPE accurately recovers reference posteriors across various population sizes and transmission regimes.
Study shows current image classification models lack robustness to real-world dataset shifts.
problem Robustness of current image classification models to natural distribution shifts in real datasets.
method Evaluation of 204 ImageNet models in 213 different test conditions.
result Little to no transfer of robustness from synthetic to natural distribution shifts.
Study finds real-world datasets contain natural experiments that can improve model performance.
problem Detecting natural experiments in real-world datasets for causal inference.
method Synthetic graph simulation and feature selection based on causal links.
result Real-world datasets contain natural experiments that can be exploited for improved model performance.
Generative Adversarial Networks simulate realistic market interactions.
problem Lack of agent-level historical data limits market simulation realism.
method Conditional Generative Adversarial Networks (CGANs) trained on real data.
result CGAN-based synthetic market generator outperforms previous methods in market responsiveness and realism.
RoPE framework calibrates misspecified simulators for reliable inference.
problem Misspecification compromises reliability of simulation-based inference.
method Data-driven calibration using optimal transport and a small calibration set.
result RoPE framework improves inference accuracy and uncertainty calibration.
Implicit probabilistic models are a flexible class of models defined by a simulation process for data. They form the basis for theories which encompass our understanding of the physical world. Despite this fundamental nature, the use of implicit models remains limited due to challenges in specifying complex latent stru…
Versatile model for High Energy Physics events.
problem Modeling complex interactions in high-energy physics data.
method Energy-based probabilistic model with multi-purpose architecture.
result Achieves success in diverse applications like simulation, anomaly detection, and particle identification.
By building on a recently introduced genetic-inspired attribute-based conceptual framework for safety risk analysis, we propose a novel methodology to compute construction univariate and bivariate construction safety risk at a situational level. Our fully data-driven approach provides construction practitioners and aca…
Neural likelihood approximates integer time series data efficiently.
problem Inference of parameters for integer-valued stochastic processes is challenging.
method Constructs a neural likelihood approximation for inference of parameters from time series data.
result Accurately approximates the true posterior with significant computational speed-ups.
Reshef et al. recently proposed a new statistical measure, the "maximal information coefficient" (MIC), for quantifying arbitrary dependencies between pairs of stochastic quantities. MIC is based on mutual information, a fundamental quantity in information theory that is widely understood to serve this need. MIC, howev…
Simulation-based inference speeds up gravitational wave data analysis.
problem High-dimensional parameter spaces and complex noise in gravitational wave data.
method Simulation-based inference methods using machine learning techniques.
result Simulation-based inference methods improve speed over traditional methods.
Enhances clustering for functional data, robust to outliers.
problem Challenges of clustering infinite-dimensional functional data and outlier sensitivity.
method Extends OCLUST algorithm to handle functional data, trimming outliers.
result Strong performance in clustering and outlier identification on simulated and real-world datasets.
Novel method for Bayesian model comparison using deep learning.
problem Comparing complex models in science with intractable likelihood functions.
method Simulation-based, purely deep learning approach that amortizes model fitting costs.
result Achieves excellent results in accuracy, calibration, and efficiency.
We examine the out-of-equilibrium phase reported by Plerou {\it et. al.} in Nature, {\bf 421}, 130 (2003) using the data of the New York stock market (NYSE) between the years 2001 --2002. We find that the observed two phase phenomenon is an artifact of the definition of the control parameter coupled with the nature of …
Extends Hawkes process for flexible residual modeling in point processes.
problem Modeling high-frequency financial data with complex residual distributions.
method Introduces self and mutually exciting point process with discretely Markovian dynamics.
result Flexible residual distributions improve intensity modeling and high-frequency data estimation.
A discrete method approximates hyperbolic curvature flow in the plane.
problem Modeling wave phenomena in solid-liquid interfaces.
method Semidiscrete finite difference method for hyperbolic curvature flow.
result Error bounds for natural discrete norms are proven.
Contextual bandit algorithms have become popular for online recommendation systems such as Digg, Yahoo! Buzz, and news recommendation in general. \emph{Offline} evaluation of the effectiveness of new algorithms in these applications is critical for protecting online user experiences but very challenging due to their "p…
Deep reinforcement learning has recently shown many impressive successes. However, one major obstacle towards applying such methods to real-world problems is their lack of data-efficiency. To this end, we propose the Bottleneck Simulator: a model-based reinforcement learning method which combines a learned, factorized …
The statistical dependencies which independent component analysis (ICA) cannot remove often provide rich information beyond the linear independent components. It would thus be very useful to estimate the dependency structure from data. While such models have been proposed, they usually concentrated on higher-order corr…
Study models opaque financial markets using multi-agent simulation.
problem Challenges in financial markets with obscured data availability.
method Multi-agent simulation with small-scale meta-heuristic methods.
result Captures bilateral market dynamics of OTC trading.
New method uses path signatures for efficient likelihood estimation in time-series data.
problem Intractable likelihood functions in complex dynamic models.
method Kernel classifier based on path signatures for sequential data.
result Path signatures yield highly performant classifiers, even with low sample numbers.
While current benchmark reinforcement learning (RL) tasks have been useful to drive progress in the field, they are in many ways poor substitutes for learning with real-world data. By testing increasingly complex RL algorithms on low-complexity simulation environments, we often end up with brittle RL policies that gene…
Flow Matching for count data improves sample quality and efficiency.
problem Mapping between count distributions across batches or time points in high-dimensional count data.
method count-FM, a flow-matching framework based on a continuous-time birth-death process with local unit jumps.
result count-FM achieves better sample quality than representative baselines while using fewer parameters.
Inverse reinforcement learning has proved its ability to explain state-action trajectories of expert agents by recovering their underlying reward functions in increasingly challenging environments. Recent advances in adversarial learning have allowed extending inverse RL to applications with non-stationary environment …
Proposes BSI for valid statistical inference on bandit algorithms.
problem Valid statistical inference on bandit algorithms' performance.
method Fits a simulator of the bandit environment from observed data and uses it to estimate mean reward under any policy.
result Proves asymptotically valid confidence intervals and maintains nominal coverage.
Probabilistic active meta-learning improves data efficiency in robotics.
problem Data-efficient learning in robotics where data collection is expensive.
method Conceptualizing meta-learning with a probabilistic latent variable model for sequential task selection.
result Improves data efficiency compared to baselines on simulated robotic experiments.
A Python tool generates synthetic data for cluster analysis from high-level descriptions.
problem Creating synthetic data for cluster analysis is laborious and requires detailed geometric parameters.
method Proposes natural language-based synthetic data generation and implements it in a Python package.
result Makes it easy to set up interpretable and reproducible benchmarks for cluster analysis.
Accurate statistical models of neural spike responses can characterize the information carried by neural populations. But the limited samples of spike counts during recording usually result in model overfitting. Besides, current models assume spike counts to be Poisson-distributed, which ignores the fact that many neur…
A2-SBNN models spatial data with copulas for non-Gaussian dependencies.
problem Capturing complex spatial relationships and extreme dependencies in non-Gaussian data.
method Embedding A2 copula into a Bayesian neural network, trained with Wasserstein loss and moment matching.
result A2-SBNN consistently delivers high accuracy across various dependency strengths.
Stochastic gradient methods enable learning probabilistic models from large amounts of data. While large step-sizes (learning rates) have shown to be best for least-squares (e.g., Gaussian noise) once combined with parameter averaging, these are not leading to convergent algorithms in general. In this paper, we conside…