New RL approach uses resets to learn complex MDPs efficiently.
problem Learning complex MDPs with high-dimensional states and function approximation.
method Local simulator access and recursive value function search.
result Proven sample-efficient learning for MDPs with low coverability.
Efficient local planning with linear approximations for agents with limited simulator access.
problem Planning with limited simulator access in reinforcement learning.
method Confident Monte Carlo Least Square Policy Iteration (Confident MC-LSPI) and Politex (Confident MC-Politex) algorithms.
result The algorithms can learn the optimal policy with local simulator access, even for linear Q-functions.
TASID learns policies in high-dimensional settings with abstract simulator knowledge.
problem RL in high-dimensional settings with limited observation knowledge.
method TASID algorithm for transfer RL from abstract simulator with bounded perturbations.
result Sample complexity polynomial in horizon, independent of number of states.
ParaMonte simplifies Monte Carlo simulations for various scientific fields.
problem Efficiently performing Monte Carlo simulations for complex models.
method Unified, high-performance, parallelized library for C, C++, Fortran.
result Automates and streamlines Monte Carlo sampling for arbitrary-dimensional functions.
Study agnostic RL in large state spaces with weak function approximation.
problem Statistical intractability of agnostic policy learning in various environments.
method Investigates agnostic policy learning with different forms of environment access.
result Agnostic policy learning remains statistically intractable with certain forms of environment access.
Generative models accelerate molecular dynamics by four orders of magnitude.
problem Femtosecond time steps limit access to slow molecular processes.
method Deep generative modeling framework that accelerates sampling.
result Quantitative characterization of equilibrium ensembles and dynamical relaxation processes.
Simulators enable learning with generalization guarantees in computationally bounded worlds.
problem Generalization in learning theory
method Simulatable processes
result Recovery of PAC learning guarantees with VC dimension
Paper uses machine learning to optimize UAV deployment for traffic offloading.
problem Optimizing UAV deployment for efficient traffic offloading from ground BSs.
method LSTM for traffic prediction, KEG algorithm for service area determination, multi-access techniques comparison.
result RSMA reduces up to 24% total power consumption compared to conventional methods.
This paper proposes a new estimation algorithm for the parameters of an HMM as to best account for the observed data. In this model, in addition to the observation sequence, we have \emph{partial} and \emph{noisy} access to the hidden state sequence as side information. This access can be seen as "partial labeling" of …
GrateTile optimizes CNN feature map storage for efficient data access.
problem Efficient storage and access of sparse CNN feature maps.
method Divides feature maps into uneven-sized subtensors, compresses and stores them in a compressed yet accessible format.
result Average 55% DRAM bandwidth reduction with minimal indexing overhead.
Generative neural network simulates characteristic functions.
problem Simulating from characteristic functions inaccessible in closed form.
method Generative neural network with Maximum-Mean-Discrepancy loss.
result Universal algorithm independent of dimensionality and function properties.
In this paper, the distributed edge caching problem in fog radio access networks (F-RANs) is investigated. By considering the unknown spatio-temporal content popularity and user preference, a user request model based on hidden Markov process is proposed to characterize the fluctuant spatio-temporal traffic demands in F…
This paper tackles belief-state selection in simulators with latent states.
problem Selecting among approximate belief-state samplers for simulators with latent variables.
method Reduces belief-state selection to conditional distribution selection, develops algorithms and analyses.
result Different formulations of belief-state selection have varying guarantees under different roll-out methods.
Deep generative networks can simulate from a complex target distribution, by minimizing a loss with respect to samples from that distribution. However, often we do not have direct access to our target distribution - our data may be subject to sample selection bias, or may be from a different but related distribution. W…
Sparse code multiple access (SCMA) has been one of non-orthogonal multiple access (NOMA) schemes aiming to support high spectral efficiency and ubiquitous access requirements for 5G wireless communication networks. Conventional SCMA approaches are confronting remarkable challenges in designing low complexity high accur…
SBBO optimizes complex spaces using sampling-based models.
problem Optimizing complex spaces with discrete variables.
method Simulation Based Bayesian Optimization (SBBO) using sampling-based surrogate models.
result Empirical effectiveness of SBBO in combinatorial optimization.
Efficiently estimates marginal posteriors for complex simulations.
problem Bayesian inference in high-dimensional, intractable likelihood scenarios.
method Simulates and estimates low-dimensional marginal posteriors, using truncated indicators.
result Simulator efficiency and robustness testing of inference results.
An electrocardiogram (ECG) is a time-series signal that is represented by one-dimensional (1-D) data. Higher dimensional representation contains more information that is accessible for feature extraction. Hidden variables such as frequency relation and morphology of segment is not directly accessible in the time domain…
PDSim simulates and estimates commodity futures prices using polynomial diffusion models.
problem Simulating and estimating commodity futures prices using polynomial diffusion models.
method Developed an R package with a Shiny app for simulation and estimation of commodity futures prices using polynomial diffusion models.
result PDSim is the only package specifically designed for the simulation and estimation of the polynomial diffusion model.
Single model learns physics from diverse data.
problem Lack of universal physics models for diverse applications.
method General Physics Transformer (GPhyT) trained on diverse physics data.
result Single model achieves superior performance across multiple physics domains.
Estimates stationary distribution from batch transitions without access to the underlying process.
problem Estimating stationary distribution from batch transitions without access to the underlying process.
method Proposes a consistent estimator based on a correction ratio function and variational power method (VPM).
result VPM provides significantly better estimates across various problems.
Study validates low latency's impact on trading profits.
problem Determining the impact of low latency on trading profits.
method Agent-based simulation of trading strategies in a controlled environment.
result Latency inversely affects trading profits; latency rank is key.
This work studies reinforcement learning in the Sim-to-Real setting, in which an agent is first trained on a number of simulators before being deployed in the real world, with the aim of decreasing the real-world sample complexity requirement. Using a dynamic model known as a rich observation Markov decision process (R…
Molecular dynamics simulations are an important tool for describing the evolution of a chemical system with time. However, these simulations are inherently held back either by the prohibitive cost of accurate electronic structure theory computations or the limited accuracy of classical empirical force fields. Machine l…
Bayesian method learns from aggregated quantile data.
problem Limited access to sensitive personal data.
method Bayesian quantile matching estimation based on order statistics.
result Correctly reflects uncertainty of empirical quantiles.
DREAM learns optimal strategies in imperfect games without needing a simulator.
problem Learning optimal strategies in imperfect-information games with multiple agents.
method DREAM is a deep reinforcement learning algorithm that converges to Nash Equilibria and coarse correlated equilibria.
result DREAM achieves state-of-the-art performance in benchmark games and is competitive with simulator-based algorithms.
Paper speeds up IoT device detection and data decoding.
problem Efficiently detect and decode massive IoT devices in grant-free random access.
method Develops multi-armed bandit approaches for more efficient detection via coordinate descent.
result Proposed bandit based algorithms achieve faster convergence rates with lower time complexity.
We use neural networks to estimate complex model posteriors efficiently.
problem Intractable likelihood functions in complex models.
method Train a neural network to map data to posterior distributions of model parameters.
result Our method converges to true posteriors in Kullback-Leibler divergence.
SBI uses neural networks to infer model parameters from simulators.
problem Computational infeasibility of Bayesian inference for complex models.
method Training neural networks on simulator-generated data.
result Efficient Bayesian inference without likelihood evaluations.
A new inference method using regression and batched discrepancies.
problem Simulating parameters from simulator outputs.
method Regression-based projection and batched discrepancy weighting.
result Method produces a self-normalized pseudo-posterior.
New method achieves near-optimal regret without simulator.
problem Adversarial linear contextual bandits with unknown loss vectors.
method Near-optimally reduces regret to sqrt(T) without simulator.
result Achieves regret of sqrt(T) without simulator, improving existing methods.
A new method simulates a lazy version of a Markov chain for empirical inference.
problem Estimating and testing unknown Markov chains with limited data.
method Simulates an α-lazy version of an unknown Markov chain, making it ergodic.
result The pseudo spectral gap can be applied to non-ergodic Markov chains.
High-fidelity quantum simulations demonstrated on short-coherence hardware.
problem Short coherence times limit the depth of quantum algorithms.
method Fixed State Variational Fast Forwarding (fsVFF) algorithm.
result Simulations of 600 time steps possible, 150x longer than previous methods.
Efficiently simulates slow dynamics of high-dimensional stochastic systems.
problem Simulating high-dimensional stochastic systems with slow dynamics and fast modes.
method Designs an algorithm to estimate an invariant manifold and its dynamics, averaging out fast modes.
result Efficient simulator of effective dynamics on low-dimensional invariant manifold.
MARL improves LBM stability and accuracy across scales.
problem Stability and accuracy issues in under-resolved LBM simulations.
method Multi-Agent Reinforcement Learning (MARL) to dynamically control local relaxation parameters.
result MARL closures stabilize simulations and recover spectra of fully resolved models.
Imitation from observation (IfO) is the problem of learning directly from state-only demonstrations without having access to the demonstrator's actions. The lack of action information both distinguishes IfO from most of the literature in imitation learning, and also sets it apart as a method that may enable agents to l…
A persistent challenge in practical classification tasks is that labeled training sets are not always available. In particle physics, this challenge is surmounted by the use of simulations. These simulations accurately reproduce most features of data, but cannot be trusted to capture all of the complex correlations exp…
MF-GLaM models improve stochastic simulator emulation with multifidelity data.
problem Challenging to emulate stochastic simulators' full conditional probability distribution.
method Proposes MF-GLaMs to efficiently emulate HF stochastic simulators using LF data.
result MF-GLaMs achieve improved accuracy or comparable performance at reduced cost.
Paper presents efficient IS for tail risk estimation with machine learning features.
problem Estimating Value at Risk and Conditional Value at Risk with black-box access.
method Efficient Importance Sampling algorithm with self-structuring transformation.
result Asymptotically optimal variance reduction in logarithmic scale.
While reinforcement learning (RL) has the potential to enable robots to autonomously acquire a wide range of skills, in practice, RL usually requires manual, per-task engineering of reward functions, especially in real world settings where aspects of the environment needed to compute progress are not directly accessibl…
This work shows how to use simulators to learn efficient exploration in real-world RL.
problem Sample complexity of real-world reinforcement learning.
method Coupling exploratory policies learned in simulators with practical approaches.
result Polynomial sample complexity in real world, exponential improvement over direct sim2real transfer.
This paper speeds up PDV model calibration by learning SPX and VIX prices.
problem Slow calibration of the 4-factor PDV model due to expensive outer simulation.
method Learning SPX and VIX prices with neural networks to reduce outer simulation time.
result Calibration times reduced to just a few seconds.
LSS learns molecular trajectories from MD data.
problem Limited integration time steps in MD simulations.
method Three deep learning networks for slow collective variables, dynamics, and configuration reconstruction.
result Generates ultra-long synthetic folding trajectories.
Proposes MCLLO for assessing and recalibrating multiclass probability predictions.
problem Limited multicategory recalibration methods for assessing and comparing model calibration.
method MCLLO recalibration method that assesses calibration without model access and is easy to interpret.
result MCLLO outperforms other methods in simulations and real-world case studies.
Sepsis is a life-threatening condition caused by the body's response to an infection. In order to treat patients with sepsis, physicians must control varying dosages of various antibiotics, fluids, and vasopressors based on a large number of variables in an emergency setting. In this project we employ a "world model" m…
Reinforcement learning (RL) methods have been shown to be capable of learning intelligent behavior in rich domains. However, this has largely been done in simulated domains without adequate focus on the process of building the simulator. In this paper, we consider a setting where we have access to an ensemble of pre-tr…
When simulating a complex stochastic system, the behavior of output response depends on input parameters estimated from finite real-world data, and the finiteness of data brings input uncertainty into the system. The quantification of the impact of input uncertainty on output response has been extensively studied. Most…
New algorithm infers reward function from agent's learning trajectories.
problem Inferring reward function from agent's learning data.
method Gradient-based approach to recover reward function.
result Improved performance compared to state-of-the-art methods.