New findings show Markov models often miss past-future dependencies.
problem Markov models fail to capture dependencies between past and future.
method Investigates how much past-future information is hidden in the present.
result Markov models often miss dependencies between past and future.
The paper identifies overfitting as the main bottleneck in efficient deep reinforcement learning.
problem Improving sample efficiency in deep reinforcement learning.
method Empirical analysis on DMC tasks to identify overfitting as the main issue and developing a hill-climbing method targeting validation TD error.
result Overfitting is the primary bottleneck in sample-efficient deep RL, and regularization techniques can control this.
Model predicts patient trajectories and interventions from EMR data.
problem Forecasting patient outcomes from EMR data.
method Deep state space generative model capturing latent state dynamics.
result Model outperforms state-of-the-art methods on real EMR data.
RAD enhances RL algorithms with data augmentations.
problem Challenges in RL learning from visual observations.
method RAD is a simple plug-and-play module for RL algorithms.
result RAD improves data-efficiency and final performance.
For arbitrary quantizable compact Kaehler manifolds, relations between the geometry given by the coherent states based on the manifold and the algebraic (projective) geometry realised via the coherent state mapping into projective space, are studied. Polar divisors, formulas relating the scalar products of coherent vec…
New framework for efficient query-based imitation learning.
problem Aligning agent policy with human expert behavior without prior knowledge.
method Adversarial reward query with successor representation.
result Significantly outperforms uncertainty-based methods in query efficiency.
New model for insurance states using Markov jump processes with non-countable state space.
problem Modeling insurance states with non-countable state spaces.
method Developed a new Thiele's differential equation for continuous time rehabilitation rates.
result Allows for consistent calculation of reserves in disability insurance.
CURL uses contrastive learning to improve reinforcement learning performance.
problem Improving reinforcement learning performance on complex tasks.
method Contrastive learning to extract high-level features from raw pixels, followed by off-policy control.
result CURL outperforms prior methods on DeepMind Control Suite and Atari Games.
The Riemannian Bures metric on the space of (normalized) complex positive matrices is used for parameter estimation of mixed quantum states based on repeated measurements just as the Fisher information in classical statistics. It appears also in the concept of purifications of mixed states in quantum physics. Here we d…
Develops inverse unscented Kalman filter for non-linear systems.
problem Estimating defender's state in adversarial settings.
method Formulated inverse unscented Kalman filter (I-UKF) and reproducing kernel Hilbert space-based UKF (RKHS-UKF).
result Proposed filters are conservative estimators with upper-bounded error covariance.
Improved stock volume prediction using Kalman Filters with various hidden states.
problem Improving accuracy of intraday trading volume prediction.
method Extended Kalman Filter with various hidden states for different stocks, using cross-validation to determine optimal state number.
result Demonstrated improved accuracy through comparison experiments and numerical analysis.
We introduce an event based framework of directional changes and overshoots to map continuous financial data into the so-called Intrinsic Network - a state based discretisation of intrinsically dissected time series. Defining a method for state contraction of Intrinsic Network, we show that it has a consistent hierarch…
Agents learn state ambiguity from non-linear sensor data using Gaussian approximations.
problem Learning state representation from non-linear sensor data.
method Second-order Taylor approximation of Gaussian distribution for non-linear measurement functions.
result Induces a preference for states based on inferability from observations.
The study uses DCC for financial market analysis, revealing hidden correlations.
problem Identifying hidden nonlinear correlations in financial markets.
method Agglomerative hierarchical clustering with distance correlation coefficient.
result DCC reveals more information than Pearson correlation for financial data.
In this work, we introduce a deep-structured conditional random field (DS-CRF) model for the purpose of state-based object silhouette tracking. The proposed DS-CRF model consists of a series of state layers, where each state layer spatially characterizes the object silhouette at a particular point in time. The interact…
DCTN uses tensor networks for image classification, achieving state-of-the-art results.
problem Improving image classification accuracy with deep neural networks.
method Developed a novel deep convolutional tensor network (DCTN) based on Entangled plaquette states (EPS).
result DCTN achieves state-of-the-art results on MNIST and FashionMNIST but overfits on CIFAR10.
RECODE uses clustering and embedding to track state visitation counts in RL.
problem Efficient novelty-based exploration in nonstationary RL environments.
method Non-parametric clustering, online density estimation, inverse dynamics loss.
result RECODE achieves state-of-the-art performance in challenging RL tasks.
Improves deep RL for partially observable environments.
problem Handling partially observable environments in deep RL.
method Action-specific Deep Recurrent Q-Network (ADRQN) architecture.
result Demonstrates effectiveness in partially observable domains.
Paper proposes method for optimal control of unknown systems with latent states.
problem Jointly estimating dynamics and latent states in systems with unmeasurable states.
method Combination of particle Markov chain Monte Carlo methods and scenario theory.
result Probabilistic performance guarantees for optimal input trajectories.
Transformers learn random walks optimally with gradient descent.
problem Transformer interpretability and learning random walks.
method Theoretical analysis and gradient descent training.
result Transformers can predict random walks optimally with gradient descent.
Develops HMRL for sparse reward RL problems, improving meta policy efficiency and transferability.
problem Difficulty in learning meta policies for sparse reward RL problems.
method Hyper-Meta RL framework with cross-environment meta state embedding and shaped meta reward.
result Improves meta policy generalization and efficiency for sparse reward RL problems.
PLUMAGE improves large model training efficiency and stability.
problem Accelerator memory and networking constraints during large model training.
method Probabilistic Low rank Unbiased Minimum Variance Gradient Estimator (PLUMAGE) that resolves bias and variance issues.
result PLUMAGE reduces training loss by 28% on average across the GLUE benchmark.
Dual model combines HMM and neural networks for energy trading during volatile periods.
problem Optimizing energy trading performance during market volatility.
method Integrates Hidden Markov Models and neural networks with Black-Litterman portfolio optimization.
result Achieved 83% return with Sharpe ratio 0.77 during COVID period.
Algorithm learns user's reward function from hypothetical behaviors.
problem Aligning agent behavior with unknown user objectives.
method Synthesizes hypothetical behaviors, asks user for rewards, trains neural network.
result Significantly outperforms prior methods in learning reward models.
Deep learning improves bank distress prediction using news data.
problem Enhance bank distress prediction using news and financial data.
method Doc2vec for text analysis, supervised neural network combining text and financial data.
result News data improves bank distress prediction accuracy.
Modeling complex systems with multi-resolution data and causal dependencies.
problem Accurate prediction of complex systems with varying causal dependencies and multi-resolution data.
method Score-based Variational Graphical Diffusion Model (Temporal-SVGDM) that constructs individual SDEs for each variable at its native resolution and couples them through a causal score mechanism.
result Improved prediction accuracy and causal understanding compared to existing methods, especially in temporal scenarios.
This work improves policy optimization by maximizing entropy of state distribution, leading to better exploration.
problem Lack of exploration in state space when maximizing policy entropy.
method Proposes maximizing the entropy of a lower bound approximation to the state weighting distribution, based on latent space representation.
result Entropy regularization based on marginal state distribution achieves superior state space coverage and better performance in various domains.
Study predicts synchronization state of financial time series using cross-recurrence plots.
problem Predicting the state of synchronization of financial time series.
method Cross-correlation analysis and deep learning framework for predicting synchronization state based on cross-recurrence plots.
result Satisfactory performance in predicting synchronization state for certain pairs of stocks.
We decode latent states in Block MDPs and learn near-optimal policies.
problem Model estimation and reward-free learning in Block MDPs.
method Information-theoretical lower bound and efficient model estimation algorithm.
result Our algorithm approaches the information-theoretical limit for latent state decoding and converges to optimal policies.
NDI aims to forecast future natural disasters risk for insurers.
problem Increasing intensity and frequency of natural disasters.
method Develops a Natural Disasters Index (NDI) based on NOAA data.
result NDI forecasts future natural disasters risk for insurers.
New RL approach tackles constrained Markov decision processes.
problem Applying RL to physical systems with safety constraints.
method Formulated as a Constrained Markov Decision Process (CMDP), introduced a safe policy improvement method.
result Agent learns to maximize returns while satisfying constraints.
Online learning algorithms are designed to perform in non-stationary environments, but generally there is no notion of a dynamic state to model constraints on current and future actions as a function of past actions. State-based models are common in stochastic control settings, but commonly used frameworks such as Mark…
Paper proposes a novel optimization method for disaggregating smart meter data.
problem Energy disaggregation, inferring appliance-specific energy consumption from aggregate meter data.
method Two-stage optimization approach: first phase uses mixed integer programming, second phase binary quadratic optimization with penalty terms and appliance constraints.
result Proposed method successfully reconstructs appliance signatures, overcoming previous optimization-based methods' limitations.
Enhances portfolio construction with tailored regime forecasts for individual assets.
problem Traditional portfolio construction methods fail to account for asset-specific market conditions.
method Hybrid framework combining unsupervised and supervised learning for regime identification and forecasting.
result Outperforms traditional portfolio models across various asset classes.
Clusters of financial market states identified over 2006-2019.
problem Understanding the statistical properties of financial markets.
method Clustering analysis of correlation matrices constructed from sliding epochs.
result Financial markets can be classified into distinct states with transitions indicating precursors to catastrophic events.
ContraBAR uses contrastive learning to learn Bayes-optimal policies in RL.
problem Learning optimal policies for unknown tasks sampled from a known distribution.
method Proposes ContraBAR, a meta RL algorithm using contrastive predictive coding (CPC) for belief inference.
result ContraBAR achieves comparable performance to state-of-the-art methods and is computationally efficient.
This paper improves Gaussian process predictions by integrating prior knowledge.
problem Gaussian processes lack predictive power when prior information is ignored.
method Derive mean and covariance functions from previous data using weighted sums of basis functions.
result Integrating prior knowledge significantly increases look-ahead time and accuracy.
A framework combining HSMM and survival analysis for lifecycle-oriented mobility analysis.
problem Understanding individual metro usage dynamics over multi-year horizons.
method A state-based lifecycle modeling framework integrating HSMM and discrete-time survival analysis.
result Identification of interpretable mobility states, transition dynamics, and state-dependent exit and re-entry processes.
Paper tackles NAT translation issues with auxiliary regularization.
problem Improves NAT translation quality by addressing repeated and incomplete translations.
method Improves decoder hidden representations via two auxiliary regularization terms.
result Significant improvement in NAT model accuracy with better inference efficiency.
One-shot path planning for multiple agents using neural networks.
problem Efficiently generating optimal or near-optimal paths for multiple agents in robotics.
method Utilizes fully convolutional neural networks for one-shot multi-agent path planning.
result Demonstrates successful generation of optimal or near-optimal paths in over 85% of cases for multi-path planning.
Agents learn truth without recalling priors via random walks on graph.
problem Agents cannot distinguish true state based on private signals alone.
method Randomly select a neighbor, refine opinion using private signal and neighbor's prior.
result Agents learn truth exponentially fast, rate depends on graph structure.
ARL bridges non-Markovian decision processes with reinforcement learning, improving foresight and stability.
problem Inaccurate foresight in non-Markovian environments due to state-based methods' limitations.
method Lifted state space into a signature-augmented manifold, using a self-consistent field approach to anticipate future path-law.
result ARL achieves deterministic evaluation of expected returns with reduced computational complexity and variance.
DL-Droid detects Android malware using deep learning and real devices.
problem Sophisticated Android malware detection challenges traditional methods.
method Deep learning system with stateful input generation on real devices.
result DL-Droid achieves up to 99.6% detection rate with dynamic + static features.
This study analyzes economic policy uncertainty indices using visibility graphs.
problem Understanding the role of economic policy uncertainty in global economies.
method Visibility graph algorithm applied to economic policy uncertainty indices.
result The economic policy uncertainty indices exhibit persistent behavior and scale-free networks.
Paper classifies economic states and optimizes portfolios for stagflationary environments.
problem Economic uncertainty and stagflationary conditions.
method Mathematical techniques for analyzing multivariate time series, economic driver analysis, self-similarity identification, and portfolio optimization.
result Constructs economic state classifications and computes economic state integrals.
Variational autoencoders improve state representation for hard quantum systems.
problem Simulating and storing quantum states is computationally infeasible.
method Introduced variational autoencoders for quantum state representation.
result Deep networks better represent hard quantum states, suggesting compositional structure.
Paper derives constraints for Bayesian Knowledge Tracing parameters.
problem Issues with EM algorithm in BKT parameter estimation.
method From first principles, derives constraints on BKT parameter space.
result Novel algorithm respects derived constraints for parameter estimation.
Quantum systems are viewed as emergent systems from the fundamental degrees of freedom. The laws and rules of quantum mechanics are understood as an effective description, valid for the emergent systems and specially useful to handle probabilistic predictions of observables. After introducing the geometric theory of Ha…