ATLAS separates invariant and transferable latent factors across diverse environments.
problem Transfer learning and robust prediction in heterogeneous environments.
method ATLAS leverages invariance principle to disentangle latent factors and uses auxiliary labels for robust prediction.
result Near-oracle performance and robust transferable prediction in new environments.
Agent learns independently controllable factors by interacting with its environment.
problem Discovering independent controllable factors in environments.
method Proposes a training framework for an agent to interact with its environment and learn independently controllable factors.
result Agent can disentangle independently controllable aspects of the environment without extrinsic reward.
Agent learns independently controllable factors by interacting with environment.
problem Discovering independently controllable factors of variation in interactions with the world.
method Proposed a training framework for an agent to interact with the environment, hypothesizing that independently controllable factors correspond to aspects of the environment that can be manipulated.
result Agent can disentangle independently controllable aspects of the environment without extrinsic reward signals.
Proposes a method to learn policies from offline data with reduced bias.
problem Learning policies from offline data with reduced bias and complexity constraints.
method Cross-fitted debiasing device for policy learning from offline data.
result Achieves N \sqrt N N regret for complex policy classes with a product-of-errors nuisance remainder. Bayes Factor Surprise enables rapid adaptation to changing environments.
problem Learning in volatile, non-stationary stochastic environments.
method Bayesian inference in a hierarchical model with a Bayes Factor Surprise probability ratio.
result Novel surprise-based algorithms improve parameter estimation and performance.
Enhances speech in noisy environments using neural networks and NMF.
problem Speaker-independent multichannel speech enhancement in unknown noisy conditions.
method Uses variational autoencoders for supervised speech modeling and NMF for unsupervised noise modeling.
result The proposed approach outperforms NMF-based methods in noisy environments.
Improved acoustic scene classification with factorized CNN.
problem Acoustic scene classification in varying environments.
method Large-margin factorized CNN with triplet loss.
result Improved performance and better generalization on unseen data.
Factored contextual policy search improves data efficiency in robot learning.
problem Scalable robot learning with limited data.
method Factoring contexts into target-type and environment-type contexts, using Bayesian optimization.
result Experience can be generalized over target-type contexts, leading to faster policy generalization.
Method learns shared and specific factors in multi-study gene expression data.
problem Understanding shared and specific factors in high-dimensional multi-study data.
method Nonlinear multi-study factor model with sparse variational autoencoder.
result Method recovers meaningful shared and specific factors in platelet gene expression data.
We identify which latent factors change between environments in linear causal models.
problem Identify latent factors that change between environments in linear causal models with fewer than d d d interventions. method Propose a method to identify shifted nodes in a smaller number of environments with coarser interventions.
result It is possible to identify the set of shifted nodes under mild assumptions.
CriticSMC improves planning efficiency in constrained environments.
problem Planning with hard constraints in dynamic environments.
method Sequential Monte Carlo with learned heuristic factors.
result CriticSMC reduces collision rates with low computational cost.
Improved model predicts future environment changes efficiently.
problem Efficiently predicting future changes in environments for agents.
method Recurrent neural networks for high-dimensional pixel observations, reducing computational load.
result Model can predict hundreds of time-steps into the future, improving exploration and adaptability.
Bayesian optimization for contextual policy search improves robot learning.
problem Scalable robot learning with limited data.
method Factored contextual representation using target and environment contexts.
result Experience can be generalized over target contexts, leading to faster learning and better generalization.
StockAgent uses AI to simulate real-world stock trading, analyzing external factors and profitability.
problem Investors need to understand how external factors affect stock trading.
method Developed StockAgent, a multi-agent system driven by large language models.
result Identified how external factors impact trading behavior and profitability.
Paper uses deep reinforcement learning for optimal stock portfolio management.
problem Optimizing stock portfolio choices in complex market environments.
method Direct deep reinforcement learning to learn factor representations and make optimal decisions.
result Deep learning outperforms average market performance in portfolio allocation.
This study compares Matlab and OpenCV for machine learning algorithms.
problem Comparing execution speeds of Matlab and OpenCV for machine learning.
method 20 real datasets, 20 different machine learning algorithms.
result OpenCV is significantly faster than Matlab in execution.
PSRL extension for continuing environments reduces regret.
problem Formalizing and analyzing resampling approach for reinforcement learning.
method Continuing PSRL maintains a model of the environment and replaces it with samples from the posterior distribution.
result Established an i l d e O ( τ S A T ) ilde{O}(τS \sqrt{A T}) i l d e O ( τ S A T ) bound on Bayesian regret. New method learns robust representations by modeling environment variation.
problem Learning invariant representations across varying environments.
method Explicitly modeling variation across environments and marginalizing it out.
result Proposed method outperforms invariant-learning methods in various settings.
Paper uses RL for high-level character control in 3D environments.
problem Creating intelligent characters with generalizable behavior.
method Combines traditional animations, heuristics, and reinforcement learning.
result Demonstrates learning of complex behaviors in a 3D environment.
The article proposes a dynamic model for a company's life cycle under competitive influence.
problem Modeling a company's life cycle in a competitive environment.
method Utilized Markov model with known action costs and transition probabilities, affected by outside factors.
result Demonstrates the usefulness of the model in determining future actions of a company.
New study finds environment significantly suppresses star formation in galaxies, contrary to previous beliefs.
problem Understanding the role of environment in galaxy formation and evolution.
method Applied causal inference framework to IllustrisTNG simulations.
result Environment suppresses star formation by a factor of ~100, contrary to previous beliefs.
Unified model improves speech enhancement in unseen environments.
problem Robustness against unknown environments in speech enhancement.
method Probabilistic integration of VAE and NMF.
result Outperforms conventional DNN-based method in unseen environments.
This paper improves route choice models by incorporating contextual factors.
problem Existing route choice models lack consideration of dynamic contextual conditions.
method Knowledge distillation from Stated Choice Experiments in Immersive Virtual Environment.
result High-fidelity route choice models with increased predictive power.
GANs improve building performance model accuracy by integrating occupant behaviors.
problem Discrepancies between design and operation performance in buildings.
method Generative Adversarial Networks (GANs) to learn mixture models combining existing BPMs with occupant behaviors.
result Augmented BPMs significantly outperform existing BPMs in achieving specified performance targets.
Aiming at quantifying and evaluating the regional commercial environment along with the level of economic development among cities in mainland China, the concept of China City Commercial Environment Credit Index(CEI) was first introduced and established in 2010. In this manuscript, a historical review and detailed intr…
Neural production systems learn visual dynamics by applying rule templates to entities.
problem Modeling interactions among entities in structured visual environments.
method Inspired by production systems, the paper uses rule templates to bind placeholder variables to specific entities, scoring and applying the best fitting rules to update entity properties.
result The architecture achieves robust future-state prediction and extrapolation from simple to complex environments, outperforming GNNs.
New algorithm learns efficiently in multi-agent settings.
problem Efficient learning in multi-agent Markov decision processes.
method Cooperative Prioritized Sweeping: model-based reinforcement learning with sample efficiency.
result Outperforms state-of-the-art on SysAdmin and randomized environments.
FinRL-Meta offers market environments and benchmarks for financial reinforcement learning.
problem Challenges in creating high-quality market environments and benchmarks for financial reinforcement learning.
method DataOps paradigm, automatic pipeline, community-wise competitions, Jupyter/Python demos.
result Openly accessible FinRL-Meta library for data-driven financial reinforcement learning.
To meet the Basel II regulatory requirements for the Advanced Measurement Approaches, the bank's internal model must include the use of internal data, relevant external data, scenario analysis and factors reflecting the business environment and internal control systems. Quantification of operational risk cannot be base…
Intelligent control for greenhouses using deep reinforcement learning.
problem Uncertain nonlinear system of greenhouse environment control.
method Model Embedded Deep Reinforcement Learning (MEDRL) with computer vision and crop growth models.
result Precision and convenience in precise control of greenhouse environment.
This paper learns actionable representations for reinforcement learning.
problem Learning comprehensive representations in reinforcement learning.
method Focuses on goal-conditioned policies to learn salient, actionable representations.
result Actionable representations improve exploration and hierarchical reinforcement learning.
Empowerment quantifies the influence an agent has on its environment. This is formally achieved by the maximum of the expected KL-divergence between the distribution of the successor state conditioned on a specific action and a distribution where the actions are marginalised out. This is a natural candidate for an intr…
The paper shows how to learn causal representations with few environments and finite samples.
problem Learning causal representations from limited data and environments.
method Explicit, finite-sample guarantees with a logarithmic number of interventions.
result Consistent recovery of latent causal graph, mixing matrix, and unknown intervention targets.
Optimistic algorithms and Thompson sampling use info-theory for better reinforcement learning.
problem Designing algorithms that balance exploration and exploitation in reinforcement learning.
method Integrating information-theoretic concepts into optimistic algorithms and Thompson sampling.
result Cumulative regret bound depends on uncertainty and quantifies prior information value.
SMEs provide a transparent testbed for RL evaluation.
problem Lack of precise, white-box diagnostics in RL environments.
method Synthetic Monitoring Environments (SMEs) with fully configurable task characteristics and known optimal policies.
result SMEs allow for precise evaluation of RL algorithms, revealing the impact of specific environmental properties.
Proposes CC-NMDF for analyzing manifold-valued data.
problem Nonlinear structure in manifold-valued data requires new analysis methods.
method Curvature-corrected nonnegative manifold data factorization (CC-NMDF) with an iterative algorithm.
result Demonstrates CC-NMDF on real-world diffusion tensor MRI data.
Introduces LTQL for factored policies in cooperative MARL.
problem Learning optimal joint policies in collaborative MARL scenarios.
method Logical Team Q-learning (LTQL) as a stochastic approximation to dynamic programming.
result LTQL provides factored policies for optimal joint behavior in cooperative MARL.
Jigsaw-VAE tackles feature imbalance in VAE latent variables, improving generalization across environments.
problem Feature imbalance in VAE latent variables leads to poor generalization and biased sample generation.
method Proposes a regularization scheme to balance features in VAE latent variables and introduces a metric to measure balance.
result The regularization scheme substantially addresses feature imbalance, leading to improved generalization and diverse sample generation.
Agents struggle with solving tasks in new environments, but new models improve performance.
problem Rapid task-solving in novel environments.
method Developed EPNs to enable deep RL agents to plan over gathered knowledge.
result EPNs enable deep RL agents to excel at RTS, outperforming baselines by factors of 2-3.
New model predicts urban sprawl sensitivity from remote sensing data.
problem Forecasting urban sprawl sensitivity to economic factors.
method Physics-constrained conditional GANs for image-to-image translation.
result Model accurately predicts urban sprawl sensitivity without detailed data.
Q-Learning overestimation bias influenced by learning rate, discount factor, and reward signal.
problem Overestimation bias in Q-Learning algorithm.
method Investigated the influence of learning rate, discount factor, and reward signal on Q-Learning's overestimation bias. Tuned parameters and used an exponential moving average of reward signal.
result Q-Learning can achieve more accurate value estimates by tuning parameters and using an exponential moving average of reward signal.
Optimizes investment portfolios with multiple correlated volatility factors.
problem Maximizing utility in a stochastic environment with multiple correlated volatility factors.
method Perturbation technique around perfectly correlated factors, reducing to single factor problem; numerical solution of linear equations.
result Approximation method reduces complexity of fully non-linear HJB equation to linear equations in lower dimension.
We propose a novel receiver for orthogonal frequency division multiplexing (OFDM) transmissions in impulsive noise environments. Impulsive noise arises in many modern wireless and wireline communication systems, such as Wi-Fi and powerline communications, due to uncoordinated interference that is much stronger than the…
GCRL learns causal factors for motion forecasting, improving out-of-distribution prediction.
problem Sensitivity to out-of-distribution data in conventional supervised learning methods.
method Generative Causal Representation Learning (GCRL) leveraging causality for knowledge transfer.
result Significantly outperforms prior models on out-of-distribution prediction.
Paper proposes a reinforcement learning method for trading using expert trajectories.
problem Inability of existing methods to handle long-term goals and delayed rewards in futures trading.
method Modeling futures trading as MDP, using reinforcement learning with expert trajectories and multiple short-term alpha factors.
result The proposed method outperforms traditional and deep learning methods in trading performance.
RL methods applied to option pricing using modified QLBS and RLOP models.
problem Applying reinforcement learning to price options accurately.
method Developed modified QLBS and RLOP models, implemented RL learning algorithm with neural networks.
result Optimal hedging strategies learned by RL outperform baseline models.
The paper explains what affects the generalization gap in visual RL with and without distractors.
problem Understanding what affects the generalization gap in visual reinforcement learning.
method Theoretical analysis and empirical evidence.
result Minimizing representation distance between training and testing environments reduces the generalization gap.
Paper presents a deep learning method for estimating asset return precision matrices in noisy financial markets.
problem Estimating precision matrices of asset returns in low signal-to-noise ratio environments.
method Non-linear factor model within deep learning framework, consistent estimator with error covariance estimator.
result Superior accuracy in simulations and empirical data.