JPS improves joint policies for multi-agent collaboration in imperfect information games.
problem Learning good joint policies for multi-agent collaboration with imperfect information.
method Decomposes global changes to localized policy changes, iteratively improving joint policies without re-evaluating the entire game.
result JPS improves solutions provided by unilateral approaches and outperforms algorithms designed for collaborative policy learning.
Bringing transparency to black-box decision making systems (DMS) has been a topic of increasing research interest in recent years. Traditional active and passive approaches to make these systems transparent are often limited by scalability and/or feasibility issues. In this paper, we propose a new notion of black-box D…
This study improves stock price prediction by incorporating anticipated macroeconomic policy changes.
problem Improving accuracy in stock price prediction.
method Incorporates future expected macroeconomic policy changes and historical stock prices.
result Our method outperforms conventional approaches with an RMSE of 1.61 compared to 1.75.
Experience replay (ER) is a fundamental component of off-policy deep reinforcement learning (RL). ER recalls experiences from past iterations to compute gradient estimates for the current policy, increasing data-efficiency. However, the accuracy of such updates may deteriorate when the policy diverges from past behavio…
Paper constructs a CRRIX index to assess cryptocurrency market risks from regulatory changes.
problem Lack of indices quantifying regulatory risks in cryptocurrencies.
method CRRIX index based on news coverage frequency, using Latent Dirichlet Allocation and Hellinger distance.
result CRRIX successfully captures major policy-changing moments and synchronizes with market volatility.
CODA resolves coordination issues in offline multi-agent reinforcement learning.
problem Coordination failure in offline multi-agent reinforcement learning.
method Diffusion-based multi-agent trajectory generator for data augmentation.
result CODA resolves coordination pathologies in continuous polynomial games and complex benchmarks.
The Australian Government uses the means-test as a way of managing the pension budget. Changes in Age Pension policy impose difficulties in retirement modelling due to policy risk, but any major changes tend to be `grandfathered' meaning that current retirees are exempt from the new changes. In 2015, two important chan…
Study examines market response to concentrated policy communication using entropy measures.
problem Characterizing market response under concentrated policy communication.
method Jointly examines dispersion and information complexity (entropy) using sliding window cumulative entropy.
result Entropy captures both market volatility and narrative constraints, signaling coherent policy-driven moves.
In batch reinforcement learning (RL), one often constrains a learned policy to be close to the behavior (data-generating) policy, e.g., by constraining the learned action distribution to differ from the behavior policy by some maximum degree that is the same at each state. This can cause batch RL to be overly conservat…
Study finds stock selection ability of Chinese mutual funds is better than asset allocation ability.
problem Evaluating the performance of actively managed mutual funds in China.
method Developed performance measures for asset allocation and selection using holding-based models and compared them with Fama-French and Treynor-Mazuy models.
result Stock selection ability from holding-based models is positively correlated with Fama-French model, while industry allocation is positively correlated with Treynor-Mazuy model.
Study on rapid policy changes in reinforcement learning.
problem Rapid change of greedy policy in reinforcement learning.
method Empirical study and ablation analysis.
result Policy churn is a beneficial form of implicit exploration.
We explain a persistent cost-of-carry spread in EUA market and suggest ECB policy change.
problem Persistent cost-of-carry spread in EUA market.
method Cointegration analysis of EUA spread with credit spread and risk-free rate.
result Cointegration found between EUA spread, credit spread, and risk-free rate.
New framework forecasts both supply and demand in rental markets.
problem Booking models ignore supply, leading to regime-specific ceilings.
method Three-part coupling framework (behavioral, informational, intervention).
result Booking models learn a regime-specific ceiling and become fragile.
A new method for estimating causal effects using synthetic controls.
problem Evaluating causal effects of policy changes in settings with observational data.
method Distributional Synthetic Controls method.
result Allows construction of entire synthetic distributions for the treated unit.
In this paper, we propose an offline counterfactual policy estimation framework called Genie to optimize Sponsored Search Marketplace. Genie employs an open box simulation engine with click calibration model to compute the KPI impact of any modification to the system. From the experimental results on Bing traffic, we s…
The paper optimizes DIA purchase policies using lifecycle models and asset allocation.
problem Determining the optimal allocation to Deferred Income Annuities (DIAs).
method Employed a lifecycle model with utility of consumption and bequest, formalized optimization process, analyzed results, and extended model to include asset allocation.
result Optimal DIA allocation varies based on refundability, asset allocation, and perceived longevity.
We consider a two-agent MDP framework where agents repeatedly solve a task in a collaborative setting. We study the problem of designing a learning algorithm for the first agent (A1) that facilitates a successful collaboration even in cases when the second agent (A2) is adapting its policy in an unknown way. The key ch…
A new algebraic framework models LOBs with physics and stochastic processes.
problem Capturing the dynamics of limit order books (LOBs).
method Algebraic framework using Dirac notation and generating functions.
result Exact simulations of market scenarios using the Gillespie algorithm.
We describe two recently proposed machine learning approaches for discovering emerging trends in fatal accidental drug overdoses. The Gaussian Process Subset Scan enables early detection of emerging patterns in spatio-temporal data, accounting for both the non-iid nature of the data and the fact that detecting subtle p…
When multiple agents learn in a decentralized manner, the environment appears non-stationary from the perspective of an individual agent due to the exploration and learning of the other agents. Recently proposed deep multi-agent reinforcement learning methods have tried to mitigate this non-stationarity by attempting t…
This study analyzes EU ETS literature trends using bibliometric methods.
problem Understanding the evolving research landscape of EU ETS.
method Bibliometric analysis of Scopus database, focusing on publication trends, themes, influential authors, and journals.
result Notable increase in research activity over two decades, particularly during policy changes and economic events.
We propose a policy improvement algorithm for Reinforcement Learning (RL) which is called Rerouted Behavior Improvement (RBI). RBI is designed to take into account the evaluation errors of the Q-function. Such errors are common in RL when learning the Q-value from finite past experience data. Greedy policies or even …
Batch Reinforcement Learning (Batch RL) consists in training a policy using trajectories collected with another policy, called the behavioural policy. Safe policy improvement (SPI) provides guarantees with high probability that the trained policy performs better than the behavioural policy, also called baseline in this…
Success conditioning optimizes policies by imitating successful trajectories, solving a trust-region optimization problem.
problem Improving policies through random actions that lead to desired outcomes.
method Success conditioning, which involves collecting and updating policies based on successful trajectories.
result Success conditioning solves a trust-region optimization problem, maximizing policy improvement with a χ2 divergence constraint. The aim of this work is to address the description of hyperinflation regimes in economy. The spirals of hyperinflation developed in Brazil, Israel, and Nicaragua are revisited. This new analysis of data indicates that the episodes occurred in Brazil and Nicaragua can be understood within the frame of the model availabl…
LHIEM model predicts health, income, and employment over years.
problem Lack of path dependency in health policy simulations.
method Discrete-time microsimulation with Markov chain modules.
result Validates health care financing proposal through detailed modeling.
SARSA is an on-policy algorithm to learn a Markov decision process policy in reinforcement learning. We investigate the SARSA algorithm with linear function approximation under the non-i.i.d.\ data, where a single sample trajectory is available. With a Lipschitz continuous policy improvement operator that is smooth eno…
This work shifts focus from prediction to intervention in social systems.
problem The limitations of focusing solely on prediction in automated decision systems.
method Shift from prediction-focused paradigm to intervention-oriented approach.
result A new perspective unifies statistical frameworks and tools for ADS design, implementation, and evaluation.
MCD reformulates conditional density estimation into binary classification.
problem Conditional density estimation in statistical and machine learning.
method Marginal Contrastive Discrimination, reformulating into marginal and ratio density functions for binary classification.
result Significantly outperforms existing methods on most density models and regression datasets.
Paper proposes MMC to avoid high-density bias in clustering.
problem High-density bias in density-based clustering.
method Introduces mass distribution as a better foundation for clustering, proposing mass-maximization clustering (MMC).
result MMC avoids high-density bias and discovers clusters of arbitrary shapes, sizes, and densities.
New method minimizes robust density power-based divergences for general parametric densities.
problem Computational complexity of minimizing DPD for general parametric densities.
method Stochastic approach to minimize DPD for general parametric density models.
result Proposed method can be applied to minimize other density power-based γ-divergences.
Paper presents an efficient algorithm for linear MDP with low switching cost.
problem Large state space reinforcement learning problems with low switching cost.
method First algorithm for linear MDP with low switching cost, achieving near-optimal regret and switching cost.
result Regret bound of $\widetilde{O}\left(\sqrt{d^3H^4K}
ight)$ and near-optimal switching cost of $O\left(d H\log K
ight)$.
This study evaluates prewar Japanese financial market efficiency using time-varying models.
problem Determining when prewar Japanese financial market lost its price formation function.
method Time-varying parameter model, generalized least squares-based time-varying vector autoregressive model.
result The prewar Japanese financial market lost its price formation function in 1932.
Modes and ridges of the probability density function behind observed data are useful geometric features. Mode-seeking clustering assigns cluster labels by associating data samples with the nearest modes, and estimation of density ridges enables us to find lower-dimensional structures hidden in data. A key technical cha…
Normalizing flows improve density estimation from noisy data.
problem Estimating underlying density from noisy samples.
method Use normalizing flows for density estimation with arbitrary noise distributions, using amortized variational inference.
result Normalizing flows can outperform Gaussian mixtures for density deconvolution.
Study exact minimax rates for density estimation over convex classes, extending previous work.
problem Deriving minimax rates for density estimation over convex density classes.
method Building on Le Cam's work, determine exact minimax rates using local metric entropy.
result Exact minimax rates derived for any convex density class, including nonparametric and parametric cases.
New method uses SoS densities and α-divergences for efficient sequential transport maps.
problem Efficiently generating samples from approximated densities.
method Sequential transport maps using Sum-of-Squares (SoS) densities and α-divergences.
result Convex optimization problems with efficient semidefinite programming solutions.
The volume density of a hyperbolic link is defined as the ratio of hyperbolic volume to crossing number. We study its properties and a closely-related invariant called the determinant density. It is known that the sets of volume densities and determinant densities of links are dense in the interval [0,v_{oct}]. We cons…
TAKDE optimizes kernel density estimation for real-time dynamic processes.
problem Real-time density estimation in applications like computer vision and signal processing.
method Derives asymptotic mean integrated squared error (AMISE) upper bound for 'sliding window' kernel density estimator and proposes TAKDE as a novel, theoretically optimal estimator.
result TAKDE outperforms other dynamic density estimators in terms of test log-likelihood and runtime.
Most density-based clustering methods largely rely on how well the underlying density is estimated. However, density estimation itself is also a challenging problem, especially the determination of the kernel bandwidth. A large bandwidth could lead to the over-smoothed density estimation in which the number of density …
Study shows Lula's Zero Hunger program reduced income inequality in Brazil.
problem Income inequality in Brazil during Lula's administration.
method Breakpoint regression analysis using detailed descriptive statistics.
result The Zero Hunger program substantially reduced income inequality and provided income security for the poor.
Optimizes kernel density ratios for better predictions and information measures.
problem Improving accuracy of kernel density estimates for density ratios.
method Derives an optimal weight function using calculus of variations.
result Reduces bias in kernel density estimates, leading to improved prediction posteriors and information-theoretic measures.
Study finds a linear lower bound on conformal dimension for random hyperbolic groups.
problem Understanding conformal dimension in random hyperbolic groups.
method Building undistorted round trees from lower density groups.
result Achieves a linear lower bound in l at all densities 0<d<1/2. Chia and Nakano (2009) introduced the concept of M-decomposability of probability densities in one-dimension. In this paper, we generalize M-decomposability to any dimension. We prove that all elliptical unimodal densities are M-undecomposable. We also derive an inequality to show that it is better to represent an M-de…
FACMAC combines deep policy gradients with factored critic for multi-agent reinforcement learning.
problem Cooperative multi-agent reinforcement learning in discrete and continuous action spaces.
method FACMAC uses a centralised but factored critic, combining per-agent utilities into a joint action-value function.
result FACMAC outperforms MADDPG and other baselines on multi-agent particle environments and StarCraft II tasks.
We introduce a novel conditional density estimation model termed the conditional density operator (CDO). It naturally captures multivariate, multimodal output densities and shows performance that is competitive with recent neural conditional density models and Gaussian processes. The proposed model is based on a novel …
Quantum method improves neural density estimation in high dimensions.
problem High-dimensional density estimation with poor performance and high computational complexity.
method Adaptive Fourier features based on quantum density matrices, integrated with neural networks.
result Competitive performance compared to state-of-the-art methods in various datasets.
Roundtrip uses deep generative models for flexible density estimation.
problem Density estimation in statistics and machine learning.
method Roundtrip is a deep generative neural density estimator that uses flexible mappings.
result Roundtrip achieves state-of-the-art performance in density estimation tasks.