We study locally differentially private algorithms for reinforcement learning to obtain a robust policy that performs well across distributed private environments. Our algorithm protects the information of local agents' models from being exploited by adversarial reverse engineering. Since a local policy is strongly bei…
A new framework for adaptive behavior using reusable value profiles.
problem Adaptive behavior in changing environments requires switching among value-control regimes, but maintaining separate parameters for each situation is impractical.
method Introduces value profiles: reusable bundles of parameters assigned to hidden states, allowing for state-conditional strategy recruitment without independent parameters for each context.
result Profile-based models outperform simpler alternatives in probabilistic reversal learning, suggesting belief-dependent control of adaptive behavior.
Investigates Q value evolution in Stable Baselines for DQL in simple vs complex environments.
problem DQL in Stable Baselines struggles with simple non-game environments.
method Comparison of TrafficLight and FrozenLake environments; Q value decomposition analysis.
result Q values meander far from optimal in complex relationships between states.
Model-free reinforcement learning (RL) requires a large number of trials to learn a good policy, especially in environments with sparse rewards. We explore a method to improve the sample efficiency when we have access to demonstrations. Our approach, Backplay, uses a single demonstration to construct a curriculum for a…
We consider an SPDE description of a large portfolio limit model where the underlying asset prices evolve according to certain stochastic volatility models with default upon hitting a lower barrier. The asset prices and their volatilities are correlated via systemic Brownian motions, and the resulting SPDE is defined o…
Sparse reward is one of the biggest challenges in reinforcement learning (RL). In this paper, we propose a novel method called Generative Exploration and Exploitation (GENE) to overcome sparse reward. GENE automatically generates start states to encourage the agent to explore the environment and to exploit received rew…
Framework improves self-play for cooperative multi-agent learning.
problem Evolutionary learning converges to bad local optima in multi-agent RL.
method Add imaginary rewards using peer prediction method to elicit truthful signals.
result State-of-the-art performance in predator prey, traffic junction and StarCraft tasks.
While there are convergence guarantees for temporal difference (TD) learning when using linear function approximators, the situation for nonlinear models is far less understood, and divergent examples are known. Here we take a first step towards extending theoretical convergence guarantees to TD learning with nonlinear…
Optimal trading strategy using LQR framework with price mean-reversion.
problem Developing a dynamic trading strategy in a market with linear and quadratic costs.
method Model Predictive Control (MPC) approach to optimize trading curve with positivity constraints.
result Optimal trading curve reacts opportunistically to price changes while satisfying constraints.
Classifies reversible and strongly reversible elements in quaternionic groups.
problem Classifying reversible and strongly reversible elements in quaternionic groups.
method Proves elements are reversible if and only if they are products of skew-involutions (resp. involutions).
result Proves elements are reversible if and only if they are products of skew-involutions (resp. involutions).
The paper classifies reversible and strongly reversible elements in Hermitian isometry groups.
problem Classifying reversible and strongly reversible elements in Hermitian isometry groups.
method Classification through group theory and algebraic manipulation.
result New classification of strongly reversible elements in Sp(n).
We give a detailed account of correlations between credit sector/quality and treasury curve factors, using the robust framework of the Barclays POINT Global Risk Model. Consistent with earlier studies, we find a strong negative correlation between sector spreads and rate shifts. However, we also observe that the correl…
This paper classifies reversible and strongly reversible elements in affine groups.
problem Classifying reversible and strongly reversible elements in affine groups.
method Identifying affine transformations and using conjugacy by involutions.
result Classification of reversible and strongly reversible elements in affine groups.
Hybrid AI system combines technical, sentiment analysis for adaptive equity trading.
problem Traditional trading strategies fail during high volatility and regime shifts.
method Combines trend-following, mean-reversion, sentiment analysis, machine learning, and market regime filtering.
result Hybrid model achieved 135.49% return on investment over 24 months.
New neural net learns time-reversible symplectic dynamics.
problem Lack of time-reversibility in neural networks for symplectic systems.
method Proposes a new neural network architecture for time-reversible symplectic systems.
result Demonstrates learning of time-reversible symplectic dynamics from data.
A new trading strategy using reinforcement learning for statistical arbitrage.
problem Traditional statistical arbitrage models rely on model assumptions and price deviations from a long-term mean.
method Empirical reversion time metric, reinforcement learning framework, and state space optimization.
result Optimal mean reversion strategy identified through reinforcement learning.
Algebraic method reveals criterion for quaternionic Möbius group reversibility.
problem Characterizing reversibility in quaternionic Möbius group elements.
method Purely algebraic approach using matrix entries and conjugacy invariants.
result Explicit criterion for reversibility in terms of matrix entries.
Recent studies have shown that online portfolio selection strategies that exploit the mean reversion property can achieve excess return from equity markets. This paper empirically investigates the performance of state-of-the-art mean reversion strategies on real market data. The aims of the study are twofold. The first…
A Finsler space is said to be geodesically reversible if each oriented geodesic can be reparametrized as a geodesic with the reverse orientation. A reversible Finsler space is geodesically reversible, but the converse need not be true. In this note, building on recent work of LeBrun and Mason, it is shown that a geodes…
Study compares employers with and without anticipating strategic labor force responses.
problem Understanding and optimizing strategic interactions in labor markets.
method Formulation of causal strategic classification, theory, and experiments.
result Performatively optimal hiring policies improve employer and labor outcomes, but can also harm labor force utility.
Trend-following strategies outperform in a noisy financial market, mirroring ancient wisdom.
problem Navigating the complex, noisy financial market environment.
method Agent-based model with 10,000 agents representing different trading strategies.
result Trend-following strategies are structurally more robust than mean-reversion strategies.
Sharp stability results for reverse isoperimetric inequalities in 2D.
problem Reverse isoperimetric inequalities in the plane.
method Stability analysis of λ-convex bodies and convex bodies with smooth boundaries. result Sharp stability results for reverse isoperimetric inequalities, including inradius and Cheeger inequalities.
Let G be a group. An element g in G is called reversible if it is conjugate to g−1 within G, and called strongly reversible if it is conjugate to its inverse by an order two element of G. Let HHn be the n-dimensional quaternionic hyperbolic space. Let PSp(n,1) be the i…
Boundary rigidity proven for non-reversible Finsler metrics.
problem Recovering non-reversible Finsler metrics from boundary distance data.
method Sum of reversible Finsler norm and closed 1-form, boundary rigidity results.
result 1-form can be uniquely recovered from boundary distance data.
On-line portfolio selection has attracted increasing interests in machine learning and AI communities recently. Empirical evidences show that stock's high and low prices are temporary and stock price relatives are likely to follow the mean reversion phenomenon. While the existing mean reversion strategies are shown to …
Policy shifts between Trump and Biden impact ESG investments, creating volatility.
problem Dramatic policy shifts between Trump and Biden administrations affect ESG investments.
method Analyzes contrasting policies of Trump and Biden administrations and their impacts on ESG investments.
result Policy changes significantly influence ESG investments, leading to volatility and portfolio reassessment.
Conditional diffusion models can approximate target distributions well with Gaussian-mixture reverse kernels.
problem Approximating target distributions in conditional diffusion models.
method Using finite Gaussian mixtures with ReLU-network logits as reverse kernels, reducing the problem to static conditional density approximation.
result The resulting neural reverse-kernel class is dense in conditional KL divergence under exact terminal matching.
New knots not rationally concordant to their reverses found.
problem Identifying knots not rationally concordant to their reverses.
method Infinite family of knots constructed, rational knot concordance group analyzed.
result Infinite rank subgroup in rational knot concordance group.
The paper classifies reversible elements in Seifert-fibered spaces and braid groups.
problem Classifying reversible elements in Seifert-fibered spaces and braid groups.
method Classification of reversible elements in Fuchsian groups, application to Seifert-fibered groups, and analysis of 3-torsion elements.
result Classification and analysis of reversible and 3-torsion elements in Seifert-fibered spaces and braid groups.
Reverse annealing boosts quantum matrix factorization performance.
problem Improving quantum matrix factorization performance.
method Combining forward and reverse annealing for nonnegative/binary matrix factorization.
result Combination of forward and reverse annealing significantly improves performance.
TRS-ODENs learn dynamics with time-reversal symmetry for more efficient learning.
problem Learning dynamics with time-reversal symmetry for more efficient learning.
method Proposed a loss function and a new framework (TRS-ODENs) to learn dynamics efficiently.
result TRS-ODENs can learn dynamics from noisy and complex trajectories efficiently.
Paper proves rigidity theorems for geodesically reversible Finsler metrics.
problem Understanding geodesically reversible Finsler metrics in closed manifolds.
method Applied theory of volumes and areas on Finsler spaces to establish rigidity theorems.
result Partial explanation of the scarcity of geodesically reversible Finsler metrics in closed manifolds.
Reduces identity testing of reversible Markov chains to simpler symmetric chain tests.
problem Testing identity of reversible Markov chains from a single trajectory.
method Using lumping-congruent Markov embeddings, the problem is simplified to testing symmetric chains over a larger state space.
result Achieves state-of-the-art sample complexity for identity testing.
Proof of reverse isoperimetric inequality for black holes.
problem Reverse isoperimetric inequality for black holes in Einstein gravity.
method Geometric-analytic approach.
result Reversal of the usual isoperimetric inequality is explained by curved backgrounds governed by Einstein's equations.
New algorithm learns bridged diffusion processes without time-reversals.
problem Learning bridged diffusion processes efficiently and accurately.
method Score matching with Doob's h-transform, avoiding time-reversals.
result Outperforms existing methods in learning bridged diffusion processes.
Consider the problem of pricing options on forwards in energy markets, when spot prices follow a geometric multi-factor model in which several rates of mean reversion appear. In this paper we investigate the role played by slow mean reversion when pricing and hedging options. In particular, we determine both upper and …
Characterizes isometries between non-reversible Finsler manifolds.
problem Understanding isometries in non-reversible Finsler manifolds.
method Generalization of Myers-Nakai Theorem for Riemannian manifolds, modification of function spaces to accommodate asymmetric structure.
result Functional characterization of isometries between non-reversible Finsler manifolds.
New method trains neural samplers to sample from multi-modal distributions efficiently.
problem Mode-seeking behavior of reverse KL divergence hinders effective sampling from multi-modal target distributions.
method Minimizing reverse diffusive KL divergence along diffusion trajectories of model and target densities.
result Demonstrated enhanced sampling performance across various multi-modal distributions.
New ODE solvers improve training efficiency and accuracy.
problem Training Neural ODEs requires efficient and accurate gradient calculation.
method Presented algebraically reversible ODE solvers that are time and memory efficient, calculate exact gradients, and are numerically stable.
result Reversible solvers strictly improve upon previous architectures in efficiency and accuracy.
This paper describes an improvement in Deep Q-learning called Reverse Experience Replay (also RER) that solves the problem of sparse rewards and helps to deal with reward maximizing tasks by sampling transitions successively in reverse order. On tasks with enough experience for training and enough Experience Replay mem…
A new sampler speeds up Bayesian mixture models.
problem Sampling from Bayesian finite mixture models is slow and hard.
method Introduces a non-reversible sampling scheme for Bayesian finite mixture models.
result The new sampler outperforms classical samplers in many scenarios, especially during convergence.
Classifies orientation-reversing homeomorphisms of even periods on surfaces.
problem Classifying orientation-reversing homeomorphisms of even periods on surfaces.
method Following the approach of [1] and correcting errors in [1] for the case of periods multiple of 4.
result Classification for orientation-reversing homeomorphisms of periods multiple of 4.
We continue our study of geometric analysis on (possibly non-reversible) Finsler manifolds, based on the Bochner inequality established by the author and Sturm. Following the approach of the Γ-calculus a la Bakry et al, we show the dimensional versions of the Poincare--Lichnerowicz inequality, the logarithmic Sobolev…
A quandle orbit's orientation is problematic when reversed.
problem The natural orientation-reversal of quandle orbits is unsuitable for medial quandles.
method Defined the orientation-reversal of a quandle orbit by inverting translations, observed it's unsuitable for medial quandles.
result The natural orientation-reversal of quandle orbits is unsuitable for medial quandles.
PETRA enables parallel training of deep models with reversible architectures.
problem Challenges in parallelizing deep model training.
method Introduces PETRA, a novel approach for parallelizing gradient computations in reversible architectures.
result Achieves competitive accuracies on CIFAR-10, ImageNet32, and ImageNet using ResNet models.
New method detects inconsistencies in AHP matrices using triadic preference reversals.
problem Challenges in assessing consistency in AHP pairwise comparison matrices.
method Triadic preference reversals to detect inconsistencies between pairs of elements.
result 97% accuracy in detecting inconsistencies, significantly surpassing traditional methods.
Enhances RL in partially observable, noisy environments by uncovering causal states.
problem Making decisions based on incomplete and noisy observations in partially observable Markov decision processes (P2OMDPs). method Causal State Representation under Asynchronous Diffusion Model (CaDiff) framework, incorporating a novel asynchronous diffusion model (ADM) and a new bisimulation metric.
result Enhances returns by at least 14.18% compared to baselines on Roboschool tasks.
RER improves sample complexity by updating in reverse order.
problem Theoretical analysis limits RER's convergence rate.
method Tighter analysis for larger learning rates and longer sequences.
result RER converges faster with larger learning rates and longer sequences.