Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

24487296 · Jun 202019922001200920172026
48 results for unbounded rewards

This paper studies poisoning attacks in episodic RL and discovers their effectiveness depends on reward bounds.

problem Understanding security threats to RL algorithms through poisoning attacks.
method Examined two types of poisoning attacks: reward and action manipulation, in bounded and unbounded reward settings.
result The effectiveness of poisoning attacks depends on reward bounds, with different attack costs and success rates.

Algorithm learns diffusion processes with high-dimensional state spaces.

problem Stochastic control of unbounded diffusion processes with high-dimensional state spaces.
method Adaptive partitioning and learning algorithm that refines discretization based on estimation bias and statistical confidence.
result Established regret bounds that depend on problem parameters, extending to unbounded diffusion processes.

CRIMED optimizes regret in bandits with unbounded stochastic corruption.

problem Minimizing regret in bandits with arbitrary unbounded corruptions.
method Introduces CRIMED, an asymptotically-optimal algorithm for Gaussian distributions with known variance.
result Achieves exact lower bound on regret for Gaussian distributions with high corruption probability.

In this paper, we propose a novel perturbation-based exploration method in bandit algorithms with bounded or unbounded rewards, called residual bootstrap exploration (\texttt{ReBoot}). The \texttt{ReBoot} enforces exploration by injecting data-driven randomness through a residual-based perturbation mechanism. This nove…

2020-02-19abs ↗pdf ↗

The paper proves the convergence of Q-value for Gaussian rewards.

problem Existing proofs cannot guarantee convergence of the Q-function for Gaussian rewards.
method Using the central limit theorem and relaxing the condition to E[r(s,a)2]<E[r(s,a)^2]<\infty.
result Proves the convergence of the Q-function under the condition of E[r(s,a)2]<E[r(s,a)^2]<\infty.

In this paper long-run risk sensitive optimisation problem is studied with dyadic impulse control applied to continuous-time Feller-Markov process. In contrast to the existing literature, focus is put on unbounded and non-uniformly ergodic case by adapting the weight norm approach. In particular, it is shown how to com…

2019-06-14abs ↗pdf ↗

Study optimal stopping for variable annuity contracts with discontinuous rewards.

problem Optimal timing to surrender a variable annuity contract with guaranteed minimum benefit.
method Analytical study of an optimal stopping problem with a discontinuous reward function, considering general fee and surrender charge functions.
result Characterization of the surrender region and its interrelation with fee and surrender charge functions.

We study a variant of the stochastic KK-armed bandit problem, which we call "bandits with delayed, aggregated anonymous feedback". In this problem, when the player pulls an arm, a reward is generated, however it is not immediately observed. Instead, at the end of each round the player observes only the sum of a number…

2017-09-20abs ↗pdf ↗

New algorithms for efficient return distribution approximation in reinforcement learning.

problem Efficiently approximating unknown return distributions in reinforcement learning.
method Introduced novel distributional dynamic programming algorithms for arbitrary probabilistic reward mechanisms.
result Proved error bounds for the algorithms in Wasserstein and Kolmogorov--Smirnov distances.

Two new algorithms improve performance in adversarial bandits with unbounded losses.

problem Adversarial Multi-Armed Bandits with unbounded losses.
method Developed UMAB-NN and UMAB-G for non-negative and general unbounded losses respectively.
result UMAB-NN achieves the first adaptive and scale-free regret bound for non-negative unbounded losses.

We design and study a Contextual Memory Tree (CMT), a learning memory controller that inserts new memories into an experience store of unbounded size. It is designed to efficiently query for memories from that store, supporting logarithmic time insertion and retrieval operations. Hence CMT can be integrated into existi…

2018-07-17abs ↗pdf ↗

This memoir presents a systematic study of the utility maximization problem of an investor in a constrained and unbounded financial market. Building upon the work of Hu et al. (2005) [Ann. Appl. Probab., 15, 1691--1712] in a bounded framework, we extend our analysis to the more challenging unbounded case. Our methodolo…

2017-07-01abs ↗pdf ↗

It is shown that the compactly supported identity component of the diffeomorphism group of the 2-dimensional punctured torus Tp2\mathbb T^2_p is an unbounded group. It follows that the fragmentation norm of Tp2\mathbb T^2_p is unbounded.

2011-03-18abs ↗pdf ↗

New inequalities for unbounded functions improve denoising score matching.

problem Statistical error bounds for denoising score matching with unbounded objective functions.
method Derive new concentration inequalities using McDiarmid's inequality and Rademacher complexity bounds.
result Improved statistical error bounds for denoising score matching.

Study focal surfaces of wave fronts with unbounded curvatures.

problem Characterizing singularities of focal surfaces near non-degenerate singular points.
method Characterizations based on types of singularities and geometrical properties of initial fronts.
result Investigation of Gaussian curvature behavior of focal surfaces.

Study shows unbounded Pontryagin numbers on curved manifolds.

problem Understanding unbounded Pontryagin numbers on curved manifolds.
method Analyzing rational linear combinations of Pontryagin numbers and their relation to the universal elliptic genus.
result Proves existence of unbounded Pontryagin numbers on nonnegatively curved spin manifolds.

Paper tackles online control of linear systems with unbounded noise.

problem Online control of linear systems under unbounded noise with unknown convex cost functions.
method Developed an algorithm achieving ildeO(T) ilde{O}(\sqrt{T}) high-probability regret under unbounded noise, and established O(mpoly(logT)) O({ m poly} (\log T)) regret bound for strongly convex costs and sub-Gaussian noise.
result Achieved ildeO(T) ilde{O}(\sqrt{T}) high-probability regret under unbounded noise, and O(mpoly(logT)) O({ m poly} (\log T)) regret bound for specific noise and cost conditions.

New approach finds solutions to games with unbounded controls.

problem Existence of equilibrium in mean-field games with unbounded controls.
method Weak formulation and new existence/stability results for quadratic-growth generalized McKean-Vlasov BSDEs.
result Existence of equilibrium result for non-Markovian mean-field games with unbounded control space.

The paper provides gradient estimates for Neumann semigroups on manifolds with boundary under unbounded curvature conditions.

problem Gradient estimates for Neumann semigroups on manifolds with boundary under unbounded curvature conditions.
method Establishes Bismut-type formulas and gradient estimates for Feynman--Kac semigroups on Riemannian manifolds with boundary, under geometric conditions formulated in terms of Ricci curvature and second fundamental form.
result Derives pointwise gradient estimates for the Neumann semigroup under variable, possibly unbounded, lower curvature bounds.

Improved sampling from Gaussian distributions with privacy constraints.

problem Sampling from unbounded Gaussian distributions with differential privacy.
method First $\widetilde{\mathcal{O}}\left(d ight)$-sample algorithm for unbounded Gaussians under $\left(\varepsilon, δ ight)$-differential privacy.
result A quadratic improvement over previous results, settling an open question.

In this paper, we derive Li-Yau inequality for unbounded Laplacian on complete weighted graphs with the assumption of the curvature-dimension inequality CDE(n,K)CDE'(n,K), which can be regarded as a notion of curvature on graphs. Furthermore, we obtain some applications of Li-Yau inequality, including Harnack inequality, hea…

2018-01-18abs ↗pdf ↗

Study unbounded sl3\mathfrak{sl}_3-laminations around punctures.

problem Classify and understand structures of sl3\mathfrak{sl}_3-laminations at punctures.
method Relate to root data, classify signed webs, describe tropicalization, clarify relationships with other approaches.
result Clarify the relationship between sl3\mathfrak{sl}_3-laminations and other approaches.

We consider the problem of minimizing the relative perimeter under a volume constraint in an unbounded convex body CRn+1C\subset \mathbb{R}^{n+1}, without assuming any further regularity on the boundary of CC. Motivated by an example of an unbounded convex body with null isoperimetric profile, we introduce the concept of…

2016-06-13abs ↗pdf ↗

Solves open problem on universally consistent online learning with unbounded losses.

problem Open problem on universally consistent online learning with unbounded losses.
method Constructs random measurable partitions of the instance space.
result Simple memorization rule is optimistically universal for any unbounded loss.