Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,738 papers · 148 categories

Trend · papers per month

68135203270 · Jun 202019922001200920172026
48 results for Reflective Exploration

Bayesian RL enhances LLMs to reflectively explore and correct errors.

problem LLMs trained via RL lack reflective behaviors like rethinking and error correction.
method Bayesian RL framework that optimizes expected return under posterior distribution over Markov decision processes.
result BARL algorithm improves LLM performance in reasoning tasks.

Proposes r2SGLD for efficient constrained exploration in non-convex learning.

problem Stagnation in high-temperature chains of reSGLD in distribution tails.
method r2SGLD: replica exchange with reflection steps in a bounded domain.
result Reflection steps enhance mixing rates with quadratic improvement in domain diameter.

Survey explores interactions between four conformal dynamics branches.

problem Understanding complex dynamics through different mathematical concepts.
method Examples and general results with technical tools.
result Dynamical relations between Schwarz reflection parameter spaces and anti-rational maps/ reflection groups.

The paper explores reflection principles for lightlike line segments on maximal surfaces.

problem Reflection property does not hold for lightlike line segments on maximal surfaces.
method Analyzes reflection properties for lightlike line segments connecting shrinking singularities.
result Shows a kind of reflection principle for lightlike line segments on maximal surfaces.

Paper develops a new method for optimal stopping in American options.

problem Optimal stopping in American options with singular generators.
method Entropy-regularized penalization scheme for reflected BSDEs with singular generators.
result Limit of the penalization scheme solves a reflected BSDE with a logarithmically singular generator.

New algorithm enhances generative modeling for bounded domains.

problem Ad-hoc thresholding techniques for boundary enforcement in diffusion models.
method Reflected Schrödinger Bridge algorithm for entropy-regularized optimal transport.
result Generative modeling in diverse bounded domains with optimal transport properties.

In this paper we sketch some reflections on the pitfalls and inconsistencies of the research program - currently dominant among the profession - aimed at providing microfoundations to macroeconomics along a Walrasian perspective. We argue that such a methodological approach constitutes an unsatisfactory answer to a wel…

2006-08-14abs ↗pdf ↗

Copula models have become popular in different applications, including modeling shocks, in view of their ability to describe better the dependence concepts in stochastic systems. The class of maxmin copulas was recently introduced by Omladič and Ružić. It extends the well known classes of Marshall-Olkin and Marshall co…

2018-08-23abs ↗pdf ↗

Intrinsically motivated goal exploration processes enable agents to autonomously sample goals to explore efficiently complex environments with high-dimensional continuous actions. They have been applied successfully to real world robots to discover repertoires of policies producing a wide diversity of effects. Often th…

2018-07-04abs ↗pdf ↗

Exploration strategy design is one of the challenging problems in reinforcement learning~(RL), especially when the environment contains a large state space or sparse rewards. During exploration, the agent tries to discover novel areas or high reward~(quality) areas. In most existing methods, the novelty and quality in …

2019-06-06abs ↗pdf ↗

The paper explores rigidity and proximality in dynamical systems, proving new results about CC^*-algebras.

problem Understanding rigidity and proximality in dynamical systems and their algebraic counterparts.
method Analyzing crossed products of dynamical systems and their CC^*-algebras, focusing on uniform rigidity and proximality.
result Uniformly rigid systems are almost reflecting, and certain crossed products are reflecting.

The paper connects machine learning interpretability with learning theory.

problem Performance and explanation generalization in local machine learning models.
method Theoretical analysis and empirical validation of local approximation explanations.
result Theoretical bounds on test-time accuracy and explanation generalization.

PRISM integrates diverse rewards in MORL, improving sample efficiency and Pareto coverage.

problem Heterogeneous MORL where dense objectives dominate, leading to poor sample efficiency.
method PRISM uses reflectional symmetry and ReSymNet to reconcile temporal-frequency mismatches and accelerate exploration.
result PRISM consistently outperforms sparse-reward baselines and oracles, achieving significant Pareto gains.

Study of multidimensional control problems with reflection controls.

problem Solving control problems with reflection controls in multidimensional settings.
method Gradient descent algorithm for polytope approximations, data-driven domain estimator, episodic learning algorithm.
result Data-driven solutions for unknown diffusion dynamics with sublinear regret.

New algorithm improves graph-based active learning by identifying unexplored regions.

problem Improving graph-based active learning by identifying unexplored regions.
method Poisson Reweighted Laplacian Uncertainty Sampling (PWLL) with a diagonal perturbation.
result PWLL effectively identifies unexplored regions in graph-based data.

New method uses hindsight to make exploration robust in stochastic environments.

problem Exploration in sparse-reward or reward-free environments, especially in stochastic settings.
method Learn representations of the future that capture unpredictable aspects, using them to predict and reward only the predictable parts of the world.
result Improves exploration in Atari games and Montezuma's Revenge, robust to stochasticity.

Transfer learning borrows knowledge from a source domain to facilitate learning in a target domain. Two primary issues to be addressed in transfer learning are what and how to transfer. For a pair of domains, adopting different transfer learning algorithms results in different knowledge transferred between them. To dis…

2017-08-18abs ↗pdf ↗

Optimal Transport has recently gained interest in machine learning for applications ranging from domain adaptation, sentence similarities to deep learning. Yet, its ability to capture frequently occurring structure beyond the "ground metric" is limited. In this work, we develop a nonlinear generalization of (discrete) …

2017-12-17abs ↗pdf ↗

A discrete subgroup of the group of isometries of the hyperbolic space is called reflective if up to a finite index it is generated by reflections in hyperplanes. The main result of this paper is a complete classification of the reflective (and quasi-reflective) subgroups among the Bianchi groups and their extensions.

2012-10-09abs ↗pdf ↗

Minimal surfaces in 3-sphere created by reflections from polygons, with new examples based on pentagons.

problem Constructing minimal surfaces in 3-sphere using reflections.
method Minimal nn-gon solves free boundary problem; curvature lines combinatorics investigated.
result New examples of minimal reflection surfaces based on pentagons.

Study improves chiral photonic metasurface design using neural networks and genetic algorithms.

problem Optimizing chiral photonic metasurfaces for high chiral dichroism and reflectivity.
method Combines neural network and genetic algorithm approaches with improved fitness functions and data augmentation.
result Demonstrates a significant increase in chiral dichroism and reflectivity.

Study values and optimizes forestry leases under risk and uncertainty.

problem Valuing and optimizing forestry leases in the presence of catastrophe risk and parameter uncertainty.
method Stochastic bio-economic models, Kalman filter, maximum likelihood estimation, RBSDEs, Monte Carlo simulations.
result Conservative strategy is recommended due to parameter uncertainty.

This study compares SPX and VIX options and quantifies their relationship.

problem Understanding the relationship between SPX and VIX options markets.
method Uses moment formulas in a model-free approach to compare implied volatilities.
result SPX options reflect the extreme-strike asymptotics of VIX options and vice versa.

Variational inference lies at the core of many state-of-the-art algorithms. To improve the approximation of the posterior beyond parametric families, it was proposed to include MCMC steps into the variational lower bound. In this work we explore this idea using steps of the Hamiltonian Monte Carlo (HMC) algorithm, an e…

2016-09-26abs ↗pdf ↗

FinReflectKG - EvalBench benchmarks financial KG extraction from SEC 10-K filings.

problem Lack of universal benchmark and evaluation framework for financial KG construction.
method Agentic and holistic evaluation principles, deterministic commit-then-justify judging protocol, binary and ordinal evaluations.
result Reflection-based extraction outperforms single-pass extraction in comprehensiveness, precision, and relevance.

The abstract explores connections between reinforcement learning, scaling, and diffusion.

problem Aligning reinforcement learning with human feedback and scaling techniques.
method Clarifying connections between reinforcement learning, scaling, and diffusion.
result Introducing a resampling approach for alignment and reward-directed diffusion models.

A hyperbolic lattice is called \textit{1.21.2-reflective} if the subgroup of its automorphism group generated by all 11- and 22-reflections is of finite index. The main result of this article is a complete classification of 1.21.2-reflective maximal anisotropic lattices of rank 44.

2016-10-19abs ↗pdf ↗