Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

66131197262 · Jun 202019922001200920172026
48 results for Markov Transition Kernel

New method detects changes in high-dimensional Markov processes without explicit likelihood evaluation.

problem Quickest change detection in Markov processes with unknown transition kernels.
method Learn conditional score from sample pairs, develop score-based CUSUM procedure.
result Exponential lower bounds on mean time to false alarm and asymptotic upper bounds on detection delay.

We introduce a new geometric approach that constructs a transition kernel of Markov chain. Our method always minimizes the average rejection rate and even reduce it to zero in many relevant cases, which cannot be achieved by conventional methods, such as the Metropolis-Hastings algorithm or the heat bath algorithm (Gib…

2011-06-17abs ↗pdf ↗

The existence of stationary Markov perfect equilibria in stochastic games is shown under a general condition called "(decomposable) coarser transition kernels". This result covers various earlier existence results on correlated equilibria, noisy stochastic games, stochastic games with finite actions and state-independe…

2013-11-07abs ↗pdf ↗

Improved KSD test for better detection of differences in distributions.

problem Low power of KSD test when distributions have same modes but different mixing proportions.
method Perturb the observed sample using Markov transition kernels to improve KSD test power.
result Perturbed KSD test can lead to substantially higher power than the original KSD test.

New algorithm achieves data-dependent regret bounds in MDPs with unknown transitions.

problem Achieving best-of-both-worlds guarantees with data-dependent regret bounds in MDPs with unknown transitions.
method Optimistic follow-the-regularized-leader algorithm with new optimistic Q-function estimators and transition bonus.
result First-order, second-order, and path-length bounds with polylog(T) regret in the stochastic regime.

We develop algorithms with low regret for learning episodic Markov decision processes based on kernel approximation techniques. The algorithms are based on both the Upper Confidence Bound (UCB) as well as Posterior or Thompson Sampling (PSRL) philosophies, and work in the general setting of continuous state and action …

2019-11-04abs ↗pdf ↗

We consider online learning for minimizing regret in unknown, episodic Markov decision processes (MDPs) with continuous states and actions. We develop variants of the UCRL and posterior sampling algorithms that employ nonparametric Gaussian process priors to generalize across the state and action spaces. When the trans…

2018-05-21abs ↗pdf ↗

We study optimal solutions to an abstract optimization problem for measures, which is a generalization of classical variational problems in information theory and statistical physics. In the classical problems, information and relative entropy are defined using the Kullback-Leibler divergence, and for this reason optim…

2010-12-02abs ↗pdf ↗

A multi-task GP model tracks time-varying transition probabilities between two states.

problem Tracking time-varying transition probabilities between 'moves' and 'pauses' states.
method Kernel-based multi-task Gaussian Process model with time-variability and constraints.
result Enforces constraints while learning transition probabilities.

Researchers approximate conditional expectation operators using kernel methods.

problem Statistical approximation of conditional expectation operators under minimal assumptions.
method Modifying the domain of the operator, approximating it by Hilbert-Schmidt operators in a reproducing kernel Hilbert space.
result The nonparametric estimate of the operator converges to a specific limiting object.

HDT improves MCMC on graphs with history-dependent sampling.

problem Efficient sampling from target distributions on general graphs with low computational overhead.
method History-driven target (HDT) framework that replaces the original target distribution with a history-dependent one.
result Near-zero variance performance and scalability to large graphs with memory-efficient implementation.

Paper tackles matrix estimation under arbitrary noise, achieving minimax optimality.

problem Noisy low-rank-plus-sparse matrix recovery under arbitrary dependence.
method Incoherent-constrained least-square estimator, novel energy spreading result.
result Achieves minimax optimality in estimating structured Markov transition kernels.

The paper studies convergence of kernel autocovariance operators for stationary processes.

problem Estimating autocovariance operators of stationary processes on Polish spaces.
method Investigates convergence of empirical estimates of autocovariance operators under various conditions.
result Provides consistency results for kernel PCA and spectral analysis methods.

Extended elliptical slice sampling for infinite-dimensional spaces, proving reversibility.

problem Proving reversibility of elliptical slice sampling in infinite-dimensional spaces.
method Extended elliptical slice sampling to infinite-dimensional separable Hilbert spaces, providing an alternative proof of reversibility.
result The approach yields a positive semi-definite Markov operator, proving reversibility.

DenseHMM improves HMMs by learning dense representations that enable gradient-based optimization.

problem Learning dense representations for hidden states and observables in HMMs.
method DenseHMM uses kernelized transition probabilities and two optimization schemes.
result DenseHMM achieves superior performance and expressiveness compared to standard HMMs.

Consider a reference Markov process with initial distribution π0π_{0} and transition kernels {Mt}t[1:T]\{M_{t}\}_{t\in[1:T]}, for some TNT\in\mathbb{N}. Assume that you are given distribution πTπ_{T}, which is not equal to the marginal distribution of the reference process at time TT. In this scenario, Schrödinger addressed t…

2019-12-31abs ↗pdf ↗

Bootstrap method for Markov chains in reinforcement learning.

problem Distributional consistency in finite controlled Markov chains with unknown control policies.
method Model-based bootstrap with novel LLN and CLT for visitation counts and transition increments.
result Asymptotically valid confidence intervals for value and QQ-functions in offline RL.

This study analyzes convergence and stability of reinforcement learning algorithms.

problem Understanding the conditions under which reinforcement learning algorithms converge and remain stable.
method Theoretical analysis of convergence and stability of Episodic Upside-Down Reinforcement Learning, Goal-Conditioned Supervised Learning, and Online Decision Transformers.
result The algorithms can achieve near-optimal behavior if the transition kernel is close to a deterministic kernel.

Risk measures applied to dynamic Markov processes with varying risk aversion.

problem Investigating dynamic risk measures in Markov decision processes with varying risk aversion.
method Distributional viewpoint on law-invariant convex risk measures, applied to Markov decision processes with latent costs and random actions.
result Existence of optimal policies in finite and infinite time horizons under mild assumptions.

New insights into Markov chain geometry via positive transition measures.

problem Lack of statistical meaning in the space of transition probabilities.
method Constructing an extension of the space of transition probabilities using Amari's theory of positive measures.
result Introduction of a new dually flat structure for the space of positive transition measures.

Develops CLTs for Markov chain transition probabilities and policies.

problem Estimating transition probabilities and policies in controlled Markov chains.
method Non-parametric estimator for transition matrices; CLTs for value, Q-, and advantage functions; goodness-of-fit tests.
result Asymptotic normality of estimators under specific logging policies.

Paper learns meaningful state and action representations from MDP trajectories.

problem Learning good state and action representations from MDP trajectories.
method Tensor decomposition, kernelization, importance sampling, low-Tucker-rank approximation.
result The learned state/action abstractions provide accurate approximations to latent block structures.

Model detects market anomalies using a Hawkes process with hidden Markov chain.

problem Detecting high-frequency market manipulation in cryptocurrency trades.
method Developed a Markov-modulated Hawkes process with piecewise constant excitation kernels.
result Demonstrated the model's effectiveness in detecting suspicious trading activities.

New method links covariates to CTMCs using RKHS, improving state transitions modeling.

problem Traditional multistate models rely on linear relationships, limiting flexibility.
method Nonparametric approach using RKHS, with Frequentist and Bayesian versions.
result Effective in identifying nonlinear transition functions and predicting long-term behaviors.

Deeptime simplifies learning dynamical models from time series data.

problem Understanding complex systems through dynamical models from time series data.
method Various tools for estimating dynamical models including conventional and kernel/deep learning methods.
result Estimates dynamical models from time series data efficiently and with rich analysis methods.

Scalable Bayesian sampling is playing an important role in modern machine learning, especially in the fast-developed unsupervised-(deep)-learning models. While tremendous progresses have been achieved via scalable Bayesian sampling such as stochastic gradient MCMC (SG-MCMC) and Stein variational gradient descent (SVGD)…

2018-11-21abs ↗pdf ↗

New algorithms reduce regret in reinforcement learning with MNL approximations.

problem Efficient reinforcement learning with MNL function approximation for MDPs.
method Proposed randomized exploration algorithms with frequentist regret guarantees.
result Achieved improved regret bounds for MNL transition models.

We develop efficient and sharp bounds on policy value under perturbations in MDPs.

problem Evaluating policies under best- and worst-case perturbations in MDPs with transition observations.
method Proposed a perturbation model for MDPs, developed semiparametrically efficient estimator with asymptotic normality.
result Semiparametrically efficient and asymptotically normal estimator for policy value bounds.

In most sampling algorithms, including Hamiltonian Monte Carlo, transition rates between states correspond to the probability of making a transition in a single time step, and are constrained to be less than or equal to 1. We derive a Hamiltonian Monte Carlo algorithm using a continuous time Markov jump process, and ar…

2015-09-13abs ↗pdf ↗

Improved algorithm detects changes in RL environments with non-stationary MDPs.

problem Learning in non-stationary reinforcement learning environments.
method R-BOCPD-UCRL2 algorithm for MDPs with multinomial state transitions.
result Near-optimal theoretical guarantees in terms of false-alarm rate and detection delay.

Proposes MIVI for efficient posterior estimation and design of MCMC transitions.

problem Efficiently estimating posterior distributions in constrained time.
method Combines variational inference and MCMC with a variational distribution and optimized Markov chain.
result Optimized Markov chain improves variational distribution and vice versa, leading to more accurate posteriors.