Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

2985968931,191 · Jun 202019922001200920172026
48 results for optimizer state

Characterizes optimal-speed quantum state evolution Hamiltonians.

problem Optimal-speed unitary time evolution of pure and quasi-pure quantum states.
method Construction of the manifold of pure states and isometry with flag manifold, characterization of equigeodesic vectors.
result Hamiltonians generating optimal-speed time evolution are fully characterized by equigeodesic vectors of the flag manifold.

Proposes a new framework for optimizing utility with state-dependent benchmarks.

problem Various interpretations of benchmarks in utility functions.
method General framework of state-dependent utility optimization with stochastic benchmarks.
result Provides optimal solutions and addresses issues of well-definedness and feasibility.

Optimizes wireless network resource management with state-augmented policies.

problem Optimizing network-wide utility with user performance constraints.
method State-augmented parameterization of RRM policy, using dual variables.
result Superior trade-off between mean, minimum, and 5th percentile rates.

Optimizes electric field to control molecule states in Hartree-Fock theory.

problem Optimizing electric field to drive molecule from initial to target state.
method Trust region optimization with gradients from adjoint state method.
result Achieves desired target states with minimal control effort.

Efficiently plans large MDPs with weak function approximations.

problem Planning in large MDPs with limited function approximation capabilities.
method Uses linear value function approximation with weak requirements and a generative oracle.
result Produces almost-optimal actions for any state with polynomial computation time.

Proposes ICC method for dynamic portfolio optimization.

problem Non-stationarity in market conditions makes traditional portfolio optimization ineffective.
method Inverse Covariance Clustering (ICC) to identify market states and integrate into dynamic optimization.
result ICC-PO generates portfolios with higher Sharpe Ratios and greater robustness.

AE-LSVI identifies near-optimal policies in complex systems with minimal data.

problem Identifying near-optimal policies in complex, costly data acquisition systems.
method Combines optimism and pessimism for active exploration in a generative model setting.
result Proves near-optimal policy identification over entire state spaces with polynomial sample complexity.

Paper proposes method for optimal control of unknown systems with latent states.

problem Jointly estimating dynamics and latent states in systems with unmeasurable states.
method Combination of particle Markov chain Monte Carlo methods and scenario theory.
result Probabilistic performance guarantees for optimal input trajectories.

Robust RL with learned optimal adversary improves agent performance under adversarial state observations.

problem Ensuring reinforcement learning agents' robustness against adversarial perturbations of state observations.
method Proposed a framework of alternating training with learned adversaries (ATLA) to find optimal adversarial policies and enhance agent robustness.
result ATLA achieves state-of-the-art performance under strong adversaries in continuous control environments.

Portfolio turnpikes state that, as the investment horizon increases, optimal portfolios for generic utilities converge to those of isoelastic utilities. This paper proves three kinds of turnpikes. In a general semimartingale setting, the abstract turnpike states that optimal final payoffs and portfolios converge under …

2011-01-05abs ↗pdf ↗

The paper calculates optimal trading turnover in terms of asset liquidity and alpha autocorrelation.

problem Understanding optimal trading turnover in the context of asset liquidity and alpha autocorrelation.
method Developed a Gaussian process model to compute steady-state turnover explicitly, relating it to asset liquidity and alpha autocorrelation.
result Steady-state optimal turnover is given by γn+1γ\sqrt{n+1}, where γγ is a liquidity-adjusted risk-aversion and nn is the mean-reversion speed ratio.

Most decision theories, including expected utility theory, rank dependent utility theory and cumulative prospect theory, assume that investors are only interested in the distribution of returns and not in the states of the economy in which income is received. Optimal payoffs have their lowest outcomes when the economy …

2013-08-29abs ↗pdf ↗

Study efficient algorithms for nonconvex optimization with state-dependent Markov data.

problem Stochastic optimization with Markovian data and state-dependent transition kernels.
method Projection-based and projection-free algorithms for constrained nonconvex problems.
result The number of oracle calls to achieve an εε-stationary point is O(1/ε2.5)\mathcal{O}(1/ε^{2.5}).

Recurrent networks learn beliefs from history in partially observable environments.

problem Learning optimal policies in partially observable environments.
method Trained recurrent neural networks to approximate value functions, measuring mutual information between hidden states and beliefs.
result Recurrent networks' hidden states correlate with beliefs of relevant state variables, improving expected return.

Optimal persuasion involves projecting state vectors onto lower-dimensional 'optimal information manifolds'.

problem Optimal persuasion of another agent observing multi-dimensional data.
method Performing non-linear dimension reduction by projecting state vectors onto the 'optimal information manifold'.
result Optimal information design splits information into 'good' and 'bad' components, revealing only the direction of good information.

New RL method handles large state-action spaces with complex models.

problem Complex models and large state-action spaces in reinforcement learning.
method π-KRVI, an optimistic modification of least-squares value iteration using kernel ridge regression.
result First order-optimal regret guarantees under general settings, improving over state of the art.

Second-order optimizers retain residual information after data deletion, affecting machine unlearning.

problem Residual information in second-order optimizers after data deletion.
method Comparison of first-order and second-order learners, eigendecomposition analysis.
result Second-order optimizers retain residual information, not detectable by first-order analysis.

We decode latent states in Block MDPs and learn near-optimal policies.

problem Model estimation and reward-free learning in Block MDPs.
method Information-theoretical lower bound and efficient model estimation algorithm.
result Our algorithm approaches the information-theoretical limit for latent state decoding and converges to optimal policies.

We propose an algorithm for deterministic continuous Markov Decision Processes with sparse rewards that computes the optimal policy exactly with no dependency on the size of the state space. The algorithm has time complexity of O(R3×A2)O( |R|^3 \times |A|^2 ) and memory complexity of O(R×A)O( |R| \times |A| ), where R|R| is the…

2018-05-17abs ↗pdf ↗

Sample inefficiency is a long-lasting problem in reinforcement learning (RL). The state-of-the-art estimates the optimal action values while it usually involves an extensive search over the state-action space and unstable optimization. Towards the sample-efficient RL, we propose ranking policy gradient (RPG), a policy …

2019-06-24abs ↗pdf ↗

Imitation learning is a control design paradigm that seeks to learn a control policy reproducing demonstrations from expert agents. By substituting expert demonstrations for optimal behaviours, the same paradigm leads to the design of control policies closely approximating the optimal state-feedback. This approach requ…

2019-01-07abs ↗pdf ↗

Counterexample shows state-constrained optimal control problems can have Young measure gaps.

problem Existence of Young measure gaps in state-constrained optimal control problems.
method Provided a counterexample for smooth controllable systems state-constrained to the unit ball.
result Gap occurs in a regular setting with non-convex Lagrangian density.

Novel method for shape optimization of non-smooth PDEs.

problem Optimizing shapes governed by non-smooth PDEs.
method Functional variational approach and sensitivity analysis.
result Necessary conditions for locally optimal shapes.

In this paper, we introduce a novel form of value function, Q(s,s)Q(s, s'), that expresses the utility of transitioning from a state ss to a neighboring state ss' and then acting optimally thereafter. In order to derive an optimal policy, we develop a forward dynamics model that learns to make next-state predictions that…

2020-02-21abs ↗pdf ↗

Study designs neural networks for fault localization, state estimation, and optimal PMU placement in power systems.

problem Fault localization, state estimation, and optimal PMU placement in power systems.
method Designs and compares various neural networks for fault localization, builds machine learning schemes for state estimation and parameter estimation, and designs an algorithm for optimal PMU placement.
result Comprehensive comparison of neural networks for fault localization shows that Graphical Convolutional NN and Neural Graph-based ODE perform best.

This paper deals with discrete-time Markov control processes on a general state space. A long-run risk-sensitive average cost criterion is used as a performance measure. The one-step cost function is nonnegative and possibly unbounded. Using the vanishing discount factor approach, the optimality inequality and an optim…

2007-04-03abs ↗pdf ↗

This paper optimizes MDP policies for efficient state aggregation.

problem Optimizing policies in aggregated Markov chains while preserving optimal performance.
method Homomorphic mappings to establish optimal policy equivalence and derive performance bounds.
result Developed HPG and EBHPG methods for efficient aggregation and policy optimization.