Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

25.0%50.0%75.0%100.0% · Feb 199419922001200920182026
48 results for A3C path finding

Paper tackles adversarial attacks on A3C path finding, proposing Gradient Band-based Adversarial Training.

problem Adversarial attacks on A3C path finding.
method Gradient Band-based Adversarial Training with CDG method.
result Gradient Band-based Adversarial Training achieves high attack immunity.

Paper proposes Terminal Prediction to improve deep RL performance.

problem Sample inefficiency and convergence to locally optimal policies in deep reinforcement learning.
method Introduces a self-supervised auxiliary task, Terminal Prediction, to help representation learning.
result A3C-TP outperforms standard A3C in most domains and provides significant improvement in Pommerman.

IRL addresses weaknesses in DDPG and A3C for continuous reinforcement learning.

problem Theoretical weaknesses in DDPG and A3C for continuous reinforcement learning.
method IRL based on stochastic differential equations, ensuring action continuity and variance control.
result IRL method guarantees action continuity and variance control, allowing positive interaction with the environment.

This work combines GANs and A3C for high-resolution image compression.

problem Image compression for high-resolution images without loss of quality.
method Hybrid approach using GANs and A3C for end-to-end learning.
result Improves PSNR for high-resolution images through end-to-end learning.

Rogue is a famous dungeon-crawling video-game of the 80ies, the ancestor of its gender. Rogue-like games are known for the necessity to explore partially observable and always different randomly-generated labyrinths, preventing any form of level replay. As such, they serve as a very natural and challenging task for rei…

2018-04-23abs ↗pdf ↗

We introduce NoisyNet, a deep reinforcement learning agent with parametric noise added to its weights, and show that the induced stochasticity of the agent's policy can be used to aid efficient exploration. The parameters of the noise are learned with gradient descent along with the remaining network weights. NoisyNet …

2017-06-30abs ↗pdf ↗

VUSFA improves transfer learning for target-driven navigation in AI2THOR.

problem Improving transfer reinforcement learning for complex visual navigation tasks.
method Introducing SFDP and Variational Information Bottlenecks to A3C agent.
result VUSFA achieves state-of-the-art performance and generalizability.

Deep reinforcement learning has shown its success in game playing. However, 2.5D fighting games would be a challenging task to handle due to ambiguity in visual appearances like height or depth of the characters. Moreover, actions in such games typically involve particular sequential action orders, which also makes the…

2018-05-05abs ↗pdf ↗

The ability to use a 2D map to navigate a complex 3D environment is quite remarkable, and even difficult for many humans. Localization and navigation is also an important problem in domains such as robotics, and has recently become a focus of the deep reinforcement learning community. In this paper we teach a reinforce…

2017-11-20abs ↗pdf ↗

RADIAL-RL improves deep RL agents' robustness against adversarial attacks.

problem Vulnerability of deep reinforcement learning agents to small adversarial perturbations.
method RADIAL-RL, a principled framework for training robust reinforcement learning agents.
result RADIAL-RL-trained agents consistently outperform prior methods in robustness tests.

Deep RL optimizes processing paths to desired material structures.

problem Optimizing processing paths to achieve desired material properties.
method Deep reinforcement learning guided by structure representations and reward signals.
result Algorithm learns to find optimal paths to target structures in material space.

Given two points on a soup can or conical cup with lid, we find and classify all paths of minimal length connecting them. When the number of minimal paths is finite, there are at most four on a can and three on a cup. At worst, minimal paths are piece-wise smooth with three components, each of which is a classical geod…

2004-01-09abs ↗pdf ↗

Motivated by recent advance of machine learning using Deep Reinforcement Learning this paper proposes a modified architecture that produces more robust agents and speeds up the training process. Our architecture is based on Asynchronous Advantage Actor-Critic (A3C) algorithm where the total input dimensionality is halv…

2018-04-13abs ↗pdf ↗

Paper proposes a new RL approach combining IL and RL methods to improve decision-making.

problem Challenges in RL with large state and action spaces, and difficulty in reward determination.
method Combines Imitation Learning and RL methods (SARSA and A3C) to learn sequential decision-making policies.
result Significantly decreases human effort and exploration time in learning decision-making policies.

The path probability of a particle undergoing stochastic motion is studied by the use of functional technique, and the general formula is derived for the path probability distribution functional. The probability of finding paths inside a tube/band, the center of which is stipulated by a given path, is analytically eval…

2016-02-13abs ↗pdf ↗

Sparse neural networks training is difficult due to optimization failures and energy landscape issues.

problem Training sparse neural networks leads to suboptimal solutions and optimization failures.
method Investigated optimization dynamics and energy landscape in sparse neural networks.
result Sparse neural networks have a linear path with a monotonically decreasing objective from initialization to a good solution, but not from a bad solution.

Novel approach to financial derivatives pricing using rough path theory.

problem No-arbitrage conditions in financial markets necessitating precise integration methods.
method Developed a polynomial-based approximation class for rough path functionals, extending to non-geometric rough paths.
result Motivated a hypothesis for payoff functionals in financial markets, facilitating analysis.

Solves infinite horizon portfolio problem with path-dependent labor income.

problem Infinite horizon portfolio choice with path-dependent labor income.
method Solves an infinite dimensional stochastic optimal control problem using explicit solutions to the HJB equation.
result Explicit solutions to the optimal controls in feedback form are found.

Study finds a limiting distribution for free path lengths on flat surfaces with circular obstacles.

problem Understanding free path lengths on flat surfaces with circular obstacles.
method Proved the existence of a limiting distribution using radius of obstacles as a parameter.
result Relates the limiting distribution to heights of zippered rectangle decompositions.

New methods improve Monte Carlo estimation of partition functions.

problem Estimating the normalization constant of complex distributions.
method Annealing through paths of distributions to estimate partition functions.
result Optimal path for estimation is arithmetic, improving efficiency.

Q-learning is a simple and powerful tool in solving dynamic problems where environments are unknown. It uses a balance of exploration and exploitation to find an optimal solution to the problem. In this paper, we propose using four basic emotions: joy, sadness, fear, and anger to influence a Qlearning agent. Simulation…

2016-09-06abs ↗pdf ↗

Develops methods to find most probable paths on complex manifolds.

problem Identifying optimal paths for manifold-valued processes, especially those with non-trivial structures.
method Constructs a general approach to defining and identifying most probable paths by measuring the Onsager-Machlup function on the anti-development of such processes.
result Derives explicit equations for development most probable paths that encompass various manifold-valued processes.

The paper reveals surprising star-shaped connectivity in neural networks.

problem Understanding mode connectivity in neural network landscapes.
method Fine-grained analysis of connectivity in overparameterized and finite minima cases.
result Star-shaped connectivity exists in neural network landscapes, suggesting near convexity.

The study examines price formation in complex networks and finds efficiency varies by network structure.

problem Understanding price formation and efficiency in complex networks.
method Price formation experiments with human subjects in large networks, agent-based model construction.
result Prices are higher and trade less efficient in small-world networks compared to random networks.

Flow Matching enables robust training of CNFs with various probability paths.

problem Training Continuous Normalizing Flows (CNFs) at large scales.
method Flow Matching (FM) is a simulation-free approach for training CNFs by regressing vector fields of conditional probability paths.
result Flow Matching with diffusion paths yields more robust and stable training compared to diffusion-based methods.

Selecting important features in non-linear or kernel spaces is a difficult challenge in both classification and regression problems. When many of the features are irrelevant, kernel methods such as the support vector machine and kernel ridge regression can sometimes perform poorly. We propose weighting the features wit…

2009-06-24abs ↗pdf ↗

The study examines insurance demand under rough volatility and path-dependent shocks.

problem Optimal insurance and investment strategies under rough volatility and path-dependent shocks.
method Rough volatility model and Hawkes process with power kernel, Functional Ito formula extension.
result Individuals demand more catastrophe insurance when path-dependent effects are considered.

Domain adaptation is an important open problem in deep reinforcement learning (RL). In many scenarios of interest data is hard to obtain, so agents may learn a source policy in a setting where data is readily available, with the hope that it generalises well to the target domain. We propose a new multi-stage RL agent, …

2017-07-26abs ↗pdf ↗

The paper sparsifies networks by finding efficient paths in their functional space.

problem Sparsifying neural networks to improve performance and efficiency.
method The authors use the geometry of weight spaces and functional manifolds to find efficient paths (geodesics) in the functional space of neural networks.
result The proposed framework can sparsify networks and improve performance on various tasks.

Path-independent equilibrium models improve network performance on harder problems.

problem Improving network performance on harder problem instances.
method Investigated path-independent equilibrium models and their impact on network performance.
result Path independence correlates with better performance on harder problem instances.

The relaxed maximum entropy problem is concerned with finding a probability distribution on a finite set that minimizes the relative entropy to a given prior distribution, while satisfying relaxed max-norm constraints with respect to a third observed multinomial distribution. We study the entire relaxation path for thi…

2013-11-07abs ↗pdf ↗

Reinforcement learning improves trading performance on stock exchanges.

problem Optimizing trading strategies on stock exchanges using machine learning.
method Markov model, asynchronous advantage actor-critic method, neural networks, recurrent layers.
result Best trading strategy for RTS Index futures achieved a 66% annual profit.

We consider the problem of packing node-disjoint directed paths in a directed graph. We consider a variant of this problem where each path starts within a fixed subset of root nodes, subject to a given bound on the length of paths. This problem is motivated by the so-called kidney exchange problem, but has potential ot…

2016-03-18abs ↗pdf ↗