Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

3857711,1561,541 · Jun 202019922001200920182026
48 results for return-based learning

Paper combines DQN and return-based RL for improved policy performance.

problem Improving policy performance in reinforcement learning.
method Integrates DQN and return-based reinforcement learning, introduces two measurements to quantify policy discrepancy.
result The proposed measurements accurately express trace coefficient and improve approximation to return.

New self-imitation learning method improves performance in continuous control tasks.

problem Improving off-policy learning in continuous control tasks.
method Proposes a n-step lower bound to generalize lower-bound Q-learning and introduces a new family of self-imitation learning algorithms.
result n-step lower bound Q-learning achieves a better trade-off between bias and contraction rate, leading to improved performance.

In this work, we take a fresh look at some old and new algorithms for off-policy, return-based reinforcement learning. Expressing these in a common form, we derive a novel algorithm, Retrace(λλ), with three desired properties: (1) it has low variance; (2) it safely uses samples collected from any behaviour policy, wha…

2016-06-08abs ↗pdf ↗

Enhances traditional MV model for socially responsible investors.

problem Traditional MV models ignore ESG scores relevant to socially responsible investors.
method Implemented an amended MV model considering ESG scores.
result SR investors can achieve competitive SR portfolios with a trade-off between Sharpe Ratio and ESG scores.

In their activity, the traders approximate the rate of return by integer multiples of a minimal one. Therefore, it can be regarded as a quantized variable. On the other hand, there is the impossibility of observing the rate of return and its instantaneous forward time derivative, even if we consider it as a continuous …

2012-11-08abs ↗pdf ↗

Neural Markov models improve time series analysis by balancing deep learning and classical models.

problem Modeling non-stationary time series with high data sparsity.
method Hybrid approach using neural networks to parameterize stochastic matrices, estimating time-inhomogeneous Markov chains.
result Reduction of Chapman-Kolmogorov discrepancy and superior likelihood in financial markets.

Optimizes impression allocation for e-commerce platforms using reinforcement learning.

problem Short-term and long-term returns are not optimized in current e-commerce platform allocation mechanisms.
method Formal lifecycle model of products, reinforcement learning framework, first principal component based permutation, novel experiences generation method.
result Significant improvement in platform and participant health with optimized impression allocation.

This paper examines the possibility of using derivative-implied risk premia to explain stock returns. The rapid development of derivative markets has led to the possibility of trading various kinds of risks, such as credit and interest rate risk, separately from each other. This paper uses credit default swaps and equi…

2010-05-30abs ↗pdf ↗

We empirically test predictability on asset price by using stock selection rules based on maximum drawdown and its consecutive recovery. In various equity markets, monthly momentum- and weekly contrarian-style portfolios constructed from these alternative selection criteria are superior not only in forecasting directio…

2014-03-31abs ↗pdf ↗

Optimizes trading policies using future price forecasts.

problem Static reinforcement learning agents lack mechanisms for using price forecasts at inference time.
method FPILOT framework inspired by Model Predictive Control (MPC). Uses a predictive model to construct an allocation-based imagined return objective at each decision step.
result Consistent improvements in total return and risk-adjusted metrics across various policy learning algorithms.

Paper develops new spot regression estimators using candlesticks for asset pricing.

problem Estimation of spot betas in asset pricing and risk management.
method Develops a new estimation and inference framework for spot regressions using high-frequency candlesticks.
result The proposed candlestick-based estimators reduce estimation risk and achieve higher power in hypothesis testing.

GMADL loss function improves model performance and reduces transaction costs.

problem Overfitting and high transaction costs in high-frequency algorithmic trading models.
method Introduces GMADL loss function for better optimization and feature selection.
result GMADL produces superior results and reduces transaction costs compared to standard loss functions.

The paper explains how to predict returns based on firm characteristics.

problem Predicting returns based on firm characteristics in equilibrium models.
method Reverse-engineering equilibrium construction process with linear demands in characteristics.
result Linear expressions for returns are derived from scaled net aggregate demands and their variations.

Paper resolves the debate on process vs. outcome supervision in reinforcement learning.

problem Distinguishing between process and outcome supervision in reinforcement learning.
method Developed a technical tool (Change of Trajectory Measure Lemma) to show equivalence between outcome and process supervision under standard data coverage assumptions.
result Reinforcement learning through outcome supervision is statistically equivalent to process supervision, up to polynomial factors in horizon.

A new model uses firm characteristics to predict asset covariances.

problem Risk models are noisy and dependent on historical returns.
method Characteristic-Driven Dynamic Factor Model (CD-DFM) that learns latent representations from firm characteristics.
result CD-DFM produces interpretable factor portfolios and competitive covariance forecasts.

Paper proposes SPO paradigm for better portfolio optimization in real markets.

problem Real-world trading frictions and constraints affect portfolio optimization quality.
method SPO paradigm with decision-focused training using surrogate loss and linear predictors.
result Decision-focused training improves risk-adjusted performance and robustness.

We establish the existence of anomalous excess returns based on trend following strategies across four asset classes (commodities, currencies, stock indices, bonds) and over very long time scales. We use for our studies both futures time series, that exist since 1960, and spot time series that allow us to go back to 18…

2014-04-12abs ↗pdf ↗

In our previous studies we have investigated the structural complexity of time series describing stock returns on New York's and Warsaw's stock exchanges, by employing two estimators of Shannon's entropy rate based on Lempel-Ziv and Context Tree Weighting algorithms, which were originally used for data compression. Suc…

2014-08-16abs ↗pdf ↗

The aim of this paper is to examine the time scaling of the semivariance when returns are modeled by various types of jump-diffusion processes, including stochastic volatility models with jumps in returns and in volatility. In particular, we derive an exact formula for the semivariance when the volatility is kept const…

2013-11-05abs ↗pdf ↗

Study examines Indian equity mutual funds' investment style and risk-shifting.

problem Understanding how Indian equity mutual funds' investment styles affect their returns.
method Estimating size and style beta coefficients, identifying breakpoints, analyzing investment styles, and assessing risk-shifting intensity.
result Funds can enhance returns by shifting to high-return styles like Small Value and Small Blend.

We study the statistical properties of the recurrence intervals ττ between successive trading volumes exceeding a certain threshold qq. The recurrence interval analysis is carried out for the 20 liquid Chinese stocks covering a period from January 2000 to May 2009, and two Chinese indices from January 2003 to April 2…

2010-02-06abs ↗pdf ↗

We study tick-by-tick financial returns belonging to the FTSE MIB index of the Italian Stock Exchange (Borsa Italiana). We can confirm previously detected non-stationarities. However, scaling properties reported in the previous literature for other high-frequency financial data are only approximately valid. As a conseq…

2012-12-03abs ↗pdf ↗

DRL improves ESG financial portfolio management by regulating returns based on ESG scores.

problem Improving ESG financial portfolio management through market regulation.
method Used Advantage Actor-Critic (A2C) agent and adapted OpenAI Gym environments for comparative analysis.
result DRL agent outperforms standard market conditions in ESG-regulated market.

This paper optimizes cryptocurrency portfolios by clustering price correlations and improving risk-return profiles.

problem Volatility and regulatory uncertainty in cryptocurrency markets make portfolio construction challenging.
method The paper combines network analysis, price forecasting, and portfolio theory to identify stable groups of correlated cryptocurrencies.
result Predictive consensus-clustering portfolios maintain positive and stable performance up to a 14-day horizon, with favourable gain-loss asymmetry and tighter tail-risk control.

The study explains stock return distributions using reaction functions.

problem Stock return distributions often deviate from normal distributions.
method Assumes normal event/information effects, financial over/underreaction, proposes reaction function model.
result Financial markets often underreact to minor events, overreact to significant ones, and react stronger to positive events.

The paper develops models for asset returns based on market conditions and uses them to construct a trading policy.

problem Developing a robust trading strategy based on market conditions.
method The authors create stratified models of asset return mean and covariance, fit these models using Laplacian regularization, and combine them with a Markowitz optimization method.
result The trading policy performs well out of sample and can be scaled to larger problems.

We present and discuss a stochastic model of financial assets dynamics based on the idea of an inverse renormalization group strategy. With this strategy we construct the multivariate distributions of elementary returns based on the scaling with time of the probability density of their aggregates. In its simplest versi…

2013-05-14abs ↗pdf ↗

We extend and test empirically the multifractal model of asset returns based on a multiplicative cascade of volatilities from large to small time scales. The multifractal description of asset fluctuations is generalized into a multivariate framework to account simultaneously for correlations across times scales and bet…

2000-08-04abs ↗pdf ↗

A new method cleans and analyzes stock return correlation matrices.

problem Improving the accuracy of covariance/correlation matrices in financial data.
method Constrained principal component analysis using financial data and optimal portfolios.
result Identified stylized patterns in correlation matrix eigenvalues and weights.

MDS selects assets by combining daily returns and intraday risk curves, improving portfolio performance.

problem High estimation error in large-scale asset selection.
method Metric Dependence Screening (MDS) incorporating high frequency information as object valued data.
result MDS improves portfolio performance over benchmarks by preserving intraday risk dynamics.