This paper balances short-term and long-term rewards in policy learning.
problem Balancing short-term and long-term rewards in policy learning.
method Formalizes a new framework to balance rewards, identifies rewards under mild assumptions, deduces efficiency bounds, and develops a policy learning approach.
result The proposed method improves the estimator of long-term reward and reduces regret.
Kernel method estimates long-term effects from short-term data.
problem Estimating long-term effects from short-term data in continuous actions.
method Kernel ridge regression to embed and extrapolate long-term effects.
result Uniform consistency and nonasymptotic error bounds for the estimator.
New framework estimates long-term outcomes from short-term data.
problem Estimating long-term outcomes from short-term data.
method Reward function decomposition-based framework (LOPE).
result LOPE outperforms existing methods, especially when surrogacy is violated.
New algorithm optimizes for long-term user satisfaction in delayed reward settings.
problem Optimizing for long-term user satisfaction in delayed reward settings.
method Developed a predictive model of delayed rewards and a bandit algorithm that combines rewards and surrogate outcomes.
result Our algorithm significantly outperforms methods that optimize for short-term proxies or rely solely on delayed rewards.
We propose and study the known-compensation multi-arm bandit (KCMAB) problem, where a system controller offers a set of arms to many short-term players for T steps. In each step, one short-term player arrives to the system. Upon arrival, the player aims to select an arm with the current best average reward and receiv…
Model user preferences for conversational LLMs using weak rewards.
problem Lack of persistent user models in conversational LLMs leading to repeated user restatements.
method Vector-Adapted Retrieval Scoring (VARS) framework that updates user vectors online from weak scalar rewards.
result Full VARS agent achieves strongest overall performance, matches strong Reflection baseline in task success, and reduces user effort.
Research predicts cryptocurrency staking rewards with high accuracy.
problem Predicting cryptocurrency staking rewards.
method Two predictive methodologies: sliding-window average and linear regression models.
result ETH staking rewards can be forecasted with RMSE within 0.7% and 1.1% of the mean value for 1-day and 7-day look-aheads respectively.
New algorithm optimizes long-term user satisfaction in recommendation systems.
problem Optimizing long-term user satisfaction in recommendation systems with delayed rewards.
method Developed a predictive model of delayed rewards and a bandit algorithm that balances exploration and exploitation.
result Our approach results in substantially better performance compared to short-term or delayed optimization.
Paper optimizes recommendation systems for long-term business metrics.
problem Short-term reward optimization ignores long-term business metrics.
method Introduced a framework for modeling long-term rewards in RecoGym.
result Proposed a simple extension leading to state-of-the-art results.
This paper proposes a new algorithm for learning guidance rewards in RL.
problem Long-term temporal credit assignment in sparse or delayed reward environments.
method Surrogate RL objective with trajectory-space smoothing to learn guidance rewards.
result Guidance rewards can be learned without additional neural networks and have intuitive interpretation.
New algorithm tackles nonstationary linear bandits with latent dynamics.
problem Nonstationary bandit problem with latent states and unknown dynamics.
method Explore-then-commit algorithm with exploration and commitment phases.
result Achieves ildeO(T2/3) regret. Quantum model outperforms classical in training but underperforms in real-world metrics.
problem Mismatch between proxy reward signals and true investment objectives in financial domains.
method Hybrid quantum-classical reinforcement learning framework with automated feature engineering.
result Quantum models achieve higher training rewards but underperform in real-world metrics.
The problem of multi-armed bandits (MAB) asks to make sequential decisions while balancing between exploitation and exploration, and have been successfully applied to a wide range of practical scenarios. Various algorithms have been designed to achieve a high reward in a long term. However, its short-term performance m…
Reinforcement learning aids decision-making in economics and finance.
problem Optimal decision-making in dynamic, uncertain environments.
method Reinforcement learning algorithms to learn optimal policies.
result Deep learning enhances solving complex behavioral problems.
Goal-oriented reinforcement learning has recently been a practical framework for robotic manipulation tasks, in which an agent is required to reach a certain goal defined by a function on the state space. However, the sparsity of such reward definition makes traditional reinforcement learning algorithms very inefficien…
With the breakthrough of computational power and deep neural networks, many areas that we haven't explore with various techniques that was researched rigorously in past is feasible. In this paper, we will walk through possible concepts to achieve robo-like trading or advising. In order to accomplish similar level of pe…
New algorithm models satiation in recommender systems.
problem Satiation effects in user preferences not modeled by existing algorithms.
method Rebounding bandits, modeling satiation as time-invariant linear dynamical systems.
result Greedy policy optimal for identical deterministic dynamics; EEP algorithm for stochastic dynamics.
Algorithm improves wildlife protection patrols.
problem Balancing exploration and exploitation in patrolling vast protected areas.
method Formulated as a stochastic multi-armed bandit problem, leveraging smoothness and decomposability.
result Algorithm LIZARD improves performance on real-world poaching data.
Paper proposes a reinforcement learning method for trading using expert trajectories.
problem Inability of existing methods to handle long-term goals and delayed rewards in futures trading.
method Modeling futures trading as MDP, using reinforcement learning with expert trajectories and multiple short-term alpha factors.
result The proposed method outperforms traditional and deep learning methods in trading performance.
Model combines long-term and short-term memory using conceptors.
problem Transfer between long-term and short-term memory.
method Recurrent neural network with gated reservoir for short-term memory and conceptors for long-term memory.
result Standard operations on conceptors allow combining long-term memories and describing their effect on short-term memory.
New algorithm for traffic routing in congested conditions.
problem Optimal routing in congested traffic networks.
method Congested Bandits model, UCB algorithm, iterative least squares planner.
result No-regret algorithms for congestion-aware routing.
Cash management is concerned with optimizing the short-term funding requirements of a company. To this end, different optimization strategies have been proposed to minimize costs using daily cash flow forecasts as the main input to the models. However, the effect of the accuracy of such forecasts on cash management pol…
QLSTM outperforms LSTM in predicting KSE 100 index movements.
problem Predicting stock market movement in uncertain economic conditions.
method Used LSTM and QLSTM models on monthly data of economic indicators.
result QLSTM provided more accurate predictions of KSE 100 index values.
The paper presents the comparative study of the nature of stock markets in short-term and long-term time scales with and without structural break in the stock data. Structural break point has been identified by applying Zivot and Andrews structural trend break model to break the original time series (TSO) into time ser…
Study shows risk-averse investors have consistent ranking of risky assets.
problem Ranking of risky assets in short-term investments.
method Analyzes various decision problems regarding risky assets with continuous returns.
result Risk-averse decision makers have the same ranking over risky assets.
Predicts short-term futures contract direction using neural networks and order flow data.
problem Challenges in predicting short-term directional movement of futures contracts.
method Engineering features from technical analysis, order flow, and order-book data; training a Tabnet neural network.
result Achieved an accuracy of 0.601 in predicting directional change on the Silver Futures Contract.
This paper proposes a framework to predict long-term trends and short-term fluctuations in multivariate time series.
problem Existing prediction methods often ignore the distinction between long-term trends and short-term fluctuations.
method The paper introduces a MTS forecasting framework that uses both original time series and its first difference to capture long-term trends and short-term fluctuations.
result The proposed method improves forecasting performance by using more supervision information.
A new model for pricing ultra-short-term options with complex volatility patterns.
problem Complex pricing of ultra-short-term options due to oscillations in implied volatility.
method Edgeworth++ model with nonparametric stochastic volatility and deterministic shift extension.
result Fast and accurate closed-form option pricing for ultra-short-term options.
Statistical models outperform mechanistic models in short-term COVID-19 incidence forecasts.
problem Comparing accuracy of mechanistic vs statistical models for short-term COVID-19 incidence forecasts.
method Empirical comparison of forecasts from mechanistic and statistical models using daily incidence data from six US states.
result Statistical models are at least as accurate as mechanistic models and better capture volatility.
TimeMixer predicts global financial asset volatility, excelling in short-term forecasts.
problem Predicting volatility in global financial markets is challenging due to complexity and non-linear dynamics.
method Uses TimeMixer, a multiscale-mixing model for forecasting across different scales.
result TimeMixer performs exceptionally well in short-term volatility forecasting but less so in longer-term predictions.
Deep RL optimizes US stock allocations with better performance.
problem Optimizing asset allocation in US equities markets.
method Reinforcement learning applied to asset allocation problems.
result Deep RL models outperform traditional methods in asset allocation.
FinRLlama wins FinRL Challenge 2024 by fine-tuning LLMs with market data.
problem Lack of contextual alignment for financial market applications in traditional LLMs.
method Fine-tuning LLaMA-3.2-3B-Instruct model with custom RLMF prompt design integrating historical data and reward feedback.
result RLMF-tuned FinRLlama framework outperforms baseline methods in signal consistency and trading outcomes.
Comparative study of neural networks for short-term FOREX forecasting.
problem Simulating expert judgment in foreign exchange market forecasting.
method Implemented and compared LSTM and ANN architectures for short-term FOREX forecasting.
result ANN custom architecture outperforms LSTM in prediction quality and resource efficiency.
Long short-term memory network outperforms seasonal model in JSE Top 40 forecasting.
problem Comparing neural network performance to traditional models in financial forecasting.
method Used long short-term memory network for JSE Top 40 return data forecasting.
result Long short-term memory network outperforms seasonal model in forecasting.
TimeBridge addresses non-stationarity in long-term time series forecasting.
problem Non-stationarity in multivariate time series leads to spurious regressions and obscures long-term relationships.
method TimeBridge segments series into patches, applying Integrated Attention for short-term non-stationarity and Cointegrated Attention for long-term cointegration.
result TimeBridge achieves state-of-the-art performance in both short-term and long-term forecasting.
Proposes LSR-IGRU for improved stock trend prediction.
problem Challenges in stock price prediction due to complex relationships and nonlinear dynamics.
method Long short-term relationships matrix and improved GRU input for better temporal and relationship integration.
result Significantly improved accuracy in predicting stock trend changes.
Study finds short-term wage increases due to COVID-19, contrary to expectations.
problem Impact of COVID-19 on wages over time.
method Empirical analysis controlling for GDP as a demand proxy.
result Short-term positive wage effect, contrary to expectations.
Study finds short-term trading signals can enhance alpha in U.S. S&P 500 portfolios.
problem Traditional factor investing misses real-time market dislocations.
method Double-selection LASSO framework to control for fundamental factors and isolate trading signals.
result 17 distinct trading signals capture significant risk premiums and enhance portfolio diversification.
Model predicts short-term Amazon rainforest fires with high accuracy.
problem Accurate short-term forecasting of Amazon rainforest fires is challenging.
method Used Seasonal and Trend decomposition based on Loess combined with multi-month-ahead load forecasting algorithms.
result Proposed decomposition-ensemble models provide more accurate forecasts than other models.
The study improves load forecasting for electricity consumers using advanced machine learning models.
problem Improving short-term load forecasting for effective scheduling and decision-making.
method Proposes and evaluates statistical nonlinear models, including LSTM and GRU, for 15-min frequency electricity load forecasting.
result Advanced models outperform other models in out-of-sample forecasting accuracy, as shown by the Diebold-Mariano test.
Paper proposes a method for predicting any quantile of short-term electricity demand.
problem Uncertainty in power systems due to multiple factors.
method Proposes a novel general approach for distributional forecasting of short-term electricity demand.
result Demonstrates state-of-the-art distributional forecasting results for short-term electricity demand.
Paper presents LSTM models for short-term stock price prediction.
problem Accurately predicting short-term stock prices is challenging.
method Univariate and multivariate LSTM models using historical data.
result Multivariate LSTM model with technical indicators outperforms univariate model.
We find a sharp local maximum in cross-correlation of EUR/USD and BTC/USD pairs, indicating short-term momentum trading.
problem The Epps effect is observed in various markets but deviates in foreign exchange and cryptocurrency markets.
method We document and analyze the cross-correlation function of EUR/USD and BTC/USD pairs to identify the Epps effect deviation.
result The sharp local maximum in cross-correlation function reveals the activity of short-term momentum traders.
This study compares machine learning models for short-term stock price forecasting.
problem Accurate short-term stock price prediction in the NYSE.
method Compared four machine learning models (XGBoost, Random Forest, Multi-layer Perceptron, Support Vector Regression) on NYSE stocks.
result XGBoost model outperformed others with highest accuracy.
Recurrent neural networks (RNN) are at the core of modern automatic speech recognition (ASR) systems. In particular, long-short term memory (LSTM) recurrent neural networks have achieved state-of-the-art results in many speech recognition tasks, due to their efficient representation of long and short term dependencies …
In this paper we analyze an extension of the Jeanblanc and Valchev (2005) model by considering a short-term uncertainty model with two noises. It is a combination of the ideas of Duffie and Lando (2001) and Jeanblanc and Valchev (2005): share quotations of the firm are available at the financial market, and these can b…
Reward hacking exploits misspecified rewards, affecting agent capabilities and true performance.
problem Reward hacking in RL models exploiting reward misspecifications.
method Constructed four RL environments with misspecified rewards; analyzed agent capabilities and behavior.
result More capable agents exploit reward misspecifications, achieving higher proxy reward but lower true reward.
Paper introduces PRMs to learn non-Markovian stochastic rewards for reinforcement learning.
problem Lack of structured representation for non-Markovian stochastic rewards in reinforcement learning.
method Introduces probabilistic reward machines (PRMs) and presents an algorithm to learn them from decision processes.
result Algorithm proves correct and convergent for learning PRMs from decision processes.