ArtificialReplay improves data efficiency in bandits using historical data.
problem Data inefficiency in warm-starting bandit algorithms.
method ArtificialReplay, a meta-algorithm for incorporating historical data into any bandit algorithm.
result ArtificialReplay uses only a fraction of historical data compared to a full warm-start approach, achieving identical regret.
Unified HS and related methods with explicit modeling assumptions.
problem Lack of clear assumptions in HS methods for Value-at-Risk.
method Explicitly defined parametric model for asset returns and extraction of innovation process.
result HS and related methods require more assumptions than commonly acknowledged.
Proposes dynamic borrowing method for historical data in clinical trials.
problem Insufficient statistical power in rare and pediatric disease clinical trials.
method Dynamic borrowing method based on frequentist approach using similarity measures.
result Demonstrates usefulness of dynamic borrowing in reanalyzing clinical trial data.
The link between different psychophysiological measures during emotion episodes is not well understood. To analyse the functional relationship between electroencephalography (EEG) and facial electromyography (EMG), we apply historical function-on-function regression models to EEG and EMG data that were simultaneously r…
Improved Bayesian inference using power priors with historical data.
problem Improving Bayesian inference with historical data.
method Generalized power priors that adapt to the α parameter of Amari's α-divergence. result Improved performance through appropriate choices of the α parameter. Combines experimental and historical data for robust policy evaluation.
problem Policy evaluation with mixed data sources, especially experimental vs historical.
method Linear integration of estimators from experimental and historical data, optimized for MSE minimization.
result Proposed estimators outperform traditional methods in ridesharing company data.
Data describing historical economic growth are analysed. Included in the analysis is the world and regional economic growth. The analysis demonstrates that historical economic growth had a natural tendency to follow hyperbolic distributions. Parameters describing hyperbolic distributions have been determined. A search …
Typically flat filling, linear or polynomial interpolation methods to generate missing historical data. We introduce a novel optimal method for recreating data generated by a diffusion process. The results are then applied to recreate historical data for stocks.
RL improves market making with historical data time travel.
problem Limited ability to simulate and fully appraise the impact of actions in competitive systems.
method Introduces 'consistent data time travel' to adjust historical data time index.
result Significant improvement in agent's gain with data time travel.
Industry datasets used for text classification are rarely created for that purpose. In most cases, the data and target predictions are a by-product of accumulated historical data, typically fraught with noise, present in both the text-based document, as well as in the targeted labels. In this work, we address the quest…
New algorithm reduces online learning regret by exploiting historical invariances.
problem Stochastic non-stationary linear bandits with changing reward models.
method ISD-linUCB algorithm that learns invariances in reward model.
result Significant regret improvements in fast-changing environments with historical data.
In this paper we look at the efficacy of different risk measures on energy markets and across several different stock market indices. We use both the Value at Risk and the Tail Conditional Expectation on each of these data sets. We also consider several different durations and levels for historical risk measures. Throu…
The paper develops a method to forecast financial risk multiple steps ahead using quantile time series and historical simulation.
problem Forecasting financial risk multiple steps ahead with accurate estimation of Value-at-Risk (VaR) and Expected Shortfall (ES).
method Quantile-based, semi-parametric historical simulation estimation of VaR and ES models, using quantile loss function and resampling.
result The proposed method accurately forecasts VaR and ES one and multiple steps ahead, superior to existing methods.
Econophysics embodies the recent upsurge of interest by physicists into financial economics, driven by the availability of large amount of data, job shortage in physics and the possibility of applying many-body techniques developed in statistical and theoretical physics to the understanding of the self-organizing econo…
New method detects and mitigates historical bias in data.
problem Detecting and explaining historical bias in data.
method Developed a sample bias criterion and algorithms to measure and counter sample bias.
result Derived bias score provides sample-level attribution and explanation of historical bias.
This research predicts stock market movements using Vision-Language models.
problem Predicting future stock market direction using historical data.
method Utilizing image and byte-based representations of stock data processed with Vision-Language models.
result The proposed approach significantly outperforms deep learning baselines.
Paper addresses off-policy evaluation and learning with covariate shift.
problem Evaluating and training a new policy using historical data with a covariate shift.
method Derives efficiency bounds and proposes doubly robust estimators for OPE and OPL under covariate shift.
result Proposes estimators for off-policy evaluation and learning under covariate shift.
Low-frequency historical data, high-frequency historical data and option data are three major sources, which can be used to forecast the underlying security's volatility. In this paper, we propose two econometric models, which integrate three information sources. In GARCH-Itô-OI model, we assume that the option-implied…
The paper explores using historical data to improve clinical trial analysis by optimizing covariate weights.
problem Limited covariates in small clinical trials reduce the effectiveness of analysis.
method Leverage historical data to pre-specify covariate weights as a composite covariate.
result A composite covariate improves the cost/benefit ratio and reduces overfitting in small clinical trials.
Combines historical and market data for better portfolio selection.
problem Improving portfolio selection through diverse information integration.
method Bayesian learning via Gaussian mixture model to harmonize historical and market data.
result The method enhances forecasting accuracy and robustness across various capital markets.
Data-driven method for option pricing using historical asset prices.
problem Tackling the gap between historical asset prices and risk-neutral option pricing.
method Identifying a pricing kernel process, solving utility maximization and functional optimization problems using deep learning.
result Demonstrated the efficiency of the data-driven option pricing methodology.
A new GNN model predicts stock trends by learning historical and future correlations.
problem Limited improvement in stock trend prediction models due to ignoring future patterns.
method DishFT-GNN framework that trains a teacher and student model to capture historical and future data correlations.
result State-of-the-art performance on real-world datasets.
The paper evaluates criteria for selecting cryptocurrencies based on historical data.
problem High risk of cryptocurrencies due to volatility.
method Characterized returns and risks using historical data in short time windows (7 and 15 days). Analyzed the importance of criteria using various methods.
result Importance of criteria for selecting cryptocurrencies is analyzed and evaluated.
Historical returns depend on historical closing prices and distributions. We describe how to compute adjusted closing prices from closing price/distribution data with an emphasis on spreadsheet implementation. Then the growth of a security from one date to another (1 + total return) is just the ratio of the correspondi…
Improves trial efficiency by adjusting for historical prognostic scores.
problem Reducing statistical uncertainty in randomized trial estimates.
method Linear covariate adjustment using a prognostic model trained on historical data.
result Prognostic covariate adjustment achieves minimum variance and reduces mean-squared error.
Study optimal product assortment using historical data, proving item coverage suffices.
problem Offline assortment optimization under MNL model with limited historical data.
method Pessimistic Rank-Breaking (PRB) algorithm combining rank-breaking and pessimistic estimation.
result Optimal item coverage is both sufficient and necessary for efficient offline learning.
Improves RL from historical data by stitching trajectories.
problem Lack of high-quality data for offline RL.
method Trajectory Stitching (TS) to augment historical data with synthetic actions.
result Improves RL policy performance over baseline.
Bayesian framework improves variance component estimation in MET data.
problem Inaccurate estimation of variance components in MET data.
method Proposes a Bayesian updating framework using historical data.
result Stabilizes variance component estimation and quantifies uncertainty.
Paper proposes a new method to simulate realistic markets from data.
problem Lack of accurate market simulators leading to misleading conclusions.
method Proposes a world agent model trained on historical data without agent calibration.
result Models consistently outperform previous methods in realism and responsiveness.
The study uses historical revenue data to forecast music catalog cashflows and multipliers.
problem Valuation of music catalogs based on historical revenue data.
method Risk-neutral approach using discounted cashflows formula.
result Ask prices are close to multipliers justified by median song cashflows, while best bids are near multipliers justified by bottom decile cashflows.
This paper reviews and compares deep generative models for financial time series and VaR.
problem Forecasting risk factor distribution in financial markets.
method Apply multiple deep generative models (CGAN, CWGAN, Diffusion, Signature WGAN) and propose new methods for conditional time series generation.
result Top performing models are Historical Simulation, GARCH, and CWGAN.
Algometrics analyzes how predictive models affect their own forecasts in algorithmic markets.
problem How predictive models affect their own forecasts in algorithmic markets.
method Introduces algometrics, a framework for time series with feedback, proving three results on deployment risk.
result Deployment risk cannot be identified from passive historical data alone, and historical rankings can invert under crowding.
To meet the Basel II regulatory requirements for the Advanced Measurement Approaches, the bank's internal model must include the use of internal data, relevant external data, scenario analysis and factors reflecting the business environment and internal control systems. Quantification of operational risk cannot be base…
How do we learn from biased data? Historical datasets often reflect historical prejudices; sensitive or protected attributes may affect the observed treatments and outcomes. Classification algorithms tasked with predicting outcomes accurately from these datasets tend to replicate these biases. We advocate a causal mode…
Identifying the type of font (e.g., Roman, Blackletter) used in historical documents can help optical character recognition (OCR) systems produce more accurate text transcriptions. Towards this end, we present an active-learning strategy that can significantly reduce the number of labeled samples needed to train a font…
Paper presents a method for geographic ratemaking using spatial embeddings.
problem Lack of historical loss data in areas with high exposures.
method Construct spatial features within a complex representation model and use them as inputs to a predictive model.
result Predictions have smaller bias and variance than other spatial interpolation models.
Paper proposes a model to predict stock prices using historical and sentiment data.
problem Improving accuracy in predicting stock prices.
method Integrates historical and sentiment data to predict stock prices using LSTM.
result Improved accuracy in predicting stock prices.
It is well known that the historical logs are used for evaluating and learning policies in interactive systems, e.g. recommendation, search, and online advertising. Since direct online policy learning usually harms user experiences, it is more crucial to apply off-policy learning in real-world applications instead. Tho…
ADR helps LLMs find and use historical analogies for foresight analysis.
problem LLMs struggle to find relevant historical analogies due to surface-level matching.
method Proposes CANA framework with mechanism alignment and cross-analogy confirmation.
result CANA improves historical analogy generation by up to 10%.
Paper introduces a new method for calibrating ESGs to both historical and forward-looking data.
problem Lack of a generally accepted methodology for calibrating ESGs to forward-looking information.
method Conditional Scenario Simulator framework for consistent calibration of economic and financial variables.
result Framework can embed various financial and macroeconomic models and demonstrate practical examples in frequentist and Bayesian settings.
This paper studies an application of machine learning in extracting features from the historical market implied corporate bond yields. We consider an example of a hypothetical illiquid fixed income market. After choosing a surrogate liquid market, we apply the Denoising Autoencoder (DAE) algorithm to learn the features…
Machine learning automates digitization of historical data.
problem Manual transcription is costly and difficult for large, detailed datasets.
method Apply machine learning techniques for unsupervised layout classification and attention-based neural networks.
result Machine learning can automate the digitization process for historical data.
Mathematical properties of the historical GDP/cap distributions are discussed and explained. These distributions are frequently incorrectly interpreted and the Unified Growth Theory is an outstanding example of such common misconceptions. It is shown here that the fundamental postulates of this theory are contradicted …
This study reviews techniques to estimate volatility and price Variance Swaps.
problem Estimating historical volatility and pricing Variance Swaps.
method Review of existing techniques.
result Discussion of various methods to estimate volatility and price Variance Swaps.
The paper proposes an asset allocation strategy using the Sortino ratio for better performance.
problem Traditional asset allocation methods like the Sharpe ratio do not penalize negative returns adequately.
method The Sortino ratio is used to maximize asset allocation, penalizing only negative return variances.
result The Sortino ratio-based strategy outperforms traditional methods like the Kelly criterion.
New algorithm combines new and historical data with different input dimensions for linear regression.
problem Combining new and historical data with different input dimensions for improved accuracy.
method Proposes a transfer learning algorithm with rigorous theoretical robustness analysis.
result Achieves state-of-the-art performance on 9 real-life datasets.
This paper investigates the impact of pre-existing offline data on online learning, in the context of dynamic pricing. We study a single-product dynamic pricing problem over a selling horizon of T periods. The demand in each period is determined by the price of the product according to a linear demand model with unkn…
In this paper a highly abstracted view on the historical development of Genetic Algorithms for the Traveling Salesman Problem is given. In a meta-data analysis three phases in the development can be distinguished. First exponential growth in interest till 1996 can be observed, growth stays linear till 2011 and after th…