Average Oracle outperforms DCC+NLS in portfolio optimization.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Improved average distance classifier for HDLSS settings with multiple population differences.
We consider stochastic gradient descent algorithms for minimizing a non-smooth, strongly-convex function. Several forms of this algorithm, including suffix averaging, are known to achieve the optimal convergence rate in expectation. We consider a simple, non-uniform averaging strategy of Lacoste-Julien et al. …
New Ricci curvature means derived from plane curvatures.
Distributed statistical inference has recently attracted enormous attention. Many existing work focuses on the averaging estimator. We propose a one-step approach to enhance a simple-averaging based distributed estimator. We derive the corresponding asymptotic properties of the newly proposed estimator. We find that th…
We prove a sharp estimate on the expected value of the integral of the index of a simple random walk on the square or triangular lattice. This gives new lower bounds on the averaged Dehn function, which measures the expected area needed to fill a random curve with a disc.
We give a short and simple proof of Cauchy's surface area formula, which states that the average area of a projection of a convex body is equal to its surface area up to a multiplicative constant in the dimension.
Efficient algorithm for global optimization of multivariate Lipschitz functions.
Study the averaging estimator on graphs with labeled nodes.
AverageTime uses simple averaging to enhance long-term time series forecasting.
Stochastic Gradient Descent (SGD) is one of the simplest and most popular stochastic optimization methods. While it has already been theoretically studied for decades, the classical analysis usually required non-trivial smoothness assumptions, which do not apply to many modern applications of SGD with non-smooth object…
Learning sentence vectors from an unlabeled corpus has attracted attention because such vectors can represent sentences in a lower dimensional and continuous space. Simple heuristics using pre-trained word vectors are widely applied to machine learning tasks. However, they are not well understood from a theoretical per…
OMD and DA perform similarly in static settings but OMD is inferior under dynamic learning rates.
This paper improves forecasts for diverse time series by averaging similar ones.
Paper proposes a fair grading method for randomized exams.
Understanding consumption dynamics and its impact on the whole economy and welfare within the present economic crisis is not an easy task. Indeed the level of consumer demand for different goods varies with the prices, consumer incomes and demographic factors. Furthermore crisis may trigger different behaviors which re…
BEMA reduces bias in EMA, leading to faster convergence and better performance.
We introduce a simple algorithm, True Asymptotic Natural Gradient Optimization (TANGO), that converges to a true natural gradient descent in the limit of small learning rates, without explicit Fisher matrix estimation. For quadratic models the algorithm is also an instance of averaged stochastic gradient, where the par…
The study calculates average crosscap numbers for 2-bridge knots.
Instability and variability of Deep Reinforcement Learning (DRL) algorithms tend to adversely affect their performance. Averaged-DQN is a simple extension to the DQN algorithm, based on averaging previously learned Q-values estimates, which leads to a more stable training procedure and improved performance by reducing …
In his seminal work, Schapire (1990) proved that weak classifiers could be improved to achieve arbitrarily high accuracy, but he never implied that a simple majority-vote mechanism could always do the trick. By comparing the asymptotic misclassification error of the majority-vote classifier with the average individual …
This paper proposes a simple but effective graph-based agglomerative algorithm, for clustering high-dimensional data. We explore the different roles of two fundamental concepts in graph theory, indegree and outdegree, in the context of clustering. The average indegree reflects the density near a sample, and the average…
Exponential smoothers are a simple and memory efficient way to compute running averages of time series. Here we define and describe practical properties of exponential smoothers for signals observed at constant and variable intervals.
Theory and methods to mitigate omitted variable bias in causal machine learning.
In this work, we addressed the issue of combining linear classifiers using their score functions. The value of the scoring function depends on the distance from the decision boundary. Two score functions have been tested and four different combination strategies were investigated. During the experimental study, the pro…
A new method for averaging data on manifolds is proposed, offering simplicity and efficiency.
In this paper, we derive a new model of synaptic plasticity, based on recent algorithms for reinforcement learning (in which an agent attempts to learn appropriate actions to maximize its long-term average reward). We show that these direct reinforcement learning algorithms also give locally optimal performance for the…
LASSO-PCA combines LASSO and PCA for automated forecast averaging.
Group averaging boosts model accuracy without training cost.
Introduces robust and decomposable AP for image retrieval.
We examine two different techniques for parameter averaging in GAN training. Moving Average (MA) computes the time-average of parameters, whereas Exponential Moving Average (EMA) computes an exponentially discounted sum. Whilst MA is known to lead to convergence in bilinear settings, we provide the -- to our knowledge …
Simple mode exploration methods do not improve performance in neural networks.
We adapt the optimization's concept of momentum to reinforcement learning. Seeing the state-action value functions as an analog to the gradients in optimization, we interpret momentum as an average of consecutive -functions. We derive Momentum Value Iteration (MoVI), a variation of Value Iteration that incorporates …
Simplified analysis of SGD for linear regression with weight averaging.
This paper studies robust estimation methods in high dimensions, comparing model-averaged and composite quantile estimators.
Paper proposes a new method for conditional coverage in conformal prediction.
Deep neural networks are typically trained by optimizing a loss function with an SGD variant, in conjunction with a decaying learning rate, until convergence. We show that simple averaging of multiple points along the trajectory of SGD, with a cyclical or constant learning rate, leads to better generalization than conv…
In this article we present new results for the pricing of arithmetic Asian options within a Black-Scholes context. To derive these results we make extensive use of the local scale invariance that exists in the theory of contingent claim pricing. This allows us to derive, in a natural way, a simple PDE for the price of …
We consider the problem of designing models to leverage a recently introduced approximate model averaging technique called dropout. We define a simple new model called maxout (so named because its output is the max of a set of inputs, and because it is a natural companion to dropout) designed to both facilitate optimiz…
Gas demand is made of three components: Residential, Industrial, and Thermoelectric Gas Demand. Herein, the one-day-ahead prediction of each component is studied, using Italian data as a case study. Statistical properties and relationships with temperature are discussed, as a preliminary step for an effective feature s…
We propose SWA-Gaussian (SWAG), a simple, scalable, and general purpose approach for uncertainty representation and calibration in deep learning. Stochastic Weight Averaging (SWA), which computes the first moment of stochastic gradient descent (SGD) iterates with a modified learning rate schedule, has recently been sho…
We prove that the binary classifiers of bit strings generated by random wide deep neural networks with ReLU activation function are biased towards simple functions. The simplicity is captured by the following two properties. For any given input bit string, the average Hamming distance of the closest input bit string wi…
The study introduces anytime learning schedules for large language models without fixed horizons.
The average economic agent is often used to model the dynamics of simple markets, based on the assumption that the dynamics of many agents can be averaged over in time and space. A popular idea that is based on this seemingly intuitive notion is to dampen electric power fluctuations from fluctuating sources (as e.g. wi…
New kernels from neural networks show better performance than traditional methods.
New Q-learning method achieves optimal sample complexity for average-reward problems.
This work improves Q-learning for average-reward MDPs, reducing sample and communication complexities in federated settings.
High-dimensional models can outperform simpler ones in causal inference.