We show Vector Autoregressive Moving Average models with scalar Moving Average components could be estimated by generalized least square (GLS) for each fixed moving average polynomial. The conditional variance of the GLS model is the concentrated covariant matrix of the moving average process. Under GLS the likelihood …
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We present a geometric proof of the averaging theorem for perturbed dynamical systems on a Riemannian manifold, in the case where the flow of the unperturbed vector field is periodic and the -action associated to this vector field is not necessarily trivial. We generalize the averaging procedure \cite{A…
This work proposes a model averaging method for SVM that avoids redundant covariates and achieves asymptotic optimality.
Learning sentence vectors from an unlabeled corpus has attracted attention because such vectors can represent sentences in a lower dimensional and continuous space. Simple heuristics using pre-trained word vectors are widely applied to machine learning tasks. However, they are not well understood from a theoretical per…
We generalise the average asymptotic linking number of a pair of divergence-free vector fields on homology three-spheres by considering the linking of a divergence-free vector field on a manifold of arbitrary dimension with a codimension two foliation endowed with an invariant transverse measure. We prove that the aver…
Optimal algorithms for Riemannian optimization with reduced complexity.
New algorithms for collaborative learning in uncertain, decentralized environments.
New Ricci curvature means derived from plane curvatures.
UCRL-CMDP algorithm optimizes RL with constraints on average costs.
FedNNNN improves FL by adjusting model update vector norms.
Study on averaging geometric structures in Finsler spaces with Lorentzian signature.
We study a variation of Bagchi and Datta's -vector of a simplicial complex , whose entries are defined as weighted averages of Betti numbers of induced subcomplexes of . We show that these invariants satisfy an Alexander-Dehn-Sommerville type identity, and behave nicely under natural operations on triangulated…
Andreas Maurer in the paper "A vector-contraction inequality for Rademacher complexities" extended the contraction inequality for Rademacher averages to Lipschitz functions with vector-valued domains; He did it replacing the Rademacher variables in the bounding expression by arbitrary idd symmetric and sub-gaussian var…
The contraction inequality for Rademacher averages is extended to Lipschitz functions with vector-valued domains, and it is also shown that in the bounding expression the Rademacher variables can be replaced by arbitrary iid symmetric and sub-gaussian variables. Example applications are given for multi-category learnin…
A new meta-algorithm for estimating the conditional average treatment effects is proposed in the paper. The main idea underlying the algorithm is to consider a new dataset consisting of feature vectors produced by means of concatenation of examples from control and treatment groups, which are close to each other. Outco…
SVM used for estimating treatment effects without confounding.
Consider vector valued harmonic maps of at most linear growth, defined on a complete non-compact Riemannian manifold with non-negative Ricci curvature. For the norm square of the pull-back of the target volume form by such maps, we report a strong maximum principle, and equalities among its supremum, its asymptotic ave…
New theory for nonsmooth systems helps optimize and control complex functions.
Stochastic variance reduction algorithms have recently become popular for minimizing the average of a large, but finite number of loss functions. The present paper proposes a Riemannian stochastic quasi-Newton algorithm with variance reduction (R-SQN-VR). The key challenges of averaging, adding, and subtracting multipl…
We consider learning the principal subspace of a large set of vectors from an extremely small number of compressive measurements of each vector. Our theoretical results show that even a constant number of measurements per column suffices to approximate the principal subspace to arbitrary precision, provided that the nu…
We present an iterative technique for finding zeroes of vector fields on Riemannian manifolds. As a special case we obtain a ``nonlinear averaging algorithm'' that computes the centroid of a mass distribution supported in a set of small enough diameter D in a Riemannian manifold M. We estimate the convergence rate of o…
We design a new sparse projection method for a set of vectors that guarantees a desired average sparsity level measured leveraging the popular Hoyer measure (an affine function of the ratio of the and norms). Existing approaches either project each vector individually or require the use of a regulariz…
NLE embeds labels for domain adaptation with neural networks.
Let be an -dimensional manifold and finite-dimensional vector spaces with Euclidean metric. We assign to each a Finsler ellipsoid, i.e., a family of ellipsoids in the fibers of the cotangent bundle of . We prove that the average number of isolated common…
Learning an encoding of feature vectors in terms of an over-complete dictionary or a information geometric (Fisher vectors) construct is wide-spread in statistical signal processing and computer vision. In content based information retrieval using deep-learning classifiers, such encodings are learnt on the flattened la…
In the standard setting of approachability there are two players and a target set. The players play repeatedly a known vector-valued game where the first player wants to have the average vector-valued payoff converge to the target set which the other player tries to exclude it from this set. We revisit this setting in …
We provide non-asymptotic convergence rates of the Polyak-Ruppert averaged stochastic gradient descent (SGD) to a normal random vector for a class of twice-differentiable test functions. A crucial intermediate step is proving a non-asymptotic martingale central limit theorem (CLT), i.e., establishing the rates of conve…
SNAP improves robust computation by emphasizing trustworthy items and downweighting outliers.
Transformationally invariant processors constructed by transformed input vectors or operators have been suggested and applied to many applications. In this study, transformationally identical processing based on combining results of all sub-processes with corresponding transformations at one of the processing steps or …
Dictionaries are collections of vectors used for representations of random vectors in Euclidean spaces. Recent research on optimal dictionaries is focused on constructing dictionaries that offer sparse representations, i.e., -optimal representations. Here we consider the problem of finding optimal dictionaries …
Generalized Berwald manifolds are Finsler manifolds admitting linear connections such that the parallel transports preserve the Finslerian length of tangent vectors. By the fundamental result of the theory \cite{V5} such a linear connection must be metrical with respect to the averaged Riemannian metric given by integr…
Learning the parameters of a (potentially partially observable) random field model is intractable in general. Instead of focussing on a single optimal parameter value we propose to treat parameters as dynamical quantities. We introduce an algorithm to generate complex dynamics for parameters and (both visible and hidde…
A machine learning approach to record fusion with high accuracy.
Introduces a new quantile regression method for financial and wage data analysis.
The original k-means clustering method works only if the exact vectors representing the data points are known. Therefore calculating the distances from the centroids needs vector operations, since the average of abstract data points is undefined. Existing algorithms can be extended for those cases when the sole input i…
Improved bounds for Monte Carlo Rademacher Averages using self-bounding functions.
This paper investigates the average-case time complexity of certifying RIP matrices.
In recent years, stochastic variance reduction algorithms have attracted considerable attention for minimizing the average of a large but finite number of loss functions. This paper proposes a novel Riemannian extension of the Euclidean stochastic variance reduced gradient (R-SVRG) algorithm to a manifold search space.…
MILDA uses unlabelled data to compute LDA projections.
SVR analyzed within RQ framework for risk management.
Gas demand is made of three components: Residential, Industrial, and Thermoelectric Gas Demand. Herein, the one-day-ahead prediction of each component is studied, using Italian data as a case study. Statistical properties and relationships with temperature are discussed, as a preliminary step for an effective feature s…
We consider random vectors drawn from a multivariate normal distribution and compute the sample statistics in the presence of non-stationary correlations. For this purpose, we construct an ensemble of random correlation matrices and average the normal distribution over this ensemble. The resulting distribution contains…
Dictionaries are collections of vectors used for representations of elements in Euclidean spaces. While recent research on optimal dictionaries is focussed on providing sparse (i.e., -optimal,) representations, here we consider the problem of finding optimal dictionaries such that representations of samples of …
ALP outperforms other data descriptors in one-class classification.
Stochastic Gradient Descent (SGD) is one of the simplest and most popular stochastic optimization methods. While it has already been theoretically studied for decades, the classical analysis usually required non-trivial smoothness assumptions, which do not apply to many modern applications of SGD with non-smooth object…
Speaker embeddings are continuous-value vector representations that allow easy comparison between voices of speakers with simple geometric operations. Among others, i-vector and x-vector have emerged as the mainstream methods for speaker embedding. In this paper, we illustrate the use of modern computation platform to …
A quantum model classifies financial sentiment by mapping text chunks to quantum circuits.
The nearest neighbor method together with the dynamic time warping (DTW) distance is one of the most popular approaches in time series classification. This method suffers from high storage and computation requirements for large training sets. As a solution to both drawbacks, this article extends learning vector quantiz…