Generative model uses random weighted support points for interpretable data sampling.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New ensemble SVM model reduces prediction error without choosing best kernel.
The least-squares support vector machine is a frequently used kernel method for non-linear regression and classification tasks. Here we discuss several approximation algorithms for the least-squares support vector machine classifier. The proposed methods are based on randomized block kernel matrices, and we show that t…
Study recovers tree structure in noisy MRFs with support size 3 or more.
New algorithm recovers model coefficients and supports from noisy data.
Algorithm reduces support of discrete measures by integrating against functions.
Dictionary learning is a popular approach for inferring a hidden basis or dictionary in which data has a sparse representation. Data generated from the dictionary A (an n by m matrix, with m > n in the over-complete setting) is given by Y = AX where X is a matrix whose columns have supports chosen from a distribution o…
In this paper, we study randomized reduction methods, which reduce high-dimensional features into low-dimensional space by randomized methods (e.g., random projection, random hashing), for large-scale high-dimensional classification. Previous theoretical results on randomized reduction methods hinge on strong assumptio…
New bounds on continuous random variables' right-tail probabilities.
Study predicts academic achievement using students' support networks.
KSG mutual information estimator, which is based on the distances of each sample to its k-th nearest neighbor, is widely used to estimate mutual information between two continuous random variables. Existing work has analyzed the convergence rate of this estimator for random variables whose densities are bounded away fr…
A new method for causal inference in high-dimensional data using machine learning.
The paper analyzes sparse PCA for incomplete data and proves support recovery conditions.
In this paper we propose a unified framework for structured prediction with latent variables which includes hidden conditional random fields and latent structured support vector machines as special cases. We describe a local entropy approximation for this general formulation using duality, and derive an efficient messa…
Study on using random subspaces for ERM with various loss functions.
We consider random walks on the mapping class group whose support generates a non-elementary subgroup and contains a pseudo-Anosov map whose invariant Teichmüller geodesic is in the principal stratum. For such random walks, we show that mapping classes along almost every infinite sample path are eventually pseudo-Anoso…
Study on recovering supports of multiple sparse vectors from mixed linear measurements.
Proposes a link between randomness and compression in deep learning.
UniNet efficiently learns network representations from large graphs.
New method reduces variance and bias in approximating indefinite kernels.
We show that the probability that a finitely supported random walk on a non-elementary subgroup of the the mapping class group gives a non-pseudo-Anosov element decays exponentially in the length of the random walk. More generally, we show that if R is a set of mapping class group elements with an upper bound on their …
Study examines large deviations in random walks on hyperbolic spaces.
Theoretical study of random forests for nonlinear time series.
Orthogonal random features approximate a Bessel kernel, offering sharper bounds than random Fourier features.
Kaimanovich and Masur showed that a random walk on the mapping class group for an initial distribution with finite first moment and whose support generates a non-elementary subgroup, converges almost surely to a point in the space PMF of projective measured foliations on the surface. This defines a harmonic measure on …
A new method for efficient inference in probabilistic programs with mixed support.
TEC combines multiple RPSTMs to classify big tensors efficiently.
We consider selection of random predictors for high-dimensional regression problem with binary response for a general loss function. Important special case is when the binary model is semiparametric and the response function is misspecified under parametric model fit. Selection for such a scenario aims at recovering th…
We study the volume distribution of nodal domains of random band-limited functions on generic manifolds, and find that in the high energy limit a typical instance obeys a deterministic universal law, independent of the manifold. Some of the basic qualitative properties of this law, such as its support, monotonicity and…
RFX-Fuse combines Breiman and Cutler's Random Forest with modern ML capabilities.
Study on random surfaces in hyperbolic 3-manifolds, focusing on geometric and topological properties.
Improvement of statistical learning models in order to increase efficiency in solving classification or regression problems is still a goal pursued by the scientific community. In this way, the support vector machine model is one of the most successful and powerful algorithms for those tasks. However, its performance d…
The field of few-shot learning has been laboriously explored in the supervised setting, where per-class labels are available. On the other hand, the unsupervised few-shot learning setting, where no labels of any kind are required, has seen little investigation. We propose a method, named Assume, Augment and Learn or AA…
We present Random Partition Kernels, a new class of kernels derived by demonstrating a natural connection between random partitions of objects and kernels between those objects. We show how the construction can be used to create kernels from methods that would not normally be viewed as random partitions, such as Random…
We systematically investigate the problem of representing Markov chains by families of random maps, and which regularity of these maps can be achieved depending on the properties of the probability measures. Our key idea is to use techniques from optimal transport to select optimal such maps. Optimal transport theory a…
Maximizes stock portfolio predictability using machine learning.
We consider random walks on the mapping class group that have finite first moment with respect to the word metric, whose support generates a non-elementary subgroup and contains a pseudo-Anosov map whose invariant Teichmuller geodesic is in the principal stratum of quadratic differentials. We show that a Teichmuller ge…
We propose a general formalism of iterated random functions with semigroup property, under which exact and approximate Bayesian posterior updates can be viewed as specific instances. A convergence theory for iterated random functions is presented. As an application of the general theory we analyze convergence behaviors…
A well-known problem in data science and machine learning is {\em linear regression}, which is recently extended to dynamic graphs. Existing exact algorithms for updating the solution of dynamic graph regression require at least a linear time (in terms of : the size of the graph). However, this time complexity might…
We define two non-linear operations with random (not necessarily closed) sets in Banach space: the conditional core and the conditional convex hull. While the first is sublinear, the second one is superlinear (in the reverse set inclusion ordering). Furthermore, we introduce the generalised conditional expectation of r…
Randomly chosen support makes sparse linear regression easy.
We show that the error probability of reconstructing kernel matrices from Random Fourier Features for the Gaussian kernel function is at most , where is the number of random features and is the diameter of the data domain. We also provide an information-theoretic method-independen…
We consider a random walk on the mapping class group of a surface of finite type. We assume that the random walk is determined by a probability measure whose support is finite and generates a non-elementary subgroup . We further assume that is not consisting only of lifts with respect to any one covering. Then w…
Researchers prove hitting measure singularity for most Fuchsian and Kleinian groups.
The overall equipment effectiveness (OEE) is a performance measurement metric widely used. Its calculation provides to the managers the possibility to identify the main losses that reduce the machine effectiveness and then take the necessary decisions in order to improve the situation. However, this calculation is done…
We prove that, under low noise assumptions, the support vector machine with random features (RFSVM) can achieve the learning rate faster than on a training set with samples when an optimized feature map is used. Our work extends the previous fast rate analysis of random features method from…
In this paper we study the problem of maximizing expected utility from the terminal wealth with proportional transaction costs and random endowment. In the context of the existence of consistent price systems, we consider the duality between the primal utility maximization problem and the dual one, which is set up on t…
Least Squares Estimators are suboptimal for 5D convex functions.