BDSG generates samples on distribution boundaries, improving anomaly detection.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Support vector regression (SVR) is one of the most popular machine learning algorithms aiming to generate the optimal regression curve through maximizing the minimal margin of selected training samples, i.e., support vectors. Recent researchers reveal that maximizing the margin distribution of whole training dataset ra…
Paper tackles distributed quantile regression with improved efficiency and support recovery.
A new method for support vector regression using a data-driven insensitive parameter.
KSG mutual information estimator, which is based on the distances of each sample to its k-th nearest neighbor, is widely used to estimate mutual information between two continuous random variables. Existing work has analyzed the convergence rate of this estimator for random variables whose densities are bounded away fr…
Kaimanovich and Masur showed that a random walk on the mapping class group for an initial distribution with finite first moment and whose support generates a non-elementary subgroup, converges almost surely to a point in the space PMF of projective measured foliations on the surface. This defines a harmonic measure on …
Dictionary learning is a popular approach for inferring a hidden basis or dictionary in which data has a sparse representation. Data generated from the dictionary A (an n by m matrix, with m > n in the over-complete setting) is given by Y = AX where X is a matrix whose columns have supports chosen from a distribution o…
KALE flow approximates KL divergence for distributions with disjoint support.
New method calibrates reference distributions for bounded support.
We stabilize the Kumaraswamy distribution for efficient sampling and differentiation.
We obtain sufficient conditions exlcuding the existence of non-trivial distribution sections of bundles over the boundary of symmetric spaces of negative curvature which are invariant with respect to a geometrically finite group of isometries and are supported on the limit set in a strong sense.
The principal support vector machines method (Li et al., 2011) is a powerful tool for sufficient dimension reduction that replaces original predictors with their low-dimensional linear combinations without loss of information. However, the computational burden of the principal support vector machines method constrains …
We present a stochastic algorithm to compute the barycenter of a set of probability distributions under the Wasserstein metric from optimal transport. Unlike previous approaches, our method extends to continuous input distributions and allows the support of the barycenter to be adjusted in each iteration. We tackle the…
When optimizing against the mean loss over a distribution of predictions in the context of a regression task, then even if there is a distribution of targets the optimal prediction distribution is always a delta function at a single value. Methods of constructing generative models need to overcome this tendency. We con…
Paper proposes a method to solve log-optimal portfolios under ambiguous return distributions.
Federated learning supports exact support recovery with minimal communication.
Given only positive (P) and unlabeled (U) data, PU learning can train a binary classifier without any negative data. It has two building blocks: PU class-prior estimation (CPE) and PU classification; the latter has been well studied while the former has received less attention. Hitherto, the distributional-assumption-f…
Maximum entropy distributions with discrete support in dimensions arise in machine learning, statistics, information theory, and theoretical computer science. While structural and computational properties of max-entropy distributions have been extensively studied, basic questions such as: Do max-entropy distributio…
Visual analytics tool detects and corrects concept drift in data streams.
We propose Radial Bayesian Neural Networks (BNNs): a variational approximate posterior for BNNs which scales well to large models while maintaining a distribution over weight-space with full support. Other scalable Bayesian deep learning methods, like MC dropout or deep ensembles, have discrete support-they assign zero…
PBM mechanism improves privacy and accuracy in federated learning.
We propose an efficient distributed online learning protocol for low-latency real-time services. It extends a previously presented protocol to kernelized online learners that represent their models by a support vector expansion. While such learners often achieve higher predictive performance than their linear counterpa…
This paper presents a kernel-based discriminative learning framework on probability measures. Rather than relying on large collections of vectorial training examples, our framework learns using a collection of probability distributions that have been constructed to meaningfully represent training data. By representing …
To keep up with increasing dataset sizes and model complexity, distributed training has become a necessity for large machine learning tasks. Parameter servers ease the implementation of distributed parameter management---a key concern in distributed training---, but can induce severe communication overhead. To reduce c…
Distributed-OMP recovers sparse vectors with low communication costs.
New Stein identity for q-Gaussians reduces gradient variance in machine learning.
New algorithm for sampling from distributions with thin tails.
New method disentangles correlated factors without independence assumption.
The support vector machines (SVM) algorithm is a popular classification technique in data mining and machine learning. In this paper, we propose a distributed SVM algorithm and demonstrate its use in a number of applications. The algorithm is named high-performance support vector machines (HPSVM). The major contributio…
Efficient algorithms for sparse parameter recovery in mixture models.
Estimates support in distributions with sampling artifacts and errors.
The Lasso performs well in ultra-sparse linear models with finite support size.
Optimal learning for parametric prophet inequalities with exponential-type distributions
Research shows guidance in diffusion models does not sample from intended distribution, affecting boundary sampling.
Paper studies regularized KKL divergence for distributions with disjoint supports.
Unique continuation property for measures in high dimensions.
The paper proves consistency of archetypal analysis for multivariate data.
The paper explores how to extrapolate from limited data points using causal mechanisms.
It is shown that bootstrap approximations of support vector machines (SVMs) based on a general convex and smooth loss function and on a general kernel are consistent. This result is useful to approximate the unknown finite sample distribution of SVMs by the bootstrap approach.
k Nearest Neighbor (kNN) method is a simple and popular statistical method for classification and regression. For both classification and regression problems, existing works have shown that, if the distribution of the feature vector has bounded support and the probability density function is bounded away from zero in i…
Normalizing flows can now estimate densities on unknown manifolds.
We propose one-class support measure machines (OCSMMs) for group anomaly detection which aims at recognizing anomalous aggregate behaviors of data points. The OCSMMs generalize well-known one-class support vector machines (OCSVMs) to a space of probability measures. By formulating the problem as quantile estimation on …
We propose one-class support measure machines (OCSMMs) for group anomaly detection which aims at recognizing anomalous aggregate behaviors of data points. The OCSMMs generalize well-known one-class support vector machines (OCSVMs) to a space of probability measures. By formulating the problem as quantile estimation on …
This work develops secure distributed algorithms for machine learning to protect against data poisoning and network attacks.
Extends SW and GSW to compare heterogeneous joint distributions.
This paper presents novel Gaussian process decentralized data fusion algorithms exploiting the notion of agent-centric support sets for distributed cooperative perception of large-scale environmental phenomena. To overcome the limitations of scale in existing works, our proposed algorithms allow every mobile sensing ag…
New distributions allow greedy arm selection in sparse bandit problems.
In this expository paper we illustrate the generality of game theoretic probability protocols of Shafer and Vovk (2001) in finite-horizon discrete games. By restricting ourselves to finite-horizon discrete games, we can explicitly describe how discrete distributions with finite support and the discrete pricing formulas…