Representations of probability measures in reproducing kernel Hilbert spaces provide a flexible framework for fully nonparametric hypothesis tests of independence, which can capture any type of departure from independence, including nonlinear associations and multivariate interactions. However, these approaches come wi…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
BanditLP optimizes personalized recommendations for large-scale systems.
Paper tackles non-convex constrained DRO with a stochastic algorithm for large-scale applications.
We present an idea of unifying small scale (topology, proximity spaces, uniform spaces) and large scale (coarse spaces, large scale spaces). It relies on an analog of multilinear forms from Linear Algebra. Each form has a large scale compactification and those include all well-known compactifications: Higson corona, Gr…
String kernels are attractive data analysis tools for analyzing string data. Among them, alignment kernels are known for their high prediction accuracies in string classifications when tested in combination with SVM in various applications. However, alignment kernels have a crucial drawback in that they scale poorly du…
Scalable and robust TR decomposition for large-scale data with missing entries and outliers.
Paper proposes PPMM for fast estimation of large-scale OTM.
GGDA simplifies DA for large models, speeding up attribution by up to 50x.
0-1 knapsack is of fundamental importance in computer science, business, operations research, etc. In this paper, we present a deep learning technique-based method to solve large-scale 0-1 knapsack problems where the number of products (items) is large and/or the values of products are not necessarily predetermined but…
Method analyzes large-scale network data to detect communication pattern shifts.
LIBTwinSVM offers a free library for efficient Twin Support Vector Machines.
Scalable PnP-ADMM for large-scale imaging problems.
Many modern big data applications feature large scale in both numbers of responses and predictors. Better statistical efficiency and scientific insights can be enabled by understanding the large-scale response-predictor association network structures via layers of sparse latent factors ranked by importance. Yet sparsit…
We study in this paper a variant of Wasserstein barycenter problem, which we refer to as tree-Wasserstein barycenter, by leveraging a specific class of ground metrics, namely tree metrics, for Wasserstein distance. Drawing on the tree structure, we propose an efficient algorithmic approach to solve the tree-Wasserstein…
We consider the problem of solving a large-scale Quadratically Constrained Quadratic Program. Such problems occur naturally in many scientific and web applications. Although there are efficient methods which tackle this problem, they are mostly not scalable. In this paper, we develop a method that transforms the quadra…
Divide-and-conquer is a general strategy to deal with large scale problems. It is typically applied to generate ensemble instances, which potentially limits the problem size it can handle. Additionally, the data are often divided by random sampling which may be suboptimal. To address these concerns, we propose the $DC^…
FibeRed reduces complex data dimensions while preserving topology.
Efficient approximation lies at the heart of large-scale machine learning problems. In this paper, we propose a novel, robust maximum entropy algorithm, which is capable of dealing with hundreds of moments and allows for computationally efficient approximations. We showcase the usefulness of the proposed method, its eq…
Decomposes ultrametric spaces into scaled simplices.
This paper surveys large-scale machine learning methods for efficient data analysis.
We consider convex-concave saddle point problems with a separable structure and non-strongly convex functions. We propose an efficient stochastic block coordinate descent method using adaptive primal-dual updates, which enables flexible parallel optimization for large-scale problems. Our method shares the efficiency an…
New neural network method simplifies high-dimensional data.
We show that uniformly finite homology of products of trees vanishes in all degrees except degree , where it is infinite dimensional. Our method is geometric and applies to several large scale homology theories, including almost equivariant homology and controlled coarse homology. As an application we determine …
Hashing has been widely used for large-scale approximate nearest neighbor search because of its storage and search efficiency. Recent work has found that deep supervised hashing can significantly outperform non-deep supervised hashing in many applications. However, most existing deep supervised hashing methods adopt a …
DistPre predicts traffic speeds efficiently for large networks.
UCB algorithm adapted for large-scale, non-sub-Gaussian problems.
This work aims to create a large-scale model for critical care time series data.
Novel LRMC tackles missing data and outliers in large-scale low-rank data recovery.
New machine learning method detects quantum separability in large-scale systems.
Paper proposes data quality measures for large-scale high-dimensional data.
New method estimates bidirectional causal effects in large-scale systems.
Improved kernel Stein discrepancy for large-scale data.
Deep neural networks provide meaningful uncertainty estimates for large-scale simulations.
Majorization-minimization algorithms consist of iteratively minimizing a majorizing surrogate of an objective function. Because of its simplicity and its wide applicability, this principle has been very popular in statistics and in signal processing. In this paper, we intend to make this principle scalable. We introduc…
New Bayesian optimization models for efficient material screening.
The sum-of-correlations (SUMCOR) formulation of generalized canonical correlation analysis (GCCA) seeks highly correlated low-dimensional representations of different views via maximizing pairwise latent similarity of the views. SUMCOR is considered arguably the most natural extension of classical two-view CCA to the m…
Kernel methods provide a principled way to perform non linear, nonparametric learning. They rely on solid functional analytic foundations and enjoy optimal statistical properties. However, at least in their basic form, they have limited applicability in large scale scenarios because of stringent computational requireme…
Principal Component Analysis (PCA) is a dimension reduction technique. It produces inconsistent estimators when the dimensionality is moderate to high, which is often the problem in modern large-scale applications where algorithm scalability and model interpretability are difficult to achieve, not to mention the preval…
Paper develops robust methods for large-scale testing without tuning parameters.
A new method resolves permutation issues in shuffled linear regression for large-scale applications.
This paper extends Mirror Descent to Riemannian manifolds for optimization.
Sparse Convex Biclustering improves accuracy and robustness in high-dimensional datasets.
Study large-scale geometry of graph braid groups via cubical structures.
Internet companies are facing the need for handling large-scale machine learning applications on a daily basis and distributed implementation of machine learning algorithms which can handle extra-large scale tasks with great performance is widely needed. Deep forest is a recently proposed deep learning framework which …
Water pollution is a major global environmental problem, and it poses a great environmental risk to public health and biological diversity. This work is motivated by assessing the potential environmental threat of coal mining through increased sulfate concentrations in river networks, which do not belong to any simple …
Paper introduces slow kill for efficient large-scale variable screening.
This paper proposes a new Nystrom-based clustering algorithm for large-scale data.
We propose LOCO, an algorithm for large-scale ridge regression which distributes the features across workers on a cluster. Important dependencies between variables are preserved using structured random projections which are cheap to compute and must only be communicated once. We show that LOCO obtains a solution which …