Extends topological theorems to new mapping families.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Geodesic X-ray transform proves injective for smooth one-forms on gas giant manifolds.
Improved analysis shows milder over-parameterization for deep neural networks.
Ensemble learning improves anomaly detection for milder symptoms.
For a density on , a {\it high-density cluster} is any connected component of , for some . The set of all high-density clusters forms a hierarchy called the {\it cluster tree} of . We present two procedures for estimating the cluster tree given samples from . The first…
Paper introduces new Finsler metrics preserved under projective transformations.
Logistic regression gets a new, simpler uniform bound.
Efficiently learns a single neuron with adversarial noise, improving on prior work.
We examine the squared error loss landscape of shallow linear neural networks. We show---with significantly milder assumptions than previous works---that the corresponding optimization problems have benign geometric properties: there are no spurious local minima and the Hessian at every saddle point has at least one ne…
Study bubbling Kahler metrics using algebraic geometry.
Paper improves classification rates for private data.
This work addresses various open questions in the theory of active learning for nonparametric classification. Our contributions are both statistical and algorithmic: -We establish new minimax-rates for active learning under common \textit{noise conditions}. These rates display interesting transitions -- due to the inte…
We prove that finitely generated purely loxodromic subgroups of a right-angled Artin group fulfill equivalent conditions that parallel characterizations of convex cocompactness in mapping class groups . In particular, such subgroups are quasiconvex in . In addition, we identify a milder cond…
Gradient methods learn a single neuron under mild assumptions.
Dynamic angles estimated from noisy measurements over time with smoothness constraints.
Unified analysis of SGD and GD convergence rates.
This paper addresses a gap in the classifcation of Codazzi tensors with exactly two eigenfunctions on a Riemannian manifold of dimension three or higher. Derdzinski proved that if the trace of such a tensor is constant and the dimension of one of the the eigenspaces is , then the metric is a warped product where t…
We propose a communication-efficient distributed estimation method for sparse linear discriminant analysis (LDA) in the high dimensional regime. Our method distributes the data of size into machines, and estimates a local sparse LDA estimator on each machine using the data subset of size . After the distri…
Simplified SGD interpretation as Ito process for broader applicability.
In high-dimension, low-sample size (HDLSS) data, it is not always true that closeness of two objects reflects a hidden cluster structure. We point out the important fact that it is not the closeness, but the "values" of distance that contain information of the cluster structure in high-dimensional space. Based on this …
We analyze a class of estimators based on convex relaxation for solving high-dimensional matrix decomposition problems. The observations are noisy realizations of a linear transformation of the sum of an approximately) low rank matrix with a second matrix endowed with a complementary …
Study on fair learning under differential privacy constraints.
We analyze a negative-parameter variant of the diversity-weighted portfolio studied by Fernholz, Karatzas, and Kardaras (Finance Stoch 9(1):1-27, 2005), which invests in each company a fraction of wealth inversely proportional to the company's market weight (the ratio of its capitalization to that of the entire market)…
We analyze the convergence behaviour of a recently proposed algorithm for regularized estimation called Dual Augmented Lagrangian (DAL). Our analysis is based on a new interpretation of DAL as a proximal minimization algorithm. We theoretically show under some conditions that DAL converges super-linearly in a non-asymp…
Ensemble models struggle with detecting mild faults.
Improved anomaly detection for incipient faults using ensemble learning.
We consider the problem of clustering data points in high dimensions, i.e. when the number of data points may be much smaller than the number of dimensions. Specifically, we consider a Gaussian mixture model (GMM) with non-spherical Gaussian components, where the clusters are distinguished by only a few relevant dimens…
In the mixture models problem it is assumed that there are distributions and one gets to observe a sample from a mixture of these distributions with unknown coefficients. The goal is to associate instances with their generating distributions, or to identify the parameters of the hidden distribu…
Paper corrects Max-Margin loss for multi-label tasks.
We demonstrate an equivalence between reproducing kernel Hilbert space (RKHS) embeddings of conditional distributions and vector-valued regressors. This connection introduces a natural regularized loss function which the RKHS embeddings minimise, providing an intuitive understanding of the embeddings and a justificatio…
The choice of admissible trading strategies in mathematical modelling of financial markets is a delicate issue, going back to Harrison and Kreps (1979). In the context of optimal portfolio selection with expected utility preferences this question has been a focus of considerable attention over the last twenty years. We…
Complex performance measures, beyond the popular measure of accuracy, are increasingly being used in the context of binary classification. These complex performance measures are typically not even decomposable, that is, the loss evaluated on a batch of samples cannot typically be expressed as a sum or average of losses…
Deep learning framework for kernel methods using RKHM and Perron-Frobenius operators.
The paper computes torsion invariants for groups acting on complexes.
We theoretically investigate the convergence rate and support consistency (i.e., correctly identifying the subset of non-zero coefficients in the large sample limit) of multiple kernel learning (MKL). We focus on MKL with block-l1 regularization (inducing sparse kernel combination), block-l2 regularization (inducing un…
1-Lipschitz networks are as accurate as classical networks and offer robustness.
In this paper, we address the problem of learning the structure of a pairwise graphical model from samples in a high-dimensional setting. Our first main result studies the sparsistency, or consistency in sparsity pattern recovery, properties of a forward-backward greedy algorithm as applied to general statistical model…
New method identifies Gaussian SEMs with varying error variances.
Improved analysis of accelerated noisy power method for PCA.
In topic modeling, many algorithms that guarantee identifiability of the topics have been developed under the premise that there exist anchor words -- i.e., words that only appear (with positive probability) in one topic. Follow-up work has resorted to three or higher-order statistics of the data corpus to relax the an…
Paper develops KMS Wasserstein for high-dimensional data reduction.
Develops a Best-of-Both-Worlds algorithm for linear contextual bandits with Tsallis entropy.
T-SCI improves Cox-MLP's guaranteed coverage for censored data.
New bounds on generalization error using information density moments.
New method identifies shared components from unpaired multimodal mixtures.
Paper proposes efficient offline neural bandit method.
New algorithm trains ReLU gates provably in linear time.
Improves SGM convergence bounds in W2-distance without strict assumptions.