Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

326496128 · Jun 202019922001200920182026
48 results for unknown subspaces

Develops methods to construct confidence regions for singular subspaces in low-rank matrix regression.

problem Recovering the singular subspace of a low-rank matrix from noisy measurements.
method Two-step procedure involving de-biasing and empirical singular vector calculation.
result Asymptotically normal joint projection distance for confidence regions of the true singular subspace.

Agents collaborate to reduce regret in a multi-agent linear bandit problem with side information.

problem Reducing regret in a multi-agent stochastic linear bandit with side information.
method A decentralized algorithm where agents communicate subspace indices and each plays a projected LinUCB on the corresponding low-dimensional subspace.
result Per-agent finite-time regret is much smaller when agents communicate compared to non-communicating case.

We consider the problem of clustering noisy high-dimensional data points into a union of low-dimensional subspaces and a set of outliers. The number of subspaces, their dimensions, and their orientations are unknown. A probabilistic performance analysis of the thresholding-based subspace clustering (TSC) algorithm intr…

2013-05-15abs ↗pdf ↗

New research shows SSC fails when points on the same subspace are mislabeled.

problem Failure of SSC when points on the same subspace are mislabeled.
method Analyzed the effect of different distributions of points on the same subspace.
result SSC fails to infer correct labels when points on the same subspace fall into more than one cluster.

We consider the problem of clustering a set of high-dimensional data points into sets of low-dimensional linear subspaces. The number of subspaces, their dimensions, and their orientations are unknown. We propose a simple and low-complexity clustering algorithm based on thresholding the correlations between the data po…

2013-03-15abs ↗pdf ↗

Subspace clustering refers to the problem of clustering high-dimensional data points into a union of low-dimensional linear subspaces, where the number of subspaces, their dimensions and orientations are all unknown. In this paper, we propose a variation of the recently introduced thresholding-based subspace clustering…

2014-03-13abs ↗pdf ↗

Unified framework for structured principal subspace estimation with bounds and rates.

problem Structured principal subspace estimation problems.
method Unified framework, minimax lower and upper bounds, information-geometric complexity.
result Minimax rates of convergence for specific settings, including optimal rates for non-negative PCA/SVD.

The paper finds non-Gaussian directions in high-dimensional data using Wasserstein distance.

problem Locating interesting non-Gaussian features in high-dimensional data.
method Projection pursuit using 2-Wasserstein distance to maximize the difference from Gaussian.
result Statistical guarantees for accurately approximating an unknown low-dimensional non-Gaussian subspace.

Subspace clustering refers to the problem of clustering unlabeled high-dimensional data points into a union of low-dimensional linear subspaces, assumed unknown. In practice one may have access to dimensionality-reduced observations of the data only, resulting, e.g., from "undersampling" due to complexity and speed con…

2014-04-27abs ↗pdf ↗

BO method identifies sparse subspaces for efficient high-dimensional optimization.

problem Efficient optimization of high-dimensional black-box functions.
method Sparse Gaussian process surrogate models on axis-aligned subspaces with Hamiltonian Monte Carlo inference.
result SAASBO achieves excellent performance on synthetic and real-world problems.

The problem of clustering noisy and incompletely observed high-dimensional data points into a union of low-dimensional subspaces and a set of outliers is considered. The number of subspaces, their dimensions, and their orientations are assumed unknown. We propose a simple low-complexity subspace clustering algorithm, w…

2013-07-18abs ↗pdf ↗

Bayesian approach for online subspace learning from incomplete data.

problem Handling incomplete large-scale datasets for subspace learning.
method Online variational Bayes subspace learning with low-rank and sparsity constraints.
result The proposed algorithm outperforms state-of-the-art methods in estimation accuracy.

A parameter-free method clusters data points from multiple subspaces.

problem Subspace clustering with unknown number of clusters and parameters.
method Clusters data points based on angle differences between subspaces; merges clusters until final clustering is obtained.
result Parameter-free approach for clustering data points from multiple subspaces.

Paper proposes clustering algorithms for data from Union of Polyhedral Cones model.

problem Clustering data from multiple convex polyhedral cones.
method Sparse Subspace Clustering, Least squares approximation, K-nearest neighbor, Spectral Clustering.
result KNN outperforms NCL and LSA in clustering data from UOPC model.

Subspace clustering refers to the problem of clustering unlabeled high-dimensional data points into a union of low-dimensional linear subspaces, whose number, orientations, and dimensions are all unknown. In practice one may have access to dimensionality-reduced observations of the data only, resulting, e.g., from unde…

2015-07-25abs ↗pdf ↗

Paper analyzes SSC for data with missing entries, improving performance.

problem Theoretical analysis of SSC with missing data entries.
method Analyzes theoretical guarantees for SSC with incomplete data, projecting zero-filled data onto observation pattern.
result Improves performance of SSC with incomplete data by projecting zero-filled data onto observation pattern.

New method identifies latent components in PNL mixtures without strong assumptions.

problem Identifying latent components in PNL mixtures under unknown nonlinear functions.
method Carefully designed UML criterion to identify a null space associated with the mixing system.
result Identification/removal of unknown nonlinearity under minimal conditions.

A new method splits surface flow discretizations into streamfunctions and harmonic fields.

problem Discretizing incompressible flows on surfaces with pressure and saddle-point structure.
method Discrete Helmholtz-Hodge decomposition for BDM elements on surfaces.
result Eliminates pressure and saddle-point structure, ensuring exact tangentiality and divergence-freeness.

Model improves covariance estimation from shared and distinct datasets.

problem Limited sample sizes and shared covariance structure across related datasets.
method Spiked covariance model with shared subspace, closed-form pooling weight, and asymptotic guarantees.
result Improves estimation of high-dimensional covariance matrices from related datasets.

Regularized M-estimators are used in diverse areas of science and engineering to fit high-dimensional models with some low-dimensional structure. Usually the low-dimensional structure is encoded by the presence of the (unknown) parameters in some low-dimensional model subspace. In such settings, it is desirable for est…

2013-05-31abs ↗pdf ↗

AGOP from KRR recovers central subspace in fewer samples than needed for prediction.

problem Recovering low-dimensional structure in multi-index polynomial functions.
method Fit kernel ridge regression and compute AGOP from the fitted predictor.
result AGOP's top rr eigenspace recovers the central subspace in ndp+δn \asymp d^{p+δ} samples.

ProSub uses angles in feature space to classify data as in- or out-of-distribution.

problem Open-set semi-supervised learning with unknown classes.
method Probabilistic approach based on angles in feature space, estimating conditional distributions of scores.
result ProSub achieves state-of-the-art performance on benchmark problems.

Algorithm learns polynomials in Gaussian inputs with reduced sample complexity.

problem Learning polynomials of few relevant dimensions in high-dimensional data.
method Filtered PCA for warm start, geodesic SGD for accuracy.
result Sample complexity roughly N=Or,d(nlog2(1/ε)(logn)d)N = O_{r,d}(n \log^2(1/ε) (\log n)^d), runtime Or,d(Nn2)O_{r,d}(N n^2).

CobBO optimizes expensive functions in high dimensions by using a two-stage kernel approach.

problem Bayesian optimization struggles in high dimensions due to computational inefficiency.
method Coordinate backoff Bayesian Optimization with two-stage kernels.
result CobBO finds solutions comparable to or better than other methods in high dimensions.

This work quantifies the space of transferable adversarial examples.

problem Understanding and quantifying the space of transferable adversarial examples.
method Estimating the dimensionality of the space of adversarial inputs and analyzing the similarity of decision boundaries.
result Adversarial examples span a contiguous subspace of large (~25) dimensionality, and higher dimensionality subspaces are more likely to intersect.

The stable 4-genus of a knot K in 3-space is the limiting value of g_4(nK)/n, where g_4 denotes the 4-genus and n goes to infinity. This induces a seminorm on CQ, the concordance group tensored with the rational numbers. Basic properties of the stable genus are developed, as are examples focused on understanding the un…

2009-04-20abs ↗pdf ↗

This paper improves diffusion models for low-dimensional data.

problem Theoretical foundations of diffusion models are lacking for low-dimensional data.
method Score approximation, estimation, and distribution recovery of diffusion models on low-dimensional data.
result Sample complexity bounds for distribution estimation using diffusion models are provided.

New algorithm tackles bilinear bandit problem with low-rank structure.

problem Finding the optimal action in a bilinear bandit problem with low-rank reward matrix.
method Two-stage algorithm: subspace exploration followed by linear bandit refinement.
result Regret bound of ESTR is O~((d1+d2)3/2rT)\widetilde{\mathcal{O}}((d_1+d_2)^{3/2} \sqrt{r T}).

EAGC boosts GCD by regulating gradient entanglement, improving known and novel category separability.

problem Gradient entanglement distorts supervised gradients and overlaps known and novel class representations.
method EAGC uses AGA and EEP to align and project gradients, reducing entanglement and overlap.
result EAGC consistently boosts GCD performance, setting new state-of-the-art results.

We consider the problem of joint estimation of structured inverse covariance matrices. We perform the estimation using groups of measurements with different covariances of the same unknown structure. Assuming the inverse covariances to span a low dimensional linear subspace in the space of symmetric matrices, our aim i…

2015-11-20abs ↗pdf ↗

The paper develops Kalman filters for unknown systems with sample complexity bounds.

problem Designing Kalman filters for systems with unknown parameters and noise.
method Combines system identification with Kalman filter design, ensuring robustness and sub-optimality guarantees.
result Proves sub-optimality guarantees for both Certainty Equivalent and robust Kalman filters with sample complexity bounds.

Efficiently recovers data corrupted by adversarial noise in structured settings.

problem Recovering clean data points from corrupted Gaussian data with low-rank noise and adversarial coordinate corruptions.
method Developed an efficient algorithm using a combinatorial approach to analyze Basis Pursuit (BP) method.
result Achieved nearly-optimal recovery of data points up to a ildeO(ks/d) ilde O(ks/d) error bound.

Study of symplectic Stiefel and Grassmann manifolds with geodesics and applications.

problem Understanding symplectic bases and subspaces for data processing.
method Lie group approach to derive geodesics and retractions for pseudo-Riemannian and Riemannian metrics.
result Efficient formulas for geodesics and retractions on symplectic manifolds.

In many applications, such as economics, operations research and reinforcement learning, one often needs to estimate a multivariate regression function f subject to a convexity constraint. For example, in sequential decision processes the value of a state under optimal subsequent decisions may be known to be convex or …

2011-09-01abs ↗pdf ↗