Anomaly detection for high-dimensional data using large deviations principle.
problem Challenges in anomaly detection for high-dimensional data.
method Large Deviations Anomaly Detection (LAD) algorithm.
result Outperforms state-of-the-art methods on high-dimensional data sets.
Proposes MamBO for efficient high-dimensional large-scale optimization.
problem High-dimensional and large-scale optimization problems in machine learning and simulation.
method Combines subsampling and subspace embeddings with model aggregation to address uncertainty in surrogate models.
result Improves robustness of Bayesian optimization algorithm and achieves superior performance.
Minimal surfaces with negative curvature found in large spheres.
problem Existence of minimal surfaces with negative curvature in large dimensional spheres.
method Applied Song's strategy to closed Riemann surfaces with large automorphism groups, resulting in almost hyperbolic minimal surfaces.
result Existence of closed minimal surfaces with negative induced curvature in any sphere of large dimension.
An extra large metric is a spherical cone metric with all cone angles greater than 2 pi and every closed geodesic longer than 2pi. We show that every two-dimensional extra large metric can be triangulated with vertices at cone points only. The argument implies the same result for Euclidean and hyperbolic cone metrics, …
FibeRed reduces complex data dimensions while preserving topology.
problem Hard embedding of topologically complex datasets in low-dimensional Euclidean space.
method Modeling datasets with vector bundles, reducing fibers while preserving topology.
result FibeRed learns topologically faithful embeddings in lower dimensions than existing methods.
This paper explores saturation effects in spectral algorithms over large dimensions.
problem Saturation effects in spectral algorithms over large dimensions.
method Improved minimax lower bound and gradient flow with early stopping strategy.
result Exact convergence rates of spectral algorithms in large dimensional settings.
Paper proposes data quality measures for large-scale high-dimensional data.
problem Lack of practical data quality measures for large-scale high-dimensional data.
method Proposes two data quality measures: class separability and in-class variability. Efficient algorithms based on random projections and bootstrapping are provided.
result Efficient algorithms for computing data quality measures on large-scale high-dimensional data.
We propose a feature selection method that finds non-redundant features from a large and high-dimensional data in nonlinear way. Specifically, we propose a nonlinear extension of the non-negative least-angle regression (LARS) called N3LARS, where the similarity between input and output is measured through the norm…
For very large datasets, random projections (RP) have become the tool of choice for dimensionality reduction. This is due to the computational complexity of principal component analysis. However, the recent development of randomized principal component analysis (RPCA) has opened up the possibility of obtaining approxim…
Large scale online kernel learning aims to build an efficient and scalable kernel-based predictive model incrementally from a sequence of potentially infinite data points. A current key approach focuses on ways to produce an approximate finite-dimensional feature map, assuming that the kernel used has a feature map wit…
We obtain an improved pseudolocality result for Ricci flows on two-dimensional surfaces that are initially almost-hyperbolic on large hyperbolic balls. We prove that, at the central point of the hyperbolic ball, the Gauss curvature remains close to the hyperbolic value for a time that grows exponentially in the radius …
New method uses tensor decompositions to overcome the curse of dimensionality for large-scale learning.
problem Large-scale machine learning problems with kernel methods.
method Deterministic Fourier features combined with low-rank tensor decomposition for tensor product structure.
result Demonstrated consistent performance and superior results compared to random Fourier features.
Machine learning-based analysis of medical images often faces several hurdles, such as the lack of training data, the curse of dimensionality problem, and the generalization issues. One of the main difficulties is that there exists computational cost problem in dealing with input data of large size matrices which repre…
The paper extends kernel ridge regression to product kernels and reveals new convergence behaviors.
problem Understanding kernel ridge regression in large dimensions with various kernels.
method Established a broad family of large dimensional kernels and derived convergence rates.
result Revealed new phenomena including minimax optimality, saturation effect, and multiple descent behavior.
A hierarchical approach improves classification accuracy in large datasets.
problem Improving classification accuracy in large datasets with high dimensionality.
method Hierarchical subspace learning to scale manifold learning methods.
result Average 5% increase in classification accuracy.
FSL-Net detects and localizes feature shifts in large, high-dimensional datasets.
problem Feature shifts between data sources lead to erroneous features in various applications.
method FSL-Net is a neural network trained on multiple datasets to localize feature shifts.
result FSL-Net accurately localizes feature shifts from unseen datasets without re-training.
Paper proposes PPMM for fast estimation of large-scale OTM.
problem Estimation of large-scale optimal transport maps (OTM) is challenging due to the curse of dimensionality.
method Combines projection pursuit regression and sufficient dimension reduction to adaptively select projection directions.
result PPMM consistently estimates the most informative projection direction and weakly converges to the target OTM.
Paper proposes a robust test for high-dimensional models with large covariates and instruments.
problem Testing high-dimensional linear instrumental variable models with large covariates and instruments.
method Introduces a test based on the maximum norm of multiple parameters and a power-enhanced test.
result The proposed test is robust to heteroskedastic errors and has higher power than existing tests.
Characterizes kernel interpolation in large dimensions, revealing optimal and sub-optimal regions.
problem Understanding the phase diagram of kernel interpolation in large dimensions.
method Characterization of variance and bias under various source conditions.
result Determined the (s,γ)-phase diagram of large-dimensional kernel interpolation. Supervised dimensionality reduction strategies have been of great interest. However, current supervised dimensionality reduction approaches are difficult to scale for situations characterized by large datasets given the high computational complexities associated with such methods. While stochastic approximation strateg…
We obtain an estimate for the volume of neighbourhoods of sets of large curvature in three-dimensional Kähler-Einstein manifolds.
CCP clusters correlated features and projects them to 1D for efficient dimensionality reduction.
problem Efficiency in handling large datasets with high intrinsic dimensions.
method CCP partitions features into correlated clusters and projects them to 1D based on sample correlations.
result CCP achieves efficient dimensionality reduction without matrix diagonalization.
New algorithm selects variables from large datasets.
problem Automatic selection of variables from large datasets.
method Uses Graphical Models and combines with OLS method.
result Outperforms LASSO method in forecasting models.
A tutorial on variational inference for high-dimensional models.
problem Approximating marginal likelihood and posterior in Bayesian models.
method Parametric approach to variational inference.
result Variational inference is now preferred for high-dimensional models and large datasets.
Paper develops robust methods for large-scale testing without tuning parameters.
problem Heavy-tailed data in high-dimensional settings.
method Revisits Hodges-Lehmann estimator for robust inference without tuning parameters.
result Develops confidence intervals and controls false discovery proportion.
This paper tackles the curse of dimensionality in semi-supervised learning using Laplacian regularization.
problem The curse of dimensionality in semi-supervised learning with Laplacian regularization.
method Statistical analysis and spectral filtering methods using kernel methods.
result The paper provides a method to overcome the curse of dimensionality in semi-supervised learning.
New method filters large networks from financial data to reveal key subnetworks.
problem Filtering large dimensional networks to isolate key constituents.
method Exploits spectral properties of high-dimensional data networks, tuning for sparsity and consistency.
result Shows method can interpolate between zero and maximal filtering, preserving spectral properties.
In this paper, we study randomized reduction methods, which reduce high-dimensional features into low-dimensional space by randomized methods (e.g., random projection, random hashing), for large-scale high-dimensional classification. Previous theoretical results on randomized reduction methods hinge on strong assumptio…
Paper proposes efficient methods for forecasting with large datasets.
problem Forecasting with large, high-dimensional economic data sets.
method Bayesian hierarchical priors, factor graphs, message passing algorithms, Generalized Approximate Message Passing (GAMP).
result The proposed methods outperform traditional approaches in forecasting U.S. price inflation.
Interpolating models can have heavy-tailed risk, leading to rare but severe errors.
problem Interpolating models' tail risk is poorly understood, affecting rare but impactful errors.
method Large-deviation methods to study the fragility of high-dimensional linear interpolators.
result Ridgeless regression exhibits heavy-tailed risk, while ridge-regularized estimators have better tail behavior.
We study the large scale geometry of the upper triangular subgroup of PSL(2,Z[1/n]), which arises naturally in a geometric context. We prove a quasi-isometry classification theorem and show that these groups are quasi-isometrically rigid with infinite dimensional quasi-isometry group. We generalize our results to a lar…
Two new scalable K-means initialization methods proposed for large-scale clustering.
problem Efficient initialization for large-scale clustering problems.
method Divide-and-conquer approach and random projection method for multiple lower-dimensional subspaces.
result The proposed methods outperform state-of-the-art in large-scale clustering tasks.
Modern large-scale datasets are frequently said to be high-dimensional. However, their data point clouds frequently possess structures, significantly decreasing their intrinsic dimensionality (ID) due to the presence of clusters, points being located close to low-dimensional varieties or fine-grained lumping. We test a…
We show that the Morse index of a closed minimal hypersurface in a four-dimensional Riemannian manifold cannot be bound in terms of the volume and the topological invariants of the hypersurface itself by presenting a method for constructing Riemannian metrics on S^4 that admit embedded minimal hyperspheres of uniformly…
The success of modern Artificial Intelligence (AI) technologies depends critically on the ability to learn non-linear functional dependencies from large, high dimensional data sets. Despite recent high-profile successes, empirical evidence indicates that the high predictive performance is often paired with low robustne…
Aggregates predictions from multiple regression models using random projections and kernel methods.
problem Combining predictions from multiple regression models to improve accuracy.
method Random projection of high-dimensional feature space, followed by kernel-based consensual aggregation.
result The aggregation scheme performs similarly to using the original high-dimensional features, with high probability.
Deep BSDE method for pricing and hedging complex financial portfolios.
problem Simultaneous pricing and delta-gamma hedging of large portfolios of multi-asset Bermudan options.
method Discretely reflected BSDEs, One Step Malliavin scheme, neural network regression Monte Carlo method.
result Efficient and accurate pricing and hedging strategies for high-dimensional portfolios.
ConMeZO speeds up zeroth-order optimization for large language models.
problem Slow convergence in high-dimensional parameter spaces of large language models.
method Adaptive directional sampling in a cone centered around a momentum estimate.
result Achieves the same convergence rate as MeZO but up to 2X faster.
Multi-label classification has received considerable interest in recent years. Multi-label classifiers have to address many problems including: handling large-scale datasets with many instances and a large set of labels, compensating missing label assignments in the training set, considering correlations between labels…
Large neural networks learn low-dimensional representations that balance complexity and regularity.
problem Understanding the tradeoff between low-dimensional representations and complexity in deep neural networks.
method Computed finite depth corrections to reveal a measure of regularity that bounds the pseudo-determinant of the Jacobian.
result Proved the conjectured bottleneck structure in learned features as network depth increases, showing almost all hidden representations are approximately low-dimensional and weight matrices have singular values close to 1.
Modern methods for data visualization via dimensionality reduction, such as t-SNE, usually have performance issues that prohibit their application to large amounts of high-dimensional data. In this work, we propose NCVis -- a high-performance dimensionality reduction method built on a sound statistical basis of noise c…
A new, fast kernel test for large data.
problem Efficient kernel two-sample tests for high-dimensional, large-scale data.
method A new kernel-based test that is computationally efficient and robust to high dimensions.
result The new test performs well across various alternatives and dimensions.
Study high-dimensional Bayesian linear regression using variational inference.
problem High-dimensional Bayesian linear regression with product priors.
method Non-linear large deviations theory and variational inference.
result Unique optimizer in variational problem governs posterior distribution under separation condition.
TSRGA scales multivariate linear regression for feature-distributed data.
problem Multivariate linear regression for feature-distributed data with high dimensions and many computing nodes.
method Two-stage relaxed greedy algorithm (TSRGA) for multivariate linear regression.
result TSRGA is highly scalable and can yield low-rank coefficient estimates.
Study shows how 3+1D cosmologies can evolve to de Sitter space under certain conditions.
problem Understanding the evolution of 3+1D cosmologies with specific symmetry constraints.
method Mean Curvature Flow methods applied to cosmologies with positive cosmological constant and specific symmetry groups.
result Asymptotically, 3+1D cosmologies evolve to de Sitter space under certain conditions.
We introduce a novel systematic construction for integrable (3+1)-dimensional dispersionless systems using nonisospectral Lax pairs that involve contact vector fields. In particular, we present new large classes of (3+1)-dimensional integrable dispersionless systems associated to the Lax pairs which are polynomial and …
Improved MTL-LSSVM for better multi-task learning performance.
problem Improving multi-task learning performance in high-dimensional data.
method Large dimensional analysis of Least Square Support Vector Machine (LSSVM) for MTL.
result Standard MTL-LSSVM is suboptimal and can lead to negative transfer, but can be corrected.
We introduce a new class of possibly noncompact n-dimensional manifolds without boundary associated to finite data which we call topological automata. This class is large enough to contain many interesting examples of open 2-dimensional and 3-dimensional manifolds of interest to low-dimensional topologists. Our main re…