Anomaly detection for high-dimensional data using large deviations principle.
problem Challenges in anomaly detection for high-dimensional data.
method Large Deviations Anomaly Detection (LAD) algorithm.
result Outperforms state-of-the-art methods on high-dimensional data sets.
Paper proposes data quality measures for large-scale high-dimensional data.
problem Lack of practical data quality measures for large-scale high-dimensional data.
method Proposes two data quality measures: class separability and in-class variability. Efficient algorithms based on random projections and bootstrapping are provided.
result Efficient algorithms for computing data quality measures on large-scale high-dimensional data.
NCVis speeds up data visualization for large datasets.
problem Performance issues in t-SNE for large datasets.
method Noise contrastive estimation for scalable visualization.
result NCVis outperforms state-of-the-art techniques in speed and quality.
FibeRed reduces complex data dimensions while preserving topology.
problem Hard embedding of topologically complex datasets in low-dimensional Euclidean space.
method Modeling datasets with vector bundles, reducing fibers while preserving topology.
result FibeRed learns topologically faithful embeddings in lower dimensions than existing methods.
ivhd tool efficiently visualizes large high-dimensional data.
problem Interactive visualization of large high-dimensional datasets.
method Minimal memory and computational time requirements for nearest neighbor visualization.
result Significantly reduces computational and memory loads.
FSL-Net detects and localizes feature shifts in large, high-dimensional datasets.
problem Feature shifts between data sources lead to erroneous features in various applications.
method FSL-Net is a neural network trained on multiple datasets to localize feature shifts.
result FSL-Net accurately localizes feature shifts from unseen datasets without re-training.
We propose a feature selection method that finds non-redundant features from a large and high-dimensional data in nonlinear way. Specifically, we propose a nonlinear extension of the non-negative least-angle regression (LARS) called N3LARS, where the similarity between input and output is measured through the norm…
Paper introduces a method to process medical images efficiently.
problem High computational cost in processing large medical image data.
method Framelet-pooling aided deep learning method to reduce complexity.
result Significant reduction in computational costs with comparable performance.
Analyzes large-margin classifiers under high-dimensional data.
problem Selecting the best classifier among various margin-based methods.
method Investigates asymptotic performance of large-margin classifiers under two component mixture models.
result Analytical results closely match with Monte Carlo simulations.
A new, fast kernel test for large data.
problem Efficient kernel two-sample tests for high-dimensional, large-scale data.
method A new kernel-based test that is computationally efficient and robust to high dimensions.
result The new test performs well across various alternatives and dimensions.
New method filters large networks from financial data to reveal key subnetworks.
problem Filtering large dimensional networks to isolate key constituents.
method Exploits spectral properties of high-dimensional data networks, tuning for sparsity and consistency.
result Shows method can interpolate between zero and maximal filtering, preserving spectral properties.
TSRGA scales multivariate linear regression for feature-distributed data.
problem Multivariate linear regression for feature-distributed data with high dimensions and many computing nodes.
method Two-stage relaxed greedy algorithm (TSRGA) for multivariate linear regression.
result TSRGA is highly scalable and can yield low-rank coefficient estimates.
FLRML efficiently handles large-scale, high-dimensional data for metric learning.
problem Metric learning for large-scale, high-dimensional datasets is computationally expensive and memory-intensive.
method FLRML reformulates low-rank metric learning as an unconstrained optimization problem on the Stiefel manifold, enabling efficient mini-batch processing.
result FLRML achieves high accuracy with significantly reduced computational and memory costs.
A novel outlier detection method for high-dimensional data.
problem Challenges in outlier detection in high-dimensional data.
method Principal component analysis and kernel density estimation.
result The proposed method outperforms benchmark methods in F1-score and execution time. This paper tackles the curse of dimensionality in semi-supervised learning using Laplacian regularization.
problem The curse of dimensionality in semi-supervised learning with Laplacian regularization.
method Statistical analysis and spectral filtering methods using kernel methods.
result The paper provides a method to overcome the curse of dimensionality in semi-supervised learning.
Paper analyzes CKRR for large data, showing risks converge to deterministic values.
problem Analyzing risks of kernel ridge regression with large data.
method Large dimensional analysis using centered kernels and random matrix theory.
result Empirical and prediction risks converge to deterministic values under specific conditions.
Supervised dimensionality reduction strategies have been of great interest. However, current supervised dimensionality reduction approaches are difficult to scale for situations characterized by large datasets given the high computational complexities associated with such methods. While stochastic approximation strateg…
Quantum-inspired CCA improves correlation analysis for high-dimensional data.
problem High-dimensional data limits conventional CCA due to time complexity.
method Developed a quantum-inspired CCA (qiCCA) with logarithmic time complexity.
result qiCCA extracts more correlations than linear CCA and is comparable to deep and kernel CCA.
The immense amount of daily generated and communicated data presents unique challenges in their processing. Clustering, the grouping of data without the presence of ground-truth labels, is an important tool for drawing inferences from data. Subspace clustering (SC) is a relatively recent method that is able to successf…
Proposes MamBO for efficient high-dimensional large-scale optimization.
problem High-dimensional and large-scale optimization problems in machine learning and simulation.
method Combines subsampling and subspace embeddings with model aggregation to address uncertainty in surrogate models.
result Improves robustness of Bayesian optimization algorithm and achieves superior performance.
Suppose that two large, multi-dimensional data sets are each noisy measurements of the same underlying random process, and principle components analysis is performed separately on the data sets to reduce their dimensionality. In some circumstances it may happen that the two lower-dimensional data sets have an inordinat…
In this paper, we study randomized reduction methods, which reduce high-dimensional features into low-dimensional space by randomized methods (e.g., random projection, random hashing), for large-scale high-dimensional classification. Previous theoretical results on randomized reduction methods hinge on strong assumptio…
For low-dimensional data sets with a large amount of data points, standard kernel methods are usually not feasible for regression anymore. Besides simple linear models or involved heuristic deep learning models, grid-based discretizations of larger (kernel) model classes lead to algorithms, which naturally scale linear…
Estimates intrinsic dimensionality of biological datasets using Fisher separability.
problem High-dimensional biological datasets with complex structures.
method Fisher separability analysis to estimate intrinsic dimensionality.
result The method performs competitively with state-of-the-art measures and is robust to noise.
This paper proposes a new method for learning covers of geometric datasets to improve topological inference and visualization.
problem Improving topological inference and visualization of large-scale geometric datasets.
method Proposes a method for learning topologically-faithful covers of geometric datasets using optimization.
result Simplicial complexes obtained from learned covers outperform standard methods in terms of size and representation of large-scale topology.
CCP clusters correlated features and projects them to 1D for efficient dimensionality reduction.
problem Efficiency in handling large datasets with high intrinsic dimensions.
method CCP partitions features into correlated clusters and projects them to 1D based on sample correlations.
result CCP achieves efficient dimensionality reduction without matrix diagonalization.
Paper develops robust methods for large-scale testing without tuning parameters.
problem Heavy-tailed data in high-dimensional settings.
method Revisits Hodges-Lehmann estimator for robust inference without tuning parameters.
result Develops confidence intervals and controls false discovery proportion.
Proposes two-stage robust and sparse distributed inference for large-scale data.
problem Statistical inference in large-scale, high-dimensional, and outlier-contaminated data.
method Two-stage approach: model selection with robust Lasso, fusion of local selections, and bootstrap methods for inference.
result Robust and computationally efficient inference procedures for variable selection, confidence intervals, and standard deviation approximations.
Detects anomalies and locates their causes in large, high-dimensional data.
problem Locating hidden issues in complex systems with high-dimensional data.
method Copula-based model for multivariate probability distributions.
result Can identify and localize anomalies in large, high-dimensional data.
New neural network method simplifies high-dimensional data.
problem Scalability issues in nonlinear sufficient dimension reduction.
method Stochastic neural network with adaptive gradient algorithm.
result Proposed method outperforms existing methods on large-scale data.
The success of modern Artificial Intelligence (AI) technologies depends critically on the ability to learn non-linear functional dependencies from large, high dimensional data sets. Despite recent high-profile successes, empirical evidence indicates that the high predictive performance is often paired with low robustne…
Improved MTL-LSSVM for better multi-task learning performance.
problem Improving multi-task learning performance in high-dimensional data.
method Large dimensional analysis of Least Square Support Vector Machine (LSSVM) for MTL.
result Standard MTL-LSSVM is suboptimal and can lead to negative transfer, but can be corrected.
Dimensionality-reduction methods are a fundamental tool in the analysis of large data sets. These algorithms work on the assumption that the "intrinsic dimension" of the data is generally much smaller than the ambient dimension in which it is collected. Alongside their usual purpose of mapping data into a smaller dimen…
Subspace clustering aims to find groups of similar objects (clusters) that exist in lower dimensional subspaces from a high dimensional dataset. It has a wide range of applications, such as analysing high dimensional sensor data or DNA sequences. However, existing algorithms have limitations in finding clusters in non-…
FsNet selects features for high-dimensional biological data efficiently.
problem Efficient feature selection for high-dimensional biological data.
method FsNet combines selection and reconstruction layers with tiny networks for weight prediction.
result FsNet outperforms standard DNNs on high-dimensional biological datasets.
Proposes a method to visualize finer cluster structures in high-dimensional data.
problem Visualization of high-dimensional data with complex cluster structures.
method Introduces a generalized sigmoid function with a parameter b to adjust the tail heaviness for better visualization.
result The method can generate visualization results comparable to UMAP, revealing finer cluster structures.
Paper solves globally optimal k-means for low dimensional data.
problem Finding globally optimal k-means solutions for low dimensional data.
method Formulates as a concave assignment problem, iteratively solving small concave and large linear programming problems.
result Solves k-means to global optimality for large data sets with several clusters.
Optimized online learning with kernels for large-scale adversarial data.
problem Efficient online learning for large-scale, potentially adversarial datasets.
method Online variations of kernel Ridge regression using approximated basis functions.
result Optimal regret for a wide range of kernels with low per-round complexity.
A new kernel, Isolation Kernel, simplifies large scale online kernel learning without sacrificing accuracy.
problem Building efficient and scalable kernel-based models from large datasets with high accuracy.
method Introducing Isolation Kernel, which creates an exact, sparse, and finite-dimensional feature map of a kernel, allowing for efficient large scale online kernel learning without accuracy loss.
result Large scale online kernel learning can be achieved efficiently and accurately using Isolation Kernel.
A powerful approach for understanding neural population dynamics is to extract low-dimensional trajectories from population recordings using dimensionality reduction methods. Current approaches for dimensionality reduction on neural data are limited to single population recordings, and can not identify dynamics embedde…
Analyzes SGD dynamics on multi-class problems with exact expressions.
problem Analyzing SGD dynamics on multi-class problems.
method Developed a framework for analyzing training and learning rate dynamics using exact expressions.
result Exact expressions for risk and overlap with true signal in terms of ODEs.
This article carries out a large dimensional analysis of standard regularized discriminant analysis classifiers designed on the assumption that data arise from a Gaussian mixture model with different means and covariances. The analysis relies on fundamental results from random matrix theory (RMT) when both the number o…
LR-GLM speeds Bayesian GLM inference for high-dimensional data.
problem Bayesian inference in high-dimensional GLMs is computationally expensive.
method Low-rank data approximation to reduce computational time and memory costs.
result LR-GLM provides a full Bayesian posterior approximation with reduced computational time.
Sparse Convex Biclustering improves accuracy and robustness in high-dimensional datasets.
problem Challenges in clustering rows and columns of large-scale datasets due to noise and computational complexity.
method Sparse Convex Biclustering (SpaCoBi) using convex optimization and stability-based tuning.
result Significantly outperforms state-of-the-art methods in accuracy for high-dimensional datasets.
Corrected whitening restores orthogonality in high-dimensional spherical Gaussian mixtures.
problem In high-dimensional data, standard whitening fails to preserve orthogonality of mixture means.
method Derived exact limits for whitened means dot products using random matrix theory, constructed a corrected whitening matrix.
result Corrected whitening allows for improved estimation of spherical Gaussian mixtures in the large-dimensional regime.
Develops efficient algorithms for data science, tackling the curse of dimensionality.
problem Tackles the curse of dimensionality in large datasets.
method Focuses on feature extraction techniques and meta-heuristic algorithms, including evolutionary algorithms.
result Evolutionary algorithms are effective in solving optimization problems with a curse of dimensionality.
In this article, a large dimensional performance analysis of kernel least squares support vector machines (LS-SVMs) is provided under the assumption of a two-class Gaussian mixture model for the input data. Building upon recent advances in random matrix theory, we show, when the dimension of data p and their number $…
Study on optimal rate of kernel regression for large-dimensional data.
problem Characterizing the upper and lower bounds of kernel regression for large-dimensional data.
method Using Mendelson complexity and metric entropy, the study characterizes the upper and lower bounds of kernel regression for large-dimensional data.
result The minimax rate of the excess risk of kernel regression is \( n^{-1/2} \) for \( n \asymp d^γ \) with \( γ=2, 4, 6, 8, \cdots \).