This work refines Cover's theory for binary classification on low-dimensional data.
problem The challenge of analyzing how low-dimensional data structures affect classification models.
method Refines Cover's function-counting theory to account for low-dimensional data structure.
result Derives dichotomy counts and analyzes the impact of data structure on classification models.
We propose a low-rank approach to learning a Mahalanobis metric from data. Inspired by the recent geometric mean metric learning (GMML) algorithm, we propose a low-rank variant of the algorithm. This allows to jointly learn a low-dimensional subspace where the data reside and the Mahalanobis metric that appropriately f…
Algorithm recovers multiple low-rank matrices from unlabeled data.
problem Learning mixtures of low-rank models from unlabelled data.
method Three-stage meta-algorithm that copes with non-convexity and noise.
result Near-optimal sample and computational complexities under Gaussian designs.
A guide to using low-pass graph filters for network data.
problem Understanding and processing graph data with low-pass filters.
method Definition and application of low-pass graph filters to graph data.
result Low-pass filters effectively retain lower frequency graph data contents.
Proposes tensor Q-rank for better tensor rank recovery in complex data.
problem Improving tensor rank recovery for complex data with low sampling rate.
method Introduces tensor Q-rank and two selection methods for Q \mathbf{Q} Q , proposing VMTQN and MOTQN models. result Demonstrates superior performance in tensor completion problems compared to TNN-based methods.
Generative models learn complex data from low-dimensional manifolds.
problem Theoretical justification for generative models on manifold structures.
method Prove statistical guarantees of generative networks under Wasserstein-1 loss, considering intrinsic dimensionality.
result Generative networks converge to zero at a fast rate depending on intrinsic dimensionality, not ambient data dimension.
New methods combine low and high-fidelity data for accurate surrogate modeling.
problem Challenges in surrogate modeling for high-dimensional outputs with limited training data.
method Projection-based multifidelity linear regression methods integrating low-fidelity and high-fidelity data.
result Multifidelity methods achieve up to 12% improvement in median accuracy compared to single-fidelity methods.
This work improves surrogate models using low-fidelity data to enhance accuracy and efficiency.
problem Limited training data makes high-fidelity models unreliable.
method Uses low-fidelity data to augment input space and condition high-fidelity models.
result Increased predictive accuracy and reduced computational cost compared to existing methods.
Enhances classification accuracy on low data sets using synthetic data.
problem Low sample size in data augmentation.
method Variational Autoencoder and manifold sampling.
result Significant improvement in classification accuracy (e.g., 88.6% vs 80.7%).
Real world data often exhibit low-dimensional geometric structures, and can be viewed as samples near a low-dimensional manifold. This paper studies nonparametric regression of Hölder functions on low-dimensional manifolds using deep ReLU networks. Suppose n n n training data are sampled from a Hölder function in $\mathc…
TempBalance boosts model performance with low data.
problem Low-data training and fine-tuning in model alignment.
method Inspired by HT-SR theory, TempBalance balances training quality across layers.
result TempBalance improves model performance as data decreases.
Adding uninformative labels improves tumor segmentation in low-data mammography.
problem Improving tumor segmentation in mammography with limited data.
method Used seemingly uninformative labels from non-expert annotators to turn a multi-label task into a multi-class problem.
result Performance gains in tumor segmentation are achieved in low-data settings with additional uninformative labels.
The input data features set for many data driven tasks is high-dimensional while the intrinsic dimension of the data is low. Data analysis methods aim to uncover the underlying low dimensional structure imposed by the low dimensional hidden parameters by utilizing distance metrics that consider the set of attributes as…
Diffusion models adapt to low-dimensional data regardless of coefficient choices.
problem Understanding how diffusion models adapt to low-dimensional data structures.
method Analysis of diffusion models with flexible coefficient choices.
result Proven that O ~ ( k / ε ) \widetilde{O}(k/\varepsilon) O ( k / ε ) iterations suffice for accurate sampling in total variation distance. Principal components analysis (PCA) is a well-known technique for approximating a tabular data set by a low rank matrix. Here, we extend the idea of PCA to handle arbitrary data sets consisting of numerical, Boolean, categorical, ordinal, and other data types. This framework encompasses many well known techniques in da…
Scattering networks maximize separation on low-dimensional data.
problem Maximizing separation capacity on low-dimensional datasets.
method Characterize and bound separation capacity for feature extractors, then apply to scattering networks with specific criteria.
result Design criteria for scattering networks to maximize separation on low-dimensional data.
We propose a method to infer stochastic low-rank RNNs from neural data.
problem Fitting low-rank RNNs to noisy, stochastic neural data.
method Variational sequential Monte Carlo methods for stochastic low-rank RNNs.
result Lower dimensional latent dynamics compared to state-of-the-art methods.
Paper studies Einstein vacuum equations with low regularity data.
problem Einstein vacuum equations with low regularity initial data.
method Combines Klainerman-Szeftel-Rodnianski curvature theorem, Czimek's extension procedure, and global elliptic estimates.
result Time of existence controlled by low regularity bounds on curvature in L 2 L^2 L 2 . New method improves tensor completion for weakly-dependent spatiotemporal data.
problem Improving tensor completion for weakly-dependent data on graphs.
method Introducing L 1 L_{1} L 1 -norm and Graph Laplacian penalties for low-rank tensor decomposition and completion. result Improved performance in metro passenger flow prediction.
Unified framework for statistical inference of low-rank tensors.
problem Statistical inference for tensors in high-dimensional data.
method Unified framework using debiasing and tangent space projection.
result Achieves asymptotic normality and minimax-optimal confidence intervals.
The paper reveals low-rank structure in neural network gradients, influenced by data and model parameters.
problem Investigating low-rank structure in gradients of neural networks under relaxed assumptions.
method Spiked data model, relaxation of isotropy assumptions, analysis of mean-field and neural-tangent-kernel scalings.
result Gradient of input weights is approximately low rank, dominated by two rank-one terms.
Matrices of (approximate) low rank are pervasive in data science, appearing in recommender systems, movie preferences, topic models, medical records, and genomics. While there is a vast literature on how to exploit low rank structure in these datasets, there is less attention on explaining why the low rank structure ap…
Bayesian model improves image completion accuracy by automatically learning low rank structure.
problem Improving image completion accuracy with limited data and avoiding overfitting.
method Developed a Bayesian low rank tensor ring model with multiplicative interaction and Student-T distribution for sparse core factors.
result The proposed method outperforms state-of-the-art image completion techniques, especially in recovery accuracy.
New spectral methods improve matrix estimation in RL with low-rank structure.
problem Estimating matrices with low-rank structure in reinforcement learning.
method Spectral-based matrix estimation approaches.
result Spectral methods efficiently recover singular subspaces and minimize entry-wise error.
New method detects if data points were used in training models with low cost and high power.
problem Detecting if a particular data point was used in training a model.
method Fine-grained modeling of null hypothesis in likelihood ratio tests, leveraging reference models and population data.
result RMIA has superior test power compared to prior methods, even at extremely low false positive rates.
A new BO framework reduces costs by using low-fidelity data.
problem Optimizing expensive experiments with low-fidelity data.
method Developed a multi-fidelity cost-aware Bayesian optimization framework.
result Significantly outperforms state-of-the-art BO methods.
ScaledGD algorithm estimates low-rank tensors efficiently from corrupted data.
problem Estimating meaningful information from corrupted tensor data.
method Scaled gradient descent (ScaledGD) algorithm with tailored spectral initializations.
result ScaledGD achieves linear convergence at a constant rate independent of condition number.
The paper analyzes side effects of learning from low-dimensional data embedded in a Euclidean space.
problem Learning from data distributed in a linear subspace of high-dimensional space.
method Derives estimates on the variation of the learning function and studies regularization effects.
result Potential regularization effects associated with network depth and noise in codimension of data manifold.
This paper improves diffusion models for low-dimensional data.
problem Theoretical foundations of diffusion models are lacking for low-dimensional data.
method Score approximation, estimation, and distribution recovery of diffusion models on low-dimensional data.
result Sample complexity bounds for distribution estimation using diffusion models are provided.
A composite loss framework is proposed for low-rank modeling of data consisting of interesting and common values, such as excess zeros or missing values. The methodology is motivated by the generalized low-rank framework and the hurdle method which is commonly used to analyze zero-inflated counts. The model is demonstr…
UAPCA projects uncertain data to low dimensions using GMMs.
problem Uncertain multidimensional data not well described by normal distributions.
method Model data with Gaussian mixture models, derive UAPCA projection from general formulation.
result Low-dimensional projections better represent multidimensional distributions.
New method improves robust low-rank matrix completion for computer vision.
problem Robust low-rank matrix completion for partially observed data.
method Formulated as a nonsmooth Riemannian optimization problem over Grassmann manifold, solved with an alternating manifold proximal gradient continuation method.
result Demonstrated advantages over existing approaches in background extraction from surveillance videos.
Proposes a multi-fidelity machine learning strategy integrating low-fidelity deterministic and high-fidelity Bayesian models.
problem Addressing the accuracy-efficiency trade-off in machine learning with scarce high-fidelity data.
method Integrates a non-probabilistic regression model for low-fidelity with a Bayesian model for high-fidelity, trained in a staggered scheme.
result Achieves comparable performance in mean and uncertainty estimation with reduced training time and effective mitigation of overfitting.
Paper uses low-dimensional sensor data analysis for better fault detection.
problem Fault detection in critical equipment using multivariate, nonlinear sensor data.
method Exploits t-SNE and KPCA for nonlinear dimension reduction and anomaly detection.
result Low-dimensional representations improve interpretability and edge processing in IoT.
We tackle anomaly detection in sparse time series data.
problem Sparse time series with low signal-to-noise ratios and non-uniform performance.
method We introduce a novel generative procedure for benchmark datasets and demonstrate how anomaly score smoothing improves performance.
result Anomaly score smoothing consistently improves performance in low-count time series anomaly detection.
Multivariate binary data is becoming abundant in current biological research. Logistic principal component analysis (PCA) is one of the commonly used tools to explore the relationships inside a multivariate binary data set by exploiting the underlying low rank structure. We re-expressed the logistic PCA model based on …
Two methods using low-discrepancy points improve data compression for neural networks.
problem Efficiently compress large datasets for neural network training.
method Two methods based on low-discrepancy points: digital nets with averaging and clustering.
result Second method outperforms supercompress in compression error and neural network accuracy.
Paper improves tensor approximation for streaming data.
problem Challenges in finding accurate low-tubal-rank tensor approximations in streaming settings.
method Extends Frequent Directions for efficient low-tubal-rank tensor approximation.
result The new algorithm achieves arbitrarily small approximation error with linear sketch size growth.
This paper analyzes air pollution trends in Rwanda using low-cost sensors and machine learning.
problem Lack of reliable air pollution data in Rwanda due to high costs of equipment.
method Analysis of existing data and development of forecasting models using low-cost sensors and machine learning.
result Proposes forecasting models for air pollution data collected by low-cost sensors.
Novel low-rank neural decoder improves μ μ μ -ECoG neural decoding.
problem Challenging neural decoding from high-dimensional μ μ μ -ECoG data. method Low-rank structure in neural network decoder.
result Low-rank decoder outperforms standard PCA.
Paper solves globally optimal k-means for low dimensional data.
problem Finding globally optimal k-means solutions for low dimensional data.
method Formulates as a concave assignment problem, iteratively solving small concave and large linear programming problems.
result Solves k-means to global optimality for large data sets with several clusters.
New algorithms extract low-dimensional representations from sequential data, revealing insights into complex processes.
problem Challenges in extracting low-dimensional representations from sequential, high-dimensional, sparse, and noisy data.
method Developed new clustering algorithms based on Block Markov Chains theory, validated on real-world data.
result These algorithms can successfully extract low-dimensional representations from real-world sequential data, revealing insights into complex processes.
Deep learning speeds up whole heart MRI to 30 seconds.
problem Long acquisition times in whole heart MRI.
method Deep learning, specifically a 3D residual U-Net, to reconstruct high-resolution images from low-resolution data.
result Super-resolution images show better edge sharpness and fewer artefacts than low-resolution images.
Generic low-entropy hypersurfaces in 4-6D flow with only generic singularities.
problem Analyzing mean curvature flow of low-entropy hypersurfaces.
method Proving flow encounters only generic singularities for specific entropy conditions.
result Proves flow encounters only generic singularities for low-entropy initial data.
Improves ML efficiency for vast, rapidly growing data.
problem Low latency and cost in ML with distributed, growing data.
method Designs ML systems exploiting ML characteristics, data structures, and data distribution.
result Improves ML latency and cost by 1-2 orders of magnitude.
Deep ensembles don't necessarily improve calibration in low data regimes.
problem Calibration issues in deep learning models, especially in low data regimes.
method Examination of data-augmentation, ensembling, and post-processing calibration methods.
result Standard ensembling techniques can lead to less calibrated models in low data regimes.
Robust low-rank matrix estimation is a topic of increasing interest, with promising applications in a variety of fields, from computer vision to data mining and recommender systems. Recent theoretical results establish the ability of such data models to recover the true underlying low-rank matrix when a large portion o…
ConvResNets approximate Besov functions and classify on low-dimensional manifolds.
problem Lack of statistical theories for deep learning on high-dimensional data.
method Exploits low-dimensional geometric structures of real-world data sets using ConvResNets.
result ConvResNets can approximate Besov functions and learn classifiers with optimal excess risk.