Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

4188371,2551,673 · Jun 202019922001200920182026
48 results for Large data sets

DALES offers a large annotated aerial LiDAR dataset for 3D deep learning.

problem Lack of large-scale annotated aerial LiDAR datasets for deep learning.
method Collection and annotation of over half a billion hand-labeled points from an ALS scanner.
result DALES is the most extensive publicly available ALS data set with improved resolution and coverage.

A new method reduces the number of features needed for kernel approximation from cubic to logarithmic.

problem Large datasets make kernel methods computationally expensive and impractical.
method Combines random feature maps with data-dependent feature selection to achieve Nystrom-like performance with fewer features.
result Achieves small kernel matrix approximation error and better test set accuracy with fewer features than state-of-the-art methods.

There has been a surge in the number of large and flat data sets - data sets containing a large number of features and a relatively small number of observations - due to the growing ability to collect and store information in medical research and other fields. Hierarchical clustering is a widely used clustering tool. I…

2014-09-02abs ↗pdf ↗

Paper analyzes Nyström and column-sampling methods for PCA of large data sets.

problem Computational and storage challenges in PCA for large data sets.
method Nyström and column-sampling methods for approximating PCA of large matrices.
result Comparison of methods and trade-off between accuracy and computational complexity.

New method estimates multivariate Gaussian fields using sparse precision matrix.

problem Estimating covariance matrices for large multivariate Gaussian fields.
method Sparse Precision Matrix Selection (SPS) algorithm for multivariate GRFs.
result Theoretical rates of convergence for estimated covariance and parameters validated.

New statistic κκ-profile helps monitor weather, soundscapes, and dynamical systems.

problem Monitoring intrinsic dimensionality of large data sets.
method Optimization problem to find κκ-profile, which is the norm of the shortest projected secant.
result The κκ-profile provides a useful statistic for understanding and monitoring large data sets.

Variational inference algorithms have proven successful for Bayesian analysis in large data settings, with recent advances using stochastic variational inference (SVI). However, such methods have largely been studied in independent or exchangeable data settings. We develop an SVI algorithm to learn the parameters of hi…

2014-11-06abs ↗pdf ↗

Study finds open data sets favor Western locales, impacting classifier performance.

problem Impact of biased open data sets on classifier performance in the developing world.
method Analysis of two large, publicly available image data sets and classifiers trained on them.
result Open data sets exhibit a bias towards Western locales, affecting classifier performance.

New method compresses large sample data for faster discriminant analysis.

problem Large sample sizes in discriminant analysis increase computational burden.
method Proposes a new compression approach for reducing training samples.
result Significant computational gains and superior predictive ability compared to random sub-sampling.

Solving different types of optimization models (including parameters fitting) for support vector machines on large-scale training data is often an expensive computational task. This paper proposes a multilevel algorithmic framework that scales efficiently to very large data sets. Instead of solving the whole training s…

2014-10-13abs ↗pdf ↗

Hierarchical Softmax approximates class probabilities for large datasets efficiently.

problem Computational inefficiency of Softmax for large-scale classification tasks.
method Used Hierarchical Softmax to approximate class probabilities efficiently.
result Hierarchical Softmax performance degrades as the number of classes increases.

A new method for efficient Nystrom approximation for large datasets.

problem Generating low-rank approximations of kernel matrices for large-scale machine learning problems.
method Randomized K-means clustering on low-dimensional random projections of data.
result Significant savings in computational efficiency for high-dimensional data.

We give examples of asymptotically flat three-manifolds (M,g)(M,g) which admit arbitrarily large constant mean curvature spheres that are far away from the center of the manifold. This resolves a question raised by G. Huisken and S.-T. Yau in 1996. On the other hand, we show that such surfaces cannot exist when (M,g)(M,g) ha…

2013-03-14abs ↗pdf ↗

A new learning method uses data to learn from large model sets.

problem Learning with large sets of candidate models where uniform convergence is hard.
method Data-dependent learning that incorporates empirical data less reliant on prior assumptions.
result Demonstrates improved generalization in various learning assumptions.

Constructs foliations of critical surfaces for Hawking energy in asymptotically flat initial data sets.

problem Positivity and rigidity of Hawking quasi-local energy in asymptotically flat spacetimes.
method Lyapunov-Schmidt reduction within a Willmore-foliation framework.
result Existence and uniqueness of foliations by Hawking surfaces, positivity and large-sphere limit of Hawking energy.

Generates synthetic laparoscopic images for training deep neural networks.

problem Lack of large labeled data sets for laparoscopic image processing.
method Unpaired image-to-image translation to generate realistic synthetic images.
result Synthetic data set improves liver segmentation performance without manual labeling.

IVF k-means algorithm improves performance on large sparse data sets.

problem Efficiently clustering large-scale sparse data sets with numerous classes.
method Sparse data representation and inverted-file structure for high-speed and low-memory clustering.
result IVF achieves better performance than other algorithms on real document data sets.

Improved GP models for scalable large data sets.

problem Computational infeasibility of Gaussian process models for large datasets.
method Composite likelihood approach with recursive computation and hyper-parameter learning.
result The derived composite GP model provides accurate predictions and hyper-parameter learning.

AgEBO-Tabular combines NAS and hyperparameter tuning for fast, high-performing tabular models.

problem Developing high-performing predictive models for large tabular data sets is challenging.
method Combines aging evolution NAS and asynchronous Bayesian optimization for hyperparameter tuning in data-parallel training.
result Automatically discovered neural network models outperform state-of-the-art AutoML ensembles in inference speed by two orders of magnitude.

Wasserstein coresets improve data sparsification for continuous distributions.

problem Efficiently summarize large continuous data distributions for Bayesian inference.
method Introduce Wasserstein measure coresets, minimizing Wasserstein distance via stochastic gradient descent.
result Wasserstein coresets enable online handling of large data streams and improve inference and clustering performance.

Method predicts computational reproducibility of large population studies data analysis pipelines.

problem Difficulty in evaluating reproducibility of large population studies due to computational and storage requirements.
method Formulated as collaborative filtering process with constraints on training set construction.
result One sampling method, 'Random File Numbers (Uniform)', predicts reproducibility with good accuracy.

Proposes a method to learn distance metrics from uncertain data.

problem Challenges of learning distance metrics from large-scale data with uncertainty.
method Margin preserving metric learning framework to learn distance metric and latent examples simultaneously.
result The learned metric is robust to uncertainty and preserves large margin for original data.

Suppose that two large, multi-dimensional data sets are each noisy measurements of the same underlying random process, and principle components analysis is performed separately on the data sets to reduce their dimensionality. In some circumstances it may happen that the two lower-dimensional data sets have an inordinat…

2013-01-09abs ↗pdf ↗

sparsebn learns large Bayesian networks from high-dimensional data.

problem Learning graphical models from large, high-dimensional datasets with interventions.
method Focuses on scalability and consistency in high-dimensional settings, learning causal networks from data.
result Achieves the goal of learning a causal network from data.

New method infers causal factors from large-scale data without full graph reconstruction.

problem Inferring causal variables from large-scale systems without full causal graph reconstruction.
method Supervised learning on simulated data using a neural network and subsampled-ensemble inference.
result Efficiently identifies causal relationships in large-scale gene regulatory networks.

Efficient multi-label classifier handles missing labels and large datasets.

problem Handling large-scale datasets with many instances and labels, missing label assignments, label correlations, and unlabeled data.
method Non-linear embedding of label vectors using a stochastic approach to predict tail labels, handling missing labels, and exploiting unlabeled data.
result Our method outperforms state-of-the-art multi-label classifiers in prediction performance and training time.