We adapted the Covertype data set for unsupervised learning.
problem Lack of suitable unsupervised learning data sets.
method Transformed the Covertype data set into the Wilderness Area data set.
result The Wilderness Area data set is more suitable for unsupervised learning.
Modern machine learning systems such as image classifiers rely heavily on large scale data sets for training. Such data sets are costly to create, thus in practice a small number of freely available, open source data sets are widely used. We suggest that examining the geo-diversity of open data sets is critical before …
Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…
DALES offers a large annotated aerial LiDAR dataset for 3D deep learning.
problem Lack of large-scale annotated aerial LiDAR datasets for deep learning.
method Collection and annotation of over half a billion hand-labeled points from an ALS scanner.
result DALES is the most extensive publicly available ALS data set with improved resolution and coverage.
New IPC data set for graph learning tasks.
problem Benchmarking graph-based machine learning methods.
method Compilation from International Planning Competitions (IPC).
result Distinctly different characteristics from popular benchmarks.
Study on generalization for data-dependent hypothesis sets.
problem Understanding generalization in hypothesis sets dependent on data.
method Learning guarantee based on transductive Rademacher complexity and hypothesis set stability.
result Generalization bound for data-dependent hypothesis sets.
A new algorithm improves credit scoring accuracy for imbalanced data.
problem Poor classification of minority class in credit scoring data sets.
method Weighted-Hybrid-Sampling-Boost (WHSBoost) algorithm with balanced data sampling.
result WHSBoost outperforms other methods in credit scoring accuracy.
Model for inferring multivariate functions from areal data.
problem Inferring multivariate functions from areal data with varying granularities.
method Probabilistic model using Gaussian processes with spatial aggregation.
result Model effectively estimates spatial correlations and dependencies between areal data sets.
Paper improves conformal prediction for imprecise training data.
problem Applying conformal prediction to partially labeled data.
method Generalizes conformal prediction for set-valued training and calibration data.
result Validates the proposed method and shows it outperforms baselines.
Paper proves rigidity of initial data sets with boundary and capillary MOTS.
problem Rigidity of initial data sets with boundary and capillary MOTS.
method Estimates area of MOTS, proves rigidity for 3D, extends to high dimensions using Yamabe constant.
result Rigidity results for initial data sets with boundary and capillary MOTS.
Paper improves Bayesian network learning from related data sets.
problem Learning from heterogeneous data sets with different probabilistic structures.
method Mixed-effects models to pool information across related data sets.
result Mixed-effects models outperform traditional methods in accuracy.
Graph data sets often contain isomorphism bias, artificially inflating model performance.
problem Isomorphism bias in graph data sets causing inflated model performance.
method Analysis of 54 graph data sets, recommendations for model setup, open sourcing new data sets.
result Graph data sets commonly contain isomorphism bias, artificially inflating model performance.
Proves density and mass theorems for specific initial data sets.
problem Initial data sets with boundary in spacetime.
method Harmonic asymptotics and dominant energy condition.
result Spacetime positive mass theorem for initial data sets with apparent horizon boundary.
The two-sample hypothesis testing problem is studied for the challenging scenario of high dimensional data sets with small sample sizes. We show that the two-sample hypothesis testing problem can be posed as a one-class set classification problem. In the set classification problem the goal is to classify a set of data …
When working with asymptotically hyperbolic initial data sets for general relativity it is convenient to assume certain simplifying properties. We prove that the subset of initial data sets with such properties is dense in the set of physically reasonable asymptotically hyperbolic initial data sets. More specifically, …
We propose a probabilistic model for refining coarse-grained spatial data by utilizing auxiliary spatial data sets. Existing methods require that the spatial granularities of the auxiliary data sets are the same as the desired granularity of target data. The proposed model can effectively make use of auxiliary data set…
Constructs constant spacetime mean curvature surfaces for hyperboloidal initial data sets.
problem Creating a foliation of constant spacetime mean curvature surfaces for asymptotically hyperboloidal initial data sets.
method Long time limit of volume preserving spacetime mean curvature flow starting from a constant mean curvature foliation.
result Obtains a foliation of constant spacetime mean curvature surfaces as the long time limit.
This paper speeds up SVC clustering by compressing data while preserving key properties.
problem Efficiently clustering large-scale real-world data sets.
method Spectrum-preserving data compression for fast support vector clustering.
result Achieved 100X and 115X speedups on real-world data sets while maintaining clustering quality.
Proves critical points of ADM mass correspond to specific initial data sets.
problem Finding initial data sets with fixed Bartnik boundary data.
method Proves existence of critical points on a Banach manifold.
result Critical points of ADM mass correspond to initial data sets with generalized Killing vector fields.
A conceptually simple way to classify images is to directly compare test-set data and training-set data. The accuracy of this approach is limited by the method of comparison used, and by the extent to which the training-set data cover configuration space. Here we show that this coverage can be substantially increased u…
The potential benefits of applying machine learning methods to -omics data are becoming increasingly apparent, especially in clinical settings. However, the unique characteristics of these data are not always well suited to machine learning techniques. These data are often generated across different technologies in dif…
Set Flow models sets of data, learns dependencies, and achieves state-of-the-art likelihoods.
problem Modeling and sampling from finite, potentially high-dimensional, non-i.i.d. sets of data.
method Extends RealNVPs to handle finite sets, maintaining invertibility and exact log-likelihood evaluation.
result Achieves state-of-the-art likelihoods on 3D point clouds.
Proposes ESCA model to analyze mixed data types in multiple sets of measurements.
problem Separating common and distinct information in mixed data types from multiple sources.
method Exponential Family Simultaneous Component Analysis (ESCA) model with structured sparse loading matrix.
result The proposed method effectively disentangles global, local common and distinct information.
Deep learning and set theory improve prediction accuracy regardless of data relevance.
problem Improving prediction accuracy with limited relevant training data.
method Deep learning and set theory applied to large labeled training data.
result Exceptional prediction results achieved with irrelevant training data.
Dimensionality-reduction methods are a fundamental tool in the analysis of large data sets. These algorithms work on the assumption that the "intrinsic dimension" of the data is generally much smaller than the ambient dimension in which it is collected. Alongside their usual purpose of mapping data into a smaller dimen…
Divide-and-conquer method splits large data sets for efficient analysis.
problem Handling large data sets that exceed computational limits.
method Split data into smaller sets, analyze each separately, then combine results.
result Combined results provide statistical inference similar to analyzing entire data set.
Proposes a new model for handling missing data.
problem Nonignorable missingness in data.
method Variational autoencoder architecture with pattern-set mixtures.
result Achieves state-of-the-art imputation performance.
Proves principles and estimates for initial data sets in Einstein equations.
problem Understanding initial data sets in Einstein equations.
method Spinorial Callias operator approach.
result Proves long neck principle and width estimates.
Smooth dec initial data sets may not extend to smooth spacetimes.
problem Whether every dec initial data set can be extended to a smooth spacetime.
method Examined the converse of the dominant energy condition for initial data sets and spacelike hypersurfaces.
result Not all dec initial data sets can be extended to smooth spacetimes.
GANs generate training data for machine learning tasks.
problem Imbalanced data sets and sensitive information.
method Generative Adversarial Networks (GANs) to create artificial training data.
result A Decision Tree classifier trained on GAN-generated data achieved similar or better accuracy and recall than on original data.
New method improves OSSL by learning from all unlabeled data.
problem Handling open-set semi-supervised learning with unknown classes.
method Self-supervision and energy-based score for all unlabeled data.
result State-of-the-art results on benchmark problems.
The age of big data has produced data sets that are computationally expensive to analyze and store. Algorithmic leveraging proposes that we sample observations from the original data set to generate a representative data set and then perform analysis on the representative data set. In this paper, we present efficient a…
Expands small recommendation datasets to industrial scale.
problem Disconnection between academic and industrial data scales.
method Randomized fractal expansions using Kronecker Graph Theory.
result Generated synthetic data sets with 1.2B ratings, 2.2M users, and 855K items.
This paper addresses GE estimation in non-standard settings using various resampling methods.
problem Biased GE estimates in non-standard settings like clustered data and concept drift.
method Tailored resampling methods for clustered, spatial, unequal sampling, concept drift, and hierarchically structured outcomes.
result Standard resampling methods often yield biased GE estimates in non-standard settings.
Enhances classification accuracy on low data sets using synthetic data.
problem Low sample size in data augmentation.
method Variational Autoencoder and manifold sampling.
result Significant improvement in classification accuracy (e.g., 88.6% vs 80.7%).
This paper finds a linear relationship between t-SNE perplexity and data set size.
problem Choosing the right perplexity for t-SNE embeddings.
method Analyzed the relationship between perplexity and data set size.
result Embeddings remain structurally consistent when perplexity is adjusted accordingly.
PAC-Bayesian theory applied to data-dependent hypothesis sets yields uniform generalization bounds.
problem Proving uniform generalization bounds for data-dependent hypothesis sets.
method Applying PAC-Bayesian framework on 'random sets' and considering data-dependent hypothesis sets.
result Data-dependent uniform generalization bounds are proven, providing tighter and unified results.
New PDE systems generalize Hawking mass monotonicity.
problem Generalizing Hawking mass monotonicity to initial data sets.
method Introduced new systems of PDE on initial data sets (M,g,k). result Generalized Geroch's monotonicity formula to initial data sets.
MAGIC method optimally estimates model predictions changes.
problem Estimating how training data affects model predictions in large-scale settings.
method Combines classical methods and recent advances in metadifferentiation.
result MAGIC method nearly optimally estimates model predictions changes.
Background: High-throughput proteomics techniques, such as mass spectrometry (MS)-based approaches, produce very high-dimensional data-sets. In a clinical setting one is often interested in how mass spectra differ between patients of different classes, for example spectra from healthy patients vs. spectra from patients…
Paper proves new inequalities for Einstein-Maxwell data sets.
problem Establishing area-charge inequalities for Einstein-Maxwell initial data sets.
method Applying Gromov's μ-bubble technique in a new geometric context.
result Novel rigidity theorems for noncompact Einstein-Maxwell data sets.
Study finds rigid properties of boundary-free hypersurfaces in specific data sets.
problem Rigidity of free boundary hypersurfaces in initial data sets with boundary.
method Extending local splitting theorems and applying results on free boundary MOTS.
result Rigidity results for compact free boundary hypersurfaces in initial data sets with boundary.
Estimates bandwidth for CMC initial data sets.
problem Estimating bandwidth for constant mean curvature (CMC) initial data sets.
method Three independent proofs: stability of null expansion, spacetime harmonic function perturbation, Dirac operator.
result Generalized Gromov's band width estimate to CMC initial data sets.
Apricot selects subsets from large data sets efficiently using submodular optimization.
problem Efficiently selecting representative subsets from large data sets.
method Submodular optimization with efficient greedy algorithm.
result Strong theoretical guarantees on the quality of selected subsets.
We consider the Einstein-Maxwell-fluid constraint equations, and make use of the conformal method to construct and parametrize constant-mean-curvature hyperboloidal initial data sets that satisfy the shear-free condition. This condition is known to be necessary in order that a spacetime development admit a regular conf…
Global properties of maximal future Cauchy developments of stationary, m-dimensional asymptotically flat initial data with an outer trapped boundary are analyzed. We prove that, whenever the matter model is well posed and satisfies the null energy condition, the future Cauchy development of the data is a black hole spa…
Publish a core-set of data to protect against adversarial use.
problem Protecting datasets from adversarial use.
method Construct a fair core-set for linear and neural models.
result Core-sets improve primary task performance while hindering unwanted tasks.
TSVQR captures heterogeneous and asymmetric data using quantile regression.
problem Capturing heterogeneous and asymmetric information in modern data.
method Twin Support Vector Quantile Regression (TSVQR) with two nonparallel planes for quantile levels.
result TSVQR outperforms previous methods in capturing and learning from data.