New causal models perform poorly when evaluated on biased training sets.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Benchmark data sets are of vital importance in machine learning research, as indicated by the number of repositories that exist to make them publicly available. Although many of these are usable in the stream mining context as well, it is less obvious which data sets can be used to evaluate data stream clustering algor…
Modern machine learning systems such as image classifiers rely heavily on large scale data sets for training. Such data sets are costly to create, thus in practice a small number of freely available, open source data sets are widely used. We suggest that examining the geo-diversity of open data sets is critical before …
Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…
DALES offers a large annotated aerial LiDAR dataset for 3D deep learning.
A new algorithm improves credit scoring accuracy for imbalanced data.
We propose a probabilistic model for inferring the multivariate function from multiple areal data sets with various granularities. Here, the areal data are observed not at location points but at regions. Existing regression-based models can only utilize the sufficiently fine-grained auxiliary data sets on the same doma…
Benchmark data sets are an indispensable ingredient of the evaluation of graph-based machine learning methods. We release a new data set, compiled from International Planning Competitions (IPC), for benchmarking graph classification, regression, and related tasks. Apart from the graph construction (based on AI planning…
Paper improves conformal prediction for imprecise training data.
Paper proves rigidity of initial data sets with boundary and capillary MOTS.
Paper improves Bayesian network learning from related data sets.
Proves density and mass theorems for specific initial data sets.
When working with asymptotically hyperbolic initial data sets for general relativity it is convenient to assume certain simplifying properties. We prove that the subset of initial data sets with such properties is dense in the set of physically reasonable asymptotically hyperbolic initial data sets. More specifically, …
The two-sample hypothesis testing problem is studied for the challenging scenario of high dimensional data sets with small sample sizes. We show that the two-sample hypothesis testing problem can be posed as a one-class set classification problem. In the set classification problem the goal is to classify a set of data …
We propose a probabilistic model for refining coarse-grained spatial data by utilizing auxiliary spatial data sets. Existing methods require that the spatial granularities of the auxiliary data sets are the same as the desired granularity of target data. The proposed model can effectively make use of auxiliary data set…
Constructs constant spacetime mean curvature surfaces for hyperboloidal initial data sets.
This paper speeds up SVC clustering by compressing data while preserving key properties.
Proves critical points of ADM mass correspond to specific initial data sets.
A conceptually simple way to classify images is to directly compare test-set data and training-set data. The accuracy of this approach is limited by the method of comparison used, and by the extent to which the training-set data cover configuration space. Here we show that this coverage can be substantially increased u…
The potential benefits of applying machine learning methods to -omics data are becoming increasingly apparent, especially in clinical settings. However, the unique characteristics of these data are not always well suited to machine learning techniques. These data are often generated across different technologies in dif…
We present a study of generalization for data-dependent hypothesis sets. We give a general learning guarantee for data-dependent hypothesis sets based on a notion of transductive Rademacher complexity. Our main result is a generalization bound for data-dependent hypothesis sets expressed in terms of a notion of hypothe…
Dimensionality-reduction methods are a fundamental tool in the analysis of large data sets. These algorithms work on the assumption that the "intrinsic dimension" of the data is generally much smaller than the ambient dimension in which it is collected. Alongside their usual purpose of mapping data into a smaller dimen…
Divide-and-conquer method splits large data sets for efficient analysis.
In this paper we propose the use of Generative Adversarial Networks (GAN) to generate artificial training data for machine learning tasks. The generation of artificial training data can be extremely useful in situations such as imbalanced data sets, performing a role similar to SMOTE or ADASYN. It is also useful when t…
Multiple sets of measurements on the same objects obtained from different platforms may reflect partially complementary information of the studied system. The integrative analysis of such data sets not only provides us with the opportunity of a deeper understanding of the studied system, but also introduces some new st…
Proposes a new model for handling missing data.
Proves principles and estimates for initial data sets in Einstein equations.
Smooth dec initial data sets may not extend to smooth spacetimes.
New method improves OSSL by learning from all unlabeled data.
The age of big data has produced data sets that are computationally expensive to analyze and store. Algorithmic leveraging proposes that we sample observations from the original data set to generate a representative data set and then perform analysis on the representative data set. In this paper, we present efficient a…
Enhances classification accuracy on low data sets using synthetic data.
This paper finds a linear relationship between t-SNE perplexity and data set size.
This paper addresses GE estimation in non-standard settings using various resampling methods.
New PDE systems generalize Hawking mass monotonicity.
PAC-Bayesian theory applied to data-dependent hypothesis sets yields uniform generalization bounds.
MAGIC method optimally estimates model predictions changes.
Background: High-throughput proteomics techniques, such as mass spectrometry (MS)-based approaches, produce very high-dimensional data-sets. In a clinical setting one is often interested in how mass spectra differ between patients of different classes, for example spectra from healthy patients vs. spectra from patients…
Paper proves new inequalities for Einstein-Maxwell data sets.
Recommender system research suffers from a disconnect between the size of academic data sets and the scale of industrial production systems. In order to bridge that gap, we propose to generate large-scale user/item interaction data sets by expanding pre-existing public data sets. Our key contribution is a technique tha…
We present ChromAlignNet, a deep learning model for alignment of peaks in Gas Chromatography-Mass Spectrometry (GC-MS) data. In GC-MS data, a compound's retention time (RT) may not stay fixed across multiple chromatograms. To use GC-MS data for biomarker discovery requires alignment of identical analyte's RT from diffe…
Estimates bandwidth for CMC initial data sets.
Study finds rigid properties of boundary-free hypersurfaces in specific data sets.
In recent years there has been a rapid increase in classification methods on graph structured data. Both in graph kernels and graph neural networks, one of the implicit assumptions of successful state-of-the-art models was that incorporating graph isomorphism features into the architecture leads to better empirical per…
We consider the Einstein-Maxwell-fluid constraint equations, and make use of the conformal method to construct and parametrize constant-mean-curvature hyperboloidal initial data sets that satisfy the shear-free condition. This condition is known to be necessary in order that a spacetime development admit a regular conf…
Global properties of maximal future Cauchy developments of stationary, m-dimensional asymptotically flat initial data with an outer trapped boundary are analyzed. We prove that, whenever the matter model is well posed and satisfies the null energy condition, the future Cauchy development of the data is a black hole spa…
Study proposes initial data sets for solving gravitational equations, proving energy estimates.
Given only information in the form of similarity triplets "Object A is more similar to object B than to object C" about a data set, we propose two ways of defining a kernel function on the data set. While previous approaches construct a low-dimensional Euclidean embedding of the data set that reflects the given similar…
TSVQR captures heterogeneous and asymmetric data using quantile regression.