New algorithm speeds up causal inference for large data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New insights into tSNE for large datasets.
The MBO scheme for data clustering is analyzed in the large data limit, proving convergence to optimal partition problems.
The paper examines how kernel approximations affect Gaussian process regression in large data applications.
Deep MF extracts hierarchical features from large data sets.
We consider the porous medium equation with power-type reaction terms on negatively curved Riemannian manifolds, and solutions corresponding to bounded, nonnegative and compactly supported data. If , small data give rise to global-in-time solutions while solutions associated to large data blow up in finite t…
Model inference, such as model comparison, model checking, and model selection, is an important part of model development. Leave-one-out cross-validation (LOO) is a general approach for assessing the generalizability of a model, but unfortunately, LOO does not scale well to large datasets. We propose a combination of u…
We present a new method for estimating multivariate, second-order stationary Gaussian Random Field (GRF) models based on the Sparse Precision matrix Selection (SPS) algorithm, proposed by Davanloo et al. (2015) for estimating scalar GRF models. Theoretical convergence rates for the estimated between-response covariance…
Derives ideal train/test split for ridge regression in large data limit.
Bayesian data sketching speeds up inference for large functional data.
Dimensionality-reduction methods are a fundamental tool in the analysis of large data sets. These algorithms work on the assumption that the "intrinsic dimension" of the data is generally much smaller than the ambient dimension in which it is collected. Alongside their usual purpose of mapping data into a smaller dimen…
Bayesian Neural Networks are robust to gradient-based attacks in the large-data limit.
New classifiers converge under large data, simplifying complex models.
Gaussian process regression (GPR) is a non-parametric Bayesian technique for interpolating or fitting data. The main barrier to further uptake of this powerful tool rests in the computational costs associated with the matrices which arise when dealing with large data sets. Here, we derive some simple results which we h…
We extend Eardley and Moncrief's estimates for the conformally invariant Yang-Mills-Higgs equations to the Einstein cylinder. Our method is to first work on Minkowski space and localise their estimates, and then carry them to the Einstein cylinder by a conformal transformation. By patching local estimates to…
We propose a practical and scalable Gaussian process model for large-scale nonlinear probabilistic regression. Our mixture-of-experts model is conceptually simple and hierarchically recombines computations for an overall approximation of a full Gaussian process. Closed-form and distributed computations allow for effici…
Bayesian optimization method tackles combinatorial spaces, scalable for large data.
To scale Gaussian processes (GPs) to large data sets we introduce the robust Bayesian Committee Machine (rBCM), a practical and scalable product-of-experts model for large-scale distributed GP regression. Unlike state-of-the-art sparse GP approximations, the rBCM is conceptually simple and does not rely on inducing or …
Extracts representative scenarios from large data panels.
The proliferation of large data sets and Bayesian inference techniques motivates demand for better data sparsification. Coresets provide a principled way of summarizing a large dataset via a smaller one that is guaranteed to match the performance of the full data set on specific problems. Classical coresets, however, n…
Localized SVMs maintain SVM's consistency properties for large datasets.
Given a data set and a subset of labels the problem of semi-supervised learning on point clouds is to extend the labels to the entire data set. In this paper we extend the labels by minimising the constrained discrete -Dirichlet energy. Under suitable conditions the discrete problem can be connected, in the large da…
Proves existence and trapped surface formation for Einstein-Vlasov system without symmetry assumptions.
New methods for Markov Blanket discovery using MML outperform existing approaches.
New iterative methods improve Vecchia-Laplace approximations for large data sets.
Unified study of ridge regression structure, cross-validation, and acceleration.
ROBOT framework solves regression without correspondence for large data and complex models.
This paper introduces online algorithms to estimate robust geometric median in large data streams.
We propose a generic Markov Chain Monte Carlo (MCMC) algorithm to speed up computations for datasets with many observations. A key feature of our approach is the use of the highly efficient difference estimator from the survey sampling literature to estimate the log-likelihood accurately using only a small fraction of …
Divide-and-conquer method splits large data sets for efficient analysis.
Study robust covariance estimation in large data with concentrated vectors.
Tensor decompositions are powerful tools for large data analytics as they jointly model multiple aspects of data into one framework and enable the discovery of the latent structures and higher-order correlations within the data. One of the most widely studied and used decompositions, especially in data mining and machi…
Graph Laplacians computed from weighted adjacency matrices are widely used to identify geometric structure in data, and clusters in particular; their spectral properties play a central role in a number of unsupervised and semi-supervised learning algorithms. When suitably scaled, graph Laplacians approach limiting cont…
Stein variational gradient descent improves inference in Gaussian process models.
A novel distributed adaptive NN classifier for large data sets.
VarFA efficiently estimates student skill levels with uncertainty for adaptive testing.
Bayesian model explains and improves black-box estimators for class distribution.
In this paper we describe a new algorithm called Fast Adaptive Sequencing Technique (FAST) for maximizing a monotone submodular function under a cardinality constraint whose approximation ratio is arbitrarily close to , is adaptive, and uses a total of queries. …
In this article, a large data set containing every course taken by every undergraduate student in a major university in Canada over 10 years is analysed. Modern machine learning algorithms can use large data sets to build useful tools for the data provider, in this case, the university. In this article, two classifiers…
New method uses symmetric splitting for efficient HMC inference in large neural networks.
In this paper we analyze approximate methods for undertaking a principal components analysis (PCA) on large data sets. PCA is a classical dimension reduction method that involves the projection of the data onto the subspace spanned by the leading eigenvectors of the covariance matrix. This projection can be used either…
CCMM efficiently solves large-scale convex clustering problems.
New method for scalable learning of IRT models from large datasets.
One of the founding paradigms of machine learning is that a small number of variables is often sufficient to describe high-dimensional data. The minimum number of variables required is called the intrinsic dimension (ID) of the data. Contrary to common intuition, there are cases where the ID varies within the same data…
Many applications of machine learning involve the analysis of large data frames-matrices collecting heterogeneous measurements (binary, numerical, counts, etc.) across samples-with missing values. Low-rank models, as studied by Udell et al. [30], are popular in this framework for tasks such as visualization, clustering…
Stochastic gradient descent procedures have gained popularity for parameter estimation from large data sets. However, their statistical properties are not well understood, in theory. And in practice, avoiding numerical instability requires careful tuning of key parameters. Here, we introduce implicit stochastic gradien…
New algorithm estimates complex probabilistic models efficiently.
This paper tackles regularization parameter learning in inverse problems using data-driven bilevel optimization.