We consider the porous medium equation with power-type reaction terms on negatively curved Riemannian manifolds, and solutions corresponding to bounded, nonnegative and compactly supported data. If , small data give rise to global-in-time solutions while solutions associated to large data blow up in finite t…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Proves existence and trapped surface formation for Einstein-Vlasov system without symmetry assumptions.
Bayesian data sketching speeds up inference for large functional data.
New methods for Markov Blanket discovery using MML outperform existing approaches.
New algorithm speeds up causal inference for large data.
ROBOT framework solves regression without correspondence for large data and complex models.
New insights into tSNE for large datasets.
The MBO scheme for data clustering is analyzed in the large data limit, proving convergence to optimal partition problems.
Study robust covariance estimation in large data with concentrated vectors.
The paper examines how kernel approximations affect Gaussian process regression in large data applications.
Deep MF extracts hierarchical features from large data sets.
VarFA efficiently estimates student skill levels with uncertainty for adaptive testing.
Smoothing splines provide a powerful and flexible means for nonparametric estimation and inference. With a cubic time complexity, fitting smoothing spline models to large data is computationally prohibitive. In this paper, we use the theoretical optimal eigenspace to derive a low rank approximation of the smoothing spl…
Model inference, such as model comparison, model checking, and model selection, is an important part of model development. Leave-one-out cross-validation (LOO) is a general approach for assessing the generalizability of a model, but unfortunately, LOO does not scale well to large datasets. We propose a combination of u…
We present a new method for estimating multivariate, second-order stationary Gaussian Random Field (GRF) models based on the Sparse Precision matrix Selection (SPS) algorithm, proposed by Davanloo et al. (2015) for estimating scalar GRF models. Theoretical convergence rates for the estimated between-response covariance…
CCMM efficiently solves large-scale convex clustering problems.
Derives ideal train/test split for ridge regression in large data limit.
Dimensionality-reduction methods are a fundamental tool in the analysis of large data sets. These algorithms work on the assumption that the "intrinsic dimension" of the data is generally much smaller than the ambient dimension in which it is collected. Alongside their usual purpose of mapping data into a smaller dimen…
Bayesian Neural Networks are robust to gradient-based attacks in the large-data limit.
New classifiers converge under large data, simplifying complex models.
Gaussian process regression (GPR) is a non-parametric Bayesian technique for interpolating or fitting data. The main barrier to further uptake of this powerful tool rests in the computational costs associated with the matrices which arise when dealing with large data sets. Here, we derive some simple results which we h…
We extend Eardley and Moncrief's estimates for the conformally invariant Yang-Mills-Higgs equations to the Einstein cylinder. Our method is to first work on Minkowski space and localise their estimates, and then carry them to the Einstein cylinder by a conformal transformation. By patching local estimates to…
We propose a practical and scalable Gaussian process model for large-scale nonlinear probabilistic regression. Our mixture-of-experts model is conceptually simple and hierarchically recombines computations for an overall approximation of a full Gaussian process. Closed-form and distributed computations allow for effici…
This paper tackles distributed estimation of the top-L eigenspace in PCA for large data sets.
Bayesian optimization method tackles combinatorial spaces, scalable for large data.
To scale Gaussian processes (GPs) to large data sets we introduce the robust Bayesian Committee Machine (rBCM), a practical and scalable product-of-experts model for large-scale distributed GP regression. Unlike state-of-the-art sparse GP approximations, the rBCM is conceptually simple and does not rely on inducing or …
New method creates coresets for deep neural networks efficiently.
Extracts representative scenarios from large data panels.
The proliferation of large data sets and Bayesian inference techniques motivates demand for better data sparsification. Coresets provide a principled way of summarizing a large dataset via a smaller one that is guaranteed to match the performance of the full data set on specific problems. Classical coresets, however, n…
New iterative methods improve scalability of Gaussian process approximations for large data.
We investigate coresets - succinct, small summaries of large data sets - so that solutions found on the summary are provably competitive with solution found on the full data set. We provide an overview over the state-of-the-art in coreset construction for machine learning. In Section 2, we present both the intuition be…
Localized SVMs maintain SVM's consistency properties for large datasets.
Given a data set and a subset of labels the problem of semi-supervised learning on point clouds is to extend the labels to the entire data set. In this paper we extend the labels by minimising the constrained discrete -Dirichlet energy. Under suitable conditions the discrete problem can be connected, in the large da…
In this paper, a class of statistics named ART (the alternant recursive topology statistics) is proposed to measure the properties of correlation between two variables. A wide range of bi-variable correlations both linear and nonlinear can be evaluated by ART efficiently and equitably even if nothing is known about the…
New iterative methods improve Vecchia-Laplace approximations for large data sets.
Support Vector Data Description (SVDD) provides a useful approach to construct a description of multivariate data for single-class classification and outlier detection with various practical applications. Gaussian kernel used in SVDD formulation allows flexible data description defined by observations designated as sup…
Inference in Gaussian process (GP) models is computationally challenging for large data, and often difficult to approximate with a small number of inducing points. We explore an alternative approximation that employs stochastic inference networks for a flexible inference. Unfortunately, for such networks, minibatch tra…
This paper introduces online algorithms to estimate robust geometric median in large data streams.
We propose a generic Markov Chain Monte Carlo (MCMC) algorithm to speed up computations for datasets with many observations. A key feature of our approach is the use of the highly efficient difference estimator from the survey sampling literature to estimate the log-likelihood accurately using only a small fraction of …
By analyzing a large data set of daily returns with data clustering technique, we identify economic sectors as clusters of assets with a similar economic dynamics. The sector size distribution follows Zipf's law. Secondly, we find that patterns of daily market-wide economic activity cluster into classes that can be ide…
Divide-and-conquer method splits large data sets for efficient analysis.
Sparse pseudo-point approximations for Gaussian process (GP) models provide a suite of methods that support deployment of GPs in the large data regime and enable analytic intractabilities to be sidestepped. However, the field lacks a principled method to handle streaming data in which both the posterior distribution ov…
A new, fast kernel test for large data.
New methods for time-to-event prediction are proposed by extending the Cox proportional hazards model with neural networks. Building on methodology from nested case-control studies, we propose a loss function that scales well to large data sets, and enables fitting of both proportional and non-proportional extensions o…
In this paper, we propose a deep, globally normalized topic model that incorporates structural relationships connecting documents in socially generated corpora, such as online forums. Our model (1) captures discursive interactions along observed reply links in addition to traditional topic information, and (2) incorpor…
Universally valid ground truth is almost impossible to obtain or would come at a very high cost. For supervised learning without universally valid ground truth, a recommended approach is applying crowdsourcing: Gathering a large data set annotated by multiple individuals of varying possibly expertise levels and inferri…
Tensor decompositions are powerful tools for large data analytics as they jointly model multiple aspects of data into one framework and enable the discovery of the latent structures and higher-order correlations within the data. One of the most widely studied and used decompositions, especially in data mining and machi…
Graph Laplacians computed from weighted adjacency matrices are widely used to identify geometric structure in data, and clusters in particular; their spectral properties play a central role in a number of unsupervised and semi-supervised learning algorithms. When suitably scaled, graph Laplacians approach limiting cont…