New matrices satisfy RIP with correlated entries for various applications.
problem Constructing RIP matrices with dependent entries.
method Introduced a new ensemble of random matrices XR where X is a fixed matrix and R is a random matrix from various models. result The constructed matrices XR satisfy the RIP with high probability. New analysis proves sketching operators' RIP guarantees for mixture models without importance sampling.
problem Proving sketching operators' Restricted Isometry Property (RIP) for mixture models without assuming importance sampling.
method Proposed alternative analysis based on new deterministic bounds and concentration inequalities.
result Theoretical guarantees for sketching operators without importance sampling.
We propose a relaxation-based approximate inference algorithm that samples near-MAP configurations of a binary pairwise Markov random field. We experiment on MAP inference tasks in several restricted Boltzmann machines. We also use our underlying sampler to estimate the log-partition function of restricted Boltzmann ma…
It is natural to ask: what kinds of matrices satisfy the Restricted Eigenvalue (RE) condition? In this paper, we associate the RE condition (Bickel-Ritov-Tsybakov 09) with the complexity of a subset of the sphere in Rp, where p is the dimensionality of the data, and show that a class of random matrices with indep…
Deep learning depends on tuning layers near critical points.
problem Understanding how deep learning architectures depend on tuning parameters.
method Random energy approach to analyze statistical dependence in deep belief networks.
result Statistical dependence can propagate only if layers are tuned near critical points.
FastForest boosts Random Forest speed by 24%.
problem Efficiency in processing speed for Random Forest.
method Subsample Aggregating, Logarithmic Split-Point Sampling, Dynamic Restricted Subspacing.
result Average 24% increase in processing speed with accuracy maintained.
New bounds on random quadratic forms hold under dependence, useful for adaptive modeling.
problem Need for independence in bounds on random quadratic forms.
method Uniform bounds on random quadratic forms of conditionally independent and sub-Gaussian stochastic processes.
result Bounds hold under general dependencies and sequential design.
New method reduces variance and bias in approximating indefinite kernels.
problem Approximating non-stationary indefinite kernels with low variance and bias.
method Generalized orthogonal random features (GORF)
result GORF achieves lower variance and approximation error compared to existing methods.
We consider the problem of how to assign treatment in a randomized experiment, in which the correlation among the outcomes is informed by a network available pre-intervention. Working within the potential outcome causal framework, we develop a class of models that posit such a correlation structure among the outcomes. …
We analyze the condition number of random feature matrices and prove their well-conditioned nature.
problem Understanding the condition number of random feature matrices and its impact on generalization error.
method Established concentration bounds and derived risk bounds for regression problems using random feature matrices.
result The risk associated with random feature matrices exhibits the double descent phenomenon, improving even with noise.
New method estimates latent gene expression factors without overlap with known confounders.
problem Estimating latent variance components in gene expression data with known confounders.
method Restricted maximum-likelihood method maximizing likelihood on orthogonal subspace.
result Method reduces runtime and attains greater likelihood values than gradient-based optimizers.
New theorem improves spectral gap for sampling from mixture distributions.
problem Sampling from multimodal distributions with simulated tempering.
method Introduced a decomposition theorem for the restricted spectral gap of simulated tempering.
result Lower bound on the restricted spectral gap for mixture distributions.
Conditional restricted Boltzmann machines are undirected stochastic neural networks with a layer of input and output units connected bipartitely to a layer of hidden units. These networks define models of conditional probability distributions on the states of the output units given the states of the input units, parame…
The restricted Boltzmann machine is a graphical model for binary random variables. Based on a complete bipartite graph separating hidden and observed variables, it is the binary analog to the factor analysis model. We study this graphical model from the perspectives of algebraic statistics and tropical geometry, starti…
The paper analyzes Karcher means on restricted PSD matrices with statistical guarantees.
problem Statistical analysis of non-linear manifolds in machine learning.
method Intrinsic mean model on restricted PSD matrices, Karcher mean analysis, extrinsic signal-plus-noise model.
result Non-asymptotic statistical analysis of Karcher means with deterministic error bounds.
We present a general construction for dependent random measures based on thinning Poisson processes on an augmented space. The framework is not restricted to dependent versions of a specific nonparametric model, but can be applied to all models that can be represented using completely random measures. Several existing …
New method constructs matrices satisfying Restricted Eigenvalue condition for sparse recovery.
problem Sparse recovery in high-dimensional settings with limited data.
method Constructs matrices from a fixed deterministic matrix and a subgaussian random matrix.
result New matrices satisfy Restricted Eigenvalue condition with high probability.
Paper proves sufficient conditions for tensor recovery using t-RIP with random measurements.
problem Establish robust recovery guarantees for low-tubal-rank tensors.
method Probabilistic arguments and random sub-Gaussian distributions to ensure t-RIP conditions.
result Minimal number of linear measurements nearly optimal for tensor recovery.
Given a Gaussian Markov random field, we consider the problem of selecting a subset of variables to observe which minimizes the total expected squared prediction error of the unobserved variables. We first show that finding an exact solution is NP-hard even for a restricted class of Gaussian Markov random fields, calle…
We analyze a monetary system of random money transfer on the basis of double entry bookkeeping. Without boundary conditions, we do not reach a price equilibrium and violate text-book formulas of economists quantity theory (MV=PQ). To match the resulting quantity of money with the model assumption of a constant price, w…
The paper offers streamlined algorithms for fitting complex linear mixed models.
problem Linear mixed models with crossed random effects in large dimensions.
method Mean field variational Bayes algorithms with various relaxations and storage strategies.
result Different inference strategies have varying trade-offs between accuracy and computational demands.
The purpose of this work is to explore the role that random arbitrage opportunities play in pricing financial derivatives. We use a non-equilibrium model to set up a stochastic portfolio, and for the random arbitrage return, we choose a stationary ergodic random process rapidly varying in time. We exploit the fact that…
Given a knot K in an Euclidean space E and a finite dimensional space V of smooth functions on K, we express the expected number of critical points of a random function in V in terms of an integral-geometric invariant of K and V. When V consists of the restrictions to K of homogeneous polynomials of degree d on E, this…
New model for STSs with restricted horizontal gluings, focusing on maximal horizontal cylinders.
problem Modeling STSs with specific horizontal restrictions.
method Modified model with conjugacy classes of permutations to restrict horizontal gluings.
result Asymptotic analysis of components, genus distribution, and saddle connections.
We use the Chebyshev knot diagram model of Koseleff and Pecker in order to introduce a random knot diagram model by assigning the crossings to be positive or negative uniformly at random. We give a formula for the probability of choosing a knot at random among all knots with bridge index at most 2. Restricted to this c…
Adversarial method learns deep models from conditional moment restrictions.
problem Learning deep neural net representations of models with conditional moment restrictions.
method Formulate as a zero-sum game between modeler and adversary, using adversarial training.
result Effective adversarial training methods, including k-means and random forests.
Given an ensemble of randomized regression trees, it is possible to restructure them as a collection of multilayered neural networks with particular connection weights. Following this principle, we reformulate the random forest method of Breiman (2001) into a neural network setting, and in turn propose two new hybrid p…
The usual development of the continuous-time random walk (CTRW) proceeds by assuming that the present is one of the jumping times. Under this restrictive assumption integral equations for the propagator and mean escape times have been derived. We generalize these results to the case when the present is an arbitrary tim…
New family of triangulated 3-spheres identified from trees.
problem Challenges in enumerating triangulations of 3-spheres.
method Identifying a restricted family of triangulations and proving bijections with triples of trees.
result These triangulations are in bijection with a combinatorial family of triples of plane trees.
Deep neural nets approximate random dynamical system trajectories uniformly in time.
problem Approximating trajectories of random dynamical systems over infinite time horizons.
method Recurrent neural networks with simple feedback structures.
result Certain random trajectories can be approximated uniformly in time to any desired accuracy.
RBM models reveal how hidden unit tail behavior affects pattern reconstruction.
problem Understanding how the tail behavior of hidden units in RBMs influences pattern reconstruction.
method Identified an effective energy function for RBMs and studied its local minima.
result The ability to reconstruct patterns depends on the tail behavior of the hidden unit prior distribution.
The multilabel learning problem with large number of labels, features, and data-points has generated a tremendous interest recently. A recurring theme of these problems is that only a few labels are active in any given datapoint as compared to the total number of labels. However, only a small number of existing work ta…
A scalable method for estimating spatial data using VREML.
problem Costly computation of REML for large, sparse precision matrices in spatial data.
method Proposes VREML framework approximating marginal likelihood with Gaussian variational distribution and deriving a coordinate-ascent algorithm.
result Empirically shows VREML outperforms MLE and INLA.
Federated Learning tackles limited user participation with a new risk-aware approach.
problem Limited availability of users in federated learning environments.
method Random Access Model (RAM) and Conditional Value-at-Risk (CVaR) to design a risk-aware federated learning algorithm.
result The proposed approach achieves significantly improved performance under various setups compared to standard federated learning.
The paper develops a new algorithm for RBMs using dynamical mean-field theory.
problem Learning in Restricted Boltzmann Machines (RBMs) with complex dependencies.
method Dynamical mean-field theory applied to RBMs with rectangular coupling matrices drawn from a bi-rotation invariant ensemble.
result The algorithm converges globally under a stability criterion, with rates matching numerical simulations.
Improved RBM training using MCLV-K outperforms CD-K on MNIST.
problem Training RBMs efficiently and with statistical guarantees.
method Markov Chain Las Vegas (MCLV-K) with stopping sets.
result MCLV-K significantly outperforms CD-K on MNIST.
New algorithm for learning RBMs with sparse latent variables.
problem Learning RBMs with sparse latent variables efficiently.
method Algorithm with time complexity O(n^(2^s+1)) for sparse RBMs.
result Improves learning time for RBMs with sparse latent variables.
Study on random hyperbolic surfaces with many cusps, focusing on tight geodesics.
problem Understanding length statistics of geodesics on random hyperbolic surfaces with cusps.
method Recursion formula for tight Weil-Petersson volumes and generalization of Mirzakhani's integration formula.
result Recovery of Poisson point process in large genus limit for length statistics of tight geodesics.
ORCCA improves CCA performance with randomized features.
problem Improving CCA performance with randomized features.
method Proposes a task-specific scoring rule for selecting random features in CCA.
result ORCCA outperforms Kernel CCA in expectation.
Given a graph embedded in an orientable surface, a process consisting of random excitations and random node and face balancing is constructed and analyzed. It is shown that given a priori bounds g' on the genus and n' on the number of nodes, one can determine the genus of the surface from local observations of the proc…
This paper provides a tutorial on Boltzmann Machines and Deep Belief Networks.
problem Understanding and applying Boltzmann Machines and Deep Belief Networks.
method Explains the structures, conditional distributions, Gibbs sampling, training methods, and deep belief networks of RBMs.
result Comprehensive overview of RBMs and DBNs, useful in various fields.
New method uses generative models to estimate aleatoric uncertainty without strict data restrictions.
problem Estimating aleatoric uncertainty with limited data distribution or dimensionality.
method Conditional generative models and two metrics for measuring distributional discrepancies.
result Metrics accurately measure conditional distributional discrepancies and train competitive models.
This paper examines from an experimental perspective random forests, the increasingly used statistical method for classification and regression problems introduced by Leo Breiman in 2001. It first aims at confirming, known but sparse, advice for using random forests and at proposing some complementary remarks for both …
New method uses manifold learning to infer latent positions of 1D submanifolds in random dot product graphs.
problem Inference on latent positions of unknown 1D submanifolds in RDPGs.
method Apply Isomap for manifold learning to estimate arc lengths on the unknown submanifold.
result Test statistics based on Isomap converge to known submanifold power as auxiliary vertices increase.
Dual random fields improve mineral potential predictions.
problem Limited understanding of multi-dimensional causalities and dependencies.
method Introduces dual random fields to pool response functions across the domain.
result Spatial inference and uncertainty assessment of response models and predictions.
Improves data recovery with optimized measurements and generalized sparsity models.
problem Data recovery with optimized measurements and generalized sparsity models.
method Optimizing over families of Banach spaces, investigating preservation of difference of sparse vectors, extending RIP to group structured measurements, and extending Fourier measurement concepts to infinite dimensions.
result Optimal scaling of number of measurements for group structured measurements and improved RIP in infinite dimensions.
New stochastic gradient descent with random search directions improves efficiency and convergence.
problem Efficiency and convergence of stochastic gradient descent methods.
method Developed a new class of stochastic gradient descent algorithms with random search directions.
result Established almost sure convergence and provided Lp rates of convergence. Proposes a new clustering algorithm using random forest.
problem Density-based clustering with optimal level determination.
method Best-scored random forest algorithm.
result Guaranteed consistency and fast convergence rates.