New proof shows rapid mixing for random walks on nilmanifolds.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New Krylov subspace methods speed up mixed-effects models with crossed random effects.
GBMixed boosts mixed models for clustered data, estimating mean and variance flexibly.
NoLimits.jl: Flexible and Composable Nonlinear Mixed-Effects Modeling in Julia
QR-MIX models joint state-action values as a distribution to handle randomness in MARL.
"Mixed Data" comprising a large number of heterogeneous variables (e.g. count, binary, continuous, skewed continuous, among other data types) are prevalent in varied areas such as genomics and proteomics, imaging genetics, national security, social networking, and Internet advertising. There have been limited efforts a…
The Tweedie Compound Poisson-Gamma model is routinely used for modeling non-negative continuous data with a discrete probability mass at zero. Mixed models with random effects account for the covariance structure related to the grouping hierarchy in the data. An important application of Tweedie mixed models is pricing …
Mixed datasets consist of both numeric and categorical attributes. Various k-means-based clustering algorithms have been developed for these datasets. Generally, these algorithms use random partition as a starting point, which tends to produce different clustering results for different runs. In this paper, we propose, …
Hashing is a basic tool for dimensionality reduction employed in several aspects of machine learning. However, the perfomance analysis is often carried out under the abstract assumption that a truly random unit cost hash function is used, without concern for which concrete hash function is employed. The concrete hash f…
Consider a matrix whose rows are independent centered non-degenerate Gaussian vectors with covariance matrices . Denote by the location-dispersion ellipsoid of . We sh…
Mixed membership factorization is a popular approach for analyzing data sets that have within-sample heterogeneity. In recent years, several algorithms have been developed for mixed membership matrix factorization, but they only guarantee estimates from a local optimum. Here, we derive a global optimization (GOP) algor…
Markov chain Monte Carlo (MCMC) algorithms are simple and extremely powerful techniques to sample from almost arbitrary distributions. The flaw in practice is that it can take a large and/or unknown amount of time to converge to the stationary distribution. This paper gives sufficient conditions to guarantee that univa…
Spectral methods improve signal recovery in mixed GLMs with precise asymptotics.
Study improves maize yield prediction using BNs with mixed-effects models.
RHMC accelerates sampling from log-concave distributions.
Gradient Boosted Mixed Models estimate mean and variance components for clustered data.
We consider the problem of solving mixed random linear equations with components. This is the noiseless setting of mixed linear regression. The goal is to estimate multiple linear models from mixed samples in the case where the labels (which sample corresponds to which model) are not observed. We give a tractable a…
Proposes integrating random effects into deep neural networks for better predictive performance.
Hamiltonian Monte Carlo (HMC) is a state-of-the-art Markov chain Monte Carlo sampling algorithm for drawing samples from smooth probability densities over continuous spaces. We study the variant most widely used in practice, Metropolized HMC with the Störmer-Verlet or leapfrog integrator, and make two primary contribut…
The multivariate version of the Mixed Tempered Stable is proposed. It is a generalization of the Normal Variance Mean Mixtures. Characteristics of this new distribution and its capacity in fitting tails and capturing dependence structure between components are investigated. We discuss a random number generating procedu…
MOMENT selects and estimates mixed-effects models using moment identities.
PAPAL algorithm finds mixed Nash equilibria in continuous games.
Graph matching in noisy environments with Markovian errors.
The beta-negative binomial process (BNBP), an integer-valued stochastic process, is employed to partition a count vector into a latent random count matrix. As the marginal probability distribution of the BNBP that governs the exchangeable random partitions of grouped data has not yet been developed, current inference f…
Paper analyzes Hit-and-Run's convergence rates and applies similar methods to randomized Kaczmarz.
The likelihood for the parameters of a generalized linear mixed model involves an integral which may be of very high dimension. Because of this intractability, many approximations to the likelihood have been proposed, but all can fail when the model is sparse, in that there is only a small amount of information availab…
OMERF extends random forest for hierarchical data and ordinal responses.
We study singular stochastic control of a two dimensional stochastic differential equation, where the first component is linear with random and unbounded coefficients. We derive existence of an optimal relaxed control and necessary conditions for optimality in the form of a mixed relaxed-singular maximum principle in a…
The paper explores the relationship between joint mixability and negative dependence structures.
Clustering is fundamental for gaining insights from complex networks, and spectral clustering (SC) is a popular approach. Conventional SC focuses on second-order structures (e.g., edges connecting two nodes) without direct consideration of higher-order structures (e.g., triangles and cliques). This has motivated SC ext…
We present an affine-invariant random walk for drawing uniform random samples from a convex body that uses maximum volume inscribed ellipsoids, known as John's ellipsoids, for the proposal distribution. Our algorithm makes steps using uniform sampling from the John's ellipsoid of the …
A new method for efficient inference in sequential latent-variable models.
Paper extends LME models to allow sign constraints on coefficients with SDTN random effects.
We improve Gaussian copula models for imputing mixed data types with precise approximations.
New method for summarizing Bayesian mixture models using sliced Wasserstein distances.
New Riemannian optimization improves variance estimation in mixed models.
Bayesian optimization reduces materials design costs by 10x.
Gibbs sampling is a Markov Chain Monte Carlo sampling technique that iteratively samples variables from their conditional distributions. There are two common scan orders for the variables: random scan and systematic scan. Due to the benefits of locality in hardware, systematic scan is commonly used, even though most st…
In this paper we introduce a new parametric distribution, the Mixed Tempered Stable. It has the same structure of the Normal Variance Mean Mixtures but the normality assumption leaves place to a semi-heavy tailed distribution. We show that, by choosing appropriately the parameters of the distribution and under the conc…
ARMED models improve deep learning interpretability and generalize better on clustered data.
In this paper, we propose a novel learning method for image classification called Between-Class learning (BC learning). We generate between-class images by mixing two images belonging to different classes with a random ratio. We then input the mixed image to the model and train the model to output the mixing ratio. BC …
The paper develops new inequalities for Markov chain sums, linking them to mixing time.
Study rare-event simulation for neural networks and random forests.
PAIN network improves imputation for mixed datasets.
We develop an HMC algorithm to easily marginalize random effects in LMMs.
While all kinds of mixed data -from personal data, over panel and scientific data, to public and commercial data- are collected and stored, building probabilistic graphical models for these hybrid domains becomes more difficult. Users spend significant amounts of time in identifying the parametric form of the random va…
We study the problem of semi-supervised singing voice separation, in which the training data contains a set of samples of mixed music (singing and instrumental) and an unmatched set of instrumental music. Our solution employs a single mapping function g, which, applied to a mixed sample, recovers the underlying instrum…
New study on No-U-Turn Sampler for accelerated mixing in Hamiltonian Monte Carlo.