Adaptive sampling method improves efficiency in complex target distributions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Proposes a new model for clustering with heavier tails.
Proposes vMF distribution for skewed elliptical distributions.
Paper develops a gradient-like proposal for discrete distributions without requiring natural differentiability.
Proposes QQE for transforming and embedding data distributions.
In this paper, we propose a data collaboration analysis method for distributed datasets. The proposed method is a centralized machine learning while training datasets and models remain distributed over some institutions. Recently, data became large and distributed with decreasing costs of data collection. If we can cen…
Comparing counterfactual distributions can provide more nuanced and valuable measures for causal effects, going beyond typical summary statistics such as averages. In this work, we consider characterizing causal effects via distributional distances, focusing on two kinds of target parameters. The first is the counterfa…
Generates samples conditioned on labels using optimal transport.
In this paper, an issue of building the RRC model using probability distributions other than beta distribution is addressed. More precisely, in this paper, we propose to build the RRR model using the truncated normal distribution. Heuristic procedures for expected value and the variance of the truncated-normal distribu…
Proposes PGPS for efficient Bayesian inference.
Proposes a method to improve SLMC for multimodal distributions.
Proposes a new complex Gaussian distribution for better modeling of complex-valued signals.
Neural network MCMC sampler maximizes proposal entropy for efficient sampling.
Paper proposes a new method to learn distribution kernels via entropy maximization.
Improved image reconstruction using VAEs with Student's t-prior.
While the Matrix Generalized Inverse Gaussian () distribution arises naturally in some settings as a distribution over symmetric positive semi-definite matrices, certain key properties of the distribution and effective ways of sampling from the distribution have not been carefully studied. In this paper…
One-shot algorithm for feature-distributed kernel PCA reduces communication costs.
Counterfactual inference has become a ubiquitous tool in online advertisement, recommendation systems, medical diagnosis, and econometrics. Accurate modeling of outcome distributions associated with different interventions -- known as counterfactual distributions -- is crucial for the success of these applications. In …
We consider distributed on-device learning with limited communication and security requirements. We propose a new robust distributed optimization algorithm with efficient communication and attack tolerance. The proposed algorithm has provable convergence and robustness under non-IID settings. Empirical results show tha…
Sequential Monte Carlo (SMC), or particle filtering, is a popular class of methods for sampling from an intractable target distribution using a sequence of simpler intermediate distributions. Like other importance sampling-based methods, performance is critically dependent on the proposal distribution: a bad proposal c…
A novel distributed adaptive NN classifier for large data sets.
Variational Auto-Encoders enforce their learned intermediate latent-space data distribution to be a simple distribution, such as an isotropic Gaussian. However, this causes the posterior collapse problem and loses manifold structure which can be important for datasets such as facial images. A GAN can transform a simple…
Adaptive importance sampling is a class of techniques for finding good proposal distributions for importance sampling. Often the proposal distributions are standard probability distributions whose parameters are adapted based on the mismatch between the current proposal and a target distribution. In this work, we prese…
Efficiently updates posterior tree distributions over meta-trees.
Estimates low-rank distributional matrices from incomplete samples.
Paper proposes a new method to aggregate multiple sources with different label distributions.
A new method solves distributed optimization problems over networks.
Transfer learning has achieved promising results by leveraging knowledge from the source domain to annotate the target domain which has few or none labels. Existing methods often seek to minimize the distribution divergence between domains, such as the marginal distribution, the conditional distribution or both. Howeve…
A distributed algorithm for training graph convolutional networks.
Paper develops a new method to improve model calibration under distribution shifts.
A new generative adversarial network is developed for joint distribution matching. Distinct from most existing approaches, that only learn conditional distributions, the proposed model aims to learn a joint distribution of multiple random variables (domains). This is achieved by learning to sample from conditional dist…
The aim of this paper is to propose distributed strategies for adaptive learning of signals defined over graphs. Assuming the graph signal to be bandlimited, the method enables distributed reconstruction, with guaranteed performance in terms of mean-square error, and tracking from a limited number of sampled observatio…
EvoMSN tackles time series forecasting under distribution shifts by evolving multi-scale normalization.
Paper proposes faster adaptation to distribution shifts in online settings.
A novel method to propagate uncertainty through the soft-thresholding nonlinearity is proposed in this paper. At every layer the current distribution of the target vector is represented as a spike and slab distribution, which represents the probabilities of each variable being zero, or Gaussian-distributed. Using the p…
Paper tackles distributed linear regression with compositional covariates.
It has been pointed out by Patriarca et al. (2005) that the power-law tailed equilibrium distribution in heterogeneous kinetic exchange models with a distributed saving parameter can be resolved as a mixture of Gamma distributions corresponding to particular subsets of agents. Here, we propose a new four-parameter stat…
AIS uses a suboptimal extended target distribution, which this paper improves using SGM.
AIS algorithm improves heavy-tailed distribution estimation.
Predicting conversion rates (CVRs) in display advertising (e.g., predicting the proportion of users who purchase an item (i.e., a conversion) after its corresponding ad is clicked) is important when measuring the effects of ads shown to users and to understanding the interests of the users. There is generally a time de…
Paper proposes a new regularization method to prevent model degradation under distribution shifts.
One key use of k-means clustering is to identify cluster prototypes which can serve as representative points for a dataset. However, a drawback of using k-means cluster centers as representative points is that such points distort the distribution of the underlying data. This can be highly disadvantageous in problems wh…
Distributional (or distribution-valued) data are a new type of data arising from several sources and are considered as realizations of distributional variables. A new set of fuzzy c-means algorithms for data described by distributional variables is proposed. The algorithms use the Wasserstein distance between dist…
A new method for faster prediction in distributed Gaussian processes.
With the recently rapid development in deep learning, deep neural networks have been widely adopted in many real-life applications. However, deep neural networks are also known to have very little control over its uncertainty for unseen examples, which potentially causes very harmful and annoying consequences in practi…
The paper proposes a method to estimate treatment effects using CAR designs with additional covariates.
In this paper, we propose a novel method for generating a synthetic dataset obeying Gaussian distribution. Compared to the commonly used benchmark datasets with unknown distribution, the synthetic dataset has an explicit distribution, i.e., Gaussian distribution. Meanwhile, it has the same characteristics as the benchm…
Paper proposes a new Wasserstein distance for mixtures of radially contoured distributions.