A new Weyl prior is proposed for Bayesian statistics, offering a more canonical choice for parameter α.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study explores geometric structure and prior for beta-logistic distribution.
Parallel sampling for smooth distributions with fast convergence.
A new parallel algorithm speeds up Hawkes process estimation.
Inference of latent feature models in the Bayesian nonparametric setting is generally difficult, especially in high dimensional settings, because it usually requires proposing features from some prior distribution. In special cases, where the integration is tractable, we can sample new feature assignments according to …
Efficiently quantifies uncertainty in DeepONets for function spaces.
New method adapts DLMs to intrinsic data dependence without prior knowledge.
New sampling method improves efficiency for diffusion models.
Alternative sampling method for autoregressive models using Langevin dynamics.
New tool for parallel and private stochastic convex optimization reduces query complexity.
We study the question of whether parallelization in the exploration of the feasible set can be used to speed up convex optimization, in the local oracle model of computation. We show that the answer is negative for both deterministic and randomized algorithms applied to essentially any of the interesting geometries and…
Stochastic convex optimization algorithms are the most popular way to train machine learning models on large-scale data. Scaling up the training process of these models is crucial, but the most popular algorithm, Stochastic Gradient Descent (SGD), is a serial method that is surprisingly hard to parallelize. In this pap…
Unified parallel ADMM for high-dimensional regression with combined regularizations.
HybridSGD improves SGD performance by balancing computation and communication.
New method solves blind inverse problems by optimizing both operator and image parameters.
We describe an embarrassingly parallel, anytime Monte Carlo method for likelihood-free models. The algorithm starts with the view that the stochasticity of the pseudo-samples generated by the simulator can be controlled externally by a vector of random numbers u, in such a way that the outcome, knowing u, is determinis…
It is well-known that the distribution over functions induced through a zero-mean iid prior distribution over the parameters of a multi-layer perceptron (MLP) converges to a Gaussian process (GP), under mild conditions. We extend this result firstly to independent priors with general zero or non-zero means, and secondl…
A new generation of manycore processors is on the rise that offers dozens and more cores on a chip and, in a sense, fuses host processor and accelerator. In this paper we target the efficient training of generalized linear models on these machines. We propose a novel approach for achieving parallelism which we call Het…
Stochastic gradient descent (SGD) is a well known method for regression and classification tasks. However, it is an inherently sequential algorithm at each step, the processing of the current example depends on the parameters learned from the previous examples. Prior approaches to parallelizing linear learners using SG…
The past years have witnessed many dedicated open-source projects that built and maintain implementations of Support Vector Machines (SVM), parallelized for GPU, multi-core CPUs and distributed systems. Up to this point, no comparable effort has been made to parallelize the Elastic Net, despite its popularity in many h…
We introduce a fast model based deep learning approach for calibrationless parallel MRI reconstruction. The proposed scheme is a non-linear generalization of structured low rank (SLR) methods that self learn linear annihilation filters from the same subject. It pre-learns non-linear annihilation relations in the Fourie…
NEST optimizes deep learning training by placing devices efficiently across networks and memory.
Paper tackles robust knowledge transfer in parallel RL tasks.
We propose a generic confidence-based approximation that can be plugged in and simplify the auto-regressive generation process with a proved convergence. We first assume that the priors of future samples can be generated in an independently and identically distributed (i.i.d.) manner using an efficient predictor. Given…
ProSper is a python library containing probabilistic algorithms to learn dictionaries. Given a set of data points, the implemented algorithms seek to learn the elementary components that have generated the data. The library widens the scope of dictionary learning approaches beyond implementations of standard approaches…
Bayesian neural networks improve with summary information and Dirichlet process.
New method ensures consistent inference across different tensor parallel sizes for large language models.
FMP sampling improves model calibration without sharing data.
SHVC improves image compression with fewer parameters.
AI-driven Bayesian inference improves decision-making uncertainty.
New algorithm improves inference for flexible models with infinite latent features.
Bayesian framework uses AI-generated data to improve parameter estimation.
Motivation: With the development of droplet based systems, massive single cell transcriptome data has become available, which enables analysis of cellular and molecular processes at single cell resolution and is instrumental to understanding many biological processes. While state-of-the-art clustering methods have been…
FEM improves attention mechanisms by applying value-driven log-linear tilts.
Bayesian imaging uses neural networks to learn prior knowledge from data.
This paper proposes a multi-channel image reconstruction method, named DeepcomplexMRI, to accelerate parallel MR imaging with residual complex convolutional neural network. Different from most existing works which rely on the utilization of the coil sensitivities or prior information of predefined transforms, Deepcompl…
A new method infers causal gene regulatory networks from parallel CRISPR interventions and transcriptomic data.
Improved SVAE models enhance sequential data prediction.
Bayesian optimization improved for high-dimensional outputs using randomized priors.
We reduce the computational cost of Neural AutoML with transfer learning. AutoML relieves human effort by automating the design of ML algorithms. Neural AutoML has become popular for the design of deep learning architectures, however, this method has a high computation cost. To address this we propose Transfer Neural A…
GShard enables scaling of large neural networks with automatic sharding and lightweight APIs.
We introduce a new, high-throughput, synchronous, distributed, data-parallel, stochastic-gradient-descent learning algorithm. This algorithm uses amortized inference in a compute-cluster-specific, deep, generative, dynamical model to perform joint posterior predictive inference of the mini-batch gradient computation ti…
We introduce topological parallelisms of oriented lines (briefly called oriented parallelisms). Every topological parallelism (of lines) on PG(3,R) gives rise to a parallelism of oriented lines, but we show that even the most homogeneous parallelisms of oriented lines other than the Clifford parallelism do not necessar…
In this tutorial we explain the inference procedures developed for the sparse Gaussian process (GP) regression and Gaussian process latent variable model (GPLVM). Due to page limit the derivation given in Titsias (2009) and Titsias & Lawrence (2010) is brief, hence getting a full picture of it requires collecting resul…
We develop an automated variational method for inference in models with Gaussian process (GP) priors and general likelihoods. The method supports multiple outputs and multiple latent functions and does not require detailed knowledge of the conditional likelihood, only needing its evaluation as a black-box function. Usi…
Parallelizes MCTS for continuous domains using leaf and root parallelization.
DWTS uses observational data to improve clinical trial efficiency.
A new method tackles Bayesian inverse problems with complex PDEs.