Study on how priors affect GPR model performance.
problem Impact of prior distributions on hyperparameter estimation and GPR performance.
method Simulated and real data experiments with various priors for hyperparameters.
result Different priors for initial hyperparameters have no significant impact on GPR prediction performance.
Proposes deep weight prior for improving neural network performance.
problem Improving neural network performance with limited training data.
method Defines deep weight prior (DWP) as an implicit distribution and proposes variational inference methods.
result Improves performance of Bayesian neural networks with limited data and accelerates conventional CNN training.
Deep networks retain initial bias after training, affecting generalization.
problem Understanding how much initial bias in neural networks survives training.
method Introduced initialization memory to measure initial bias's survival.
result SGD can preserve initial bias, while Adam-family methods erase it.
Randomly initialized sentence encoders perform well on tasks, suggesting learning is key.
problem The role of sentence encoder architectures in language tasks.
method Random initialization and fixed architecture approach to evaluate sentence encoders.
result Priors do not leverage additional information, learning is necessary.
Proposes LLM-DCD for improved causal discovery from data.
problem Challenges in discovering causal relationships from observational data.
method Uses LLM to initialize DCD optimization, incorporating priors.
result Higher accuracy on benchmark datasets compared to state-of-the-art.
Randomly-initialized neural networks can capture image statistics as well as deep learning models.
problem Capturing image statistics without learning from examples.
method Using randomly-initialized neural networks as image priors.
result Randomly-initialized networks achieve similar results to deep learning models in image restoration tasks.
Combines spatial context priors with data term for better dendritic spine segmentation.
problem Poor segmentation results due to overlapping pixel intensity distributions.
method Combines nonparametric context priors with learned-intensity data term and nonparametric shape priors.
result Significant improvements in dendritic spine segmentation.
Enhances inference of spreading processes using neural-network priors.
problem Estimating initial states of graph processes from partial observations.
method Bayesian framework with single-layer perceptron neural network for initial states; hybrid BP-AMP algorithm.
result Model exhibits first-order phase transitions, creating a statistical-to-computational gap.
DAPS++ improves diffusion-based image restoration by decoupling prior and likelihood.
problem Decoupling prior and likelihood in diffusion-based inverse problems.
method Introducing DAPS++, which fully decouples diffusion-based initialization from likelihood-driven refinement.
result Achieves high computational efficiency and robust reconstruction performance.
L2F learns to forget, improving few-shot learning performance.
problem Few-shot learning challenges in adapting to unseen tasks.
method Task-and-layer-wise attenuation on compromised initialization to reduce prior knowledge influence.
result Faster adaptation and improved performance demonstrated experimentally.
Near-optimal sample complexity for phase retrieval with generative priors.
problem Phase retrieval with magnitude-only measurements and sparse signals.
method Near-optimal sample complexity with i.i.d. Gaussian measurements and generative models.
result O(k log L) samples suffice for phase retrieval with generative priors.
DAPS++ improves diffusion-based image restoration by decoupling prior and likelihood.
problem Decoupling prior and likelihood in diffusion-based inverse problems for better performance.
method Introducing DAPS++, which separates diffusion initialization from likelihood refinement.
result DAPS++ achieves high computational efficiency and robust reconstruction performance.
This paper solves quadratic systems with sparse or generative priors.
problem Recovering signals from quadratic systems with full-rank matrices.
method Thresholded Wirtinger flow (TWF) and projected gradient descent (PGD) algorithms.
result The proposed methods significantly outperform existing algorithms in signal recovery.
Study characterizes training and test risks for MAP regression with Gaussian priors.
problem Understanding high-dimensional behavior of regularized linear regression with informative priors.
method Maximum a posteriori (MAP) regression with Gaussian priors, using random matrix theory.
result Closed-form risk formulas reveal the bias-variance-prior tradeoff and explain double descent.
Study evaluates initialization strategies for infinite hidden Markov models.
problem Limited attention to initialization in infinite hidden Markov models.
method Systematically evaluated distance-based clustering, model-based, and uniform initializations.
result Distance-based clustering initializations consistently outperform other methods.
Paper proposes a method for estimating complex low-rank matrices from phase-only measurements.
problem Estimating complex low-rank matrices from magnitude-only measurements.
method A hierarchical prior model with a Gaussian-Wishart distribution is used to promote low-rankness. A variational EM algorithm is developed to solve the problem.
result The proposed method is less sensitive to initialization and performs well with random initialization.
New ReLU initialization improves network performance and dynamical isometry.
problem Improving the initialization of ReLU units for better network performance.
method Derive exact joint signal output distribution for fully-connected networks with Gaussian weights and biases, and propose a new initialization scheme for ReLU units.
result Proposed initialization scheme achieves dynamical isometry, improving network performance.
A new probabilistic method solves IVPs by inferring a latent path.
problem Solving initial value problems for ordinary differential equations.
method Formulates IVP solution as inference on a latent Gaussian process path.
result Probabilistic methods connect to classic ODE solvers.
Bayesian quadrature improves integration efficiency with invariant priors.
problem Efficient numerical integration with known structure.
method Invariance priors for bijective transformations in input domain.
result Superior performance in synthetic and real-world applications.
Study phase retrieval under misspecified models using generative priors.
problem Estimating signals from phase measurements with model misspecification.
method Two-step approach: spectral initialization followed by iterative refinement.
result Statistical rate of order ( k log L ) ⋅ ( log m ) / m \sqrt{(k\log L)\cdot (\log m)/m} ( k log L ) ⋅ ( log m ) / m under suitable conditions. Simplified Butterfly-Net2 improves CNN efficiency in solving PDEs and signal processing tasks.
problem Improving CNN efficiency in solving PDEs and signal processing tasks.
method Introducing BNet2, a simplified Butterfly-Net, and Fourier transform initialization.
result BNet2 achieves similar accuracy as CNN but with fewer parameters and improves accuracy over randomly initialized CNN.
We introduce a new method for training deep Boltzmann machines jointly. Prior methods require an initial learning pass that trains the deep Boltzmann machine greedily, one layer at a time, or do not perform well on classifi- cation tasks.
The article applies empirical Bayes to improve initial parameter choices in collaborative filtering models.
problem Improving initial parameter choices in collaborative filtering models.
method Formulated and implemented empirical Bayes to tune hyperparameters in a Bayesian collaborative filtering setup.
result Empirical Bayes can provide good initial parameter choices, especially for datasets where MCMC struggles.
AMP method reconstructs rank-one matrices from noisy data efficiently.
problem Reconstructing rank-one matrices with prior structural information from noisy observations.
method Approximate Message Passing (AMP) with random initialization.
result AMP from random initialization converges rapidly and globally.
MixTS uses a mixture prior to analyze Thompson Sampling in multi-task learning.
problem Analyzing Thompson Sampling in environments with uncertain and multi-class problems.
method Developed MixTS by incorporating a mixture prior into Thompson Sampling and using a novel proof technique for mixture distributions.
result Proved Bayes regret bounds for MixTS in linear bandits and finite-horizon reinforcement learning.
Diffusion models help in learning priors for Thompson Sampling in bandit problems.
problem Learning effective strategies for diverse bandit tasks.
method Training a denoising diffusion model to learn task distributions, combining with Thompson Sampling.
result The approach significantly improves performance across different bandit tasks.
Framework expands particle filtering to estimate states beyond prior boundaries.
problem Limitations of traditional particle filtering in estimating states outside prior support.
method Diffusion-Enhanced Particle Filtering Framework with adaptive diffusion, entropy-driven regularisation, and kernel-based perturbations.
result Framework significantly improves state estimation accuracy and success rates for out-of-boundary targets.
A novel diffusion method for Bayesian posterior sampling with theoretical guarantees.
problem Efficiently sampling from complex posterior distributions in Bayesian inversion.
method Diffusion-based posterior sampling using Langevin dynamics and PnP framework.
result The method converges even for multi-modal posterior distributions with theoretical error bounds.
We develop a scoring and classification procedure based on the PAC-Bayesian approach and the AUC (Area Under Curve) criterion. We focus initially on the class of linear score functions. We derive PAC-Bayesian non-asymptotic bounds for two types of prior for the score parameters: a Gaussian prior, and a spike-and-slab p…
Deep audio prior uses neural networks to solve audio problems without data.
problem Challenging audio problems like source separation, editing, and synthesis.
method Randomly-initialized neural network with carefully designed audio prior.
result Superior audio results on Universal-150 benchmark dataset.
EigenNoise provides a competitive word vector initialization scheme without pre-training data.
problem Improving word vector initialization without pre-training data.
method EigenNoise uses a dense, independent co-occurrence model to initialize word vectors.
result EigenNoise can approach GloVe performance without pre-training data.
Proposes a new method to initialize neural networks by estimating global curvature of weights.
problem Improving the initialization of neural networks for better training and convergence.
method Estimates the global curvature of weights across layers using the Hessian matrix norm.
result The proposed method helps in more rigorously initializing weights, leading to better performance.
VampPrior Mixture Model improves clustering in DLVMs.
problem Simplicity of standard priors in DLVMs leads to poor clustering performance.
method Leverages VampPrior concepts to fit a Bayesian GMM prior in a VAE.
result VMM achieves highly competitive clustering performance on benchmark datasets.
The paper addresses how to estimate initial conditions in kernel-based system identification.
problem The impact of initial conditions on system dynamics when few data samples are available.
method Three methods using maximum likelihood and a posteriori estimators to estimate initial conditions, with optimization via expectation-maximization.
result The proposed methods improve accuracy in reconstructing system impulse response compared to existing kernel-based schemes.
New network learns non-parametric invariances from data.
problem Modeling non-parametric invariances in data.
method Introduces PRC-NPTN networks with permanent random connectomes.
result Improves generalization and outperforms existing methods.
New method prunes neural networks at initialization, improving performance.
problem Improving neural network compression at initialization.
method Formally characterizes initialization conditions for reliable pruning based on connection sensitivity.
result Improved neural network performance on image classification tasks.
The paper analyzes how good initial guesses affect the amount of data needed for low-rank matrix recovery.
problem Theoretical guarantee of local optimization algorithms requires excessive data to prevent spurious local minima.
method Quantifies the relationship between initial guess quality and sample complexity using restricted isometry constant.
result A linear improvement in initial guess quality leads to a constant factor improvement in sample complexity.
Constructs initial data for Einstein equations and estimates Bartnik mass outside time-symmetry.
problem Estimating Bartnik mass outside time-symmetry.
method Constructs initial data for Einstein equations and connects Bartnik data to time-symmetric data.
result Obtains estimates for the Bartnik mass outside of time-symmetry.
Proposes a new Bayesian modeling framework.
problem Establishing principled priors and consolidating Bayesian analysis.
method Bayes via goodness of fit.
result Shows practical benefits of new approach.
Proposes learning task-agnostic dynamics priors for faster RL.
problem Challenges in learning accurate dynamics models for RL.
method Pre-training a frame predictor on physics videos to initialize and fine-tune dynamics models.
result Improves policy learning and convergence, outperforming competitors.
PGD algorithms solve nonlinear inverse problems with generative priors using noisy measurements.
problem Signal estimation from noisy nonlinear measurements with generative priors.
method Projected gradient descent algorithms for two cases: unknown and known nonlinearity.
result PGD algorithms converge linearly to optimal statistical rates using arbitrary initialization.
New method for initializing low-rank neural networks improves performance.
problem Training low-rank neural networks efficiently and accurately.
method Inspired by function approximation, proposes a novel low-rank initialization framework.
result Demonstrates significant gap between spectral and low-rank initialization approaches.
Gradient descent converges linearly for overparameterized linear networks.
problem Convergence of gradient descent for overparameterized neural networks.
method Local Polyak-Lojasiewicz and Descent Lemma for overparameterized linear models.
result Gradient descent achieves linear convergence for two-layer linear networks under relaxed assumptions.
Infinite arms bandit problem solved with confidence bounds.
problem Optimizing allocation in an infinite arms bandit problem with bounded rewards.
method Constructs confidence bounds for each arm and compares them against a target value to determine sampling.
result Achieves optimality for rewards with bounded means.
A new method maps high-dimensional Bayesian inverse problems to lower dimensions.
problem High-dimensional Bayesian inverse problems with complex prior information.
method Data-driven VAE prior and KRnet map for posterior approximation in latent space.
result Efficiently reduces computational cost and approximates posterior distributions.
New approach to neural networks by incorporating observation noise and arbitrary prior means.
problem Misspecification on noisy data and limitations of NTK-GP equivalence.
method Introducing a regularizer for observation noise and proposing a shifted network for arbitrary prior means.
result Removes key obstacles to practical Gaussian process modeling in neural networks.
Gradient descent and SGD achieve low test error in specific network weight regimes.
problem Optimizing two-layer ReLU networks with standard initialization.
method Gradient flow and stochastic gradient descent, analyzing margins and weight norms.
result Gradient descent and SGD can achieve globally maximal margins under certain constraints.
Maximal initial learning rate for deep ReLU networks identified.
problem Finding the optimal initial learning rate for deep neural networks.
method Simple approach to estimate maximal initial learning rate η ∗ η^{\ast} η ∗ , analyzing its behavior in constant-width fully-connected ReLU networks. result Maximal initial learning rate η ∗ η^{\ast} η ∗ is well predicted as a power of depth × width, with specific conditions for network width and input layer training.