Bayesian models' singular fluctuation is shown to be akin to specific heat, influencing model complexity and generalization.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We introduce thermodynamic response functions for singular Bayesian models.
A widely applicable Bayesian information criterion (Watanabe, 2013) is applicable for both regular and singular models in the model selection problem. This criterion tends to overestimate the log marginal likelihood. We identify an overestimating term of a widely applicable Bayesian information criterion. Adjustment of…
Bayesian networks are now being used in enormous fields, for example, diagnosis of a system, data mining, clustering and so on. In spite of their wide range of applications, the statistical properties have not yet been clarified, because the models are nonidentifiable and non-regular. In a Bayesian network, the set of …
New method corrects Laplace/BIC errors in singular models, revealing effective dimension.
Quantum statistical models with singularities are studied for state estimation and model selection.
The paper derives an equation linking WAIC and WBIC for singular models.
Advances variational Bayesian neural networks using singular learning theory.
LS improves model selection for singular statistical models.
Bayesian neural networks can be simplified by parameterizing weights as rank- matrices, reducing parameter count and improving performance.
Bayesian free energy remains bounded for deep ReLU networks in overparametrized cases.
Low-rank matrix estimation from incomplete measurements recently received increased attention due to the emergence of several challenging applications, such as recommender systems; see in particular the famous Netflix challenge. While the behaviour of algorithms based on nuclear norm minimization is now well understood…
Evaluation of the marginal likelihood plays an important role in model selection problems. The widely applicable Bayesian information criterion (WBIC) and singular Bayesian information criterion (sBIC) give approximations to the log marginal likelihood, which can be applied to both regular and singular models. When the…
We present and implement two algorithms for analytic asymptotic evaluation of the marginal likelihood of data given a Bayesian network with hidden nodes. As shown by previous work, this evaluation is particularly hard for latent Bayesian network models, namely networks that include hidden variables, where asymptotic ap…
A statistical model or a learning machine is called regular if the map taking a parameter to a probability distribution is one-to-one and if its Fisher information matrix is always positive definite. If otherwise, it is called singular. In regular statistical models, the Bayes free energy, which is defined by the minus…
Study clarifies Bayesian generalization error in CBM for 3-layered linear neural networks.
Gaussian latent tree models, or more generally, Gaussian latent forest models have Fisher-information matrices that become singular along interesting submodels, namely, models that correspond to subforests. For these singularities, we compute the real log-canonical thresholds (also known as stochastic complexities or l…
This paper studies a specific blow-up algorithm for sop polynomials and their RLCT.
The sBIC outperforms other model selection criteria in LDA topic modeling.
Singularities of a statistical model are the elements of the model's parameter space which make the corresponding Fisher information matrix degenerate. These are the points for which estimation techniques such as the maximum likelihood estimator and standard Bayesian procedures do not admit the root- parametric rate…
This paper develops a Bayesian procedure for estimation and forecasting of the volatility of multivariate time series. The foundation of this work is the matrix-variate dynamic linear model, for the volatility of which we adopt a multiplicative stochastic evolution, using Wishart and singular multivariate beta distribu…
Paper calculates the exact error of LDA models.
Paper bounds tensor decomposition's RLCT, aiding Bayesian inference.
Hierarchical learning models, such as mixture models and Bayesian networks, are widely employed for unsupervised learning tasks, such as clustering analysis. They consist of observable and hidden variables, which represent the given data and their hidden generation process, respectively. It has been pointed out that co…
PCBM improves neural network generalization by partially observing concepts.
New method estimates high-dimensional GoM models efficiently.
A Bayesian procedure is developed for multivariate stochastic volatility, using state space models. An autoregressive model for the log-returns is employed. We generalize the inverted Wishart distribution to allow for different correlation structure between the observation and state innovation vectors and we extend the…
Estimates LLC for deep linear networks up to 100M parameters.
SLT reveals how grokking occurs via basin selection in training.
New Bayesian matrix completion method using Stiefel manifolds.
Proposes a new algorithm for Sparse Bayesian Learning connected to Stepwise Regression.
A new geometric concept, the dead direction, bridges singular learning theory and information geometry.
Local mass perspective on Bayesian inference
Dead-Direction Signatures (DDS) provide a cheap, closed-form spectral reading of a network's singular complexity.
Latent Dirichlet allocation (LDA) is useful in document analysis, image processing, and many information systems; however, its generalization performance has been left unknown because it is a singular learning machine to which regular statistical theory can not be applied. Stochastic matrix factorization (SMF) is a res…
In this paper we develop a Bayesian procedure for estimating multivariate stochastic volatility (MSV) using state space models. A multiplicative model based on inverted Wishart and multivariate singular beta distributions is proposed for the evolution of the volatility, and a flexible sequential volatility updating is …
Dropout, a stochastic regularisation technique for training of neural networks, has recently been reinterpreted as a specific type of approximate inference algorithm for Bayesian neural networks. The main contribution of the reinterpretation is in providing a theoretical framework useful for analysing and extending the…
We propose a general framework for reduced-rank modeling of matrix-valued data. By applying a generalized nuclear norm penalty we can directly model low-dimensional latent variables associated with rows and columns. Our framework flexibly incorporates row and column features, smoothing kernels, and other sources of sid…
Study the geometry of matrix multiplication in deep neural networks.
In this letter, we consider two sets of observations defined as subspace signals embedded in noise and we wish to analyze the distance between these two subspaces. The latter entails evaluating the angles between the subspaces, an issue reminiscent of the well-known Procrustes problem. A Bayesian approach is investigat…
A new data-adaptive prior stabilizes kernel learning in operators.
Simple method for estimating missing panel data entries with confidence intervals.
In this paper we investigate the singularities of Lagrangian mean curvature flows in by means of smooth singularity models. Type I singularities can only occur at certain times determined by invariants in the cohomology of the initial data. In the type II case, these smooth singularity models are asympto…
A rising topic in computational journalism is how to enhance the diversity in news served to subscribers to foster exploration behavior in news reading. Despite the success of preference learning in personalized news recommendation, their over-exploitation causes filter bubble that isolates readers from opposing viewpo…
A new method for anomaly detection using random subspaces and Gaussian mixture models.
A new algorithm improves posterior sampling for linear inverse problems.
The Kähler-Ricci flow near conical singularities is described with a curvature bound.
Improved graph neural network bounds using graph diffusion matrix.