New method samples Jeffreys prior for objective Bayesian inference.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New optimal prior avoids bias in complex models with limited data.
Informative Bayesian priors are often difficult to elicit, and when this is the case, modelers usually turn to noninformative or objective priors. However, objective priors such as the Jeffreys and reference priors are not tractable to derive for many models of interest. We address this issue by proposing techniques fo…
A new Weyl prior is proposed for Bayesian statistics, offering a more canonical choice for parameter α.
Optimality of TS with noninformative priors proven for Pareto model.
Thompson Sampling has been demonstrated in many complex bandit models, however the theoretical guarantees available for the parametric multi-armed bandit are still limited to the Bernoulli case. Here we extend them by proving asymptotic optimality of the algorithm using the Jeffreys prior for 1-dimensional exponential …
We study online learning under logarithmic loss with regular parametric models. Hedayati and Bartlett (2012b) showed that a Bayesian prediction strategy with Jeffreys prior and sequential normalized maximum likelihood (SNML) coincide and are optimal if and only if the latter is exchangeable, and if and only if the opti…
We construct geometric shrinkage priors for Kählerian signal filters. Based on the characteristics of Kähler manifolds, an efficient and robust algorithm for finding superharmonic priors which outperform the Jeffreys prior is introduced. Several ansätze for the Bayesian predictive priors are also suggested. In particul…
Jeffrey guidance extends diffusion-model control to more complex applications.
Jeffreys Flow improves robustness of Boltzmann generators for rare event sampling.
This study explores how choosing noninformative priors affects Thompson Sampling in multiparameter bandit models.
Paper proves Jeffrey's update rule minimizes relative entropy.
We propose a generalized double Pareto prior for Bayesian shrinkage estimation and inferences in linear models. The prior can be obtained via a scale mixture of Laplace or normal distributions, forming a bridge between the Laplace and Normal-Jeffreys' priors. While it has a spike at zero like the Laplace density, it al…
Bayesian method improves extreme quantile estimation with zero coverage error.
Due to the success of the bag-of-word modeling paradigm, clustering histograms has become an important ingredient of modern information processing. Clustering histograms can be performed using the celebrated -means centroid-based algorithm. From the viewpoint of applications, it is usually required to deal with symm…
We use the language of uninformative Bayesian prior choice to study the selection of appropriately simple effective models. We advocate for the prior which maximizes the mutual information between parameters and predictions, learning as much as possible from limited data. When many parameters are poorly constrained by …
Develops a statistical test for IV, improving feature selection reliability.
Bayesian neural networks update beliefs with soft evidence, improving accuracy and calibration.
In this paper, we derive a Bayesian model order selection rule by using the exponentially embedded family method, termed Bayesian EEF. Unlike many other Bayesian model selection methods, the Bayesian EEF can use vague proper priors and improper noninformative priors to be objective in the elicitation of parameter prior…
We review the information geometry of linear systems and its application to Bayesian inference, and the simplification available in the Kähler manifold case. We find conditions for the information geometry of linear systems to be Kähler, and the relation of the Kähler potential to information geometric quantities such …
Consider the space of rational functions of several variables with poles on a fixed arrangement of hyperplanes. We obtain a decomposition of as a module over the ring of differential operators with constant coefficients. We generalize to the space the notions of principal part and of residue, and …
We consider the estimation of the multi-period optimal portfolio obtained by maximizing an exponential utility. Employing Jeffreys' non-informative prior and the conjugate informative prior, we derive stochastic representations for the optimal portfolio weights at each time point of portfolio reallocation. This provide…
Automated feature selection is important for text categorization to reduce the feature size and to speed up the learning process of classifiers. In this paper, we present a novel and efficient feature selection framework based on the Information Theory, which aims to rank the features with their discriminative capacity…
This study investigates self-supervised learning with Wasserstein distance on tree structures.
The paper explores how to handle uncertain evidence in probabilistic models.
MsIGN tackles high-dimensional Bayesian inference using multiscale structure.
We propose a method for recovering the structure of a sparse undirected graphical model when very few samples are available. The method decides about the presence or absence of bonds between pairs of variable by considering one pair at a time and using a closed form formula, analytically derived by calculating the post…
We announce the following result and give several applications: A Hamiltonian -space (for a torus) with isolated fixed points is cobordant to a disjoint union of weighted projective spaces which are constructed from its fixed point data. The applications concern the Duistermaat-Heckman formula, the topological J…
Jeffrey and Kirwan suggested expressions for intersection pairings on the reduced space of a Hamiltonian G-space in terms of multiple residues. In this paper we prove a residue formula for symplectic volumes of reduced spaces of a quasi-Hamiltonian SU(2)-space. The definition of quasi-Hamiltonian G-spaces was recently …
We propose a Bayesian expectation-maximization (EM) algorithm for reconstructing Markov-tree sparse signals via belief propagation. The measurements follow an underdetermined linear model where the regression-coefficient vector is the sum of an unknown approximately sparse signal and a zero-mean white Gaussian noise wi…
We show that if the connected sum of two knots with coprime Alexander polynomials is doubly slice, then the Ozsváth-Szabó correction terms as smooth double sliceness obstructions vanish for both knots. Recently, Jeffrey Meier gave smoothly slice knots that are topologically doubly slice, but not smoothly doubly slice. …
Let be a smooth manifold and a compact connected Lie group acting on by isometries. In this paper, we study the equivariant cohomology of , and relate it to the cohomology of the Marsden-Weinstein reduced space via certain residue formulae. In case that is a compact symplectic mani…
Using the notion of equivariant Kirwan map, as defined by Goldin, we prove that -- in the case of Hamiltonian torus actions with isolated fixed points -- Tolman and Weitsman's description of the kernel of the Kirwan map can be deduced directly from the residue theorem of Jeffrey and Kirwan. A characterization of the ke…
The Global Vectors for word representation (GloVe), introduced by Jeffrey Pennington et al. is reported to be an efficient and effective method for learning vector representations of words. State-of-the-art performance is also provided by skip-gram with negative-sampling (SGNS) implemented in the word2vec tool. In this…
This work is a continuation of our previous paper arXiv:1812.06473 where we have constructed supersymmetric Yang-Mills theory on 4D manifolds with a Killing vector field with isolated fixed points. In this work we expand on the mathematical aspects of the theory, with a particular focus on its nature as a …
We compare two models of corporate default by calculating the Jeffreys-Kullback-Leibler divergence between their predicted default probabilities when asset correlations are either high or low. Our main results show that the divergence between the two models increases in highly correlated, volatile, and large markets, b…
Lisa Jeffrey and Frances Kirwan developed an integration theory for symplectic reductions. That is, given a symplectic manifold with symplectic group action, they developed a way of pulling the integration of forms on the reduction back to an integration of group-equivariant forms on the original space. We seek an anal…
We prove an analogue of the Atiyah-Bott-Berline-Vergne localization formula in the setting of equivariant basic cohomology of -contact manifolds. As a consequence, we deduce analogues of Witten's nonabelian localization and the Jeffrey-Kirwan residue formula, which relate equivariant basic integrals on a contact man…
We define a moment map associated to a smooth torus action on a smooth manifold, without a two-form. We define cobordisms of such structures, allowing non compact manifolds as long as the moment maps are proper. We prove that a compact manifold with a torus action and a moment map is cobordant to the disjoint union of …
We prove that the Grothendieck-Springer simultaneous resolution viewed as a correspondence between the adjoint quotient of a Lie algebra and its maximal torus is Lagrangian in the sense of shifted symplectic structures. As Hamiltonian spaces can be interpreted as Lagrangians in the adjoint quotient, this allows one to …
Goldman parametrizes the -Hitchin component of a closed oriented hyperbolic surface of genus by parameters. Among them, coordinates are canonical. We prove that the -Hitchin component equipped with the Atiyah-Bott-Goldman symplectic form admi…
In this article we prove that iterated renormalisations of circle diffeomorphisms with breaks, , with given size of breaks, converge to an invariant family of piecewise Moebius maps, of dimension . We prove that this invariant family identifies with a \textit{relative character variety} $χ(…
We develop a Chern-Weil theory for compact Lie group action whose generic stabilizers are finite in the framework of equivariant cohomology. This provides a method of changing an equivariant closed form within its cohomological class to a form more suitable to yield localization results. This work is motivated by our w…
We study the counting function of topological Poincaré series associated with rational homology sphere plumbed 3-manifold with connected negative definite tree, interpreting as an alternating sum of coefficient functions associated with some Taylor expansions. It is motivated by a theorem of Szenes and Vergne which exp…
The word2vec software of Tomas Mikolov and colleagues (https://code.google.com/p/word2vec/ ) has gained a lot of traction lately, and provides state-of-the-art word embeddings. The learning models behind the software are described in two research papers. We found the description of the models in these papers to be some…
The space of all based loops in a compact semisimple simply connected Lie group has an action of the maximal torus (by pointwise conjugation) and of the circle (by rotation of loops). Let $μ: Ω(G)\to (\t\times i\mathbb{R})^*$ be a moment map of the resulting action. We show t…
We establish a splitting formula for the spectral flow of the odd signature operator on a closed 3-manifold M coupled to a path of SU(2) connections, provided M = S cup X, where S is the solid torus. It describes the spectral flow on M in terms of the spectral flow on S, the spectral flow on X (with certain Atiyah-Pato…
Artificial neural network training with stochastic gradient descent can be destabilized by "bad batches" with high losses. This is often problematic for training with small batch sizes, high order loss functions or unstably high learning rates. To stabilize learning, we have developed adaptive learning rate clipping (A…