Dropout improves regularization in flexible models for rare features.
problem Understanding theoretical properties of dropout in generalized linear models.
method Theoretical analysis and application to adaptive smoothing with B-splines.
result Dropout prefers rare features in mean and dispersion parameters.
A framework connects VAEs to GLMs for better model initialization and performance.
problem Understanding and optimizing loss function critical points in VAEs.
method Introducing a theoretical framework based on GLM and EDFs.
result Maximum likelihood initialization improves VAE performance.
New calibration bands for various distributions improve testing for auto-calibration.
problem Testing for auto-calibration in finite samples is challenging.
method Construct calibration bands for the exponential dispersion family using finite sample properties.
result Calibration bands allow for various tests for calibration and auto-calibration.
We describe the underlying probabilistic interpretation of alpha and beta divergences. We first show that beta divergences are inherently tied to Tweedie distributions, a particular type of exponential family, known as exponential dispersion models. Starting from the variance function of a Tweedie model, we outline how…
Deep generative models are commonly used for generating images and text. Interpretability of these models is one important pursuit, other than the generation quality. Variational auto-encoder (VAE) with Gaussian distribution as prior has been successfully applied in text generation, but it is hard to interpret the mean…
Exponential dispersion model is a useful framework in machine learning and statistics. Primarily, thanks to the additive structure of the model, it can be achieved without difficulty to estimate parameters including mean. However, tight conditions on cumulant function, such as analyticity, strict convexity, and steepne…
Novel GLMMNet model tackles high-cardinality categorical features in actuarial applications.
problem Inadequate encoding methods for high-cardinality categorical features in actuarial data.
method Generalised Linear Mixed Model Neural Network (GLMMNet) integrating a generalised linear mixed model in a deep learning framework.
result GLMMNet often outperforms or performs comparably with entity embedded neural networks, providing transparency.
sgdGMF efficiently estimates generalized matrix factorization models for single-cell RNA sequencing data.
problem Challenges in dimensionality reduction for large single-cell RNA sequencing datasets.
method Scalable adaptive stochastic gradient descent algorithm for generalized matrix factorization models.
result sgdGMF outperforms existing methods in scalability and accuracy for large datasets.
Study global solutions for Boussinesq systems on curved manifolds.
problem Global existence and uniqueness of solutions to Boussinesq systems on non-compact Riemannian manifolds with gravitational fields.
method Used dispersive and smoothing estimates of a vectorial matrix semigroup to establish global existence and uniqueness of mild solutions for linear systems. Then, applied fixed point arguments to semilinear systems. Proved exponential stability using Gronwall's inequality.
result Established global existence, uniqueness, and exponential stability of mild solutions to the Boussinesq systems on non-compact Riemannian manifolds with gravitational fields.
Probabilistic modeling is cyclical: we specify a model, infer its posterior, and evaluate its performance. Evaluation drives the cycle, as we revise our model based on how it performs. This requires a metric. Traditionally, predictive accuracy prevails. Yet, predictive accuracy does not tell the whole story. We propose…
Study proves well-posedness and scattering for wave equations on hyperbolic spaces with singular data.
problem Proving well-posedness and scattering for wave equations on hyperbolic spaces with singular initial data.
method Using weak-Lp spaces and dispersive estimates on Lorentz spaces, the study establishes global well-posedness and exponential asymptotic stability. result Developed a scattering theory and constructed wave operators in a singular framework.
The paper discusses methods for interval estimation of coefficients in penalized regression models for insurance data.
problem Valid inference on coefficients after feature selection in GLM family for insurance data.
method Proposes methodologies for constructing confidence intervals of coefficients after feature selection in GLM family.
result Valid inference on coefficients after feature selection in GLM family for insurance data.
This paper generalizes beta divergence beyond its classical form associated with power variance functions of Tweedie models. Generalized form is represented by a compact definite integral as a function of variance function of the exponential dispersion model. This compact integral form simplifies derivations of many pr…
Paper introduces kernel deformed exponential families for sparse continuous attention.
problem Creating efficient attention mechanisms for sparse data.
method Developed kernel deformed exponential families, theoretically and experimentally.
result Kernel deformed exponential families can attend to multiple compact regions of data.
Modified EAT method improves Poisson gradient estimation.
problem Challenging differentiation through Poisson-distributed latent variables.
method Exponential Arrival Time (EAT) simulation with modifications and Gumbel-SoftMax relaxation.
result Modified EAT method provides unbiased first moment and reduced second-moment bias.
Exponential family distributions are highly useful in machine learning since their calculation can be performed efficiently through natural parameters. The exponential family has recently been extended to the t-exponential family, which contains Student-t distributions as family members and thus allows us to handle noi…
New method denoises images without clean reference using Tweedie distributions.
problem Image denoising without clean reference images.
method Combining Tweedie distributions, Noise2Score, and saddle point approximation.
result General closed-form denoising formula for various noise distributions.
Skewness dispersion predicts future stock market returns, especially in months with monetary policy announcements.
problem Predicting future stock market returns using skewness dispersion.
method Cross-sectional analysis of firm-level realized skewness and stock market returns.
result Skewness dispersion is a significant predictor of future stock market returns, robust to various estimation methods.
The study explores generalized divergences and exponential families with a focus on sufficient conditions and laws of large numbers.
problem Generalization of Kullback-Leibler divergence and exponential families.
method Investigation of (h,τ)-divergence and (h,τ)-exponential families, definition of (h,τ)-dependence, proof of law of large numbers. result Sufficient condition for (h,τ)-divergence to induce Hessian structure on (h,τ)-exponential family, proof of law of large numbers. Generalizes moment-matching for exponential families with conditioning or hidden data.
problem Generalizing moment-matching conditions for exponential families with conditioning or hidden data.
method First-principles explanation and self-contained derivation of generalized moment-matching conditions.
result Derives generalized moment-matching conditions for conditional exponential families and hidden data.
New Gini indices capture more nuanced income inequality.
problem Measuring joint dispersion across multiple observations.
method Axiomatic approach to define and characterize n-th order Gini deviations.
result Higher-order Gini coefficients reveal more extreme income disparities.
Correspondence found between exponential families and affine Grassmannians.
problem Understanding the relationship between exponential families and geometric structures.
method Established a one-to-one correspondence between exponential families and affine Grassmannians.
result Found a correspondence between minimal exponential families and affine Grassmannians.
New dispersion indices based on inaccuracy and divergence introduced for information measures.
problem Measuring variability in uncertainty measures.
method Introducing new dispersion indices based on Kerridge inaccuracy and Kullback-Leibler divergence.
result Properties, bounds, and examples of new dispersion indices presented.
We provide a classification of graphical models according to their representation as subfamilies of exponential families. Undirected graphical models with no hidden variables are linear exponential families (LEFs), directed acyclic graphical models and chain graphs with no hidden variables, including Bayesian networks …
Constructing exponential families from statistical manifolds.
problem The central problem of constructing exponential families from statistical manifolds.
method Constructive approach proving every compact statistical manifold admits a foliation of Hessian manifolds.
result Compact orientable leaves are either finite quotients of flat torus or mapping torus with periodic monodromy.
Thompson Sampling has been demonstrated in many complex bandit models, however the theoretical guarantees available for the parametric multi-armed bandit are still limited to the Bernoulli case. Here we extend them by proving asymptotic optimality of the algorithm using the Jeffreys prior for 1-dimensional exponential …
New heat dispersion laws established for smooth compact manifolds.
problem Understanding heat dispersion in smooth compact manifolds.
method Established new heat dispersion laws through Theorem 1.1 and explored them further with Propositions 3.1 and 3.2.
result New heat dispersion laws for smooth compact manifolds.
Moment polytope of toric exponential families is a projection of a simplex.
problem Understanding the geometry of exponential families in finite sample spaces.
method Toric torification and projection of higher-dimensional simplices.
result Moment polytope is a projection of a higher-dimensional simplex.
New Thompson sampling algorithm reduces regret for exponential family bandits.
problem Minimizing regret in multi-armed bandit problems with exponential family rewards.
method Proposes ExpTS and ExpTS+ algorithms using novel sampling distributions. result Minimizes both finite-time and asymptotic regret for exponential family rewards.
We study online learning under logarithmic loss with regular parametric models. Hedayati and Bartlett (2012b) showed that a Bayesian prediction strategy with Jeffreys prior and sequential normalized maximum likelihood (SNML) coincide and are optimal if and only if the latter is exchangeable, and if and only if the opti…
We propose a novel approach for density estimation with exponential families for the case when the true density may not fall within the chosen family. Our approach augments the sufficient statistics with features designed to accumulate probability mass in the neighborhood of the observed points, resulting in a non-para…
The versatility of exponential families, along with their attendant convexity properties, make them a popular and effective statistical model. A central issue is learning these models in high-dimensions, such as when there is some sparsity pattern of the optimal parameter. This work characterizes a certain strong conve…
Paper proposes a surrogate model for efficient experience rating in large insurance portfolios.
problem Inexpensive and transparent computation of Bayesian premiums for large insurance portfolios.
method Surrogate modeling approach using likelihood-based summary statistics.
result Reduced computational burden and provided a transparent way of computing Bayesian premiums.
New insights into natural exponential families improve regret bounds for bandit problems.
problem Improving regret bounds for bandit problems with subexponential tails.
method Proving self-concordance for natural exponential families and applying to bandits.
result Optimistic algorithms for generalized linear bandits have second-order regret bounds that are free of an exponential dependence on problem parameters.
MallowsPO enhances LLM fine-tuning with a dispersion index of human preferences.
problem Lack of diversity in human preferences in DPO.
method Developed a dispersion index based on Mallows' theory to characterize preference diversity.
result Demonstrated improved performance in various tasks using the dispersion index.
Extends likelihood ratio exponential families to analyze various optimization methods.
problem Analyzing optimization methods like rate-distortion and information bottleneck.
method Linking geometric mixture paths to exponential families and using hypothesis testing.
result Provides a common mathematical framework for understanding these methods.
In this paper we propose a novel index to quantify and measure the flow of information on macro and micro scales. We discuss the implications of this index for knowledge management fields and also as intellectual capital that can thus be utilized by entrepreneurs. We explore different function and human oriented metric…
Language models allocate information storage, not collapsing into uniform representations.
problem Incomplete neural collapse in language model representations.
method Analyzing variance and information sharing across 14 models, proving an information floor.
result Within-class variance is allocated information storage, not collapsed into uniform representations.
Geometric focusing affects dispersive estimates for Schrödinger and wave equations.
problem Long-time decay rate in dispersive estimates for Schrödinger and wave equations on non-trapping asymptotically conic manifolds and exact metric cones.
method Classifying the long-time decay rate in dispersive estimates for the Schrödinger and wave equations on non-trapping asymptotically conic manifolds and exact metric cones in terms of the intensity of geometric focusing.
result Each multiplicity of conjugate points within distance π on Y = ∂X0 leads to a |t|1/2-loss in the long-time decay order and a half-order shift in the regularity index in the dispersive estimate for the Schrödinger equation.
Efficient method for learning continuous exponential families beyond Gaussian.
problem Learning continuous exponential families with unbounded support.
method Interaction Screening approach for scalable learning of continuous graphical models.
result Our estimator maintains similar accuracy and sample complexity scalings compared to alternative approaches, while improving run-time.
EFDA extends LDA to non-Gaussian models using exponential families.
problem Classifying non-Gaussian data with LDA's limitations.
method EFDA uses exponential families to derive closed-form estimators for natural parameters and a linear decision rule.
result EFDA matches LDA's accuracy while reducing ECE by 2-6x, proving asymptotic calibration and efficiency.
Study on stock market volatility and return dispersion during COVID-19.
problem Impact of COVID-19 on stock market volatility and return dispersion.
method Used Google index to proxy epidemic impact, modeled volatility, and analyzed influencing factors of log-return.
result Volatility significantly affected by epidemic and cross-sectional return dispersion, with positive coefficients.
Orthogonal random features approximate a Bessel kernel, offering sharper bounds than random Fourier features.
problem Approximating Gaussian kernel efficiently for large datasets.
method Use of Haar orthogonal matrices to construct orthogonal random features and analyze their bias and variance.
result Orthogonal random features approximate a Bessel kernel, not the Gaussian kernel, with sharper bounds.
We derive a new variational principle, leading to a new momentum map and a new multisymplectic formulation for a family of Euler--Poincaré equations defined on the Virasoro-Bott group, by using the inverse map (also called `back-to-labels' map). This family contains as special cases the well-known Korteweg-de Vries, Ca…
In the recent years, banks have sold structured products such as worst-of options, Everest and Himalayas, resulting in a short correlation exposure. They have hence become interested in offsetting part of this exposure, namely buying back correlation. Two ways have been proposed for such a strategy : either pure correl…
New bounds for score matching in polynomial exponential families.
problem Understanding the sample complexity of score matching for polynomial exponential families.
method Non-asymptotic sample complexity analysis for score matching.
result First finite sample bounds for score matching in polynomial exponential families.
Exponential family extensions of principal component analysis (EPCA) have received a considerable amount of attention in recent years, demonstrating the growing need for basic modeling tools that do not assume the squared loss or Gaussian distribution. We extend the EPCA model toolbox by presenting the first exponentia…
The study examines Hawkes processes and their long-term behavior.
problem Understanding the long-term behavior of Hawkes processes.
method Proving functional limit theorems under various conditions on the dispersion of child events.
result Functional limit theorems hold for Hawkes processes with different levels of child event dispersion.