Derives a new variational approach to information bottleneck.
problem Information theoretic inference and predictive modeling.
method Variational lower bound of predictive information bottleneck.
result Generalizes modern inference procedures and suggests new ones.
Technical proofs for Radon-Nikodym derivative identities.
problem Formalizing and proving theorems on Radon-Nikodym derivatives.
method Careful consideration of conditional and marginal probability measures.
result New interpretation of mutual and lattum information sums.
Batch Active Learning uses derivative information for Gaussian Process regression.
problem Efficiently selecting data batches in Gaussian Process regression models.
method Proposes using the predictive covariance matrix to select data batches, exploiting full correlation.
result Demonstrates the effectiveness of incorporating derivative information across diverse applications.
Improved price bounds for multi-asset derivatives using market option data.
problem Creating robust price bounds for multi-asset derivatives under market-implied dependence.
method Extracting inter-asset dependence information from market option prices and applying modified martingale optimal transport.
result Improved price bounds for multi-asset derivatives, demonstrating relevance and tractability.
New method scales Gaussian processes with derivatives using variational inference.
problem Scaling Gaussian processes with derivative information for high-dimensional problems.
method Introducing inducing directional derivatives to sparsify derivative information using variational inference.
result Achieves fully scalable Gaussian process regression with derivatives.
Paper derives best- and worst-case GlueVaR measures with incomplete data.
problem Risk measurement with limited information and shape constraints.
method Unified framework based on partial distribution information and shape properties.
result Characterization of extremal GlueVaR distributions with convex envelopes.
The paper establishes bounds for transductive learning using information theory.
problem Transductive learning generalization gap control.
method Information theory, PAC-Bayes, mutual information, conditional mutual information, different information measures.
result Established transductive information-theoretic and PAC-Bayesian bounds.
We study the pricing of credit derivatives with asymmetric information. The managers have complete information on the value process of the firm and on the default threshold, while the investors on the market have only partial observations, especially about the default threshold. Different information structures are dis…
Bayesian optimization has been successful at global optimization of expensive-to-evaluate multimodal objective functions. However, unlike most optimization methods, Bayesian optimization typically does not use derivative information. In this paper we show how Bayesian optimization can exploit derivative information to …
Study derives new equation for reserves in non-monotone information scenarios.
problem Modeling reserves in situations where information is not always increasing.
method Infinitesimal approach to derive generalized stochastic Thiele equation.
result New equation allows for information discarding and solves open problems.
We derive asset pricing formula for markets with incomplete information and subjective views.
problem Asset pricing in markets with informational imperfections and subjective investor beliefs.
method Closed-form market equilibrium formula based on Merton's model, non-linear system of equations, conditional posterior distribution.
result Derivation of market reference model for excess returns under random shadow-costs.
Bayesian optimization improved for nanophotonic device design.
problem Scalability and derivative information limitations in Bayesian optimization.
method Combining forward shape derivatives and iterative inversion scheme.
result Optimal designs of nanophotonic devices achieved with fewer iterations.
In this communication, we describe some interrelations between generalized q-entropies and a generalized version of Fisher information. In information theory, the de Bruijn identity links the Fisher information and the derivative of the entropy. We show that this identity can be extended to generalized versions of en…
New bounds derived using conditional f-information for machine learning models.
problem Improving generalization bounds in machine learning.
method Introducing novel information-theoretic generalization bounds via conditional f-information. result Derives generalization bounds applicable to both bounded and unbounded loss functions.
Paper derives bounds on prediction errors using information theory.
problem Understanding maximum prediction errors in sequential data.
method Information-theoretic approach focusing on conditional entropy.
result Fundamental bounds on prediction errors depend on conditional entropy.
Meta learning with information theory and Gaussian processes.
problem Few-shot learning problems.
method Information bottleneck, mutual information, variational approximations, Gaussian processes.
result Competitive accuracy on few-shot classification problems.
Bayesian optimization sped up with scalable Gaussian processes.
problem Optimizing functions with derivative information and large datasets.
method Combines derivative acceleration and scalable Gaussian process models.
result Significant speedup in optimization convergence for large datasets.
Min-cut clustering, based on minimizing one of two heuristic cost-functions proposed by Shi and Malik, has spawned tremendous research, both analytic and algorithmic, in the graph partitioning and image segmentation communities over the last decade. It is however unclear if these heuristics can be derived from a more g…
Complex dynamical systems driven by the unravelling of information can be modelled effectively by treating the underlying flow of information as the model input. Complicated dynamical behaviour of the system is then derived as an output. Such an information-based approach is in sharp contrast to the conventional mathem…
UMAP connects to Information Geometry principles.
problem None explicitly stated; focuses on connections.
method None explicitly stated; focuses on connections.
result UMAP has a natural geometric interpretation.
A pricing formula for discount bonds, based on the consideration of the market perception of future liquidity risk, is established. An information-based model for liquidity is then introduced, which is used to obtain an expression for the bond price. Analysis of the bond price dynamics shows that the bond volatility is…
In this paper, we show that feedforward and recurrent neural networks exhibit an outer product derivative structure but that convolutional neural networks do not. This structure makes it possible to use higher-order information without needing approximations or infeasibly large amounts of memory, and it may also provid…
This work improves generalisation bounds using chaining and information theory.
problem Improving generalisation bounds for supervised learning algorithms.
method Developed a theoretical framework linking generalisation bounds to their chained counterparts, derived new bounds using Wasserstein distance.
result Chained generalisation bounds can be tighter than standard bounds, especially for concentrated hypothesis distributions.
Neural networks compute bounds on multi-asset derivatives.
problem Computing precise prices of complex financial derivatives.
method Using neural networks and constrained optimal transport.
result Neural networks provide tighter bounds on derivative prices.
NGD improves multivariate Gaussian inference by optimizing Fisher information.
problem Efficiently optimizing multivariate Gaussian models.
method Natural Gradient Descent applied to multivariate Gaussian parameters.
result NGD updates are more efficient for symmetric covariance matrices.
We develop an entropic framework to model the dynamics of stocks and European Options. Entropic inference is an inductive inference framework equipped with proper tools to handle situations where incomplete information is available. The objective of the paper is to lay down an alternative framework for modeling dynamic…
Softmax emerges naturally in neural networks as a measure of conditional mutual information.
problem The artificial nature of softmax in neural networks.
method Information-theoretic perspective to derive log-softmax and evaluate conditional mutual information.
result Training deterministic neural networks through log-softmax maximises conditional mutual information.
We study information theoretic methods for ranking biomarkers. In clinical trials there are two, closely related, types of biomarkers: predictive and prognostic, and disentangling them is a key challenge. Our first step is to phrase biomarker ranking in terms of optimizing an information theoretic quantity. This formal…
Derivative-informed models improve financial surrogates for accurate hedging and risk management.
problem Developing fast surrogate models for financial derivatives and risk quantities.
method Derivative-informed operator learning framework combining neural operators, random features, and tangent sensitivity equations.
result The framework reduces hedging and risk errors by 40-76% compared to standard surrogates.
We estimate Radon-Nikodym derivatives using regularization in reproducing kernel Hilbert spaces.
problem Estimating Radon-Nikodym derivatives in various applications.
method General regularization scheme in reproducing kernel Hilbert spaces.
result High order accuracy in reconstructing Radon-Nikodym derivatives at any point.
A new algorithm speeds up neural network derivative calculations.
problem Exponential runtime of autodifferentiation for high-order derivatives in neural networks.
method n-TangentProp, a quasilinear algorithm for computing higher-order derivatives.
result Computes exact derivatives in quasilinear time, not exponential.
A typical goal of supervised dimension reduction is to find a low-dimensional subspace of the input space such that the projected input variables preserve maximal information about the output variables. The dependence maximization approach solves the supervised dimension reduction problem through maximizing a statistic…
We construct an infinite-dimensional information manifold based on exponential Orlicz spaces without using the notion of exponential convergence. We then show that convex mixtures of probability densities lie on the same connected component of this manifold, and characterize the class of densities for which this mixtur…
In this paper, we provide an information-theoretic interpretation of the Vector Quantized-Variational Autoencoder (VQ-VAE). We show that the loss function of the original VQ-VAE can be derived from the variational deterministic information bottleneck (VDIB) principle. On the other hand, the VQ-VAE trained by the Expect…
Lower bounds on Bayes risk for realizable models derived using information theory.
problem Deriving lower bounds on Bayes risk for realizable machine learning models.
method Information-theoretic analysis using rate-distortion theory and mutual information.
result Lower bounds on Bayes risk for realizable models, matching known bounds up to logarithmic factors.
The paper derives a new theorem for predicting batches of data.
problem Finding lower bounds on minimal batch regret.
method Derives a conditional version of the regret-capacity theorem.
result Reveals a connection between conditional Rényi divergence and conditional Sibson's mutual information.
This paper aims to propose a novel deep learning-integrated framework for deriving reliable simulation input models through incorporating multi-source information. The framework sources and extracts multisource data generated from construction operations, which provides rich information for input modeling. The framewor…
Bayesian active learning method improved for censored regression data.
problem Challenges in estimating BALD for censored regression data.
method Derived entropy and mutual information for censored distributions, developed C-BALD objective, proposed novel modelling approach. result Demonstrated C-BALD outperforms other methods in censored regression. Derivative-free method solves stochastic optimization problems with noisy objectives and constraints.
problem Solving nonlinear optimization problems with stochastic objectives and deterministic constraints using only zero-order information.
method Derivative-Free Stochastic Sequential Quadratic Programming (DF-SSQP) method using simultaneous perturbation stochastic approximation (SPSA) for gradient and Hessian estimation.
result Global almost-sure convergence of the DF-SSQP method under standard assumptions, with local asymptotic normality and statistical inference.
Study utility maximization with delayed information in continuous time Gaussian markets.
problem Maximizing utility with delayed information in continuous time Gaussian markets.
method Purely probabilistic approach based on Radon-Nikodym derivatives of Gaussian measures.
result Solution for optimal control and value in a specific Gaussian framework.
We show that gamma distributions provide models for departures from randomness since every neighbourhood of an exponential distribution contains a neighbourhood of gamma distributions, using an information theoretic metric topology. We derive also the information geometry of the 3-manifold of McKay bivariate gamma dist…
Upper bound derived for informed traders' gains in a model, akin to thermodynamics.
problem Informed traders' gains in a financial model with finite horizon.
method Bayesian inference and entropic inequality.
result Upper bound for expected gain, analogous to thermodynamics.
New bounds estimate learning algorithm performance using prediction information.
problem Estimating the performance of black-box learning algorithms.
method Information-theoretic bounds based on prediction information.
result Improved bounds applicable to deterministic algorithms and easier to estimate.
Optimal adversarial attacks minimize mutual information, revealing classifier vulnerabilities.
problem Designing optimal attacks to degrade machine learning performance.
method Information-theoretic approach to finding optimal perturbations.
result Optimal attacks minimize mutual information between degraded and original signals.
FisherNet extends Autoencoder using Fisher information for better data reconstruction.
problem Data reconstruction accuracy and model scalability in high-dimensional latent spaces.
method Introduces FisherNet architecture that uses Fisher information to quantify and account for latent space uncertainty.
result FisherNet produces more accurate reconstructions and scales better with latent space dimensions compared to VAE.
New bounds show limitations of sample-wise information-theoretic generalization.
problem Limitations of sample-wise information-theoretic generalization bounds.
method Analysis of existing bounds and derivation of new bounds.
result No sample-wise information-theoretic bounds exist for expected squared generalization gap.
The probability distribution function (PDF) for prices on financial markets is derived by extremization of Fisher information. It is shown how on that basis the quantum-like description for financial markets arises and different financial market models are mapped by quantum mechanical ones.
A one-factor asset pricing model with an Ornstein--Uhlenbeck process as its state variable is studied under partial information: the mean-reverting level and the mean-reverting speed parameters are modeled as hidden/unobservable stochastic variables. No-arbitrage pricing formulas for derivative securities written on a …