Most simple braids have positive topological entropy.
problem Understanding the topological entropy of simple braids.
method Reduction from simple braids to non-simple 3-strand braids.
result The proportion of simple braids with positive entropy approaches 100% as the number of strands increases.
Entropy derived from Colding's volume on Ricci-flat manifolds.
problem Deriving Perelman's entropy from Colding's monotonic volume.
method Applying Colding's monotonic volume to Perelman's N-space for harmonic functions on Ricci-flat manifolds.
result Entropy is the limit of Colding's monotonic volume.
In this paper, we first introduce the weighted forward reduced volume of Ricci flow. The weighted forward reduced volume, which related to expanders of Ricci flow, is well-defined on noncompact manifolds and monotone non-increasing under Ricci flow. Moreover, we show that, just the same as the Perelman's reduced volume…
EVODiff optimizes DM inference by reducing conditional entropy, improving image generation.
problem Slow and inaccurate inference in diffusion models.
method Entropy-aware variance optimization for efficient inference.
result Significant improvement in image generation quality and efficiency.
Improved exploration methods for reinforcement learning with reduced sample complexity.
problem Challenges in reinforcement learning exploration in unknown environments.
method Proposed game-theoretic and trajectory entropy algorithms with improved sample complexity.
result Established statistical advantage of entropy-regularized MDPs for exploration and reduced sample complexity.
Efficient approximations reduce computation of matrix-based Renyi's entropy.
problem High computational complexity of matrix-based Renyi's entropy.
method Taylor, Chebyshev, and Lanczos approximations to reduce complexity.
result Reduced complexity to significantly less than O(n2) with negligible accuracy loss. The topological entropy of a braid is the infimum of the entropies of all homeomorphisms of the disc which have a finite invariant set represented by the braid. When the isotopy class represented by the braid is pseudo-Anosov or is reducible with a pseudo-Anosov component, this entropy is positive. Fried and Kolev prov…
Efficiently constructs sparse ROMs for high-dimensional data using causation entropy.
problem Creating effective reduced-order models for high-dimensional dynamical data.
method Uses causation entropy to identify important terms and construct ROMs with varying sparsity.
result Demonstrates the effectiveness of causation entropy in constructing sparse ROMs for chaotic systems with skewed statistics.
We show that a certain entropy-like function is convex, under an optimal transport problem that is adapted to Ricci flow. We use this to reprove the monotonicity of Perelman's reduced volume.
Entropy corrections improve GBM's predictive accuracy for non-log-normal distributions.
problem Log-normal distribution limitations in GBM predictions.
method Entropy corrections to geometric Brownian motion (GBM).
result Improved predictive accuracy for non-log-normal distributions.
A new method reduces compounding errors in model-based reinforcement learning.
problem Compounding errors in long horizon predictions from model-based reinforcement learning.
method Maximum Entropy Model Rollouts (MEMR) with non-uniform sampling and prioritized experience replay.
result Significantly reduces computation requirements compared to other model-based methods.
Paper proves Jeffrey's update rule minimizes relative entropy.
problem Improving Bayesian learning algorithms.
method More concise proof of Jeffrey's update rule.
result Jeffrey's update rule reduces relative entropy.
We explore a new method for discrete-time control problems using randomization and entropy.
problem Discrete-time linear-exponential quadratic Gaussian (LEQG) control problem.
method Introduce exploration through randomization and apply duality between free energy and relative entropy.
result Reduced LEQG problem to equivalent risk-neutral LQG control problem with entropy regularization.
AR-DAE approximates entropy gradient for machine learning models.
problem Intractable computation of entropy gradient for continuous distributions.
method Amortized residual denoising autoencoder (AR-DAE) to approximate entropy gradient.
result AR-DAE provides an unbiased gradient approximation for entropy.
Fine-tuning improves information conveyance in language models by reorganizing uncertainty into more informative sequences.
problem Uncertainty reduction in large language models through fine-tuning is not fully understood, especially regarding output length.
method Proposed Canopy Entropy (CE⋆) to measure uncertainty in both output length and sequence, capturing total Shannon entropy. result Fine-tuned models exhibit stronger positive correlation between entropy rate and semantic diversity, indicating more informative and semantically meaningful generations.
At the core of any inference procedure in deep neural networks are dot product operations, which are the component that require the highest computational resources. A common approach to reduce the cost of inference is to reduce its memory complexity by lowering the entropy of the weight matrices of the neural network, …
Study uses holography to analyze entanglement entropy in deformed CFTs.
problem Analyzing entanglement entropy in $Tar{T}$-deformed CFTs.
method Holographic methods and direct bulk gravitational action evaluation.
result Agreement with known results for entanglement entropy.
The paper generalizes Bayesian Cramér-Rao inequality using information geometry of relative α-entropy.
problem Establishing a lower bound for the variance of an unbiased estimator for the α-escort distribution.
method Proposes a general Riemannian metric based on relative α-entropy to derive a generalized Bayesian Cramér-Rao inequality.
result Establishes a lower bound for the variance of an unbiased estimator for the α-escort distribution.
Deep learning speeds up protein mapping entropy calculation.
problem Efficiently calculating the mapping entropy of protein structures.
method Deep graph networks for accelerating mapping entropy computation.
result Deep graph networks achieve a speedup factor of up to 10^5.
AECF improves multimodal inference robustness and calibration.
problem Robustness and calibration issues in multimodal systems with missing inputs.
method Adaptive Entropy-Gated Contrastive Fusion (AECF) layer.
result Improves masked-input mAP by +18 pp at a 50% drop rate.
In our previous work we showed that for an ancient solution to the Ricci flow with nonnegative curvature operator, assuming bounded geometry on one time slice, bounded entropy implies noncollapsing on all scales. In this paper we prove the implication in the other direction, that for an ancient solution with bounded no…
The paper studies entropy calibration in language models and finds that miscalibration improves slowly with scale.
problem The problem is whether language model entropy calibration improves with scale and if it's possible to calibrate without reducing log loss.
method The authors study a simplified theoretical setting to characterize miscalibration scaling behavior and measure it empirically in language models ranging from 0.5B to 70B parameters.
result The observed scaling behavior of miscalibration is similar to theoretical predictions, indicating slow improvement with scale. The authors also prove theoretically that it is possible to reduce entropy while preserving log loss if access to a black box predicting future entropy is available.
ANTIDOTE reduces noisy labels influence during learning.
problem Learning with noisy labels.
method Information-divergence neighborhood relaxation and adversarial training.
result ANTIDOTE outperforms standard cross-entropy loss in noisy label settings.
The paper introduces a new intrinsic reward method for exploration in reinforcement learning.
problem Improving exploration in reinforcement learning agents.
method Intrinsic rewards proportional to the entropy of future state-action features.
result The new objective leads to improved visitation of features within individual trajectories.
New regularization method reduces support of empirical risk minimization solutions.
problem Regularization in empirical risk minimization with relative entropy.
method Introduces Type-II regularization, characterizes solutions, analyzes properties of relative entropy.
result Type-II regularization collapses solution support into reference measure's support.
Perelman has discovered two integral quantities, the shrinker entropy $\cW$ and the (backward) reduced volume, that are monotone under the Ricci flow $\pa g_{ij}/\pa t=-2R_{ij}$ and constant on shrinking solitons. Tweaking some signs, we find similar formulae corresponding to the expanding case. The {\it expanding entr…
New method uncovers zero entropy in dependent observations after finite samples.
problem Understanding uncertainty reduction in dependent observations.
method Minimum list entropy coupling, greedy algorithm.
result Zero entropy achieved with O(log(1/P_min)) samples for dependent observations.
Tent adapts models during testing by minimizing entropy of predictions.
problem Adapting models to new data during testing with limited information.
method Test entropy minimization (tent) and online channel-wise affine transformations.
result Reduces generalization error on various datasets and benchmarks.
Compressed Counting (CC) [22] was recently proposed for estimating the ath frequency moments of data streams, where 0 < a <= 2. CC can be used for estimating Shannon entropy, which can be approximated by certain functions of the ath frequency moments as a -> 1. Monitoring Shannon entropy for anomaly detection (e.g., DD…
We propose a general framework for neural network compression that is motivated by the Minimum Description Length (MDL) principle. For that we first derive an expression for the entropy of a neural network, which measures its complexity explicitly in terms of its bit-size. Then, we formalize the problem of neural netwo…
Insider trading is reduced when penalized, affecting expected penalties in a non-monotone way.
problem Reducing insider trading behavior when insiders face legal penalties.
method Characterized via a backward stochastic differential equation (BSDE) with a non-linear operator.
result The insider's expected penalties are non-monotone in the fee structure and determined by relative entropy.
The R Package CEC performs clustering based on the cross-entropy clustering (CEC) method, which was recently developed with the use of information theory. The main advantage of CEC is that it combines the speed and simplicity of k-means with the ability to use various Gaussian mixture models and reduce unnecessary cl…
We approximate differential entropy for efficient Bayesian experimental design.
problem Efficiently estimating expected information gain in large-scale inference problems.
method Approximate differential entropy using Monte Carlo or quasi-Monte Carlo surrogates.
result Our approach achieves comparable or better convergence rates than state-of-the-art methods.
We prove a Margulis' Lemma à la Besson Courtois Gallot, for manifolds whose fundamental group is a nontrivial free product A*B, without 2-torsion. Moreover, if A*B is torsion-free we give a lower bound for the homotopy systole in terms of upper bounds on the diameter and the volume entropy. We also provide examples and…
The ability of many powerful machine learning algorithms to deal with large data sets without compromise is often hampered by computationally expensive linear algebra tasks, of which calculating the log determinant is a canonical example. In this paper we demonstrate the optimality of Maximum Entropy methods in approxi…
Bayesian neural networks learn efficiently at infinite width, matching polynomial-width performance.
problem Understanding the inductive bias of infinite-width neural networks.
method Analyzing the reduced entropy and using subsampling techniques.
result The Bayesian mean-field learner generalizes exactly on polynomially-bounded targets.
GAN+VER improves GANs by regularizing entropy to reduce mode collapse.
problem Mode collapse in GANs where the generator fails to capture all modes.
method Maximizing a variational lower bound on the entropy of generated samples.
result Significant improvement in evaluation metrics for real and generated samples.
We study large-scale kernel methods for acoustic modeling and compare to DNNs on performance metrics related to both acoustic modeling and recognition. Measuring perplexity and frame-level classification accuracy, kernel-based acoustic models are as effective as their DNN counterparts. However, on token-error-rates DNN…
Proposes a method to solve deep neural networks' local minimum problem.
problem Local minimum problem in deep neural networks training.
method Transforms cross-entropy loss into risk-averse error criterion, adjusts RSI, and uses convexity region.
result Trained deep learning machine is expected to be inside a global minimum's attraction basin.
This paper provides efficient algorithms for computing entropy and KL divergence in Bayesian networks.
problem Computing entropy and KL divergence for Bayesian networks efficiently.
method Leveraging the graphical structure of Bayesian networks, the paper provides computationally efficient algorithms.
result Reduces computational complexity of KL divergence from cubic to quadratic for Gaussian BNs.
W-entropy and reduced volume for the Ricci flow were introduced by Perelman, which had proved their importance in the study of the Ricci flow. L. Ni studied the analogous concepts for the linear heat equation on the static manifolds, and established an equation which links the large time behavior of these t…
This paper generalizes BO uncertainty measures using decision-theoretic entropies.
problem Efficiently inferring optima of expensive black-box functions.
method Introduces a generalized entropy measure from statistical decision theory to optimize Bayesian optimization.
result Demonstrates strong empirical performance across various sequential decision-making tasks.
In this communication, we describe some interrelations between generalized q-entropies and a generalized version of Fisher information. In information theory, the de Bruijn identity links the Fisher information and the derivative of the entropy. We show that this identity can be extended to generalized versions of en…
Debiased Wasserstein barycenters improve on entropy regularization in OT.
problem Entropy regularization in OT introduces bias, leading to blurred barycenters.
method Propose debiased Wasserstein barycenters using Sinkhorn iterations.
result Debiased barycenters preserve fast Sinkhorn-like iterations without entropy smoothing bias.
AdaDEM decouples EM into two parts to improve class overlap and uncertainty.
problem Improper EM limits its effectiveness in various machine learning tasks.
method Decouple EM into CADF and GMC, and AdaDEM normalizes CADF reward and uses MEC.
result AdaDEM outperforms classical EM and improves performance in noisy and dynamic environments.
Unified framework connects EI and information-theoretic acquisition functions.
problem Distinguish between Expected Improvement and information-theoretic acquisition functions.
method Introduces Variational Entropy Search (VES) to unify EI and information-theoretic approaches.
result EI can be seen as a variational inference approximation of Max-value Entropy Search (MES).
Algorithm learns Nash equilibria in stochastic games using entropy-regularized policies.
problem Learning Nash equilibria in zero-sum stochastic games is computationally expensive.
method Entropy-regularized soft policies for Q-function updates.
result Algorithm converges to Nash equilibrium under certain conditions.
Trust-region methods have yielded state-of-the-art results in policy search. A common approach is to use KL-divergence to bound the region of trust resulting in a natural gradient policy update. We show that the natural gradient and trust region optimization are equivalent if we use the natural parameterization of a st…