Entropy-minimal measure calculated for a stochastic volatility model.
problem Calculating the entropy-minimal equivalent martingale measure in a stochastic volatility model.
method Revised related theory, calculated entropy-minimal measure.
result Entropy-minimal measure for the exponential Ornstein-Uhlenbeck model.
Proposes MEDM to balance entropy minimization and diversity maximization for better domain adaptation.
problem Trivial solutions in entropy minimization for unsupervised domain adaptation.
method Introduces diversity maximization to balance with entropy minimization, controlled by deep embedded validation.
result MEDM outperforms state-of-the-art methods on four domain adaptation datasets.
The study proves compactness and existence of entropy minimizers for self-shrinking surfaces.
problem Understanding entropy in higher-codimension mean curvature flow.
method Measure-theoretical techniques and rigidity results for self-shrinkers.
result Existence of entropy minimizers and improved rigidity results.
Tent adapts models during testing by minimizing entropy of predictions.
problem Adapting models to new data during testing with limited information.
method Test entropy minimization (tent) and online channel-wise affine transformations.
result Reduces generalization error on various datasets and benchmarks.
Constructs entropy-minimizing pseudo-Anosov diffeomorphisms on K3 surfaces.
problem Finding minimal entropy diffeomorphisms on K3 surfaces.
method Constructs pseudo-Anosov diffeomorphisms minimizing entropy.
result Obtains infinitely many entropy-minimizing diffeomorphisms.
Solves ambiguity in incomplete markets by minimizing price measure entropy.
problem Ambiguity in pricing incomplete markets.
method Minimizes the entropy of the price measure from the economic measure, subject to mark-to-market constraints.
result Resolves ambiguity and provides a consistent pricing measure.
Study shows rigidity for entropy minimizers in non-monotone cases.
problem Rigidity of entropy minimizers in non-monotone settings.
method Elementary proofs in non-monotone situations.
result Showed rigidity for minimizers of generalized Colding-Minicozzi entropies.
AdaDEM decouples EM into two parts to improve class overlap and uncertainty.
problem Improper EM limits its effectiveness in various machine learning tasks.
method Decouple EM into CADF and GMC, and AdaDEM normalizes CADF reward and uses MEC.
result AdaDEM outperforms classical EM and improves performance in noisy and dynamic environments.
The resolution and calibration of pure spectra of minority components in measurements of chemical mixtures without prior knowledge of the mixture is a challenging problem. In this work, a combination of band target entropy minimization (BTEM) and target partial least squares (T-PLS) was used to obtain estimates for sin…
Method determines credit transition matrix from cumulative default probabilities.
problem Quantifying changes in bond credit ratings.
method Setup an ill-posed, linear inverse problem with entropy minimization.
result Method successfully determines CTM from cumulative default probabilities.
In this paper, we propose a general framework to learn a robust large-margin binary classifier when corrupt measurements, called anomalies, caused by sensor failure might be present in the training set. The goal is to minimize the generalization error of the classifier on non-corrupted measurements while controlling th…
In this paper, we propose a general framework to learn a robust large-margin binary classifier when corrupt measurements, called anomalies, caused by sensor failure might be present in the training set. The goal is to minimize the generalization error of the classifier on non-corrupted measurements while controlling th…
Study on entropy stability in product spaces of negatively curved symmetric spaces.
problem Stability of minimal entropy rigidity in product spaces of negatively curved symmetric spaces.
method Analysis of minimal entropy sequences and proof of intrinsic uniqueness of spherical Plateau solutions.
result Entropy-minimizing sequences converge to the model space after removing subsets whose n-volume converges to zero.
Self-training avoids spurious features in domain adaptation.
problem Domain shift with large differences between source and target domains.
method Entropy minimization on unlabeled target data, initialized with a source classifier.
result Entropy minimization avoids using spurious features in large domain shifts.
COME replaces entropy minimization to prevent model collapse.
problem Overconfidence in entropy minimization leads to model collapse.
method COME explicitly models uncertainty with a Dirichlet prior distribution.
result COME achieves state-of-the-art performance on various TTA settings.
The paper constructs diffeomorphisms on a G2-manifold achieving entropy bounds.
problem Achieving Yomdin's homological lower bound for topological entropy on G2-manifolds. method Constructs diffeomorphisms mimicking Farb-Looijenga's for K3 surfaces and acts freely on Teichmüller space.
result The homotopy moduli space of G2 structures on the manifold has an infinite fundamental group. This paper improves test-time adaptation for distribution shifts using confidence maximization and input transformation.
problem Improving deep networks' performance on data shifted from the training distribution.
method Proposes a novel loss function combining confidence maximization and batch-wise entropy maximization with an input transformation module.
result Significantly improves robustness of pretrained networks to corruptions on benchmarks like ImageNet-C.
We propose a general framework for neural network compression that is motivated by the Minimum Description Length (MDL) principle. For that we first derive an expression for the entropy of a neural network, which measures its complexity explicitly in terms of its bit-size. Then, we formalize the problem of neural netwo…
Boosting with tempered exponential measures improves AdaBoost's convergence rate.
problem Improving the convergence rate of AdaBoost.
method Introducing tempered exponential measures (TEMs) to generalize AdaBoost's approach.
result t-AdaBoost achieves an improved convergence rate compared to AdaBoost, especially for t∈[0,1). A new method for generative modeling of discrete data using geometric latent subspaces.
problem Learning generative models for discrete data with statistical dependencies.
method Geometric latent-subspace framework in exponential parameter space of product manifolds of categorical distributions.
result Low-dimensional latent space encodes statistical dependencies and accurately models high-dimensional discrete data.
This work improves deep learning from noisy crowdsourced labels.
problem Learning label correction and neural classifier from noisy crowdsourced data.
method Coupled Cross-Entropy Minimization (CCEM) with identifiability and regularization.
result The CCEM criterion correctly identifies annotators' confusion and neural classifier under realistic conditions.
The paper improves semi-supervised learning using f-divergences and α-Rényi divergences.
problem Improving semi-supervised learning with noisy pseudo-labels.
method Inspired by f-divergences and α-Rényi divergences, the paper develops new empirical risk functions and regularization techniques. result The new methods show better performance than traditional self-training methods, especially in noisy pseudo-label scenarios.
We explain SSL objectives as log-likelihoods in a data curation model.
problem Lack of understanding of SSL objectives as log-likelihoods.
method Formulate SSL objectives as a log-likelihood in a generative model of data curation.
result SSL methods can be understood as lower-bounds on a principled log-likelihood.
In this short note, we analyze geometric properties of orbit spaces of certain involutions in dimensions four, five, and six. We consider constructions of F-structures on manifolds of dimension at least four that allows us to study minimal entropy, minimal volume, collapse with bounded curvature, and sign o…
New method restores source features for SFDA without source data.
problem Domain adaptation without access to source data.
method Feature Restoration (FR) and Bottom-Up Feature Restoration (BUFR).
result BUFR outperforms existing SFDA methods in accuracy, calibration, and data efficiency.
We propose a new regularization method based on virtual adversarial loss: a new measure of local smoothness of the conditional label distribution given input. Virtual adversarial loss is defined as the robustness of the conditional label distribution around each input data point against local perturbation. Unlike adver…
We make use of F-structures and technology developed by Paternain - Petean to compute minimal entropy, minimal volume, and Yamabe invariant of symplectic 4-manifolds, as well as to study their collapse with sectional curvature bounded from below. À la Gompf, we show that these invariants vanish on symplecti…
Develops a mean-field theory for multi-head self-attention under cross-entropy training.
problem Mean-field analysis of multi-head self-attention under cross-entropy training.
method Mean-field theory for a simplified single-layer causal multi-head self-attention model.
result Proves a static finite-head approximation bound for the optimal risk.
We introduce Negative Sampling in Semi-Supervised Learning (NS3L), a simple, fast, easy to tune algorithm for semi-supervised learning (SSL). NS3L is motivated by the success of negative sampling/contrastive estimation. We demonstrate that adding the NS3L loss to state-of-the-art SSL algorithms, such as the Virtual Adv…
Paper proposes a new hierarchical attention mechanism for multi-scale data.
problem Challenges in applying neural attention to multi-scale, multi-modal data.
method Developed a mathematical framework for multi-modal, multi-scale data and derived optimal neural attention mechanics.
result Proposed hierarchical attention mechanism improves transformer performance in multi-scale, multi-modal settings.
New theorems on compactness and finiteness for specific types of self-shrinkers.
problem Characterizing rotationally symmetric self-shrinkers with constraints.
method Compactness and finiteness theorems for self-shrinkers with specific symmetries and constraints.
result Existence of entropy minimizing self-shrinkers diffeomorphic to S1imesSn−1 for each n≥2. Study examines uncertainty in adversarially trained models and proposes improved AT methods.
problem Uncertainty quantification in adversarially trained models for safety-critical applications.
method Investigates conformal prediction (CP) for adversarial attacks, proposes Beta-weighting loss with entropy minimization for improved prediction set size (PSS).
result Proposed AT-UR method improves CP efficiency and prediction set size.
The need to reason about uncertainty in large, complex, and multi-modal datasets has become increasingly common across modern scientific environments. The ability to transform samples from one distribution P to another distribution Q enables the solution to many problems in machine learning (e.g. Bayesian inference…
In 2004, Taubes introduced the space of minimal hyperbolic germs with elements consisting of the first and second fundamental form of an equivariant immersed minimal disk in hyperbolic 3-space. Herein, we initiate a further study of this space by studying the behavior of a dynamically defined function which records the…
Data-driven anomaly detection methods suffer from the drawback of detecting all instances that are statistically rare, irrespective of whether the detected instances have real-world significance or not. In this paper, we are interested in the problem of specifically detecting anomalous instances that are known to have …
New approach for semi-supervised learning under covariate shifts.
problem Semi-supervised learning under covariate shifts where labeled and unlabeled data distributions differ.
method Information-theoretical approach, addressing covariate shifts.
result Improved performance compared to previous methods.
In this short note, exploits of constructions of F-structures coupled with technology developed by Cheeger-Gromov and Paternain-Petean are seen to yield a procedure to compute minimal entropy, minimal volume, Yamabe invariant and to study collapsing with bounded sectional curvature on inequivalent smooth st…
The paper proves asymptotic normality for multinomial logistic regression on null covariates.
problem Classical asymptotic normality results fail in high-dimensional multinomial logistic models.
method Developed asymptotic normality and chi-square results for multinomial logistic MLE on null covariates.
result Validated new methodology to test feature significance in high-dimensional classification problems.
Regression aims at estimating the conditional mean of output given input. However, regression is not informative enough if the conditional density is multimodal, heteroscedastic, and asymmetric. In such a case, estimating the conditional density itself is preferable, but conditional density estimation (CDE) is challeng…
Adaptive importance sampling for estimating point process statistics.
problem Estimating the expected value of a statistic of a locally stable point process.
method Adaptive importance sampling with Poisson point processes and cross-entropy minimization.
result The proposed estimator converges to the target value almost surely and is asymptotically normal.
Study of focal-entropy for class-imbalanced classification.
problem Understanding the focal-loss in class-imbalanced settings.
method Distributional viewpoint and information-theoretic analysis of focal-entropy.
result Focal-entropy minimizer exists and departs from data distribution.
A new method separates data points using entropy minimization over a hypercube.
problem Classifying data points with linear or non-linear decision boundaries.
method Minimizing entropy over a hypercube to find separating parameters.
result Efficient and versatile for linear and non-linear separability.
SupSup model learns thousands of tasks without forgetting, using randomly initialized subnetworks.
problem Sequentially learning many tasks without forgetting.
method Randomly initialized base network with task-specific subnetworks (supermasks).
result Gradient-based optimization can identify the correct subnetwork for new tasks.
We study the problem of training an accurate linear regression model by procuring labels from multiple noisy crowd annotators, under a budget constraint. We propose a Bayesian model for linear regression in crowdsourcing and use variational inference for parameter estimation. To minimize the number of labels crowdsourc…
Timely detection of abrupt anomalies is crucial for real-time monitoring and security of modern systems producing high-dimensional data. With this goal, we propose effective and scalable algorithms. Proposed algorithms are nonparametric as both the nominal and anomalous multivariate data distributions are assumed unkno…
Generative model for time series using Schrödinger bridges with jumps.
problem Creating realistic synthetic time series from observed data.
method Entropic optimal transport, Schrödinger bridge framework, jump-diffusion process.
result Jump-diffusion Schrödinger bridge model generates more realistic time series.
A new approach for test-time adaptation detects and reacts to distribution shifts.
problem Improving test-time accuracy under distribution shifts.
method Online self-training with a detection tool based on entropy values and betting martingales.
result The classifier's entropy values match those of the source domain, building invariance to distribution shifts.
Gradient descent biases linear models in next-token prediction towards data entropy.
problem Optimization bias in next-token prediction models.
method Analysis of gradient descent on linear models with sparse conditional distributions.
result Gradient descent selects parameters that equate token logits differences to log-odds in the data subspace.