Data-driven anomaly detection methods suffer from the drawback of detecting all instances that are statistically rare, irrespective of whether the detected instances have real-world significance or not. In this paper, we are interested in the problem of specifically detecting anomalous instances that are known to have …
In this paper, we present a novel and general framework called {\it Maximum Entropy Discrimination Markov Networks} (MaxEnDNet), which integrates the max-margin structured learning and Bayesian-style estimation and combines and extends their merits. Major innovations of this model include: 1) It generalizes the extant …
A new algorithm learns diverse policies in reinforcement learning.
problem Learning diverse behaviors in reinforcement learning.
method Proposes Maximum Entropy Diverse Exploration (MEDE) algorithm.
result The set of policies learned by MEDE capture the same modalities as the optimal maximum entropy policy.
A new classifier updates sequentially using maximum margin principles.
problem Sequential data collection and partial labeling.
method Maximum margin classifier with Maximum Entropy Discrimination principle, kernel representation, and regularization.
result Improved performance compared to non-sequential classifiers.
Exponential models of distributions are widely used in machine learning for classiffication and modelling. It is well known that they can be interpreted as maximum entropy models under empirical expectation constraints. In this work, we argue that for classiffication tasks, mutual information is a more suitable informa…
Incorporating feature selection into a classification or regression method often carries a number of advantages. In this paper we formalize feature selection specifically from a discriminative perspective of improving classification/regression accuracy. The feature selection method is developed as an extension to the r…
Three LF training criteria improve neural network acoustic models without cross-entropy pre-training.
problem Improving purely sequence-trained neural network acoustic models.
method Comparison of three lattice-free discriminative training criteria (MMI, bMMI, sMBR) on LVCSR tasks.
result LF-bMMI models outperform plain LF-MMI models by 5% WER on Switchboard datasets.
We introduce a Maximum Entropy model able to capture the statistics of melodies in music. The model can be used to generate new melodies that emulate the style of the musical corpus which was used to train it. Instead of using the n−body interactions of (n−1)−order Markov models, traditionally used in automatic mus…
In this paper, we propose a general framework to learn a robust large-margin binary classifier when corrupt measurements, called anomalies, caused by sensor failure might be present in the training set. The goal is to minimize the generalization error of the classifier on non-corrupted measurements while controlling th…
Paper mitigates information leakage in image representations using maximum entropy.
problem Mitigating unintended leakage of user information from image representations.
method Formulates an adversarial non-zero sum game to find an embedding function that maximizes task-dependent discriminative information while minimizing entropy of sensitive attributes.
result Proposed approach learns image representations with high task performance and reduced leakage of sensitive information.
DNLL loss improves deep LDA accuracy and consistency.
problem Pathological solutions in unconstrained Deep LDA.
method Introducing Discriminative Negative Log-Likelihood (DNLL) loss.
result Deep LDA trained with DNLL produces clean latent spaces and better calibrated probabilities.
In this paper, we propose a general framework to learn a robust large-margin binary classifier when corrupt measurements, called anomalies, caused by sensor failure might be present in the training set. The goal is to minimize the generalization error of the classifier on non-corrupted measurements while controlling th…
VAEs improve collaborative filtering for implicit feedback.
problem Limited modeling capacity of linear factor models in collaborative filtering.
method Introduced a generative model with multinomial likelihood and used Bayesian inference for parameter estimation.
result Significantly outperforms state-of-the-art baselines on real-world datasets.
Paper proposes a new method to learn distribution kernels via entropy maximization.
problem Challenges in applying kernel methods to distribution regression tasks.
method Proposes a novel objective for unsupervised learning of data-dependent distribution kernels based on entropy maximization.
result Demonstrates the effectiveness of the learned kernel across different modalities.
Paper develops MRCs for supervised classification using generalized maximum entropy.
problem Developing robust classifiers for decision problems.
method Generalized maximum entropy principle applied to minimax risk classifiers.
result Learning techniques for determining MRCs with performance guarantees.
Maximum entropy modeling is a flexible and popular framework for formulating statistical models given partial knowledge. In this paper, rather than the traditional method of optimizing over the continuous density directly, we learn a smooth and invertible transformation that maps a simple distribution to the desired ma…
Supervised topic models utilize document's side information for discovering predictive low dimensional representations of documents. Existing models apply the likelihood-based estimation. In this paper, we present a general framework of max-margin supervised topic models for both continuous and categorical response var…
The well known maximum-entropy principle due to Jaynes, which states that given mean parameters, the maximum entropy distribution matching them is in an exponential family, has been very popular in machine learning due to its "Occam's razor" interpretation. Unfortunately, calculating the potentials in the maximum-entro…
MGD combines maximum entropy and diffusion methods for efficient sampling.
problem Generating samples from limited information in high dimensions.
method Moment Guided Diffusion (MGD) using stochastic differential equations.
result MGD efficiently samples maximum entropy distributions in finite time.
Researchers use Gaussian processes to approximate Lagrange multipliers for Maximum-Entropy distributions.
problem Finding Lagrange multipliers for Maximum-Entropy distributions is computationally challenging.
method Employed Gaussian processes to approximate the Lagrange multipliers as a map of moments. Optimized hyperparameters by maximizing log-likelihood.
result Data-driven Maximum-Entropy closure performs well in approximating non-equilibrium distributions.
Enhances RL by controlling policy stochasticity through trajectory entropy constraints.
problem Non-stationary Q-value estimation and short-sighted entropy tuning in maximum entropy RL.
method Proposes TECRL framework with separate Q-functions for reward and entropy, enforcing a trajectory entropy constraint.
result DSAC-E algorithm achieves higher returns and better stability on OpenAI Gym benchmarks.
The paper presents a method to estimate joint interventional distributions from marginal interventional data.
problem Estimating joint interventional distributions from marginal interventional data.
method The paper extends the Causal Maximum Entropy method to use interventional data and employs Lagrange duality to prove the solution lies in the exponential family.
result The method allows for causal feature selection and inference of joint interventional distributions.
We present a new statistical learning paradigm for Boltzmann machines based on a new inference principle we have proposed: the latent maximum entropy principle (LME). LME is different both from Jaynes maximum entropy principle and from standard maximum likelihood estimation.We demonstrate the LME principle BY deriving …
The need to estimate smooth probability distributions (a.k.a. probability densities) from finite sampled data is ubiquitous in science. Many approaches to this problem have been described, but none is yet regarded as providing a definitive solution. Maximum entropy estimation and Bayesian field theory are two such appr…
We discuss the systemic risk implied by the interbank exposures reconstructed with the maximum entropy method. The maximum entropy method severely underestimates the risk of interbank contagion by assuming a fully connected network, while in reality the structure of the interbank network is sparsely connected. Here, we…
Learning and understanding the typical patterns in the daily activities and routines of people from low-level sensory data is an important problem in many application domains such as building smart environments, or providing intelligent assistance. Traditional approaches to this problem typically rely on supervised lea…
A new IRL model recovers reward and state structure from expert demonstrations.
problem Limitation of classical maximum entropy model in capturing state structure.
method Generalized maximum causal entropy for IRL models.
result Empirically outperforms classical models in recovering reward and state structure.
Unified framework for maximum entropy RL using Tsallis entropy.
problem Generalizing maximum entropy reinforcement learning with various entropies.
method Tsallis MDPs with Tsallis entropy maximization, controlling entropic index.
result Different entropic indices lead to different optimal policies and exploration tendencies.
Bayesian models use hyperparameters to indirectly assign priors, and this work shows how these priors can be derived from maximum entropy principles.
problem Understanding the assumptions and dependencies in Bayesian hierarchical models.
method Demonstrates how canonical distributions and maximum entropy principles can be used to derive marginal priors in hierarchical models.
result Marginal priors in hierarchical models derived from maximum entropy principles have different constraints compared to the original priors.
The maximum entropy principle can be used to assign utility values when only partial information is available about the decision maker's preferences. In order to obtain such utility values it is necessary to establish an analogy between probability and utility through the notion of a utility density function. According…
The paper extends entropy maximization to multiscale settings and applies it to neural networks.
problem Achieving optimal risk bounds in neural networks using multiscale entropy.
method Generalizing maximum entropy to multiscale settings and applying it to neural networks.
result The multiscale Gibbs posterior can achieve a smaller excess risk than the single-scale Gibbs posterior in a teacher-student scenario.
Improves GAN training by guiding the discriminator to have more diverse binary activation patterns.
problem Stability and convergence issues in GAN training.
method Binarized Representation Entropy (BRE) regularization to guide the discriminator's model capacity allocation.
result Improves GAN training stability and convergence speed, higher sample quality, and higher classification accuracy.
Paper derives the maximum entropy characteristics of a rank order distribution for socio-economic applications.
problem Deriving the maximum entropy characteristics of a rank order distribution for socio-economic applications.
method Maximum entropy framework, deriving the discrete generalized beta distribution under a bivariate utility constraint.
result The discrete generalized beta distribution is a natural maximum entropy distribution under an appropriate bivariate utility constraint.
Improved exploration methods for reinforcement learning with reduced sample complexity.
problem Challenges in reinforcement learning exploration in unknown environments.
method Proposed game-theoretic and trajectory entropy algorithms with improved sample complexity.
result Established statistical advantage of entropy-regularized MDPs for exploration and reduced sample complexity.
This work compares lattice-free and lattice-based training criteria for LVCSR.
problem Improving acoustic model performance in speech recognition.
method Direct comparison of lattice-free and lattice-based sequence discriminative training criteria using GPU.
result Lattice-free MMI performance is comparable to lattice-based criteria, while lattice-based sMBR remains superior.
Study on merging predictors in causal and anticausal directions using CMAXENT.
problem Comparing merging predictors in causal and anticausal directions.
method Using CMAXENT as inductive bias, study differences in merging predictors.
result CMAXENT solution reduces to logistic regression in causal direction and LDA in anticausal direction.
Paper proposes a policy-search algorithm to learn entropy-maximizing exploration policies in reward-free environments.
problem Reward-free learning in high-dimensional, continuous-control domains.
method Maximum Entropy POLicy optimization (MEPOL) algorithm that maximizes a non-parametric state entropy estimate.
result MEPOL learns a maximum-entropy exploration policy that facilitates learning various reward-based tasks.
New method calibrates reference distributions for bounded support.
problem Lack of principled method for bounded-support statistical reference distributions.
method Formulated maximum entropy on projective space of nonnegative measures.
result Prescribed acceptance region uniquely determines deformation parameter.
Proposes a new loss function for deep neural networks.
problem Deep neural networks lack a direct method to discriminate between correct and competing classes.
method Introduces a discriminative loss function based on negative log likelihood ratio.
result Significantly outperforms cross-entropy loss on image classification tasks.
The paper introduces a new intrinsic reward method for exploration in reinforcement learning.
problem Improving exploration in reinforcement learning agents.
method Intrinsic rewards proportional to the entropy of future state-action features.
result The new objective leads to improved visitation of features within individual trajectories.
New method quantifies multivariate redundancy using maximum entropy decompositions.
problem Elusive multivariate measures of redundancy that comply with nonnegativity and axioms.
method Maximum entropy framework, rooted tree-based decompositions of mutual information.
result Quantifies different multivariate redundancy contributions.
The problem of determining the joint probability distributions for correlated random variables with pre-specified marginals is considered. When the joint distribution satisfying all the required conditions is not unique, the "most unbiased" choice corresponds to the distribution of maximum entropy. The calculation of t…
Paper finds a new principle for optimizing consumption and wealth using Tsallis entropy.
problem Optimal consumption-investment problem with recursive utility.
method Established connection to quadratic BSDE, derived stochastic maximum principle.
result Proved existence of optimal strategy and analyzed coupled system.
New results on max-entropy distributions with succinct descriptions and stability.
problem Understanding the complexity and stability of max-entropy distributions.
method Polynomial-time algorithms and bounds on bit complexity.
result Polynomial bit complexity of ε-optimal dual solutions to max-entropy convex programs.
MEP-Net uses MEP to generate solutions from limited data.
problem Generating solutions to scientific problems with incomplete information.
method Combines MEP with neural networks to learn complex distributions from moment constraints.
result Demonstrates MEP-Net's effectiveness in modeling biochemical reaction networks and generating complex distributions.
MEMe efficiently approximates large-scale ML problems with hundreds of moments.
problem Efficient approximation in large-scale machine learning.
method Maximum entropy algorithm with hundreds of moments for computationally efficient approximations.
result Superior to existing approaches in fast log determinant estimation and Bayesian optimisation.
Unified framework for network model assessment using maximum entropy.
problem Statistical inference for network models.
method Constrained entropy-maximization problem, Lagrange multipliers.
result Consistent goodness-of-fit and two-sample tests for network models.
A new metric DJP-MMD improves domain adaptation by balancing transferability and discriminability.
problem Improving domain adaptation performance by balancing transferability and discriminability.
method Discriminative Joint Probability Maximum Mean Discrepancy (DJP-MMD) replaces the traditional joint MMD.
result DJP-MMD outperforms traditional MMDs in image classification tasks.