This work tackles online memory selection in continual learning using information theory.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Feature selection aims to select the smallest feature subset that yields the minimum generalization error. In the rich literature in feature selection, information theory-based approaches seek a subset of features such that the mutual information between the selected features and the class labels is maximized. Despite …
IndiSeek learns disentangled representations by balancing independence and completeness.
Paper designs a penalty for model order selection using information criteria.
This paper introduces efficient approximations for fairness criteria in regression models.
Estimating the dependences between random variables, and ranking them accordingly, is a prevalent problem in machine learning. Pursuing frequentist and information-theoretic approaches, we first show that the p-value and the mutual information can fail even in simplistic situations. We then propose two conditions for r…
High-dimensional predictive models, those with more measurements than observations, require regularization to be well defined, perform well empirically, and possess theoretical guarantees. The amount of regularization, often determined by tuning parameters, is integral to achieving good performance. One can choose the …
In this work we present a review of the state of the art of information theoretic feature selection methods. The concepts of feature relevance, redundance and complementarity (synergy) are clearly defined, as well as Markov blanket. The problem of optimal feature selection is defined. A unifying theoretical framework i…
Density destructors simplify complex PDFs to maximize entropy, linking to information theory.
vsOED optimizes experiment design with reinforcement learning for Bayesian models.
Multi-class classification methods based on both labeled and unlabeled functional data sets are discussed. We present a semi-supervised logistic model for classification in the context of functional data analysis. Unknown parameters in our proposed model are estimated by regularization with the help of EM algorithm. A …
We propose orthogonality as a necessary condition for disentangling aleatoric and epistemic uncertainty.
In this work, we consider the problem of autonomously discovering behavioral abstractions, or options, for reinforcement learning agents. We propose an algorithm that focuses on the termination condition, as opposed to -- as is common -- the policy. The termination condition is usually trained to optimize a control obj…
Statistical shape models enhance machine learning algorithms providing prior information about deformation. A Point Distribution Model (PDM) is a popular landmark-based statistical shape model for segmentation. It requires choosing a model order, which determines how much of the variation seen in the training data is a…
Maximal correlation framework improves fairness in machine learning algorithms.
Comparing with traditional learning criteria, such as mean square error (MSE), the minimum error entropy (MEE) criterion is superior in nonlinear and non-Gaussian signal processing and machine learning. The argument of the logarithm in Renyis entropy estimator, called information potential (IP), is a popular MEE cost i…
Visual exploration of high-dimensional real-valued datasets is a fundamental task in exploratory data analysis (EDA). Existing methods use predefined criteria to choose the representation of data. There is a lack of methods that (i) elicit from the user what she has learned from the data and (ii) show patterns that she…
New bounds improve generalization in learning scenarios.
New bounds show limitations of sample-wise information-theoretic generalization.
Improved ITL descriptors using explicit inner product spaces for scalable systems.
Information-theoretic Bayesian optimisation techniques have demonstrated state-of-the-art performance in tackling important global optimisation problems. However, current information-theoretic approaches require many approximations in implementation, introduce often-prohibitive computational overhead and limit the choi…
New research shows existing information-theoretic methods can't establish minimax rates for gradient descent in stochastic convex optimization.
Information-theoretic Bayesian regret bounds of Russo and Van Roy capture the dependence of regret on prior uncertainty. However, this dependence is through entropy, which can become arbitrarily large as the number of actions increases. We establish new bounds that depend instead on a notion of rate-distortion. Among o…
In this paper we consider an information theoretic approach for the accounting classification process. We propose a matrix formalism and an algorithm for calculations of information theoretic measures associated to accounting classification. The formalism may be useful for further generalizations and computer-based imp…
We give some general criteria of being a homeomorphism for continuous mappings of topological manifolds, as well as criteria of being a diffeomorphism for smooth mappings of smooth manifolds. As an illustration, we apply these criteria to the problems arising in two- and three-dimensional grid generation.
The study reveals flaws in pruning criteria and proposes a new assumption for better filter selection.
New method finds optimal hyperparameters for multiple tasks and criteria.
New criteria for Heegaard splittings ensure strong irreducibility and finite Goeritz groups.
Multi-criteria recommender systems have been increasingly valuable for helping consumers identify the most relevant items based on different dimensions of user experiences. However, previously proposed multi-criteria models did not take into account latent embeddings generated from user reviews, which capture latent se…
Develops scenario theory for multi-criteria decision making.
We consider the problem of identifying patterns in a data set that exhibit anomalous behavior, often referred to as anomaly detection. In most anomaly detection algorithms, the dissimilarity between data samples is calculated by a single criterion, such as Euclidean distance. However, in many cases there may not exist …
Study reveals mutual information is crucial for understanding algorithm performance in stochastic convex optimization.
The paper evaluates criteria for selecting cryptocurrencies based on historical data.
New tighter bounds for learning algorithms from Steinke & Zakynthinou's supersample setting.
Optimistic algorithms and Thompson sampling use info-theory for better reinforcement learning.
The paper explores the information-theoretic nature of excess risk in machine learning.
We present criteria for establishing a triangulation of a manifold. Given a manifold M, a simplicial complex A, and a map H from the underlying space of A to M, our criteria are presented in local coordinate charts for M, and ensure that H is a homeomorphism. These criteria do not require a differentiable structure, or…
When sufficient labeled data are available, classical criteria based on Receiver Operating Characteristic (ROC) or Precision-Recall (PR) curves can be used to compare the performance of un-supervised anomaly detection algorithms. However , in many situations, few or no data are labeled. This calls for alternative crite…
The paper analyzes performance criteria for competing fund managers in Ito-diffusion markets.
Paper defines numerical criteria to test handlebody link irreducibility.
Spectral flow connects manifold geometry to rigidity criteria.
Optimizes SGLD noise structure for better generalization bounds.
Unified framework for information-theoretic bounds on learning algorithms.
Unified notation simplifies information-theoretic concepts in machine learning.
We shall give useful criteria of lips, beaks and swallowtail singularities of smooth map from the plane into the plane. As an application of criteria, we will discuss the singularities of Cauchy problem of single conservation law.
A new method for multi-criteria recommender systems using graph attention networks.
New framework for resilient bi-criteria optimization under noisy feedback.
New algorithms minimize risk in MNL bandits, achieving near-optimal performance.