Minimum Description Length prevents overfitting in noisy data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper introduces a new method for model selection and more generally hyperparameter selection in machine learning. Minimum description length (MDL) is an established method for model selection, which is however not directly aimed at minimizing generalization error, which is often the primary goal in machine learni…
Study shows LLC correlates with neural network compressibility.
We present an asymptotic criterion to determine the optimal number of clusters in k-means. We consider k-means as data compression, and propose to adopt the number of clusters that minimizes the estimated description length after compression. Here we report two types of compression ratio based on two ways to quantify t…
The study quantifies the information needed for causal queries at different levels of Pearl's hierarchy.
New method improves bivariate causal discovery by accurately estimating cause variable complexity.
New complexity measure ADL connects to classical complexity measures.
Method estimates dataset utility via minimal program length proxy.
Critical trajectories in a sphere are found for a specific bending functional.
Reformulated Markov's conjecture in combinatorial terms.
Time-invariant linear dynamical system arises in many real-world applications,and its usefulness is widely acknowledged. A practical limitation with this model is that its latent dimension that has a large impact on the model capability needs to be manually specified. It can be demonstrated that a lower-order model cla…
Neural networks generalize on simple data generated by a programming language.
DL/FBF improves GPSR solutions by selecting compact, generalising expressions.
PCA (Principal Component Analysis) and its variants areubiquitous techniques for matrix dimension reduction and reduced-dimensionlatent-factor extraction. One significant challenge in using PCA, is thechoice of the number of principal components. The information-theoreticMDL (Minimum Description Length) principle gives…
This paper studies spectral properties of spheres with one equator.
The Fisher information approximation (FIA) is an implementation of the minimum description length principle for model selection. Unlike information criteria such as AIC or BIC, it has the advantage of taking the functional form of a model into account. Unfortunately, FIA can be misleading in finite samples, resulting i…
Study improves neural network performance in sequential learning for image classification.
CDL index improves clustering validation for non-convex data.
We tackle the problem of penalty selection of regularization on the basis of the minimum description length (MDL) principle. In particular, we consider that the design space of the penalty function is high-dimensional. In this situation, the luckiness-normalized-maximum-likelihood(LNML)-minimization approach is favorab…
Paper establishes generalization bounds for representation learning using Minimum Description Length.
The Minimum Description Length (MDL) principle states that the optimal model for a given data set is that which compresses it best. Due to practial limitations the model can be restricted to a class such as linear regression models, which we address in this study. As in other formulations such as the LASSO and forward …
Kernel networks' stability edge linked to Fisher Information singularity.
A new method avoids overfitting in network reconstruction by using the minimum description length principle.
We introduce length dilatation structures on metric spaces, tempered dilatation structures and coherent projections and explore the relations between these objects and the Radon-Nikodym property and Gamma-convergence of length functionals. Then we show that the main properties of sub-riemannian spaces can be obtained f…
Given a surface of infinite topological type, there are several Teichmüller spaces associated with it, depending on the basepoint and on the point of view that one uses to compare different complex structures. This paper is about the comparison between the quasiconformal Teichmüller space and the length-spectrum Teichm…
New methods evaluate data representations by complexity of low-loss predictor learning.
APD method decomposes neural network parameters into simple, faithful components.
Fast, fully-automated histograms for large data sets.
Two trees in the boundary of outer space are said to be \emph{primitive-equivalent} whenever their translation length functions are equal in restriction to the set of primitive elements of . We give an explicit description of this equivalence relation, showing in particular that it is nontrivial. This question is …
We propose a simple, tractable lower bound on the mutual information contained in the joint generative density of any latent variable generative model: the GILBO (Generative Information Lower BOund). It offers a data-independent measure of the complexity of the learned latent variable description, giving the log of the…
New architectures improve KANs, making them more interpretable and accurate.
Robust low-rank matrix estimation is a topic of increasing interest, with promising applications in a variety of fields, from computer vision to data mining and recommender systems. Recent theoretical results establish the ability of such data models to recover the true underlying low-rank matrix when a large portion o…
New approach to learning kernels from data using AIT principles.
New model explains how concepts grow based on experience.
New method compares community detection algorithms without ground truth.
In this paper we consider flat metrics (semi-translation structures) on surfaces of finite type. There are two main results. The first is a complete description of when a set of simple closed curves is spectrally rigid, that is, when the length vector determines a metric among the class of flat metrics. Secondly, we gi…
This is an up-to-date introduction to and overview of the Minimum Description Length (MDL) Principle, a theory of inductive inference that can be applied to general problems in statistics, machine learning and pattern recognition. While MDL was originally based on data compression ideas, this introduction can be read w…
Why do deep neural networks (DNNs) benefit from very high dimensional parameter spaces? Their huge parameter complexities vs stunning performance in practice is all the more intriguing and not explainable using the standard theory of model selection for regular models. In this work, we propose a geometrically flavored …
Geometric characterization of sub-Riemannian geodesics on frame bundles.
The paper explores how smaller data sets can lead to better model selection decisions.
Bayesian networks are convenient graphical expressions for high dimensional probability distributions representing complex relationships between a large number of random variables. They have been employed extensively in areas such as bioinformatics, artificial intelligence, diagnosis, and risk management. The recovery …
S2KAN integrates symbolic primitives into neural network activations for improved interpretability.
Non-negative matrix factorization (NMF) is a dimensionality reduction technique which tends to produce a sparse representation of data. Commonly, the error between the actual and recreated matrices is used as an objective function, but this method may not produce the type of representation we desire as it allows for th…
We analyze differences between two information-theoretically motivated approaches to statistical inference and model selection: the Minimum Description Length (MDL) principle, and the Minimum Message Length (MML) principle. Based on this analysis, we present two revised versions of MML: a pointwise estimator which give…
We study generalizations of Lorentzian warped products with one-dimensional base of the form , where is an interval, is a length space and is a positive continuous function. These generalized cones furnish an important class of Lorentzian length spaces in the sense of [Kunzinger, Sämann; Ann. G…
A new complexity measure MDL-COMP for overparameterized models improves generalization performance.
Study magnetic geodesics on Heisenberg groups and manifolds.
The minimum number of self-intersection points for members of a free homotopy class of curves on the punctured torus is bounded above in terms of the number L of letters required for a minimal description of the class in terms of the generators of the fundamental group and their inverses: it is less than or equal to (L…