Model selection based on classical information criteria, such as BIC, is generally computationally demanding, but its properties are well studied. On the other hand, model selection based on parameter shrinkage by -type penalties is computationally efficient. In this paper we make an attempt to combine their st…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Proposes SNML for selecting word2vec Skip-gram dimensionality.
This work tackles online memory selection in continual learning using information theory.
This paper introduces efficient approximations for fairness criteria in regression models.
New method speeds up model selection for complex scientific tasks.
A new criterion selects models in overparameterized settings.
We test three common information criteria (IC) for selecting the order of a Hawkes process with an intensity kernel that can be expressed as a mixture of exponential terms. These processes find application in high-frequency financial data modelling. The information criteria are Akaike's information criterion (AIC), the…
Automated model assesses online health info quality using machine learning.
Feature selection aims to select the smallest feature subset that yields the minimum generalization error. In the rich literature in feature selection, information theory-based approaches seek a subset of features such that the mutual information between the selected features and the class labels is maximized. Despite …
Many statistical models are given in the form of non-normalized densities with an intractable normalization constant. Since maximum likelihood estimation is computationally intensive for these models, several estimation methods have been developed which do not require explicit computation of the normalization constant,…
The sBIC outperforms other model selection criteria in LDA topic modeling.
Selective regression allows abstention to improve fairness criteria.
High-dimensional predictive models, those with more measurements than observations, require regularization to be well defined, perform well empirically, and possess theoretical guarantees. The amount of regularization, often determined by tuning parameters, is integral to achieving good performance. One can choose the …
In this work, three lattice-free (LF) discriminative training criteria for purely sequence-trained neural network acoustic models are compared on LVCSR tasks, namely maximum mutual information (MMI), boosted maximum mutual information (bMMI) and state-level minimum Bayes risk (sMBR). We demonstrate that, analogous to L…
Framework benchmarks optimizers on multiple criteria.
The paper derives an equation linking WAIC and WBIC for singular models.
Unified perspective unites Bayesian optimization and active learning for efficient goal-oriented optimization.
Paper discusses prediction errors for penalized regressions using GAMP and LOOCV.
We propose a cost-effective framework for preference elicitation and aggregation under the Plackett-Luce model with features. Given a budget, our framework iteratively computes the most cost-effective elicitation questions in order to help the agents make a better group decision. We illustrate the viability of the fram…
IndiSeek learns disentangled representations by balancing independence and completeness.
This paper proposes new methods for ALR that consider informativeness, representativeness, and diversity.
Automatically assesses the quality of online health articles.
Estimating the dependences between random variables, and ranking them accordingly, is a prevalent problem in machine learning. Pursuing frequentist and information-theoretic approaches, we first show that the p-value and the mutual information can fail even in simplistic situations. We then propose two conditions for r…
New methods ensure fairness in noisy protected groups.
We study tick-by-tick financial returns belonging to the FTSE MIB index of the Italian Stock Exchange (Borsa Italiana). We can confirm previously detected non-stationarities. However, scaling properties reported in the previous literature for other high-frequency financial data are only approximately valid. As a conseq…
LS improves model selection for singular statistical models.
Graph embedding provides an efficient solution for graph analysis by converting the graph into a low-dimensional space which preserves the structure information. In contrast to the graph structure data, the i.i.d. node embedding can be processed efficiently in terms of both time and space. Current semi-supervised graph…
We use variational Gaussian approximations to analyze parametric models with unknown data-generating distributions.
With the wealth of information produced by social networks, smartphones, medical or financial applications, speculations have been raised about the sensitivity of such data in terms of users' personal privacy and data security. To address the above issues, Federated Learning (FL) has been recently proposed as a means t…
SIC detects elbows in error curves automatically.
Study evaluates various regularization methods for electricity price forecasting.
In recommender systems, cold-start issues are situations where no previous events, e.g. ratings, are known for certain users or items. In this paper, we focus on the item cold-start problem. Both content information (e.g. item attributes) and initial user ratings are valuable for seizing users' preferences on a new ite…
Online feature selection has been an active research area in recent years. We propose a novel diverse online feature selection method based on Determinantal Point Processes (DPP). Our model aims to provide diverse features which can be composed in either a supervised or unsupervised framework. The framework aims to pro…
We propose definitions of fairness in machine learning and artificial intelligence systems that are informed by the framework of intersectionality, a critical lens arising from the Humanities literature which analyzes how interlocking systems of power and oppression affect individuals along overlapping dimensions inclu…
Detecting and recovering labels in binomial logistic mixtures is challenging due to an information gap.
This paper presents a novel signal compression algorithm based on the Blaschke unwinding adaptive Fourier decomposition (AFD). The Blaschke unwinding AFD is a newly developed signal decomposition theory. It utilizes the Nevanlinna factorization and the maximal selection principle in each decomposition step, and achieve…
New GIC improves model selection for structured sparse models.
MOSAIC selects few informative exemplars from high-dimensional data with non-linear structures.
Meta-analysis improves personalized treatment rules across multiple sites.
vsOED optimizes experiment design with reinforcement learning for Bayesian models.
Batch Active Learning uses derivative information for Gaussian Process regression.
In this work we present a review of the state of the art of information theoretic feature selection methods. The concepts of feature relevance, redundance and complementarity (synergy) are clearly defined, as well as Markov blanket. The problem of optimal feature selection is defined. A unifying theoretical framework i…
In experimental design, we are given vectors in dimensions, and our goal is to select of them to perform expensive measurements, e.g., to obtain labels/responses, for a linear regression task. Many statistical criteria have been proposed for choosing the optimal design, with popular choices including A…
In this work we study loss functions for learning and evaluating probability distributions over large discrete domains. Unlike classification or regression where a wide variety of loss functions are used, in the distribution learning and density estimation literature, very few losses outside the dominant ar…
Bayesian active learning improves holistic educational assessments.
New method detects essential tori in mixed singularity links.
The paper analyzes how investors' wealth can decline collectively under partial information.
Paper improves feature selection accuracy using transfer learning.