Model selection based on classical information criteria, such as BIC, is generally computationally demanding, but its properties are well studied. On the other hand, model selection based on parameter shrinkage by ℓ1-type penalties is computationally efficient. In this paper we make an attempt to combine their st…
A new method selects optimal PHMM models for sequence alignment, improving accuracy.
problem Improving sequence alignment accuracy using PHMMs with optimal hidden states.
method Factorized Asymptotic Bayesian algorithm (FIC) for model selection.
result Improved alignment accuracy with more complex models than previous studies.
Paper proposes MIM-DRCFR to learn disentangled factors for better treatment effect estimation.
problem Learning disentangled factors precisely for individual-level treatment effect estimation.
method Multi-task learning framework with MI minimization criteria.
result MIM-DRCFR outperforms state-of-the-art methods in treatment effect estimation.
Assume (1) asset returns follow a stochastic multi-factor process with time-varying conditional expectations; (2) investments are linear functions of factors. This paper calculates asymptotic joint moments of the logarithm of investor's wealth and the factors. These formulas enable fast computation of a wide range of i…
A new criterion HBIC improves model selection for factor analysis with missing data.
problem Model selection for factor analysis with incomplete data.
method Proposes a novel criterion HBIC that uses actual observed information in the penalty term.
result HBIC is more accurate than BIC when missing data rates are high.
Flexible models cluster RNA sequencing data.
problem Clustering discrete data from RNA sequencing studies.
method Finite mixtures of multivariate Poisson-log normal factor analyzers with constraints.
result Models give favorable clustering performance on real and simulated data.
Novel algorithm compresses ECG signals with preserved R peaks.
problem Efficiently compressing ECG signals while preserving R peak information.
method Blaschke unwinding AFD for faster convergence and higher fidelity.
result The proposed algorithm outperforms state-of-the-art approaches in ECG signal compression.
Developed criteria for selecting non-normalized models using NCE and score matching.
problem No information criteria for non-normalized models estimated by NCE or score matching.
method Developed information criteria based on discrepancy measures for non-normalized models estimated by NCE or score matching.
result The proposed criteria enable selection of the appropriate non-normalized model in a data-driven manner.
Proposes SNML for selecting word2vec Skip-gram dimensionality.
problem Selecting optimal dimensionality for word2vec Skip-gram models.
method Information criteria (AIC, BIC, SNML) applied to SG and SG Negative Sampling models.
result SNML outperforms AIC and BIC, selecting closer optimal dimensionality.
Suggests stopping criteria for feature selection using mutual information.
problem Automatic determination of optimal feature subset size and stopping criterion.
method Monitoring conditional mutual information (CMI) among groups of variables using Renyi's α-entropy.
result Easy to implement stopping criteria for feature selection.
The study evaluates three IC for selecting Hawkes process model order in financial data.
problem Model selection for Hawkes process with financial data.
method Testing AIC, BIC, HQ on simulated data.
result Correct model selection success rate varies with sample size and IC type.
Three LF training criteria improve neural network acoustic models without cross-entropy pre-training.
problem Improving purely sequence-trained neural network acoustic models.
method Comparison of three lattice-free discriminative training criteria (MMI, bMMI, sMBR) on LVCSR tasks.
result LF-bMMI models outperform plain LF-MMI models by 5% WER on Switchboard datasets.
MIM learns joint distributions with mutual information and low divergence.
problem Learning joint distributions over observations and latent variables.
method Probabilistic auto-encoder with three design principles: low divergence, high mutual information, and low marginal entropy.
result MIM learns representations with high mutual information, consistent encoding and decoding distributions, effective latent clustering, and comparable data log likelihood to VAE.
This work tackles online memory selection in continual learning using information theory.
problem Online selection of a representative replay memory from data streams.
method Information-theoretic criteria (surprise, learnability) and Bayesian model for efficient computation.
result InfoRS improves robustness against data imbalance compared to reservoir sampling.
This paper introduces efficient approximations for fairness criteria in regression models.
problem Measuring fairness in real-valued outcomes (regression settings) is computationally challenging.
method Fast approximations of mutual information for independence, separation, and sufficiency fairness criteria.
result The method achieves state-of-the-art accuracy/fairness tradeoffs in real-world datasets.
A natural approach to analyze interaction data of form "what-connects-to-what-when" is to create a time-series (or rather a sequence) of graphs through temporal discretization (bandwidth selection) and spatial discretization (vertex contraction). Such discretization together with non-negative factorization techniques c…
New method speeds up model selection for complex scientific tasks.
problem Exhaustive model selection is computationally infeasible for large model spaces.
method Branch-and-bound algorithm with non-monotonic criteria.
result Guaranteed identification of optimal models with significant computational speedups.
Unified framework for disentangled representations using mechanistic independence.
problem Identifiability of disentangled latent factors under statistical dependencies.
method Introduces mechanistic independence to characterize latent factors by their actions on observed variables, proposing various independence criteria.
result Establishes conditions for identifiability of latent subspaces without statistical assumptions.
Cost-effective framework for eliciting and aggregating preferences.
problem Eliciting preferences efficiently under budget constraints.
method Iterative computation of cost-effective questions using Plackett-Luce model and various information criteria.
result Carefully designed information criteria lead to more accurate predictions with fewer questions.
A new criterion selects models in overparameterized settings.
problem Model selection for overparameterized models with more parameters than data.
method Establishes Bayesian duality and introduces the Interpolating Information Criterion.
result The Interpolating Information Criterion selects models in overparameterized settings.
Automated model assesses online health info quality using machine learning.
problem Low quality health information on the internet poses risks to patients.
method Used machine learning models, specifically hierarchical encoder attention-based neural networks (HEA) with BERT and BioBERT embeddings.
result HEA models outperform traditional models in evaluating health info quality.
The sBIC outperforms other model selection criteria in LDA topic modeling.
problem Selecting the optimal number of topics in Latent Dirichlet Allocation (LDA) models.
method Monte Carlo simulations comparing sBIC to other criteria.
result sBIC is superior for choosing the number of topics in LDA models.
Selective regression allows abstention to improve fairness criteria.
problem Selective regression can exacerbate disparities between subgroups.
method Proposes new fairness criteria and two approaches to mitigate performance disparity.
result Proposed fairness criteria ensures performance improvement for every subgroup with reduced coverage.
High-dimensional predictive models, those with more measurements than observations, require regularization to be well defined, perform well empirically, and possess theoretical guarantees. The amount of regularization, often determined by tuning parameters, is integral to achieving good performance. One can choose the …
Paper designs a penalty for model order selection using information criteria.
problem Selecting the correct model order from a set of candidate models.
method Designs a penalty for the generalized information criterion (GIC) to minimize underestimation.
result Optimal penalty minimizes underestimation while keeping overestimation below a specified level.
New algorithms for risk management in incomplete markets.
problem Risk management in incomplete markets with various sources of incompleteness.
method Machine-learning-based algorithms to solve hedging problems.
result One algorithm is flexible and can use multiple risk criteria.
This research simplifies PCA model selection using MDL principle.
problem Choosing the right number of principal components in PCA.
method Reduces NML problems to lower-dimension problems and bounds PCA NML.
result Bound the NML of PCA by terms of the NML of linear regression.
Framework benchmarks optimizers on multiple criteria.
problem Benchmarking optimizers across diverse test functions.
method Union-free generic depth function for partial orders/rankings.
result Identifies central and outlying rankings of optimizers.
RIC-NN predicts stock returns with deep learning, outperforming traditional methods.
problem Predicting stock returns consistently over long periods with minimal human intervention.
method Deep learning framework with nonlinear multi-factor approach, ranked IC stopping criteria, and deep transfer learning.
result RIC-NN outperforms machine learning methods and major equity funds in stock return prediction.
The paper derives an equation linking WAIC and WBIC for singular models.
problem In singular models, conventional criteria fail due to likelihood and posterior breakdown.
method Theoretical derivation linking WAIC and WBIC.
result An asymptotic equation linking WAIC and WBIC for singular models.
The paper discusses the impact of prior densities on Bayesian model selection.
problem The sensitivity of marginal likelihood to prior choice in Bayesian model selection.
method Analyzes the role of prior densities in model selection, discusses improper priors, and proposes solutions.
result Marginal likelihood can be sensitive to prior choice, but improper priors can still be used with caution.
We empirically test predictability on asset price by using stock selection rules based on maximum drawdown and its consecutive recovery. In various equity markets, monthly momentum- and weekly contrarian-style portfolios constructed from these alternative selection criteria are superior not only in forecasting directio…
We consider a problem of clustering a sequence of multinomial observations by way of a model selection criterion. We propose a form of a penalty term for the model selection procedure. Our approach subsumes both the conventional AIC and BIC criteria but also extends the conventional criteria in a way that it can be app…
Unified perspective unites Bayesian optimization and active learning for efficient goal-oriented optimization.
problem Efficiently optimize expensive engineering and scientific problems with limited data.
method Unified framework linking Bayesian infill criteria and active learning criteria.
result Unified approach formalizes Bayesian infill criteria and active learning criteria.
When performing regression or classification, we are interested in the conditional probability distribution for an outcome or class variable Y given a set of explanatoryor input variables X. We consider Bayesian models for this task. In particular, we examine a special class of models, which we call Bayesian regression…
Paper discusses prediction errors for penalized regressions using GAMP and LOOCV.
problem Prediction accuracy of penalized regression models.
method Derives prediction error estimators using GAMP and LOOCV.
result Information criteria and LOOCV error estimators differ in large parameter regions.
IndiSeek learns disentangled representations by balancing independence and completeness.
problem Learning disentangled representations with mutual information in multi-modal data.
method Combines independence-enforcing objective with a reconstruction loss that bounds conditional mutual information.
result Demonstrates effectiveness on synthetic data, CITE-seq, and real-world multi-modal benchmarks.
Study Poisson structures on fibered 5-manifolds with compatibility conditions.
problem Understanding Poisson structures on fibered 5-manifolds.
method Using almost coupling condition and bigraded factorization of the Jacobi identity.
result Describe global behavior and singularities of almost coupling Poisson tensors.
Optimal reinsurance and investment strategies are derived under mean-variance criteria with partial information.
problem Optimal reinsurance and investment strategies for an insurance firm under mean-variance criteria with partially observable market dynamics.
method Formulated as a stochastic LQ control problem, solved using separation principle and stochastic filtering theory for partial information, and viscosity solution for full information.
result Efficient strategies and efficient frontier presented in closed forms via solutions to extended stochastic Riccati equations.
Nonnegative matrix factorization (NMF) is a popular dimension reduction technique that produces interpretable decomposition of the data into parts. However, this decompostion is not generally identifiable (even up to permutation and scaling). While other studies have provide criteria under which NMF is identifiable, we…
This paper proposes new methods for ALR that consider informativeness, representativeness, and diversity.
problem Efficiently label samples for regression models with limited labeled data.
method Integrates informativeness, representativeness, and diversity in pool-based sequential active learning.
result Demonstrates effectiveness of new ALR approaches on 12 datasets.
Representation learning systems typically rely on massive amounts of labeled data in order to be trained to high accuracy. Recently, high-dimensional parametric models like neural networks have succeeded in building rich representations using either compressive, reconstructive or supervised criteria. However, the seman…
Automatically assesses the quality of online health articles.
problem Lack of automated tools to evaluate the quality of online health information.
method Data mining approach using 10 quality criteria and feature selection.
result Classifier achieved 84%-90% accuracy on 10 criteria.
We study the projected gradient descent method on low-rank matrix problems with a strongly convex objective. We use the Burer-Monteiro factorization approach to implicitly enforce low-rankness; such factorization introduces non-convexity in the objective. We focus on constraint sets that include both positive semi-defi…
A novel online feature selection method using DPP for diversity.
problem Online feature selection for diverse feature sets.
method DPP-based framework with three stages: sampling, local criteria, and global criteria.
result Demonstrated better compactness and comparable/outsuperior performance.
New optimization criteria improve variational autoencoders for clearer images and latent features.
problem Improving clarity and informativeness of variational autoencoders' latent features and samples.
method Proposed new optimization criteria and a sequential VAE model.
result New criteria help generate clearer images and more informative latent features.
Estimating the dependences between random variables, and ranking them accordingly, is a prevalent problem in machine learning. Pursuing frequentist and information-theoretic approaches, we first show that the p-value and the mutual information can fail even in simplistic situations. We then propose two conditions for r…
The paper establishes criteria for spacetime inextendibility using asymptotic volume-distance-ratio analysis.
problem Determining inextendibility of spacetimes near singularities.
method Asymptotic analysis of volume-distance-ratio (VDR) to prove inextendibility criteria.
result Failure of VDR convergence to the Minkowski value implies inextendibility of spacetime.