New framework for identifying unimportant variables in convex optimization.
problem Identifying unimportant variables in convex optimization problems.
method Two-step approach: gather information on optimal solution structure, then produce screening rules.
result New screening rules for various optimization problems.
Paper applies ANOVA decomposition for interpretable data approximation.
problem High-dimensional data interpretation and dimensionality reduction.
method ANOVA decomposition and Grouped Transformations for interpretability.
result Ability to rank variable interactions and unimportant variables.
Bayesian method selects important covariates in modal regression.
problem Bayesian modal regression with heavy-tailed responses.
method Expectation-maximization algorithm for parameter estimation; test statistic for variable selection.
result Efficacy of the proposed method in identifying important covariates.
New approach optimizes expensive black-box systems with uncertain outputs.
problem Optimizing expensive black-box systems with limited data and uncertainty.
method Flexible non-interpolating surrogate model (TK-MARS) and Smart-Replication approach.
result TK-MARS outperforms original MARS and detects important variables.
Variable selection for optimal treatment regime in a clinical trial or an observational study is getting more attention. Most existing variable selection techniques focused on selecting variables that are important for prediction, therefore some variables that are poor in prediction but are critical for decision-making…
CPI overcomes limitations of permutation importance by providing accurate variable selection.
problem Misidentification of unimportant variables in complex models due to covariate correlations.
method Developed a model agnostic and computationally lean Conditional Permutation Importance (CPI) approach.
result CPI provides accurate type-I error control and more parsimonious variable selection.
BetaExplainer improves GNN interpretability by masking unimportant edges.
problem Interpreting GNNs' predictions is difficult due to black-box behavior and lack of uncertainty quantification.
method BetaExplainer uses a sparsity-inducing prior to mask unimportant edges during training.
result BetaExplainer provides uncertainty in edge importance and improves predictive accuracy on challenging datasets.
New framework for inference with LAR, explaining variable contributions and providing stopping rules.
problem LAR's lack of well-understood termination point and basic behavioral properties.
method Developed a novel framework for inference with LAR, providing new mathematical properties and stopping rules.
result LAR estimates of non-zero population correlations have independent normal distributions for inference, and zero-valued correlations have a non-normal joint distribution.
Dirichlet pruning compresses neural networks by removing unimportant units.
problem Compressing large neural network models without sacrificing performance.
method Assigns Dirichlet distribution over network layers' units and uses variational inference to estimate parameters.
result Achieves state-of-the-art compression performance on larger architectures like VGG and ResNet.
We propose maximum likelihood estimation for learning Gaussian graphical models with a Gaussian (ell_2^2) prior on the parameters. This is in contrast to the commonly used Laplace (ell_1) prior for encouraging sparseness. We show that our optimization problem leads to a Riccati matrix equation, which has a closed form …
Proposes a neural network framework for feature selection in high-dimensional settings.
problem Challenges in feature selection and non-linear function estimation in high-dimensional settings.
method Sparse-input neural networks using group concave regularization.
result Establishes finite-sample guarantees for variable selection consistency and prediction accuracy.
Machine learning explainability limits identifying causal variables.
problem Limiting ability to identify important variables in machine learning models.
method Exploring machine learning explainability techniques and their limitations in identifying causal variables.
result Machine learning algorithms are sensitive to underlying causal structure, leading to misidentification of important variables.
Improves dialogue response model interpretability using attention and regularization.
problem Improving interpretability of dual encoder models for dialogue response suggestions.
method Integrates attention mechanism and novel regularization loss to emphasize important words.
result Improves model accuracy and interpretability compared to existing methods.
A simplified model for fixed income portfolio optimisation.
problem Modeling interest rates and credit risk in fixed income portfolios.
method Proposes a two-factor model for the time evolution of the efficient frontier.
result The efficient frontier is mainly controlled by linear constraints, with standard deviation less important.
Drop Pruning uses stochastic optimization to prune and recover weights, reducing model size and improving performance.
problem Complexity and inefficiency in pruning deep neural networks.
method Introduces stochastic optimization with 'drop away' and 'drop back' strategies to prune and recover weights.
result Achieves competitive compression performance and accuracy compared to state-of-the-art approaches.
DASNet improves neural network efficiency by dynamically pruning activations.
problem Reducing neural network size and computational complexity in embedded systems.
method Dynamic Activation Sparsity (DAS) using a winners-take-all (WTA) dropout technique.
result DASNet achieves better computation cost reduction compared to static feature map pruning methods.
A model explains credit assignment in deep learning networks.
problem Understanding how deep learning assigns credit to parameters.
method Mean-field learning model with ensembles of sub-networks.
result Synaptic connections can be categorized into three types.
Corrects gaps in a method for optimizing high-frequency trading strategies.
problem Optimizing bid and ask limit order strategies in high-frequency trading.
method Uses an approximation method based on Avellaneda and Stoikov's 2008 article, correcting gaps found in it.
result The main answer in Avellaneda and Stoikov's article remains unchanged despite corrections.
New algorithm learns abstractions from experience to simplify complex tasks.
problem Solving complex problems in environments with continuous state spaces.
method Uses MDP homomorphisms to find abstract MDPs and guides exploration.
result Demonstrates task transfer method outperforms deep Q-networks.
A new method to simplify deep neural networks by removing unnecessary parts.
problem Overly complex deep neural networks require significant resource investment for size reduction.
method A fully differentiable sparsification method that optimizes a regularized objective function with stochastic gradient descent.
result The method can learn both the sparsified structure and weights of a network in an end-to-end manner.
Two attention models improve human activity recognition by focusing on important signals and sensor modalities.
problem Noise and unimportant signal components in recurrent networks for human activity recognition.
method Temporal and sensor attention mechanisms with continuity constraints.
result State-of-the-art results on three datasets, showing improved understandability and mean F1 score.
A new method learns object representations from motion in slot representations.
problem Unsupervised object extraction from low-level visual data.
method Contrastive learning in slot representations, focusing on moving objects and distinct entities.
result Introduced a new evaluation metric to measure diversity of slot vectors.
SPP prunes CNN weights probabilistically for faster inference.
problem Efficiently accelerate Convolutional Neural Networks (CNNs) without significant accuracy loss.
method Structured Probabilistic Pruning (SPP) with adjustable pruning probabilities.
result 4x speedup with minimal accuracy loss (0.3% for AlexNet, 0.8% for VGG-16).
PSP dynamically learns structured sparsity in DNNs for better parallelism.
problem Efficiently compressing large DNNs to match hardware capabilities.
method Parameterized Structured Pruning (PSP) that dynamically learns the shape of DNNs through structured sparsity.
result PSP maintains prediction performance while creating substantial structured sparsity.
Network Implosion reduces ResNet layers without accuracy loss.
problem High computation costs in Residual Networks.
method Static layer pruning and retraining to erase unimportant layers.
result Reduces ResNet layers by 24.00-42.86% without accuracy drop.
AGS-CL selectively updates penalties based on node importance for continual learning.
problem Catastrophic forgetting in continual learning.
method Adaptive Group Sparsity (AGS) with proximal gradient descent.
result Significantly outperforms baselines on various continual learning benchmarks.
RAU integrates attention into GRU for better sequence learning.
problem Lack of attention mechanism in GRU leads to information redundancy or loss.
method RAU adds an attention gate to GRU to adaptively focus on regions of interest.
result RAU consistently outperforms GRU and other methods in various tasks.
AOFP prunes CNN filters faster and more accurately.
problem Finding optimal CNN width and pruning filters efficiently.
method Binary search, random masking, multi-path framework.
result AOFP achieves faster pruning with minimal accuracy loss.
Proposes sparsified intervals for high-dimensional regression coefficients.
problem Challenges of high-dimensional regression coefficient inference.
method Sparsified simultaneous confidence intervals.
result Intervals can shrink some coefficients to zero, indicating unimportance.
Improved deep learning models using new attribution priors and expected gradients.
problem Improving interpretability and performance of deep learning models.
method Introducing new attribution priors and expected gradients method that satisfies interpretability axioms.
result Improves model performance across various real-world tasks.
Centripetal SGD prunes deep CNNs by making filters collapse.
problem Pruning deep CNNs with complex structures.
method Centripetal SGD, a novel optimization method.
result Pruning deep CNNs without performance loss.
Linear regression can overfit without harm when data is high-dimensional.
problem Understanding when a perfect fit to noisy training data in linear regression leads to accurate predictions.
method Characterization of linear regression problems with a minimum norm interpolating prediction rule that has near-optimal prediction accuracy.
result Overparameterization is essential for benign overfitting in this setting, and the number of unimportant prediction directions must exceed the sample size.
Simplifies complex RL policies by ranking important decisions.
problem Complexity in RL policies makes them hard to analyze and interpret.
method Statistical fault localisation to rank states and prune unimportant decisions.
result Pruned policies can perform similarly to original policies, improving interpretability.
RNNs are vulnerable to adversarial attacks, especially in network traffic.
problem Adversarial attacks on RNNs for IDSs in network traffic.
method Developed new explainability techniques and ARS for comparing IDSs.
result RNNs are vulnerable to adversarial attacks, even in sequential data.
A taxonomy of saliency metrics helps in pruning deep neural networks.
problem Difficulty in separating saliency metric effectiveness from pruning algorithms.
method Proposed a taxonomy based on four orthogonal components.
result Constructed metrics can outperform existing state-of-the-art metrics.
CAP algorithm controls FCR in online selective prediction.
problem Online predictive tasks with temporal multiplicity and FCR control.
method CAP framework with adaptive pick rule and calibration set construction.
result CAP achieves exact selection-conditional coverage guarantee and FCR control.
Coarse ground truth improves semantic segmentation accuracy for some classes.
problem Efficiently preparing high-quality datasets for autonomous driving.
method Comparative analysis of fine and coarse ground truth annotations on Cityscapes dataset using PSPNet.
result Coarse ground truth annotations can improve semantic segmentation accuracy for some classes without significant loss.
Proposes MEED framework for model interpretation.
problem Improving model interpretability and avoiding undesired characteristics.
method Adversarial Infidelity Learning (AIL) for effective feature selection.
result AIL mechanism helps learn desired conditional distribution.
StaPLR selects important views for multi-view learning models.
problem Selecting important views in multi-view learning to reduce data collection costs.
method Developed stacked penalized logistic regression (StaPLR) for view selection.
result StaPLR outperforms existing methods in view selection, reducing false positives.
A new method identifies a stable subnetwork in overparameterized student models.
problem Identifying a stable subnetwork in overparameterized student models.
method Spectral representation of linear transfer of information, focusing on eigenvalues and eigenvectors.
result A stable student substructure is isolated that mirrors the true complexity of the teacher.
Graph attention improves node classification by distinguishing important edges.
problem Node classification in graph-based learning models.
method Theoretical analysis of graph attention networks for node classification.
result Graph attention can perfectly classify nodes in an 'easy' regime but fails in a 'hard' regime.
Study shows safely discarding features based on aggregate SHAP values is sound.
problem Discarding features based on aggregate SHAP values without proper justification.
method Investigated the soundness of discarding features based on aggregate SHAP values, proposing to aggregate SHAP values over the extended support.
result A small aggregate SHAP value implies safely discarding the corresponding feature.
Proposes a criterion for selecting relevant auxiliary variables in incomplete data analysis.
problem Selecting useful auxiliary variables for incomplete data analysis.
method Formulates model selection problem, proposes an information criterion based on Kullback-Leibler divergence.
result Proposed information criterion is an asymptotically unbiased estimator of Kullback-Leibler divergence.
Proposes a multi-variable LSTM for accurate time series forecasting and variable importance.
problem Current attention mechanisms in recurrent neural networks fail to characterize variable importance in time series with exogenous variables.
method Develops a multi-variable LSTM with tensorized hidden states to learn variable importance and a mixture of temporal and variable attention.
result Demonstrates superior prediction performance and variable importance quantification compared to baselines.
VC-PCR improves prediction by clustering correlated variables.
problem Decreased prediction accuracy due to cluster structure in predictor variables.
method Supervised variable selection and clustering to integrate cluster information into a sparse modeling process.
result VC-PCR achieves better prediction, variable selection, and clustering performance.
Study on inequalities for multinomial variables.
problem Understanding concentration inequalities for multinomial variables.
method Investigation of Dirichlet and Multinomial random variables.
result Results on concentration inequalities for multinomial variables.
This research proposes a variable importance cloud to assess variable importance across multiple good models.
problem Current variable importance measures are tied to a single model, limiting understanding of variable importance across different models.
method Introduces a variable importance cloud that maps every variable to its importance for every good predictive model.
result Shows how variable importance can vary significantly across different good models.
Proposes an interpretable LSTM for time series with exogenous variables.
problem Lack of variable importance characterization in recurrent neural networks.
method Develops a multi-variable LSTM with tensorized hidden states for learning variable-specific representations.
result Variable attention in real datasets is highly aligned with statistical causality.