Gradient descent biases neural networks to use an average of features, leading to non-robustness.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new method averages neural network parameters to rank features robustly.
Unordered feature sets are a nonstandard data structure that traditional neural networks are incapable of addressing in a principled manner. Providing a concatenation of features in an arbitrary order may lead to the learning of spurious patterns or biases that do not actually exist. Another complication is introduced …
Introduces joint Shapley values to measure feature importance in models.
This work analyzes benefits and limitations of data augmentation and feature averaging in deep learning models.
AGOP mechanism explains deep neural collapse in neural networks.
Proposes a new method to find features affecting treatment effect distribution.
New method for interpreting non-linear models using forward marginal effects.
Given a model that predicts a target from a vector of input features , we seek to measure the importance of each feature with respect to the model's ability to make a good prediction. To this end, we consider how (on average) some measure of goodness or badness of prediction (wh…
Random Feature (RF) models are used as efficient parametric approximations of kernel methods. We investigate, by means of random matrix theory, the connection between Gaussian RF models and Kernel Ridge Regression (KRR). For a Gaussian RF model with features, data points, and a ridge , we show that the avera…
In recent years, a large amount of model-agnostic methods to improve the transparency, trustability and interpretability of machine learning models have been developed. We introduce local feature importance as a local version of a recent model-agnostic global feature importance method. Based on local feature importance…
The volume of stroke lesion is the gold standard for predicting the clinical outcome of stroke patients. However, the presence of stroke lesion may cause neural disruptions to other brain regions, and these potentially damaged regions may affect the clinical outcome of stroke patients. In this paper, we introduce the t…
Feature-based time series representations have attracted substantial attention in a wide range of time series analysis methods. Recently, the use of time series features for forecast model averaging has been an emerging research focus in the forecasting community. Nonetheless, most of the existing approaches depend on …
Improved self-distillation reduces label noise and enhances model accuracy.
Revises Bayesian model averaging for foundation models.
Paper proves non-zero generalization boost for equivariant models.
Automatically learns summary features from time series data for likelihood-free inference.
Recursive Feature Machines show grokking in modular arithmetic without neural networks.
Feature extraction is a very crucial task in image and pixel (voxel) classification and regression in biomedical image modeling. In this work we present a machine learning based feature extraction scheme based on inception models for pixel classification tasks. We extract features under multi-scale and multi-layer sche…
Ensemble learning is a method of combining multiple trained models to improve model accuracy. We propose the usage of such methods, specifically ensemble average, inside Convolutional Neural Network (CNN) architectures by replacing the single convolutional layers with Inner Average Ensembles (IEA) of multiple convoluti…
Improves contrastive learning invariance with novel training objectives and feature averaging.
Improved TD learning with tail averaging and regularization achieves optimal convergence rates.
UCRL2-VTR achieves nearly optimal regret for learning MDPs with linear function approximation.
A new meta-algorithm for estimating the conditional average treatment effects is proposed in the paper. The main idea underlying the algorithm is to consider a new dataset consisting of feature vectors produced by means of concatenation of examples from control and treatment groups, which are close to each other. Outco…
We study least squares linear regression over uncorrelated Gaussian features that are selected in order of decreasing variance. When the number of selected features is at most the sample size , the estimator under consideration coincides with the principal component regression estimator; when , the esti…
FBMS R package simplifies Bayesian model selection and averaging.
Study compares and contrasts various ML explanation methods, highlighting their disagreements and similarities.
Kernel balancing weights are generalized as KRRR, providing better confidence intervals for treatment effects.
Federated learning is a distributed framework according to which a model is trained over a set of devices, while keeping data localized. This framework faces several systems-oriented challenges which include (i) communication bottleneck since a large number of devices upload their local updates to a parameter server, a…
Deep neural-kernel models combine neural networks and kernel machines for scalable large datasets.
Freezing of gait (FoG) is a common gait disability in Parkinson's disease, that usually appears in its advanced stage. Freeze episodes are associated with falls, injuries, and psychological consequences, negatively affecting the patients' quality of life. For detecting FoG episodes automatically, a highly accurate dete…
Rotationally equivariant convolutions improve molecular property prediction.
Predicting the risk of mortality for patients with acute myocardial infarction (AMI) using electronic health records (EHRs) data can help identify risky patients who might need more tailored care. In our previous work, we built computational models to predict one-year mortality of patients admitted to an intensive care…
No feature ranking can be faithful, stable, and complete when features are collinear.
New statistical methods improve explainability of boosting models.
This paper proposes a novel generic one-class feature learning method based on intra-class splitting. In one-class classification, feature learning is challenging, because only samples of one class are available during training. Hence, state-of-the-art methods require reference multi-class datasets to pretrain feature …
Bayesian model averaging fails under covariate shift, affecting neural networks' performance.
A machine learning approach to record fusion with high accuracy.
Paper analyzes TD() convergence rates for arbitrary features.
Max-margin classifiers' behavior is studied in high dimensions with non-Gaussian features.
New method improves domain generalization by aligning causal mechanisms across domains.
clusterBMA combines clustering results from multiple models using Bayesian model averaging.
New versions of the set-valued average value at risk for multivariate risks are introduced by generalizing the well-known certainty equivalent representation to the set-valued case. The first "regulator" version is independent from any market model whereas the second version, called the market extension, takes trading …
The paper reduces the complexity of financial market correlation matrices to a 2x2 matrix.
Optimizes LightGBM for stock market forecasting with novel feature engineering and transformation methods.
Federated learning allows edge devices to collaboratively learn a shared model while keeping the training data on device, decoupling the ability to do model training from the need to store the data in the cloud. We propose Federated matched averaging (FedMA) algorithm designed for federated learning of modern neural ne…
Motivated by modern applications in which one constructs graphical models based on a very large number of features, this paper introduces a new class of cluster-based graphical models, in which variable clustering is applied as an initial step for reducing the dimension of the feature space. We employ model assisted cl…
The assessment of energy expenditure in real life is of great importance for monitoring the current physical state of people, especially in work, sport, elderly care, health care, and everyday life even. This work reports about application of some machine learning methods (linear regression, linear discriminant analysi…