ERM performs well in feature learning with minimal feature maps.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Identifies features most relevant to concept drift in data.
Mixup helps learn both features in multi-view classification problems.
P-SE explains model decisions with minimal feature subsets and fast estimators.
Self-training avoids spurious features in domain adaptation.
NGP selects N features from P using neural networks in a greedy, iterative process.
Novel Newton method for large-scale kernel methods using random features.
Binary Stochastic Filtering (BSF), the algorithm for feature selection and neuron pruning is proposed in this work. The method defines filtering layer which penalizes amount of the information involved in the training process. This information could be the input data or output of the previous layer, which directly lead…
Feature selection problems have been extensively studied for linear estimation, for instance, Lasso, but less emphasis has been placed on feature selection for non-linear functions. In this study, we propose a method for feature selection in high-dimensional non-linear function estimation problems. The new procedure is…
Decentralized learning for GLMs with feature distribution and network connectivity.
New random forest algorithms for PU learning minimize risk directly.
FeAT improves OOD generalization by learning richer features.
A2I Transformer predicts atom energies from coordinates, avoiding heavy featurization.
Researchers quantify the relationship between feature depth and performance in deep neural networks.
Paper relaxes differential privacy for correlated features, improving privacy-utility trade-off.
Fairness is becoming a rising concern w.r.t. machine learning model performance. Especially for sensitive fields such as criminal justice and loan decision, eliminating the prediction discrimination towards a certain group of population (characterized by sensitive features like race and gender) is important for enhanci…
New features generated from kernel methods are minimally dependent on sensitive features.
Study shows exponential convergence in classification errors using random features and SGD.
We introduce new families of Integral Probability Metrics (IPM) for training Generative Adversarial Networks (GAN). Our IPMs are based on matching statistics of distributions embedded in a finite dimensional feature space. Mean and covariance feature matching IPMs allow for stable training of GANs, which we will call M…
Many efforts have been made to use various forms of domain knowledge in malware detection. Currently there exist two common approaches to malware detection without domain knowledge, namely byte n-grams and strings. In this work we explore the feasibility of applying neural networks to malware detection and feature lear…
A method to identify important features without solving the full problem.
APD method decomposes neural network parameters into simple, faithful components.
Random Fourier features model reconstructs wind fields from sparse measurements.
We study black-box attacks on machine learning classifiers where each query to the model incurs some cost or risk of detection to the adversary. We focus explicitly on minimizing the number of queries as a major objective. Specifically, we consider the problem of attacking machine learning classifiers subject to a budg…
For the problem of binary linear classification and feature selection, we propose algorithmic approaches to classifier design based on the generalized approximate message passing (GAMP) algorithm, recently proposed in the context of compressive sensing. We are particularly motivated by problems where the number of feat…
Recent unsupervised approaches to domain adaptation primarily focus on minimizing the gap between the source and the target domains through refining the feature generator, in order to learn a better alignment between the two domains. This minimization can be achieved via a domain classifier to detect target-domain feat…
For massive data sets, efficient computation commonly relies on distributed algorithms that store and process subsets of the data on different machines, minimizing communication costs. Our focus is on regression and classification problems involving many features. A variety of distributed algorithms have been proposed …
Kernel models learn low-dimensional predictive subspaces from input data.
We extend the work of Narasimhan and Bilmes [30] for minimizing set functions representable as a dierence between submodular functions. Similar to [30], our new algorithms are guaranteed to monotonically reduce the objective function at every step. We empirically and theoretically show that the per-iteration cost of ou…
Proposes IIB for domain generalization, overcoming failure modes of IRM.
EMAP finds minimal perturbations to change model predictions, combining feature weighting and counterfactuals.
Deep models learn spurious features correlated with target, but can still perform well.
Optimal AFs minimize RFR test error and sensitivity.
GRANITE unifies feature-based explanation methods to reduce disagreement.
We compute the condition of minimality of a G-structure for the Gray-Hervella class of almost hermitian manifolds and class of almost contact metric structures. We also consider class by comparison with the Grey-Hervella class . The common feature is the ex…
Scientific explanation often requires inferring maximally predictive features from a given data set. Unfortunately, the collection of minimal maximally predictive features for most stochastic processes is uncountably infinite. In such cases, one compromises and instead seeks nearly maximally predictive features. Here, …
We establish parabolicity and quadratic area growth for minimal surfaces-with-boundary contained in regions of R^3 which are within a sub-logarithmic factor of the exterior of a cone. Unlike previous work showing that these two properties hold for minimal surfaces-with-boundary contained between two catenoids, we do no…
Distributed machine learning has been widely studied in order to handle exploding amount of data. In this paper, we study an important yet less visited distributed learning problem where features are inherently distributed or vertically partitioned among multiple parties, and sharing of raw data or model parameters amo…
Label noise SGD converges to a simple model with a single linear feature.
Study finds minimizers for complex membrane models without symmetry assumptions.
New method MRI improves machine learning models' ability to generalize to unseen data.
DRSS method identifies unnecessary samples and features in DR covariate shift.
In this study, a novel feature coding method that exploits invariance for transformations represented by a finite group of orthogonal matrices is proposed. We prove that the group-invariant feature vector contains sufficient discriminative information when learning a linear classifier using convex loss minimization. Ba…
Improves convex biclustering for high-dimensional data.
In recent years, there have been significant efforts on mitigating unethical demographic biases in machine learning methods. However, very little is done for kernel methods. In this paper, we propose a new fair kernel regression method via fair feature embedding (FKR-FE) in kernel space. Motivated by prior works on…
DDLK uses deep learning to find important features in models.
Optimal biomarker combinations for treatment-selection can be derived by minimizing total burden to the population caused by the targeted disease and its treatment. However, when multiple biomarkers are present, including all in the model can be expensive and hurt model performance. To remedy this, we consider feature …
The paper proposes criteria and methods for evaluating and aggregating feature-based model explanations.