Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

229458686915 · Jun 202019922001200920182026
48 results for small feature space

GOAL algorithm reduces and rotates feature space for small data classification.

problem Challenges in identifying important features for classification in small data settings.
method GOAL algorithm reduces and rotates feature space in a lower-dimensional gauge, providing an analytically tractable solution.
result GOAL algorithm outperforms state-of-the-art ML tools in synthetic and real-world applications.

GapTV improves interpretability in small feature spaces.

problem Estimating regression functions with small feature sets and high interpretability needs.
method Divides feature space into blocks, fits values jointly using convex optimization, incorporates automatic hyperparameter tuning.
result GapTV finds a better balance between accuracy and interpretability compared to CART and CRISP.

Unified Bayesian model for multi-modal, small sample size biomedical data classification.

problem Classifying high-dimensional, multi-modal biomedical data with small sample sizes.
method Combines multi-modal data views into a latent space, prunes irrelevant features, and uses dual kernels for small sample size scenarios.
result Outperforms state-of-the-art models and identifies features aligned with existing markers.

A method for clustering small datasets in high dimensions using random projections.

problem Challenges in clustering small datasets in high-dimensional spaces.
method Random projection followed by binary clustering in one-dimensional space.
result Statistically significant clustering structures can be found with as few as 100-200 points.

Proposes a method to extract robust features that improve classifier robustness.

problem Improving classifier robustness to small perturbations in input space.
method Introduces an additional penalty term in the information bottleneck framework to minimize Fisher information, optimizing a variational bound using stochastic gradient descent.
result Optimally robust features are jointly Gaussian, and the method produces classifiers with increased robustness to perturbations.

Poor approximators found in neural networks and random feature models.

problem Understanding why certain neural networks and models perform poorly in approximating functions.
method Established a scale separation of Kolmogorov width type and applied it to neural networks and random feature models.
result Reproducing kernel Hilbert spaces and two-layer neural networks are poor L2L^2-approximators for certain functions.

Study compares deep feature methods for anomaly detection in limited data scenarios.

problem Handling limited data in industrial inspection applications.
method Three approaches (KNN, Mahalanobis, PaDiM) using pre-trained deep features with data augmentation.
result Data augmentation significantly improves performance in small data regimes.

Feature selection is frequently used as a pre-processing step to machine learning. It is a process of choosing a subset of original features so that the feature space is optimally reduced according to a certain evaluation criterion. The central objective of this paper is to reduce the dimension of the data by finding a…

2014-01-05abs ↗pdf ↗

Proposes FBFAN to defend against adversarial attacks by learning semantic features.

problem Vulnerability of deep neural networks to adversarial attacks.
method Featurized Bidirectional Generative Adversarial Networks (FBGAN) that learns semantic features and filters non-semantic perturbations.
result FBGAN effectively reconstructs adversarial data to denoised data, improving classifier performance.

Centroids Matching tackles catastrophic forgetting by matching feature vectors to class centroids.

problem Catastrophic forgetting in neural networks when learning new tasks.
method Centroids Matching operates in the embedding space of neural network features, matching these vectors to class centroids.
result Centroids Matching achieves high accuracy on all tasks without using external memory, even in realistic scenarios.

ALIEN detects multiple small objects and estimates their pixel locations and features.

problem Detecting and localizing multiple small objects in a scene.
method Deep-learning network (ALIEN) that performs simultaneous x,y pixel estimation and feature extraction in a single forward pass.
result Efficient detection and localization of hundreds to thousands of small objects.

Locality-sensitive hashing converts high-dimensional feature vectors, such as image and speech, into bit arrays and allows high-speed similarity calculation with the Hamming distance. There is a hashing scheme that maps feature vectors to bit arrays depending on the signs of the inner products between feature vectors a…

2012-12-26abs ↗pdf ↗

Random small feature subsets outperform FS in diverse datasets.

problem The significance of selected features in high-dimensional datasets is questionable.
method Analysis of 28 diverse datasets (microarray, RNA-Seq, etc.).
result Any arbitrary set of features performs as well as or better than selected features across datasets.

A new method for few-sample FS using manifold learning.

problem Few-sample supervised feature selection in high-dimensional spaces.
method Learn feature associations on manifolds, compute composite kernel, and use spectral analysis for FS score.
result Our method outperforms competitors in feature selection and classification accuracy.

Random feature model shows slow self-correction of generalization gap.

problem Slow deterioration of generalization error in random feature model.
method Examined the dynamic behavior of gradient descent in the model's resonance regime.
result Gradient descent exhibits a self-correction mechanism, reducing generalization gap over time.

Adaptive Siamese network improves local feature descriptor learning efficiency.

problem Estimating the size of neural networks for local feature descriptors.
method Adaptive pruning Siamese architecture based on neuron activation.
result Learned local feature descriptors outperform state-of-the-art methods in patch matching.

This paper improves prediction in small data sets by eliciting expert knowledge about feature similarities.

problem Improving predictive models from small high-dimensional data sets.
method Eliciting expert knowledge about pairwise feature similarities and using sequential decision making techniques.
result Improvement in predictive performance on both simulated and real data.

Large learning rates cause oscillations in NN weights that improve generalization.

problem Improving generalization of neural networks trained with large learning rates.
method Theoretical analysis and feature-noise data generation model.
result Oscillating SGD with large learning rates benefits NN generalization by effectively learning weak features.

The paper develops a method to select features from multiple kernels for efficient risk minimization.

problem Identifying promising features leading to satisfactory out-of-sample performance in nonlinear kernel approximation.
method A greedy selection process using a correlation metric to choose features from multiple kernels.
result An out-of-sample error bound capturing trade-offs between approximation and spectral errors, showing poly-logarithmic scaling with data.

IMKPL learns interpretable prototypes for better classification.

problem Efficient trade-offs between interpretability and prediction accuracy in kernel-based data.
method Local discrimination in feature space, condensed class-homogeneous neighborhoods, combined embedding.
result IMKPL achieves better interpretability and discriminative representation.

Improves MKL for multi-class classification with better feature selection and representation.

problem Real-world multi-class classification problems with non-linear separations.
method Large-margin multiple kernel learning (LMMK) with sparsity term for discriminative feature selection.
result Competitive classification accuracy and sparse non-zero kernel weights.

Generative model initializes 2-layer network weights for small datasets.

problem Approximating functions with 2-layer networks using small datasets and gradient-based training.
method Initialize hidden weights with a learned proposal distribution parameterized as a deep generative model. Refine with gradient-based post-processing and regularization.
result Demonstrates effectiveness of the approach with numerical examples.

New geometric interpretation explains over-parameterized models and adversarial perturbations.

problem Geometric understanding of over-parameterized regression and adversarial perturbations.
method Alternative geometric interpretation of regression in feature space.
result Adversarial perturbations are a natural feature of biased models due to underlying geometry.

We propose a novel model for generating graphs similar to a given example graph. Unlike standard approaches that compute features of graphs in Euclidean space, our approach obtains features on a surface of a hypersphere. We then utilize a von Mises-Fisher distribution, an exponential family distribution on the surface …

2011-05-15abs ↗pdf ↗

WeatherFormer learns robust weather features from small datasets.

problem Modeling complex weather dynamics from limited data.
method Pretrained transformer encoder on large satellite dataset, with spatiotemporal encoding.
result State-of-the-art performance in county-level soybean yield prediction and influenza forecasting.

Study of local optima in neural networks for feature interactions.

problem NNs struggle with local optima in feature interactions for small datasets.
method Proposed a node pruning and feature selection algorithm to improve NN performance.
result NNs have many non-equivalent local optima in XOR-like data with irrelevant variables.

Interactive tool improves prediction accuracy in small datasets.

problem Challenges in machine learning with small data sets and tacit expert knowledge.
method Interactive visualization and user model to elicit feature relevance.
result User model significantly improves prediction accuracy and prior knowledge elicitation.

Paper develops a novel approach for unsupervised dimension selection.

problem Tackles the combinatorial problem of identifying top-k dimensions in high-dimensional data.
method Develops a novel approach based on graph signal analysis to measure feature influence.
result Demonstrates the superiority of the proposed approach over existing techniques in capturing crucial characteristics of high-dimensional spaces using only a small subset of features.

Paper analyzes robustness of non-Lipschitz networks, proving powerful adversarial attacks but offering solutions.

problem Adversarial attacks on deep networks, especially non-Lipschitz networks.
method Developed an attack model that abstracts the challenge of adversarial robustness, proving the power of such attacks and offering solutions.
result Proves powerful adversarial attacks on non-Lipschitz networks but offers solutions with abstention.

Novel method converts time series data into functional data for high dimensional classification.

problem Small sample size problem in high dimensional time series data.
method Classwise Functional Principal Component Analysis (PCA) followed by Bayesian linear classifier.
result Demonstrated efficacy on synthetic and real data sets.

GVM replaces SVM with general project vectors and a Monte Carlo algorithm for improved feature extraction.

problem Improving feature extraction and reducing overlearning in SVM.
method Replaces support vectors with general project vectors and uses a Monte Carlo algorithm to find them.
result The GVM can achieve better performance, especially for small-set training problems.

Study shows safely discarding features based on aggregate SHAP values is sound.

problem Discarding features based on aggregate SHAP values without proper justification.
method Investigated the soundness of discarding features based on aggregate SHAP values, proposing to aggregate SHAP values over the extended support.
result A small aggregate SHAP value implies safely discarding the corresponding feature.

MVTV improves interpretability in low-dimensional regression.

problem Estimating regression functions with few features and high interpretability needs.
method MVTV divides space into blocks, fits values jointly, and optimizes automatically.
result MVTV outperforms CART and CRISP in both complexity and human interpretability studies.

CM algorithm improves MMI classifications for unseen instances.

problem Improving classification accuracy for unseen instances using MMI criterion.
method Introduces CM algorithm for MMI classifications, combining semantic and Shannon channels for matching.
result Achieves high mutual information (99%) with minimal iterations in low-dimensional feature spaces.

Improved SVM classification with interpretable features from scattered data.

problem Classification of scattered data points in high-dimensional spaces.
method Truncated ANOVA decomposition for sparse feature selection; use of trigonometric or wavelet feature maps.
result Better classification accuracy and interpretability with 1\ell_1-norm regularization.