Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

25.0%50.0%75.0%100.0% · Dec 199219922001200920182026
48 results for local feature size

Adaptive Siamese network improves local feature descriptor learning efficiency.

problem Estimating the size of neural networks for local feature descriptors.
method Adaptive pruning Siamese architecture based on neuron activation.
result Learned local feature descriptors outperform state-of-the-art methods in patch matching.

Paper uses deep learning to improve thermal-hydraulic simulations.

problem Limited credibility of thermal-hydraulic codes in real plant conditions.
method Feature Similarity Measurement (FSM) and deep learning.
result Deep learning constructs relationships between local physical features and simulation errors.

Method learns feature maps from deep CNN layers for weakly supervised chest pathology localization.

problem Localization of chest pathologies in X-ray images is challenging due to varying sizes and appearances.
method Class-aware deep multiscale feature learning using intermediate feature maps from CNN layers.
result Improves localization performance of small pathologies like nodules and masses.

New molecular descriptors improve machine learning for small and large molecules.

problem Improving machine learning models for molecular properties.
method Developed constant-size molecular descriptors combining connectivity counts and encoded distances.
result Models using these descriptors perform comparably to or better than state-of-the-art models.

MR-GNN predicts interactions between structured entities using multi-resolution and dual graph neural networks.

problem Predicting interactions between structured entities, especially considering features in substructures of different sizes and interactions between entities.
method MR-GNN uses a multi-resolution architecture and dual graph-state L-STMs to extract features from different neighborhoods and pairwise graphs, respectively.
result MR-GNN improves prediction accuracy compared to state-of-the-art methods.

A new framework SIMBA improves graph classification performance on size-imbalanced datasets.

problem Size imbalance in graph classification leads to poor model performance.
method Energy-guided structural smoothing between head and tail graphs, re-weighting based on energy propagation.
result SIMBA outperforms existing methods in size-imbalanced graph classification tasks.

Framework learns asymmetric and local features in multi-dimensional data.

problem Learning features in multi-dimensional data, especially images.
method Bayesian hierarchical modeling with recursive wavelet transforms.
result Framework achieves high computational scalability and adaptivity.

Locally sparse neural networks improve interpretability for biomedical tabular data.

problem Overfitting and lack of interpretability in neural networks for tabular biomedical data.
method Locally sparse neural network with a gating network to select relevant features.
result The method outperforms state-of-the-art models in synthetic and real-world biomedical datasets.

Local convolutions bias neural networks towards high-frequency adversarial examples.

problem High-frequency adversarial examples in neural networks.
method Analysis of different linear and nonlinear architectures, focusing on the impact of local convolution operations.
result Local convolutions induce an implicit bias towards high frequency features, leading to high-frequency adversarial examples.

New method uncovers global topology through local interactions, reducing algorithm complexity.

problem Global interaction is necessary for forming feature maps that preserve global topology.
method Competing agents engage in local interactions to form feature maps without global interaction.
result Local interactions can uncover global topology, leading to consistent map quality across diverse datasets.

New method improves Gaussian kernel approximations for high-frequency data.

problem Limited scalability of kernel-based models to large data sets.
method Local random feature approximations using Maclaurin expansions and polynomial sketches.
result Significant improvement in kernel approximations and downstream performance for high-frequency data.

GOTabPFN improves tabular model performance with compact tokenization for HDLSS data.

problem Making tabular models effective for high-dimensional, low-sample size data without retraining.
method Introducing Graph-guided Ordering with Local Refinement (GO-LR) and Neuro-Inspired Subunit Compression (NSC) to create compact meta-features.
result GOTabPFN improves stability and accuracy in tabular benchmarks with compact tokenization.

New method explains deep neural networks by ranking feature importance.

problem Limited ability to explain deep neural networks.
method Proposes a novel approach to global feature ranking in DNNs, leveraging partial covariance structures and variable dependence.
result Demonstrates improved feature ranking and interpretation in various domains.

A new test for conditional independence adapts to nonlinear dependencies efficiently.

problem Testing conditional independence in nonlinear and high-dimensional data.
method Nearest-neighbor estimator of conditional mutual information combined with local permutation scheme.
result The test reliably simulates null distribution and is better calibrated for non-smooth densities.

Adaptive batch sizes improve local gradient methods in distributed training.

problem Communication bottlenecks in distributed deep learning.
method Adaptive batch size strategies for local gradient methods.
result Adaptive batch sizes reduce minibatch gradient variance and improve training efficiency.

A framework for collaborative learning reduces communication rounds.

problem Collaborative learning with distributed features and privacy concerns.
method Federated Stochastic Block Coordinate Descent (FedBCD) algorithm.
result The algorithm achieves O(T)O(\sqrt{T}) communication rounds and O(1/T)O(1/\sqrt{T}) accuracy.

Two graph auto-encoders decouple feature propagation from graph convolution layers.

problem Designing efficient graph auto-encoders with fixed receptive fields.
method L-GAE and L-VGAE using linear matrix computation before auto-encoder input.
result Comparable performance to VGAEs with smaller, simpler networks.

Study uses machine learning to detect early COVID-19 from CT images.

problem Early detection of COVID-19 from CT images.
method Machine learning methods applied to patches of CT images, feature extraction (GLCM, LDP, GLRLM, GLSZM, DWT), SVM classification.
result Best classification accuracy of 99.68% with 10-fold cross-validation and GLSZM feature extraction.

New graph kernel scales well with graph size and number, achieving state-of-the-art performance.

problem Graph kernels lose structure information when representing graphs.
method Proposes a positive-definite global alignment graph kernel using random features and random graph embeddings.
result Achieves quasi-linear scalability with respect to graph size and number.

Algorithm segments glandular structures in colon histology images for cancer grading.

problem Manual gland segmentation is time-consuming and risky for patients.
method Local intensity and texture features, Random Forest classifier, multilevel approach.
result Fast, accurate automatic gland segmentation for clinical use.

NN-Stacking improves predictive power of regression models by adjusting stacking coefficients with features.

problem Low predictive power of linear stacking methods.
method NN-Stacking uses neural networks to estimate adaptive stacking coefficients.
result NN-Stacking leads to better predictive power, especially in large datasets.

The paper adapts step sizes in TD learning to identify relevant features.

problem Identifying which features are relevant for temporal-difference learning.
method Adapting step sizes in stochastic gradient descent for feature relevance in TD learning.
result TD IDBD effectively distinguishes relevant features in gridworld and robotic tasks.

Flexible classifier using Mahalanobis distances for non-elliptical distributions.

problem Classifying non-elliptical and multimodal distributions.
method Semiparametric classifier based on Mahalanobis distances and generalized additive models.
result The proposed classifiers outperform traditional methods in high-dimensional, low-sample-size scenarios.

Study shows DNNs often extract redundant features, influenced by network size and activation function.

problem Redundancy in deep neural network features.
method Hierarchical clustering of features based on cosine distances, varying network sizes and activation functions.
result Network size and activation function are key factors in DNN redundancy.

The paper investigates why GNNs struggle to generalize from small to large graphs.

problem Challenges in graph neural networks' ability to generalize across different graph sizes.
method Identified and studied the effect of local structure on size generalization; proposed a novel SSL task.
result GNNs can converge to non-generalizing solutions when there is a discrepancy in local structure.

Improved variational inequality algorithms using adaptive step sizes.

problem Solving monotone variational inequalities and convex-concave min-max problems efficiently.
method Adaptive step sizes that eliminate hyperparameters and global Lipschitz continuity requirements.
result Eliminated the need for the golden ratio in the algorithm and improved complexity bounds.

Study ridge ensembles in proportional feature-to-sample size regime, proving risk equivalence and GCV consistency.

problem Characterizing and optimizing ridge ensembles in proportional feature-to-sample size regimes.
method Proportional asymptotics analysis, GCV for tuning, proving risk equivalence.
result Risk of optimal full ridgeless ensemble matches optimal ridge predictor's risk.

Ensembles of random-feature models can't outperform a single large model.

problem Finding the optimal balance between model size and ensemble size.
method Deterministic equivalent risk estimates and scaling laws analysis.
result Ensembles of random-feature models achieve near-optimal performance only under specific conditions.

Geometric step decay schedules improve stochastic algorithms' convergence on sharp nonconvex problems.

problem Convergence of stochastic algorithms on sharp nonconvex problems.
method Geometric step decay schedule applied to stochastic algorithms.
result Geometric step decay schedules lead to local linear convergence rates for sharp nonconvex problems.

NIS learns optimal embedding sizes for recommendation models.

problem Finding optimal embedding sizes for large-scale recommendation models.
method Neural Input Search (NIS) uses reinforcement learning to automatically find optimal vocabulary sizes and embedding dimensions.
result NIS improves prediction accuracy by 6.8% on Recall@1 and 1.8% on ROC-AUC.

New method learns domain-invariant local feature patterns for unsupervised domain adaptation.

problem Performance degradation due to domain-shift in unsupervised domain adaptation.
method Jointly learns domain-invariant local feature patterns and holistic feature distributions.
result Superior performance on benchmark datasets compared to state-of-the-art methods.