Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

143285428570 · Jun 202019922001200920182026
48 results for local feature descriptor

Adaptive Siamese network improves local feature descriptor learning efficiency.

problem Estimating the size of neural networks for local feature descriptors.
method Adaptive pruning Siamese architecture based on neuron activation.
result Learned local feature descriptors outperform state-of-the-art methods in patch matching.

Paper introduces a method to learn optimal local feature aggregation functions.

problem Learning optimal local feature aggregation functions for classification.
method Compose local feature aggregation function with classifier cost function and backpropagate gradient to update parameters.
result Method outperforms state-of-the-art local feature aggregation functions.

New molecular descriptors improve machine learning for small and large molecules.

problem Improving machine learning models for molecular properties.
method Developed constant-size molecular descriptors combining connectivity counts and encoded distances.
result Models using these descriptors perform comparably to or better than state-of-the-art models.

Paper proposes CNN with SIFT for rotation invariant feature extraction.

problem Max-pooling layer discards rotational information, leading to rotation invariance issues.
method Uses SIFT descriptor to capture orientation and spatial relationships.
result Improves feature extraction on MNIST and fashionMNIST datasets.

Unsupervised feature learning has shown impressive results for a wide range of input modalities, in particular for object classification tasks in computer vision. Using a large amount of unlabeled data, unsupervised feature learning methods are utilized to construct high-level representations that are discriminative en…

2013-01-14abs ↗pdf ↗

iStructTab improves multimodal learning by optimizing feature sequencing.

problem Redundancy, dispersion, and generalization issues in multimodal learning of images and tabular data.
method Graph-Enhanced Descriptor Sequencing (GEDS) algorithm that refines statistical descriptors through similarity graph-based computations.
result iStructTab effectively minimizes feature dispersion, improving predictive performance and robustness.

Paper proposes a new descriptor for early trajectory characterization in matrix iterations.

problem Comparing early behavior of high-dimensional trajectories in nonlinear matrix iterations.
method Develops a two-channel fuzzy coordinate system using F-transform for compact representation.
result The descriptor achieves high R^2 values (mean = 0.6480) in approximating convergence lengths.

Efficiently extracts local features from whole images using CNNs with pooling layers.

problem Efficiently extracting local features from whole images for various tasks.
method A method to compute patch-based local feature descriptors efficiently in presence of pooling and striding layers for whole images at once, applicable to nearly all existing network architectures.
result Our approach significantly speeds up feature extraction from whole images compared to existing methods.

Proposes a new model to predict polymer properties by integrating various data types.

problem Inaccurate polymer property prediction due to separate modeling of different data types.
method Multi-modal cascade feature transfer using GCN for chemical structure and molecular descriptors.
result Empirically evaluated model shows higher predictive performance than single-feature approaches.

A framework to compare atomistic descriptors and their transformations.

problem Comparing and understanding different atomistic descriptors and their transformations.
method Introducing a framework to compare different sets of descriptors and their transformations by metrics and kernels.
result Diagnostic tools to determine equivalent information and distorted common information between feature spaces.

Optimizes atomic descriptors to reduce redundancy and improve machine learning models.

problem Redundant descriptors in atomistic machine learning models increase computational burden and limit model expressivity.
method Employing techniques from pattern recognition, we refine and augment existing atomistic representations to produce optimal sets of descriptors.
result New architectures recognize up to 5-body patterns with low computational cost and high accuracy.

Informative and discriminative feature descriptors play a fundamental role in deformable shape analysis. For example, they have been successfully employed in correspondence, registration, and retrieval tasks. In the recent years, significant attention has been devoted to descriptors obtained from the spectral decomposi…

2011-10-23abs ↗pdf ↗

Paper introduces stable vectorization for multiparameter PH using signed barcodes.

problem Lack of stable vectorization methods for multiparameter persistent homology.
method Signed barcodes as measures for stable vectorization of MPH.
result Stable feature vectors from signed barcodes improve performance in data science.

Support Vector Machines (SVMs) are powerful learners that have led to state-of-the-art results in various computer vision problems. SVMs suffer from various drawbacks in terms of selecting the right kernel, which depends on the image descriptors, as well as computational and memory efficiency. This paper introduces a n…

2013-07-19abs ↗pdf ↗

Wittgenstein's Rule Following evolves datasets by extrapolating structural descriptors.

problem Generating meaningful continuations of evolving datasets.
method Wittgenstein's Rule Following (WRF) uses structural descriptors to extrapolate trajectories and average historical descriptors.
result WRF can generate meaningful continuations of evolving datasets.

PolyGraph Discrepancy improves graph generative model evaluation.

problem Inability of existing metrics to provide an absolute performance measure and comparability across different graph descriptors.
method Approximates Jensen-Shannon distance using binary classifiers trained to distinguish between real and generated graphs.
result PGD provides a more robust and insightful evaluation compared to MMD metrics.

Method embeds numeric tabular datasets into a shared vector space for similarity and retrieval.

problem Lack of meaningful representation for numeric tabular datasets in large language models.
method Structured exploratory data analysis descriptors, sentence transformer embedding, CCA for cross-dataset alignment.
result Total P@1 score of 0.9 across 15 datasets, robust nearest-neighbor retrieval and cluster structure.

Scheme for online state discovery in financial markets using feature correlations and clustering.

problem Discovering temporal states in high-frequency financial data without human intervention.
method Unbiased Fourier estimator for feature correlations, high-speed clustering algorithm, state space enumeration.
result Feature cluster configuration is a candidate for system state representation.

Proposes a method for weakly-supervised object localization to improve few-shot learning.

problem Challenges of few-shot learning, especially with fine-grained categories.
method Introduces a Self-Attention Based Complementary Module (SAC Module) for weakly-supervised object localization.
result Significantly outperforms state-of-the-art methods on benchmark datasets, especially for fine-grained few-shot tasks.

Method infers domain-specific models without domain semantic descriptors.

problem Poor performance of standard supervised learning methods in unseen domains.
method Introduces latent domain vectors and neural networks for optimization.
result Inference of appropriate domain-specific models without semantic descriptors.

Paper classifies movie genres using multimodal data.

problem Challenging task of multi-label movie genre classification.
method Created dataset from video clips, subtitles, synopses, and posters. Extracted features using various descriptors. Evaluated using different classifiers and late fusion strategy.
result Best F-Score result of 0.628 achieved by combining LSTM on synopses and CNN on movie trailer frames.

LC-GAP uses localized Coulomb descriptors for accurate molecular potential predictions.

problem Creating accurate molecular potentials for large molecules.
method Combining localized Coulomb matrix representations with Gaussian approximation potential.
result LC-GAP generates accurate potentials for molecules larger than training data with chemical accuracy.

HSSE framework embeds single-cell RNA-seq data at multiple scales.

problem Capturing heterogeneous local structure in single-cell RNA-seq data.
method Hierarchical sheaf spectral embedding (HSSE) framework.
result HSSE achieves competitive or improved performance in single-cell RNA-seq data representation learning.

A new method interprets astrophysical spectra using geometric paths to distinguish line profiles.

problem Tackling the indistinguishability of spectral line profiles under scalar summaries.
method Introduces a geometric representation of line profiles using rough path theory, mapping profiles to a common velocity grid and defining descriptors from path properties.
result Compact descriptors separate morphologies with similar scalar summaries, revealing ordered line structures.

New method uses geometric moments for accurate machine learning potentials.

problem Creating high-dimensional potential energy surfaces efficiently.
method Feed-forward neural networks with invariant local molecular descriptors based on geometric moments.
result Accuracy comparable to established models, high efficiency.

New method uses machine learning to predict CO2 reduction catalysts without expensive ab initio calculations.

problem Predicting catalytic activity for CO2 reduction reactions using computationally expensive ab initio methods.
method Combining muffin-tin orbital theory descriptors with machine learning (ANN and KRR) for large-scale screening.
result Predicted CO adsorption energy with 0.05 eV mean absolute deviation, significantly improved over previous methods.

Proposes a new method to describe graph vertex features using characteristic functions.

problem Describing the distribution of vertex features at multiple scales on graphs.
method Introduces FEATHER, a computationally efficient algorithm to calculate characteristic functions based on random walk transition probabilities.
result Demonstrates that the proposed method creates high-quality graph representations and is robust to data corruption.

A novel feature representation method for non-image based features.

problem Inability of Convolutional Neural Networks for non-image based features or features without spatial correlations.
method REFINED: Representation of Features as Images with Neighborhood Dependencies.
result Higher prediction accuracy compared to existing methodologies.

Paper presents a method for recognizing human actions using GLAC features from motion and static images.

problem Action recognition in 3D depth videos.
method 3D Motion Trail Model (3DMTM) for MHIs and SHIs, GLAC features extraction, l2-regularized Collaborative Representation Classifier (l2-CRC) for classification.
result The method outperforms other approaches in recognizing human actions.

New method learns shape correspondences robustly from raw geometry.

problem Inaccurate and poor generalization of shape correspondences.
method Learning-based approach with feature-extraction network and functional map representation.
result Robust and accurate shape correspondence learning with less training data.

This work characterizes topological descriptors of graph products and their expressive power.

problem Capturing multiscale structural information in graph products using topological descriptors.
method Analysis of various filtrations on graph products, including Euler characteristic and persistent homology.
result Persistent homology of graph products contains more information than individual graphs.

This study benchmarks deep learning for unsupervised near-duplicate image detection.

problem Detecting near-duplicates in large image datasets with high specificity.
method Binary classification using Receiver Operating Curve (ROC) for comparison of different descriptors.
result Fine-tuning deep convolutional networks generally outperforms off-the-shelf features, with best performance on MFND dataset.

A new QSAR model selects relevant molecular descriptors for bioactivity prediction.

problem Redundant, noisy, and irrelevant descriptors in QSAR models.
method SPL-Logsum method using regularization and self-paced learning.
result SPL-Logsum method outperforms other methods in classification performance and model interpretability.

A framework separates chemical and structural contributions to aqueous solubility.

problem Merging chemical and structural information in solubility models obscures their relative contributions.
method Additive MLP-GNN framework with separate chemical and structural branches.
result Framework reveals distinct roles of chemical and structural information in solubility.

ALP outperforms other data descriptors in one-class classification.

problem Challenges in one-class classification using data descriptors.
method Determined optimal default hyperparameters for data descriptors, proposed ALP, evaluated using leave-one-dataset-out procedure.
result ALP outperforms other data descriptors, including IF and SVM.