Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

100199299398 · Jun 202019922001200920172026
48 results for feature abundance

Paper analyzes weak-to-strong generalization in CNNs, identifying data-scarce and data-abundant regimes.

problem Weak-to-strong generalization in CNNs trained on weak models.
method Formal analysis of gradient descent dynamics in data-scarce and data-abundant regimes.
result Identifies two regimes and distinct mechanisms of generalization in each.

Study of superintegrable systems linked to affine hypersurfaces.

problem Understanding superintegrable systems through geometric structures.
method Established a correspondence between superintegrable systems and affine hypersurfaces, defining conformal equivalence.
result Identified conformal classes of abundant manifolds with abundant hypersurface immersions.

We propose a novel high-performance and interpretable canonical deep tabular data learning architecture, TabNet. TabNet uses sequential attention to choose which features to reason from at each decision step, enabling interpretability and more efficient learning as the learning capacity is used for the most salient fea…

2019-08-20abs ↗pdf ↗

Method maps imperfect simulations to observed stellar spectra using unsupervised domain adaptation.

problem Mapping from large sets of imperfect simulations and observational data.
method Adversarial autoencoders, cycle-consistency constraint, and generative surrogate physics emulator network.
result Reconstructed spectra quality and discovery of new spectral features.

The paper proves conditions for the Abundance conjecture in minimal projective klt pairs.

problem Proving the Abundance conjecture for minimal klt pairs with non-zero canonical bundle.
method Analyzing asymptotic behavior of multiplier ideals and properties of supercanonical currents.
result Supercanonical currents are central to proving the Abundance conjecture.

Study shows torsion grows subexponentially in book of I-bundles but can grow exponentially in non-regular covers.

problem Growth rates of torsion in book of I-bundles.
method Analysis of torsion in homology of book of I-bundles using finite-sheeted covers.
result Torsion growth rates differ between regular and non-regular finite-sheeted covers.

Study reveals DNNs prefer easy-to-learn cues over essential ones in image recognition.

problem DNNs learn easy-to-learn features that aren't essential to the task.
method WCST-ML training setup with shortcut cues on synthetic and face datasets.
result DNNs converge to solutions focusing on preferred cues, leading to flat minima.

A novel method uses blockchain transaction graphs for Bitcoin price prediction.

problem Insufficient effectiveness of manually designed features for Bitcoin price prediction.
method Mining patterns from Bitcoin transactions using k-order transaction graphs and proposing a novel prediction method.
result The proposed method outperforms state-of-the-art Bitcoin price prediction methods.

The paper proves conditions for minimal compact Kähler manifolds with vanishing second Chern class.

problem Conditions for minimal compact Kähler manifolds with vanishing second Chern class.
method Study of the abundance conjecture and associated Iitaka fibrations.
result For a minimal compact Kähler manifold, the second Chern class vanishes if and only if the cotangent bundle is nef and the canonical bundle has numerical dimension 0 or 1.

Predictive analytics is increasingly used to guide decision-making in many applications. However, in practice, we often have limited data on the true predictive task of interest, and must instead rely on more abundant data on a closely-related proxy predictive task. For example, e-commerce platforms use abundant custom…

2018-12-28abs ↗pdf ↗

Meta-learning, or learning to learn, is a machine learning approach that utilizes prior learning experiences to expedite the learning process on unseen tasks. As a data-driven approach, meta-learning requires meta-features that represent the primary learning tasks or datasets, and are estimated traditonally as engineer…

2019-05-27abs ↗pdf ↗

We formulate and prove that there are "abundant" in nilpotent orbits in real semisimple Lie algebras, in the following sense. If S denotes the collection of hyperbolic elements corresponding the weighted Dynkin diagrams coming from nilpotent orbits, then S span the maximally expected space, namely, the (-1)-eigenspace …

2016-12-09abs ↗pdf ↗

Improves unsupervised domain adaptation by mixing source and target domains.

problem Improves unsupervised domain adaptation by mixing source and target domains.
method Enforces training constraints across domains using mixup formulation and feature-level consistency regularizer.
result Significantly improves state-of-the-art performance on image classification and human activity recognition tasks.

New product structures encode superintegrable Hamiltonian systems in Euclidean spaces.

problem Encoding superintegrable Hamiltonian systems using product structures.
method Introducing commutative and associative product structures on Euclidean spaces of dimension at least three, satisfying specific conditions.
result All abundant superintegrable Hamiltonian systems on Euclidean space of dimension at least three arise from these product structures.

PRESTO improves rare event prediction by shrinking towards proportional odds model.

problem Difficult to predict rare events due to class imbalance.
method PRESTO relaxes proportional odds model by estimating separate weights for transitions between categories, imposing L1 penalty to shrink towards proportional odds.
result PRESTO consistently estimates decision boundary weights under sparsity assumption, improving rare probability estimation.

The paper proposes a method to analyze categorical feature interactions in large datasets using graph covariance and LLMs.

problem Analyzing complex datasets with numerous categorical features and timestamps.
method Binarization of categorical features using one-hot encoding, computation of graph covariance, identifying significant feature pairs, and using LLMs to generate explanations.
result The method identifies meaningful feature pairs and potential data stories underlying categorical feature interactions.

The study finds abundant normal generators for mapping class groups.

problem Understanding normal generation in mapping class groups.
method Analyzing restrictions on invariant subsurfaces and Teichmüller spaces.
result Reducible mapping classes can normally generate mapping class groups based on their asymptotic translation lengths.

DC-SIS selects features faster than mRMR for Parkinson's vocal diagnosis.

problem Feature selection for Parkinson's disease vocal data.
method DC-SIS (Distance Correlation Sure Independence Screening) using distance correlation measure.
result 90 times faster feature selection with similar accuracy.

A new method resolves permutation issues in shuffled linear regression for large-scale applications.

problem Estimating latent features through linear transformation with unknown permutations.
method Spectral matching method to align spectral components of measurement and feature covariances.
result Achieves accurate estimates in shuffled LS and LASSO settings with sufficient samples.

Machine learning on graph structured data has attracted much research interest due to its ubiquity in real world data. However, how to efficiently represent graph data in a general way is still an open problem. Traditional methods use handcraft graph features in a tabular form but suffer from the defects of domain expe…

2019-11-11abs ↗pdf ↗

DWTS uses observational data to improve clinical trial efficiency.

problem Lack of definitive conclusions from randomized clinical trials due to insufficient patient cohorts and confounding biases.
method DWTS combines observational data with randomized clinical trials using Doubly Debiased LASSO (DDL) to identify reliable covariates.
result DWTS reduces cumulative regret in clinical trials compared to standard methods.

Study symplectically aspherical Kähler manifolds with unique properties.

problem Existence and properties of symplectically aspherical Kähler manifolds.
method Detailed study and analysis of geometric and topological features.
result Existence of symplectically aspherical Kähler manifolds with large fundamental groups.

New method uses neural networks to infer dark matter subhalo abundance from stellar streams.

problem Constrain warm dark matter mass using stellar streams.
method Amortized Approximate Likelihood Ratios (AALR) for likelihood-free Bayesian inference.
result Demonstrates effectiveness of new method for estimating dark matter subhalo abundance.

We construct examples of complete Riemannian manifolds having the property that every geodesic lies in a totally geodesic hyperbolic plane. Despite the abundance of totally geodesic hyperbolic planes, these examples are not locally homogenous.

2016-12-06abs ↗pdf ↗

Develops fair feature importance scores for tree-based models to interpret fairness.

problem Ensuring fairness in machine learning models, especially tree-based ones.
method Inspired by decision trees, proposes a novel fair feature importance score based on mean decrease in group bias.
result Valid interpretations of fairness for tree-based ensembles and surrogates of other ML systems.

LoCEC classifies user relationships in large social networks, addressing sparsity issues.

problem Sparse relationship feature and label data in real social platforms.
method Local Community-based Edge Classification (LoCEC) framework with three-phase processing.
result Effective and efficient classification of user relationships in large-scale networks.