A novel graph spectral method for mixed categorical and numerical data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new method generates mixed-type features in tabular data with improved realism and accuracy.
Proposes CDTD, a diffusion model for mixed-type tabular data.
There are time series that are amenable to recurrent neural network (RNN) solutions when treated as sequences, but some series, e.g. asynchronous time series, provide a richer variation of feature types than current RNN cells take into account. In order to address such situations, we introduce a unified RNN that handle…
Phytoplankton plays an important role in marine ecosystem. It is defined as a biological factor to assess marine quality. The identification of phytoplankton species has a high potential for monitoring environmental, climate changes and for evaluating water quality. However, phytoplankton species identification is not …
New method generates diverse EHR data types while maintaining privacy.
gOMP algorithm selects features for various types of data.
The paper explores features from orderbooks to improve intraday electricity price forecasting.
We investigate a class of feature allocation models that generalize the Indian buffet process and are parameterized by Gibbs-type random measures. Two existing classes are contained as special cases: the original two-parameter Indian buffet process, corresponding to the Dirichlet process, and the stable (or three-param…
Behavior of a malware varies with respect to malware types. Therefore,knowing type of a malware affects strategies of system protection softwares. Many malware type classification models empowered by machine and deep learning achieve superior accuracies to predict malware types.Machine learning based models need to do …
As machine learning is applied to an increasing variety of complex problems, which are defined by high dimensional and complex data sets, the necessity for task oriented feature learning grows in importance. With the advancement of Deep Learning algorithms, various successful feature learning techniques have evolved. I…
New unsupervised feature selection method for imbalanced datasets.
Proposes a new method to find features affecting treatment effect distribution.
Paper proposes BSP to find stable bimodules of cross-correlated features.
VAEM extends VAEs to handle mixed-type data heterogeneity.
In this paper, we aim to develop a unified view of causal and non-causal feature selection methods. The unified view will fill in the gap in the research of the relation between the two types of methods. Based on the Bayesian network framework and information theory, we first show that causal and non-causal feature sel…
In this paper we demonstrate predicting electroencephalograpgy (EEG) features from acoustic features using recurrent neural network (RNN) based regression model and generative adversarial network (GAN). We predict various types of EEG features from acoustic features. We compare our results with the previously studied p…
Automated feature selection is important for text categorization to reduce the feature size and to speed up the learning process of classifiers. In this paper, we present a novel and efficient feature selection framework based on the Information Theory, which aims to rank the features with their discriminative capacity…
New method selects causal features from diverse data types.
Random neural nets learn Black-Scholes PDEs without dimensionality issues.
Proposes a new model to predict polymer properties by integrating various data types.
Background: Depression has become a major health burden worldwide, and effective detection depression is a great public-health challenge. This Electroencephalography (EEG)-based research is to explore the effective biomarkers for depression recognition. Methods: Resting state EEG data was collected from 24 major depres…
In this thesis, we study the problem of feature learning on heterogeneous knowledge graphs. These features can be used to perform tasks such as link prediction, classification and clustering on graphs. Knowledge graphs provide rich semantics encoded in the edge and node types. Meta-paths consist of these types and abst…
Enhances random forest performance with exogenous randomness.
Automatic classification of epileptic seizure types in electroencephalograms (EEGs) data can enable more precise diagnosis and efficient management of the disease. This task is challenging due to factors such as low signal-to-noise ratios, signal artefacts, high variance in seizure semiology among epileptic patients, a…
Method uses aggregate crop statistics to improve satellite-based crop type mapping.
We prove a version of the Arezzo-Pacard-Singer blow-up theorem in the setting of Poincaré type metrics. We apply this to give new examples of extremal Poincaré type metrics. A key feature is an additional obstruction which has no analogue in the compact case. This condition is conjecturally related to ensuring the metr…
Two new regularization methods improve neural network performance and complexity control.
This paper compares methods for handling mixed-attribute data in GFMM neural networks.
We develop a simple and computationally efficient significance test for the features of a machine learning model. Our forward-selection approach applies to any model specification, learning task and variable type. The test is non-asymptotic, straightforward to implement, and does not require model refitting. It identif…
AdaEnsemble learns adaptive feature interactions for CTR prediction.
Pyramidal GNN combines RC and pooling for efficient graph embeddings.
The paper investigates causal relationships in heart failure prediction using machine learning.
Aggregates predictions from multiple regression models using random projections and kernel methods.
In this paper we provide a new Bennequin-type inequality for the Rasmussen- Beliakova-Wehrli invariant, featuring the numerical transverse braid invariants (the c-invariants) introduced by the author. From the Bennequin type-inequality, and a combinatorial bound on the value of the c-invariants, we deduce a new computa…
Two new Hie-TAN and Hie-TAN-Lite algorithms improve TAN for hierarchical feature spaces.
In the integrative analyses of omics data, it is often of interest to extract data representation from one data type that best reflect its relations with another data type. This task is traditionally fulfilled by linear methods such as canonical correlation analysis (CCA) and partial least squares (PLS). However, infor…
AFD learns features resistant to adversarial attacks.
DCoM uses deep neural networks to detect semantic data types from raw column values.
It has long been recognized that the invariance and equivariance properties of a representation are critically important for success in many vision tasks. In this paper we present Steerable Convolutional Neural Networks, an efficient and flexible class of equivariant convolutional networks. We show that steerable CNNs …
Motivation: State-of-the-art biomedical named entity recognition (BioNER) systems often require handcrafted features specific to each entity type, such as genes, chemicals and diseases. Although recent studies explored using neural network models for BioNER to free experts from manual feature engineering, the performan…
We introduce a method called multi-scale local shape analysis, or MLSA, for extracting features that describe the local structure of points within a dataset. The method uses both geometric and topological features at multiple levels of granularity to capture diverse types of local information for subsequent machine lea…
The paper studies multiple descent in multi-component prediction models.
GraphDINO learns neuronal morphologies from unlabeled data.
Theoretical analysis of vision transformers' performance with MAE and CL objectives.
Cold-start is a very common and still open problem in the Recommender Systems literature. Since cold start items do not have any interaction, collaborative algorithms are not applicable. One of the main strategies is to use pure or hybrid content-based approaches, which usually yield to lower recommendation quality tha…
tsflex speeds up time series processing and feature extraction.
The study analyzes decision trees on real and categorical features, deriving bounds on their VC dimension and proposing improved pruning algorithms.