Improves molecular activity prediction using graph convolutional neural networks considering graph distances.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We consider the problem of aligning a pair of databases with jointly Gaussian features. We consider two algorithms, complete database alignment via MAP estimation among all possible database alignments, and partial alignment via a thresholding approach of log likelihood ratios. We derive conditions on mutual informatio…
This paper tackles unpaired data in multi-view learning, proposing a new framework and models.
A new framework for paired-sample testing in high-dimensional data.
Proposes a new clustering algorithm for high-dimensional data.
Link prediction requires predicting which new links are likely to appear in a graph. Being able to predict unseen links with good accuracy has important applications in several domains such as social media, security, transportation, and recommendation systems. A common approach is to use features based on the common ne…
Develops framework to evaluate feature attribution methods.
Pharmaceutical targeting is one of key inputs for making sales and marketing strategy planning. Targeting list is built on predicting physician's sales potential of certain type of patient. In this paper, we present a time-sensitive targeting framework leveraging time series model to predict patient's disease and treat…
The paper proposes a method to analyze categorical feature interactions in large datasets using graph covariance and LLMs.
CURE extracts relations without supervision by clustering similar entity pairs.
Noisy Feature Mixup improves model robustness with noise-perturbed convex combinations.
Deep model learns protein interfaces from high-order interactions.
Proves accuracy guarantees for self-supervised learning with correlated positive pairs.
A principal pair consists of a holomorphic principal -bundle together with a holomorphic section of an associated Kaehler fibration. Such objects support natural gauge theoretic equations coming from a moment map condition, and also admit a notion of stability based on Geometric Invariant Theory. The Hitchin--Kobaya…
Levy copulas are the most general concept to capture jump dependence in multivariate Levy processes. They translate the intuition and many features of the copula concept into a time series setting. A challenge faced by both, distributional and Levy copulas, is to find flexible but still applicable models for higher dim…
Paper proposes BSP to find stable bimodules of cross-correlated features.
We describe the 1st place winning approach for the CIKM Cup 2016 Challenge. In this paper, we provide an approach to reasonably identify same users across multiple devices based on browsing logs. Our approach regards a candidate ranking problem as pairwise classification and utilizes an unsupervised neural feature ense…
This paper proposes an active metric learning method for clustering with pairwise constraints.
MTRGL learns temporal correlations from multi-modal data for improved pair trading.
Proposes an efficient method for ordered counterfactual explanations.
In this paper, we deal with two challenges for measuring the similarity of the subject identities in practical video-based face recognition - the variation of the head pose in uncontrolled environments and the computational expense of processing videos. Since the frame-wise feature mean is unable to characterize the po…
In this work, we ask two questions: 1. Can we predict the type of community interested in a news article using only features from the article content? and 2. How well do these models generalize over time? To answer these questions, we compute well-studied content-based features on over 60K news articles from 4 communit…
GeoLifeCLEF 2020 dataset pairs species observations with environmental data.
Consider a data set collected by (individuals-features) pairs in different times. It can be represented as a tensor of three dimensions (Individuals, features and times). The tensor biclustering problem computes a subset of individuals and a subset of features whose signal trajectories over time lie in a low-dimensiona…
Mutual Information (MI) is often used for feature selection when developing classifier models. Estimating the MI for a subset of features is often intractable. We demonstrate, that under the assumptions of conditional independence, MI between a subset of features can be expressed as the Conditional Mutual Information (…
Novel unsupervised feature selection method using multi-step Markov transition probability.
We employ a wavelet approach and conduct a time-frequency analysis of dynamic correlations between pairs of key traded assets (gold, oil, and stocks) covering the period from 1987 to 2012. The analysis is performed on both intra-day and daily data. We show that heterogeneity in correlations across a number of investmen…
PatchUp improves CNN robustness with mixed feature blocks.
Obtaining common representations from different modalities is important in that they are interchangeable with each other in a classification problem. For example, we can train a classifier on image features in the common representations and apply it to the testing of the text features in the representations. Existing m…
A new estimator, OddSHAP, simplifies Shapley value computation by focusing on odd components.
A new weighted FDA method improves face recognition accuracy.
Recently we proposed a general, ensemble-based feature engineering wrapper (FEW) that was paired with a number of machine learning methods to solve regression problems. Here, we adapt FEW for supervised classification and perform a thorough analysis of fitness and survival methods within this framework. Our tests demon…
Can neural networks learn to compare graphs without feature engineering? In this paper, we show that it is possible to learn representations for graph similarity with neither domain knowledge nor supervision (i.e.\ feature engineering or labeled graphs). We propose Deep Divergence Graph Kernels, an unsupervised method …
The paper introduces a new pairs trading model using nonlinear and non-Gaussian state-space models.
Training features used to analyse physical processes are often highly correlated and determining which ones are most important for the classification is a non-trivial tasks. For the use case of a search for a top-quark pair produced in association with a Higgs boson decaying to bottom-quarks at the LHC, we compare feat…
Details of quantum knot invariant calculations using a specific SU(3)_q-module are given which distinguish the Conway and Kinoshita-Teresaka pair of mutant knots. Features of Kuperberg's skein-theoretic techniques for SU(3)_q invariants in the context of mutant knots are also discussed.
Predicting the click-through rate of an advertisement is a critical component of online advertising platforms. In sponsored search, the click-through rate estimates the probability that a displayed advertisement is clicked by a user after she submits a query to the search engine. Commercial search engines typically rel…
A graph neural network detects beneficial feature interactions for recommender systems.
We consider the problem of learning a policy for a Markov decision process consistent with data captured on the state-actions pairs followed by the policy. We assume that the policy belongs to a class of parameterized policies which are defined using features associated with the state-action pairs. The features are kno…
Proposes Contrastive Clustering for improved clustering performance.
DRSVM uses deep learning to rank relative attributes between image pairs.
Bayesian principles improve neural additive models for better feature selection and uncertainty.
New sampling methods improve Shapley values for explaining machine learning predictions.
Sparse neural networks visualize paired transcriptomic and electrophysiological data.
Model forecasts market structure from financial networks using machine learning.
We consider analysis of relational data (a matrix), in which the rows correspond to subjects (e.g., people) and the columns correspond to attributes. The elements of the matrix may be a mix of real and categorical. Each subject and attribute is characterized by a latent binary feature vector, and an inferred matrix map…
A model for choosing crypto assets based on security and stability.
Paper investigates multimodal contrastive learning and incorporates unpaired data.