Paper proposes a method to recover accurate labels from partially valid data in multi-label learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
DS2CF-Net learns hierarchical representations with deep coupled factorization and enriched prior.
Study reduces memory needs for active learning with enriched queries.
Deep networks learn clean structure before memorizing corrupted labels, leaving a spectral signature in gradient centered scatter.
Small satellite constellations provide daily global coverage of the earth's landmass, but image enrichment relies on automating key tasks like change detection or feature searches. For example, to extract text annotations from raw pixels requires two dependent machine learning models, one to analyze the overhead image …
BioBO optimizes gene perturbation design using Bayesian optimization with biological priors.
Adding uninformative labels improves tumor segmentation in low-data mammography.
In this paper we propose Structuring AutoEncoders (SAE). SAEs are neural networks which learn a low dimensional representation of data which are additionally enriched with a desired structure in this low dimensional space. While traditional Autoencoders have proven to structure data naturally they fail to discover sema…
GitHub has become an important platform for code sharing and scientific exchange. With the massive number of repositories available, there is a pressing need for topic-based search. Even though the topic label functionality has been introduced, the majority of GitHub repositories do not have any labels, impeding the ut…
Machine learning has become pervasive in multiple domains, impacting a wide variety of applications, such as knowledge discovery and data mining, natural language processing, information retrieval, computer vision, social and health informatics, ubiquitous computing, etc. Two essential problems of machine learning are …
For sales and marketing organizations within large enterprises, identifying and understanding new markets, customers and partners is a key challenge. Intel's Sales and Marketing Group (SMG) faces similar challenges while growing in new markets and domains and evolving its existing business. In today's complex technolog…
We define a symmetric monoidal (4,3)-category with duals whose objects are certain enriched multi-fusion categories. For every modular tensor category , there is a self enriched multi-fusion category giving rise to an object of this symmetric monoidal (4,3)-category. We conjecture that the e…
In this paper we are interested in the prediction of preterm birth based on diagnosis codes from longitudinal EHR. We formulate the prediction problem as a supervised classification with noisy labels. Our base classifier is a Recurrent Neural Network with an attention mechanism. We assume the availability of a data sub…
In this paper we analyze supergeometric locally covariant quantum field theories. We develop suitable categories SLoc of super-Cartan supermanifolds, which generalize Lorentz manifolds in ordinary quantum field theory, and show that, starting from a few representation theoretic and geometric data, one can construct a f…
Cross-domain collaborative filtering (CF) aims to alleviate data sparsity in single-domain CF by leveraging knowledge transferred from related domains. Many traditional methods focus on enriching compared neighborhood relations in CF directly to address the sparsity problem. In this paper, we propose superhighway const…
This paper evaluates data enrichment techniques for rare event detection in manufacturing.
Deep learning uses ROC cost functions to improve virtual screening accuracy.
Study assesses weakly-supervised methods for rare outcomes in medical records.
KANEL combines models for early hit enrichment in virtual screening.
This note compares two recently published machine learning methods for constructing flexible, but tractable families of variational hidden-variable posteriors. The first method, called "hierarchical variational models" enriches the inference model with an extra variable, while the other, called "auxiliary deep generati…
Proposes a deep tree-ensemble model for multi-output prediction.
A new method for virtual drug screening detects top treatments.
Chinese named entity recognition (CNER) is an important task in Chinese natural language processing field. However, CNER is very challenging since Chinese entity names are highly context-dependent. In addition, Chinese texts lack delimiters to separate words, making it difficult to identify the boundary of entities. Be…
This paper connects virtual biquandles to biquandles for virtual link colorings.
Hashing techniques have been applied broadly in retrieval tasks due to their low storage requirements and high speed of processing. Many hashing methods based on a single view have been extensively studied for information retrieval. However, the representation capacity of a single view is insufficient and some discrimi…
Variational autoencoders are powerful algorithms for identifying dominant latent structure in a single dataset. In many applications, however, we are interested in modeling latent structure and variation that are enriched in a target dataset compared to some background---e.g. enriched in patients compared to the genera…
The Lax-Hopf formula simplifies the value function of an intertemporal optimization (infinite dimensional) problem associated with a convex transaction-cost function which depends only on the transactions (velocities) of a commodity evolution: it states that the value function is equal to the marginal fonction of a fin…
Algorithm recovers multiple low-rank matrices from unlabeled data.
In this study, we propose the integration of competitive learning into convolutional neural networks (CNNs) to improve the representation learning and efficiency of fine-tuning. Conventional CNNs use back propagation learning, and it enables powerful representation learning by a discrimination task. However, it require…
Geometric problems are usually formulated by means of (exterior) differential systems. In this theory, one enriches the system by adding algebraic and differential constraints, and then looks for regular solutions. Here we adopt a dual approach, which consists to enrich a plane field, as this is often practised in cont…
Improves neural network performance by enriching training dataset.
CDSSL improves representation quality by integrating linear and nonlinear dependencies.
Paper reviews neurolinguistics and language technologies, emphasizing mutual enrichment.
Flowification enriches neural networks with an inverse pass and likelihood monitoring.
Enhances LLMs for predicting stock movements by considering news dissemination and context.
Work shows hallucination detection by LLMs is impossible without expert feedback.
Paper proposes a novel GCN-based SSL algorithm to enhance node representations using contrastive and generative losses.
Enhanced tree-based classifiers use derivatives and geometry for better function classification.
Decision support tools that rely on supervised learning require large amounts of expert annotations. Using past radiological reports obtained from hospital archiving systems has many advantages as training data above manual single-class labels: they are expert annotations available in large quantities, covering a popul…
Regression, unlike classification, has lacked a comprehensive and effective approach to deal with cost-sensitive problems by the reuse (and not a re-training) of general regression models. In this paper, a wide variety of cost-sensitive problems in regression (such as bids, asymmetric losses and rejection rules) can be…
We explore 4d Yang-Mills gauge theories (YM) living as boundary conditions of 5d gapped short/long-range entangled (SRE/LRE) topological states. Specifically, we explore 4d time-reversal symmetric pure YM of an SU(2) gauge group with a second-Chern-class topological term at (SU(2) YM), by turning on backg…
In this era of digital information explosion, an abundance of data from numerous modalities is being generated as well as archived everyday. However, most problems associated with training Deep Neural Networks still revolve around lack of data that is rich enough for a given task. Data is required not only for training…
A recent trend in machine learning has been to enrich learned models with the ability to explain their own predictions. The emerging field of Explainable AI (XAI) has so far mainly focused on supervised learning, in particular, deep neural network classifiers. In many practical problems however, label information is no…
Face verification remains a challenging problem in very complex conditions with large variations such as pose, illumination, expression, and occlusions. This problem is exacerbated when we rely unrealistically on a single training data source, which is often insufficient to cover the intrinsically complex face variatio…
This work interprets supergravity as a super Cartan geometry linking it to Yang-Mills theory.
Generative model learns conditional distributions on collective variable levels.
Variational inference for latent variable models is prevalent in various machine learning problems, typically solved by maximizing the Evidence Lower Bound (ELBO) of the true data likelihood with respect to a variational distribution. However, freely enriching the family of variational distribution is challenging since…
The study improves compound selection in in silico screening by focusing on model's ability to predict desirable outcomes.