Adaptive sparseness enhances robust regression using MCC and ARD.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Proposes MCC-F1 curve for better binary classification evaluation.
A new weighted MCC measure improves classifier performance evaluation.
Robust diffusion adaptive estimation algorithms based on the maximum correntropy criterion (MCC), including adaptation to combination MCC and combination to adaptation MCC, are developed to deal with the distributed estimation over network in impulsive (long-tailed) noise environments. The cost functions used in distri…
Correntropy is a local similarity measure defined in kernel space and the maximum correntropy criterion (MCC) has been successfully applied in many areas of signal processing and machine learning in recent years. The kernel function in correntropy is usually restricted to the Gaussian function with center located at ze…
New method turns any regression model into a calibrated probabilistic model.
The maximum correntropy criterion (MCC) has recently been successfully applied in robust regression, classification and adaptive filtering, where the correntropy is maximized instead of minimizing the well-known mean square error (MSE) to improve the robustness with respect to outliers (or impulsive noises). Considerab…
New metrics improve performance in imbalanced classification problems.
Given two maps f_1, f_2 : M^m \longrightarrow N^n between manifolds of the indicated arbitrary dimensions, when can they be deformed away from one another? More generally: what is the minimum number MCC (f_1, f_2) of pathcomponents of the coincidence space of maps f'_1, f'_2 where f'_i is homotopic to f_i, i = 1, 2? Ap…
Derivative traders are usually required to scan through hundreds, even thousands of possible trades on a daily basis. Up to now, not a single solution is available to aid in their job. Hence, this work aims to develop a trading recommendation system, and apply this system to the so-called Mid-Curve Calendar Spread (MCC…
Enhanced metrics for multiclass classification improve on existing methods.
Extracts geometric information from point-clouds for multiclass classification.
FedMCC learns from distributed data to cluster and extract features.
With the emergence of diverse data collection techniques, objects in real applications can be represented as multi-modal features. What's more, objects may have multiple semantic meanings. Multi-modal and Multi-label (MMML) problem becomes a universal phenomenon. The quality of data collected from different channels ar…
Defines tensor product of profinitely many vector spaces over F2.
Study uses machine learning to predict potato clones suitable for processing.
HNet detects significant associations in mixed data types efficiently.
As an effective and efficient discriminative learning method, Broad Learning System (BLS) has received increasing attention due to its outstanding performance in various regression and classification problems. However, the standard BLS is derived under the minimum mean square error (MMSE) criterion, which is, of course…
Computational identification of promoters is notoriously difficult as human genes often have unique promoter sequences that provide regulation of transcription and interaction with transcription initiation complex. While there are many attempts to develop computational promoter identification methods, we have no reliab…
SAEs struggle with feature consistency across runs, hindering MI reliability.
Constrained adaptive filtering algorithms inculding constrained least mean square (CLMS), constrained affine projection (CAP) and constrained recursive least squares (CRLS) have been extensively studied in many applications. Most existing constrained adaptive filtering algorithms are developed under mean square error (…
Heart disease is the leading cause of death, and experts estimate that approximately half of all heart attacks and strokes occur in people who have not been flagged as "at risk." Thus, there is an urgent need to improve the accuracy of heart disease diagnosis. To this end, we investigate the potential of using data ana…
Drug-drug interaction (DDI) is a major cause of morbidity and mortality and a subject of intense scientific interest. Biomedical literature mining can aid DDI research by extracting evidence for large numbers of potential interactions from published literature and clinical databases. Though DDI is investigated in domai…
The existing automatic fingerprint verification methods are designed to work under the assumption that the same sensor is installed for enrollment and authentication (regular matching). There is a remarkable decrease in efficiency when one type of contact-based sensor is employed for enrolment and another type of conta…
Graph-based ML improves defect prediction in software development.
Principal component analysis (PCA) is recognised as a quintessential data analysis technique when it comes to describing linear relationships between the features of a dataset. However, the well-known sensitivity of PCA to non-Gaussian samples and/or outliers often makes it unreliable in practice. To this end, a robust…
As a robust nonlinear similarity measure in kernel space, correntropy has received increasing attention in domains of machine learning and signal processing. In particular, the maximum correntropy criterion (MCC) has recently been successfully applied in robust regression and filtering. The default kernel function in c…
Traditional Kalman filter (KF) is derived under the well-known minimum mean square error (MMSE) criterion, which is optimal under Gaussian assumption. However, when the signals are non-Gaussian, especially when the system is disturbed by some heavy-tailed impulsive noises, the performance of KF will deteriorate serious…
In classical fixed point and coincidence theory the notion of Nielsen numbers has proved to be extremely fruitful. Here we extend it to pairs (f_1, f_2) of maps between manifolds of arbitrary dimensions. This leads to estimates of the minimum numbers MCC(f_1, f_2) (and MC(f_1, f_2), resp.) of pathcomponents (and of poi…
Machine learning and deep learning have gained popularity and achieved immense success in Drug discovery in recent decades. Historically, machine learning and deep learning models were trained on either structural data or chemical properties by separated model. In this study, we proposed an architecture training simult…
The unscented transformation (UT) is an efficient method to solve the state estimation problem for a non-linear dynamic system, utilizing a derivative-free higher-order approximation by approximating a Gaussian distribution rather than approximating a non-linear function. Applying the UT to a Kalman filter type estimat…
Develops RES metrics for stable rare-event forecasting evaluation.
In classical fixed point and coincidence theory the notion of Nielsen numbers has proved to be extremely fruitful. We extend it to pairs (f_1,f_2) of maps between manifolds of arbitrary dimensions, using nonstabilized normal bordism theory as our main tool. This leads to estimates of the minimum numbers MCC(f_1,f_2) (a…
QA-Token improves tokenization for noisy data, boosting model performance.
As a novel similarity measure that is defined as the expectation of a kernel function between two random variables, correntropy has been successfully applied in robust machine learning and signal processing to combat large outliers. The kernel function in correntropy is usually a zero-mean Gaussian kernel. In a recent …
Framework learns to transform majority to minority samples for balanced classification.
Graph theory criterion for Hodge theory to match linearly.
Advanced AI model predicts stock movements post earnings reports.
This paper visualizes uncertainty in classifier performance metrics.
GCNET predicts stock price movements using graph convolutional networks.
Benchmark evaluates AI-generated financial QA hallucinations, highlighting system vulnerabilities.
The study of genetic variants can help find correlating population groups to identify cohorts that are predisposed to common diseases and explain differences in disease susceptibility and how patients react to drugs. Machine learning algorithms are increasingly being applied to identify interacting GVs to understand th…
There are a variety of Domain Adaptation (DA) scenarios subject to label sets and domain configurations, including closed-set and partial-set DA, as well as multi-source and multi-target DA. It is notable that existing DA methods are generally designed only for a specific scenario, and may underperform for scenarios th…
The paper discusses thresholds and bounds for accuracy in binary classification systems.