ALICE combines feature selection and inter-rater agreeability for ML model insights.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Framework assesses autograders' reliability and biases.
Quantifying the degree of atrophy is done clinically by neuroradiologists following established visual rating scales. For these assessments to be reliable the rater requires substantial training and experience, and even then the rating agreement between two radiologists is not perfect. We have developed a model we call…
Generative models have made immense progress in recent years, particularly in their ability to generate high quality images. However, that quality has been difficult to evaluate rigorously, with evaluation dominated by heuristic approaches that do not correlate well with human judgment, such as the Inception Score and …
An automated metric to evaluate dialogue quality is vital for optimizing data driven dialogue management. The common approach of relying on explicit user feedback during a conversation is intrusive and sparse. Current models to estimate user satisfaction use limited feature sets and rely on annotation schemes with low …
Study evaluates consistency of LLMs in binary text classification, providing systematic guidance.
k-Rater reliability corrects under-reporting of aggregated data reliability.
Sleep stage classification constitutes an important element of sleep disorder diagnosis. It relies on the visual inspection of polysomnography records by trained sleep technologists. Automated approaches have been designed to alleviate this resource-intensive task. However, such approaches are usually compared to a sin…
Study evaluates five LLMs for financial report analysis, revealing performance differences and variability.
We derive the mapping between two of the most pervasive utility functions, the mean square error () and the concordance correlation coefficient (CCC, ). Despite its drawbacks, is one of the most popular performance metrics (and a loss function); along with lately in many of the sequence prediction…
Machine learning (ML) methods have the potential to automate clinical EEG analysis. They can be categorized into feature-based (with handcrafted features), and end-to-end approaches (with learned features). Previous studies on EEG pathology decoding have typically analyzed a limited number of features, decoders, or bot…
Solves a 60-year-old question on agreement measures in statistics.
Infants' spontaneous and voluntary movements mirror developmental integrity of brain networks since they require coordinated activation of multiple sites in the central nervous system. Accordingly, early detection of infants with atypical motor development holds promise for recognizing those infants who are at risk for…
SMART is an open source web application designed to help data scientists and research teams efficiently build labeled training data sets for supervised machine learning tasks. SMART provides users with an intuitive interface for creating labeled data sets, supports active learning to help reduce the required amount of …
Improves reliability of medical diagnosis uncertainty estimates.
This work introduces significativity indices for agreement values between classifiers.
The problem of maximizing (or minimizing) the agreement between clusterings, subject to given marginals, can be formally posed under a common framework for several agreement measures. Until now, it was possible to find its solution only through numerical algorithms. Here, an explicit solution is shown for the case wher…
SNAP improves robust computation by emphasizing trustworthy items and downweighting outliers.
Discriminatory trade liberalization policies are becoming more popular among world economies. Countries are motivated to enter for regional trade agreements to capture faster economic growth for alleviating poverty. In developing economies like most of the member countries of the Association of South East Asian Nations…
LFD method improves text classification by making features clearer and less label-leaking.
In unsupervised machine learning, agreement between partitions is commonly assessed with so-called external validity indices. Researchers tend to use and report indices that quantify agreement between two partitions for all clusters simultaneously. Commonly used examples are the Rand index and the adjusted Rand index. …
In many machine learning problems, labeled training data is limited but unlabeled data is ample. Some of these problems have instances that can be factored into multiple views, each of which is nearly sufficent in determining the correct labels. In this paper we present a new algorithm for probabilistic multi-view lear…
In Bipartite Correlation Clustering (BCC) we are given a complete bipartite graph with `+' and `-' edges, and we seek a vertex clustering that maximizes the number of agreements: the number of all `+' edges within clusters plus all `-' edges cut across clusters. BCC is known to be NP-hard. We present a novel approx…
Proposes a strategy to train models with minimal labeled data.
Model selection is a problem that has occupied machine learning researchers for a long time. Recently, its importance has become evident through applications in deep learning. We propose an agreement-based learning framework that prevents many of the pitfalls associated with model selection. It relies on coupling the t…
A new method for combining multiple data views in supervised learning.
The retinal vascular condition is a reliable biomarker of several ophthalmologic and cardiovascular diseases, so automatic vessel segmentation may be crucial to diagnose and monitor them. In this paper, we propose a novel method that combines the multiscale analysis provided by the Stationary Wavelet Transform with a m…
Computable contracts simplify financial transactions and reduce legal costs.
Generalization and reliability of multilingual translation often highly depend on the amount of available parallel data for each language pair of interest. In this paper, we focus on zero-shot generalization---a challenging setup that tests models on translation directions they have not been optimized for at training t…
We apply an asymmetric version of Kirman's herding model to volatile financial markets. In the relation between returns and agent concentration we use the square root law proposed by Zhang. This can be derived by extending the idea of a critical mean field theory suggested by Plerou et al. We show that this model is eq…
Study finds simple model-agreement scores perform well in various error estimation scenarios.
Study optimizes deep learning models for sleep stage classification.
Researchers develop multi-agent systems for quadcopters to collaborate in missions.
The adjusted Rand index (ARI) is commonly used in cluster analysis to measure the degree of agreement between two data partitions. Since its introduction, exploring the situations of extreme agreement and disagreement under different circumstances has been a subject of interest, in order to achieve a better understandi…
Ensemble learning is a powerful approach to construct a strong learner from multiple base learners. The most popular way to aggregate an ensemble of classifiers is majority voting, which assigns a sample to the class that most base classifiers vote for. However, improved performance can be obtained by assigning weights…
Study shows more data improves model explanations, aiding reliable knowledge extraction.
Unified framework for policy learning using weak supervision.
There is considerable debate whether the domestic political institutions (specifically, the country s level of democracy) of the host developing country toward foreign investors are effective in establishing the credibility of commitments are still underway, researchers have also analyzed the effect of international in…
MAS scores cluster size consistency from points, robust to label changes.
In important applications involving multi-task networks with multiple objectives, agents in the network need to decide between these multiple objectives and reach an agreement about which single objective to follow for the network. In this work we propose a distributed decision-making algorithm. The agents are assumed …
Co-learning BO improves global optimization with limited samples.
Model explains money creation under regulatory constraints.
New index improves anomaly detection in correlated time series data.
The paper uses Black-Scholes model to analyze political support and coalition agreements.
This paper examines how regional trade agreements affect global trade relationships.
ValueBlindBench tests LLM-generated investment rationales for validity before returns are known.
New algorithms for collaborative learning in uncertain, decentralized environments.
Paper develops framework for valuing and assessing credit risk in renewable PPAs.