Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

129259388517 · Jun 202019922001200920182026
48 results for expected classification agreement

The paper proposes a method to trim Bayesian network classifiers robustly.

problem Removing costly features from Bayesian network classifiers while maintaining robustness.
method Introduces an expected classification agreement (ECA) metric and a branch-and-bound search algorithm to find optimal feature subsets and thresholds.
result The proposed method maximizes expected agreement between the original and trimmed classifiers, subject to a budgetary constraint.

Solves a 60-year-old question on agreement measures in statistics.

problem The challenge of measuring agreement between two raters or measures.
method Developed a new algorithm to minimize diagonals in contingency tables, formulated the minimum feasible agreement, and studied the lower limit of maximum feasible agreement.
result Formulated the lower limit of Cohen's kappa and two statistics for agreement analysis.

ARIMLE optimizes classifier fusion for brain-computer interface.

problem Improving ensemble classifier aggregation performance.
method ARIMLE uses agreement rate to estimate classifier accuracy, then refines a maximum likelihood estimator.
result ARIMLE outperforms majority voting and other methods in brain-computer interface applications.

LFD method improves text classification by making features clearer and less label-leaking.

problem Creating interpretable text representations that are both predictive and understandable.
method LFD method: proposes lexical and semantic features from contrastive text pairs, screens candidates using κκ, and selects features by residual gain.
result LFD features achieve higher human-human and human-LLM agreement than baseline concepts and are less label-leaking.

In many machine learning problems, labeled training data is limited but unlabeled data is ample. Some of these problems have instances that can be factored into multiple views, each of which is nearly sufficent in determining the correct labels. In this paper we present a new algorithm for probabilistic multi-view lear…

2012-06-13abs ↗pdf ↗

This paper examines how regional trade agreements affect global trade relationships.

problem The relationship between regional trade agreements and global trade purity.
method Defined and decomposed synthesized trade resistance, separated natural and artificial factors, used expectation maximization algorithm to optimize parameters, and quantified trade purity indicator.
result Regional trade agreements contribute to the relative prosperity of EU and NAFTA countries, but weaken the role of trade unions and accelerate multilateral trade liberalization.

MAS scores cluster size consistency from points, robust to label changes.

problem Desired uniformity in cluster sizes, stability under label perturbations.
method Mass Agreement Score (MAS) measures point-centric cluster size consistency, robust to label changes.
result MAS yields similar scores for partitions with similar bulk structure, sensitive to genuine redistribution of cluster mass.

Extends Yagil's model to include stochastic dividends in stock-for-stock mergers.

problem Determining exchange ratios in stock-for-stock mergers with uncertain dividend growth.
method Generalizes Yagil's deterministic model to a stochastic environment, considering both expected values and variance of dividends.
result Identifies a more complex bargaining region for exchange ratios, influenced by the mean and standard deviation of dividends' growth rate.

The operator realizing a Dehn twist in quantum Teichmuller theory is diagonalized and continuous spectrum is obtained. This result is in agreement with the expected spectrum of conformal weights in quantum Liouville theory at c>1. The completeness condition of the eigenvectors includes the integration measure which app…

2000-08-18abs ↗pdf ↗

Fuses ITRs for primary and secondary outcomes to minimize harm.

problem Learn an ITR maximizing primary outcome while minimizing harm to secondary outcomes.
method Introduces fusion penalty to encourage similar recommendations for different outcomes. Two algorithms estimate the ITR using surrogate loss functions.
result Agreement rate between primary and secondary optimal ITRs converges faster than ignoring secondary outcomes.

Leverage is strongly related to liquidity in a market and lack of liquidity is considered a cause and/or consequence of the recent financial crisis. A repurchase agreement is a financial instrument where a security is sold simultaneously with an agreement to buy it back at a later date. Repurchase agreements (repos) ma…

2010-11-01abs ↗pdf ↗

Bayesian active learning improves holistic educational assessments.

problem Gap between holistic CJ and criterion-based rubrics in education.
method Extends Bayesian CJ to handle multiple LO components, using entropy-based active learning.
result Enhanced predictive rankings with uncertainty estimates and quantified assessor agreement.

New heuristics improve genetic programming's parent selection for classification problems.

problem Improving genetic programming's parent selection for classification tasks.
method Proposed three heuristics inspired by specific classifiers' characteristics, using similarity measures.
result Combination of agreement-based selection and random selection outperforms classical and state-of-the-art schemes.

Risk-controlled post-processing optimizes decision policies under risk constraints.

problem Optimizing decision policies with risk constraints for better outcomes.
method Developed a post-processing algorithm that selects a threshold based on fitted fallback policy and score, leveraging tools from algorithmic stability and stochastic processes.
result The post-processed policy achieves precise expected risk control under exchangeability and meets or nearly meets risk budgets while preserving more agreement with the baseline.

We present a fully nonparametric method to estimate the value function, via simulation, in the context of expected infinite-horizon discounted rewards for Markov chains. Estimating such value functions plays an important role in approximate dynamic programming and applied probability in general. We incorporate "soft in…

2013-12-26abs ↗pdf ↗

We apply random matrix theory to compare correlation matrix estimators C obtained from emerging market data. The correlation matrices are constructed from 10 years of daily data for stocks listed on the Johannesburg Stock Exchange (JSE) from January 1993 to December 2002. We test the spectral properties of C against ra…

2004-02-14abs ↗pdf ↗

SNAP improves robust computation by emphasizing trustworthy items and downweighting outliers.

problem Improving robustness in computation, especially in high-dimensional settings.
method SNAP assigns weights based on mutual agreement, suppressing outlier contributions.
result SNAP ensures outliers contribute negligibly to computations, even in high-dimensional settings.

This paper shows faster convergence rates for stochastic gradient descent in binary classification.

problem Achieving faster convergence rates for stochastic gradient descent in binary classification.
method Stochastic gradient descent and averaging variant, focusing on exponential convergence rates under strong low-noise conditions.
result Exponential convergence of the expected classification error in the final phase of stochastic gradient descent and averaged stochastic gradient descent for differentiable convex loss functions.

Paper tackles zero-shot translation by encouraging consistent agreement in models.

problem Challenges of generalizing multilingual translation without parallel data.
method Reformulated as probabilistic inference, introduced consistent agreement-based training.
result Agreement-based learning improves zero-shot translation by 2-3 BLEU points.

New deep learning method validated across multiple sleep staging databases.

problem Improving automatic sleep scoring accuracy across different datasets.
method Ensemble of local models using deep learning for automatic sleep staging.
result Good general performance compared to human experts and state-of-the-art methods.

New approach combines multi-view learning for improved convergence in knowledge transfer.

problem Improving convergence in knowledge transfer settings like learning with privileged information and distillation.
method Adopting a multi-view approach under reasonable assumptions about hypothesis spaces, encouraging agreement between teacher and student.
result Improved convergence rate achieved with regularized empirical risk minimization.

Study explores how to explain deep neural networks using interactive naming.

problem Explaining deep neural networks' decisions in human-understandable terms.
method Developed an interactive naming interface for human annotators to cluster activation maps into visual concepts.
result Significant agreement among annotators about visual concepts, many activation maps have recognizable concepts.

Proposes using Wasserstein barycenters for model ensembling in multiclass/multilabel learning.

problem Finding consensus between models in multiclass/multilabel learning settings.
method Uses Wasserstein (W.) barycenters to find consensus between models, incorporating semantic side information.
result Wasserstein ensembling balances confidence and semantics in model agreement.

The paper analyzes indices based on counting object pairs for assessing partition agreement in unsupervised learning.

problem The difficulty in interpreting overall indices like Rand and adjusted Rand indices.
method Analysis of three families of indices based on counting object pairs, decomposing overall indices into cluster-level indices.
result Overall indices based on pair-counting approach are sensitive to cluster size imbalance and provide limited information on smaller clusters.

We make a precision test of a recently proposed conjecture relating Chern-Simons gauge theory to topological string theory on the resolution of the conifold. First, we develop a systematic procedure to extract string amplitudes from vacuum expectation values (vevs) of Wilson loops in Chern-Simons gauge theory, and then…

2000-04-27abs ↗pdf ↗

We introduce the formalism of generalized Fourier transforms in the context of risk management. We develop a general framework to efficiently compute the most popular risk measures, Value-at-Risk and Expected Shortfall (also known as Conditional Value-at-Risk). The only ingredient required by our approach is the knowle…

2009-09-22abs ↗pdf ↗

Gaptron algorithm reduces mistakes in online multiclass classification.

problem Online multiclass classification with limited information.
method Randomized first-order algorithm exploiting the gap between zero-one loss and surrogate losses.
result First linear time algorithm with O(KT)O(K\sqrt{T}) expected regret.

GraphCL learns node representations by maximizing similarity between perturbed node features.

problem Learning node representations in graph data without labeled data.
method Contrastive learning of node embeddings using graph neural networks and a loss function.
result Significantly outperforms state-of-the-art in unsupervised node classification benchmarks.

Forest tree species mapped with high accuracy using satellite data.

problem Classifying dominant tree species in Swedish forests.
method Extreme gradient boosting model with Bayesian optimization, combining Sentinel-1/2 satellite data and field observations.
result Overall accuracy of 85%, F1 score of 0.82, Matthews correlation coefficient of 0.81.

Unified framework for comparing clusterings from information-theoretic and pair-counting perspectives.

problem Divergent evaluations of unsupervised models due to different clustering similarity measures.
method Developed an analytical framework that unifies pair-counting and information-theoretic clustering similarity measures.
result Unified framework clarifies when and why the two regimes diverge and provides a principled basis for selecting and interpreting clustering similarity measures.

In Bipartite Correlation Clustering (BCC) we are given a complete bipartite graph GG with `+' and `-' edges, and we seek a vertex clustering that maximizes the number of agreements: the number of all `+' edges within clusters plus all `-' edges cut across clusters. BCC is known to be NP-hard. We present a novel approx…

2016-03-09abs ↗pdf ↗

Binary classification models get more efficient predictive probabilities.

problem Computing predictive probabilities in Bayesian probit models is computationally challenging.
method Use of expectation propagation (EP) to find a closed-form expression for predictive probabilities.
result Closed-form predictive probabilities improve over existing methods.