Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

4.1%8.1%12.2%16.2% · May 202619922001200920182026
48 results for Hartigan consistency

Kernel k-Groups uses Hartigan's method for clustering in metric spaces of negative type.

problem Clustering in metric spaces of negative type.
method Weighted energy statistics, quadratically constrained quadratic program, kernel k-groups, Hartigan's method.
result Improved performance in higher dimensions compared to spectral clustering and kernel k-means.

Defines hierarchical clustering axioms for various densities.

problem Defining hierarchical clustering for different types of densities.
method An axiomatic approach to piecewise constant densities, then extending to general densities.
result Our axiomatic definition results in Hartigan's cluster tree under certain conditions.

SCAMP clusters data by selecting candidate clusters that follow shape constraints, avoiding the need for tuning parameters.

problem Clustering data in high-dimensional space with unknown number of clusters.
method SCAMP formulates clustering as a search and selection problem, using shape constraints and preference functions to select clusters.
result SCAMP can be run multiple times to assess clustering uncertainty, providing a robust method for data annotation.

Following Hartigan, a cluster is defined as a connected component of the t-level set of the underlying density, i.e., the set of points for which the density is greater than t. A clustering algorithm which combines a density estimate with spectral clustering techniques is proposed. Our algorithm is composed of two step…

2010-02-11abs ↗pdf ↗

A new fuzzy clustering method using hyperbolic smoothing for large datasets.

problem Building fuzzy clusters for large data sets efficiently.
method A novel smoothing numerical approach to relax the sum-of-squares criterion, converting the problem into a differentiable optimization problem.
result The method produces better fuzzy partitions compared to traditional fuzzy CC-means.

SparseMix clusters sparse high dimensional binary data efficiently.

problem Clustering sparse high dimensional binary data.
method SparseMix is a mixture model designed for sparse data, using an on-line Hartigan optimization algorithm.
result SparseMix builds partitions with higher compatibility with reference grouping than related methods.

Unified theory for estimating and completing matrices with biclustering structures.

problem Estimating and completing matrices with biclustering structures in partially observed and noisy data.
method Developed a constrained least squares estimator achieving minimax rate-optimal performance.
result Unified high probability upper bounds and matching minimax lower bounds for various scenarios.

The level set tree approach of Hartigan (1975) provides a probabilistically based and highly interpretable encoding of the clustering behavior of a dataset. By representing the hierarchy of data modes as a dendrogram of the level sets of a density estimator, this approach offers many advantages for exploratory analysis…

2013-07-30abs ↗pdf ↗

We describe kk-MLE, a fast and efficient local search algorithm for learning finite statistical mixtures of exponential families such as Gaussian mixture models. Mixture models are traditionally learned using the expectation-maximization (EM) soft clustering technique that monotonically increases the incomplete (expec…

2012-03-23abs ↗pdf ↗

Study of loss functions for learning to defer, proving consistency.

problem Learning to defer in machine learning.
method Introduced a family of surrogate losses parameterized by ΨΨ and proved their consistency.
result Proved realizable HH-consistency and Bayes-consistency of specific surrogate losses.

In this paper we study the consistency of an empirical minimum error entropy (MEE) algorithm in a regression setting. We introduce two types of consistency. The error entropy consistency, which requires the error entropy of the learned function to approximate the minimum error entropy, is shown to be always true if the…

2014-12-17abs ↗pdf ↗

This paper improves deep learning model consistency through ensemble methods.

problem Consistency and correct-consistency issues in deep learning models.
method Formal definition of consistency and correct-consistency, proving ensemble improvement, proposing dynamic snapshot ensemble method.
result Ensemble methods can improve correct-consistency of deep learning models.

Empirical study shows consistent meta-RL algorithms adapt to OOD tasks.

problem Theoretical consistency of meta-RL algorithms and its practical implications.
method Empirical investigation of representative meta-RL algorithms, focusing on consistency and adaptation to out-of-distribution tasks.
result Theoretical consistent algorithms can adapt to OOD tasks, while inconsistent ones cannot, but can still fail for poor exploration.

Paper connects risk consistency to L_p consistency for broader loss functions.

problem Establishing risk consistency for a wider class of loss functions.
method Analyzes the connection between risk consistency and L_p-consistency for various loss functions.
result Shifted loss functions do not reduce assumptions as much as other results.

Improves GAN-based semi-supervised learning with consistency regularization.

problem Lack of consistency in class probability predictions under local perturbations.
method Introduces consistency regularization to GANs, leveraging both local and interpolation consistency.
result Significantly improves performance and achieves new state-of-the-art results.

This paper establishes a theoretical foundation for consistency training in diffusion models.

problem Lack of a comprehensive theoretical understanding of consistency training in diffusion models.
method Demonstrates the necessity of a number of steps in consistency learning exceeding d5/2/εd^{5/2}/\varepsilon for generating samples within ε\varepsilon proximity to the target distribution.
result Establishes rigorous insights into the validity and efficacy of consistency models, offering theoretical underpinnings for their utility.

Paper introduces new actuarial-consistent valuations for insurance liabilities.

problem Valuation of insurance liabilities considering both financial and actuarial risks.
method Proposes two-step actuarial valuations and actuarial-consistent procedures.
result Actuarial-consistent valuations are equivalent to two-step actuarial valuations under coherence.

New clustering method avoids problematic properties of existing algorithms.

problem Existing clustering algorithms cannot satisfy all natural clustering properties.
method Developed Morse Clustering using Morse Theory to satisfy Kleinberg's axioms with a new property, Monotonic Consistency.
result Morse Clustering satisfies Kleinberg's original axioms with Consistency replaced by Monotonic Consistency.

We improve random forest consistency and performance with DMRF, a new variant.

problem Improving the consistency and performance of random forest models.
method Developed DMRF, a data-driven multinomial random forest, by modifying proof methods and improving data utilization.
result DMRF achieves strong consistency with probability 1, surpassing previous models in classification tasks.

Fisher consistency improves class probability estimation under dataset shift.

problem Lack of Fisher consistency can lead to unreliable class probability estimates.
method Introduced Fisher consistency as a desirable property for class prior probability estimators.
result CDE-Iterate is not Fisher consistent and cannot be trusted for reliable estimates.

This paper tackles deferral learning with multiple experts, providing strong theoretical guarantees.

problem Optimizing input assignment to experts balancing accuracy and computational cost.
method Introducing new surrogate loss functions and efficient algorithms with strong theoretical learning guarantees.
result Realizable HH-consistency, HH-consistency bounds, and Bayes-consistency for deferral learning.

Unified surrogate loss framework for multi-label learning with strong consistency guarantees.

problem Improving consistency and accounting for label correlations in multi-label learning.
method Introducing multi-label logistic loss and extending it to comprehensive multi-label comp-sum losses, proving strong consistency guarantees for any multi-label loss.
result Unified surrogate loss framework benefiting from strong consistency guarantees for any multi-label loss.

Learning rule consistency tied to non-existence of real-valued measurable cardinals.

problem Consistency of k-NN learning rule in metric spaces.
method Analyzing separable subspaces and density conditions.
result The k-NN classifier's consistency depends on the absence of real-valued measurable cardinals.

New method evaluates language model forecasters by checking consistency of predictions.

problem Evaluating the performance of language model forecasters is difficult due to lack of ground truth.
method Developed a consistency check framework based on arbitrage to evaluate forecasters.
result Consistency metrics correlate with ground truth performance of LLM forecasters.

Investigates time-consistency of cash-subadditive risk measures.

problem Investigates conditions for time-consistency of cash-subadditive convex dynamic risk measures.
method Uses dual representation and generalized cocycle condition to provide sufficient conditions for strong time-consistency.
result Provides sufficient condition for strong time-consistency in cash-subadditive convex dynamic risk measures.

MTSCI uses diffusion models to impute multivariate time series data with consistency.

problem Imputation of missing values in multivariate time series data.
method MTSCI employs a contrastive complementary mask and mixup mechanism to ensure intra-consistency and inter-consistency.
result MTSCI achieves state-of-the-art performance on multivariate time series imputation tasks.

Self-consistent models improve reinforcement learning by aligning predictions with future values.

problem Improving reinforcement learning by aligning model predictions with future values.
method Proposes multiple self-consistency updates to encourage a learned model and value function to be consistent with each other.
result Self-consistency helps both policy evaluation and control in both tabular and function approximation settings.

Consistent estimation of constrained autoregressive processes.

problem Estimating autoregressive processes with coefficients constrained to an ellipsoid.
method Use of constrained and penalized estimators under different norms.
result Provide consistency results for estimation of constrained autoregressive processes.

Prefix consistency improves model reliability by weighting answers based on their reproducibility.

problem Improving the reliability of large language models' reasoning traces.
method Use prefix consistency to weight candidate answers based on their reproducibility during regeneration.
result Prefix consistency is the best correctness predictor, reaching Standard MV plateau accuracy with up to 21x fewer tokens.

The paper explores time consistency for scalar multivariate risk measures in markets with transaction costs.

problem Time consistency of scalar multivariate risk measures in markets with transaction costs.
method Presented dual representations and derived an equivalent recursive formulation for multivariate scalar risk measures.
result Developed a direct notion of a 'moving scalarization' for scalar time consistency.