This paper tackles deep clustering evaluation challenges in high-dimensional data.
problem Evaluation of deep clustering methods is problematic due to the curse of dimensionality and variations in embedding spaces.
method Develops a theoretical framework to highlight the ineffectiveness of internal validation measures and proposes a systematic approach to applying clustering validity indices in deep learning.
result The proposed framework reduces misguidance from improper use of clustering validity indices in deep learning.
The paper assesses quality measures for machine learning models using cross-validation.
problem Evaluating the accuracy and robustness of quality measures for machine learning models.
method Cross-validation approach to estimate prediction error and quantify explained variation. Confidence bounds and local quality measures derived from residuals.
result The reliability and robustness of quality measures are assessed through numerical examples and confidence bounds.
The paper challenges the validity of cluster validity measures in unsupervised learning.
problem The validity of cluster validity measures in selecting optimal clusterings.
method The authors investigate the use of cluster validity measures as objective functions in unsupervised learning and introduce a new variant of the Dunn index.
result Many cluster validity measures promote clusterings that do not match expert knowledge well.
The paper validates a centrality measure for financial networks during financial distress.
problem Systemic risk and shock propagation in financial networks.
method Statistical validation method for network centrality measures.
result The proposed centrality measure increases significantly during financial distress.
Study benchmarks 26 clustering validity measures.
problem Determining the best clustering solution from candidates.
method Enhanced revision of previous methodology with three sub-methodologies.
result Comprehensive evaluation of 26 internal validity indexes.
Study validates metrics for offline MBO using diffusion models.
problem Evaluate metrics for offline MBO without ground truth oracle.
method Propose and quantify validation metrics over datasets.
result Identify most effective validation metrics.
A new measure normalizes clustering accuracy to evaluate algorithms better.
problem Evaluation of clustering algorithms is challenging due to limitations of existing measures.
method Proposes a new, normalised clustering accuracy measure.
result The new measure identifies worst-case scenarios and is more interpretable.
Validates composite systems using discrepancy propagation.
problem Validation of industrial systems with costly real-world tests.
method Propagates bounds on distributional discrepancy measures through a composite system.
result Derives upper bound on real system failure probability from simulations.
Proposes a method to generate counterfactuals for ensemble models using entropic risk measures.
problem Finding a single counterfactual explanation for an ensemble of models.
method Incorporates entropic risk measure into a constrained optimization to generate counterfactuals valid for an adjustable fraction of models.
result Entropic risk measure allows generation of counterfactuals valid for all models in the ensemble under a limiting case.
We show how to adjust the coefficient of determination (R2) when used for measuring predictive accuracy via leave-one-out cross-validation.
Prediction of future observations is an important and challenging problem. The two mainstream approaches for quantifying prediction uncertainty use prediction regions and predictive distributions, respectively, with the latter believed to be more informative because it can perform other prediction-related tasks. The st…
New method for selecting clusters in residential electricity data.
problem Selecting useful clusters in electricity consumption data.
method Formalizing expert knowledge as external validation measures.
result Successfully reconstructed customer archetypes.
Robust variable selection for high-dimensional data with missing and measurement errors.
problem Missing data and measurement errors confound data distribution.
method Exponential loss function with inverse probability weighting and additive error models.
result The Atan punishment method improves robust variable selection.
Machine learning models trained on indirect data labels can fail on real-world examples.
problem Validity issues in machine learning when target labels are indirectly defined.
method Identification of problematic datasets and models using a general procedure.
result Machine learning models trained on indirect data labels will fail on real-world examples.
New method uses spherical convolutional Wasserstein distance to validate climate models.
problem Ensuring the accuracy of global climate models.
method Spherical convolutional Wasserstein distance to measure model differences.
result Phase 6 models show modest improvements in realistic climatologies.
Improves full conformal prediction for stochastic non-conformity measures.
problem Inability of existing conditions to guarantee full conformal prediction validity under stochastic settings.
method Introduces a new sufficient condition: Conditional Independence & Permutation Invariance in Distribution.
result Corrects the insufficient condition and provides a new sufficient condition for full conformal prediction validity.
This paper evaluates and validates cluster results using external and internal evaluation methods.
problem Evaluating and validating the quality of clustering results.
method External evaluation using Homogeneity, Correctness, and V-measure scores; internal evaluation using Silhouette Index and Sum of Square Errors.
result Validation of the number of clusters using dendrogram and statistical frequency distribution.
Scenario-based testing for the safety validation of highly automated vehicles is a promising approach that is being examined in research and industry. This approach heavily relies on data from real-world scenarios to derive the necessary scenario information for testing. Measurement data should be collected at a reason…
A new measure identifies clusters without assuming data distribution.
problem Identifying the correct number of clusters in data without distribution assumptions.
method Nonparametric interpoint distance-based approach.
result Superior to existing clustering measures, validated on synthetic and real data.
Measures neural network complexity via effective degrees of freedom.
problem Challenges in quantifying neural network complexity.
method Adapts generalized degrees of freedom (GDF) for binary outcomes and compares with cross-validation and null degrees of freedom.
result GDF provides a robust measure of model complexity for neural networks.
Inferring causal interactions from observed data is a challenging problem, especially in the presence of measurement noise. To alleviate the problem of spurious causality, Haufe et al. (2013) proposed to contrast measures of information flow obtained on the original data against the same measures obtained on time-rever…
The paper examines curvature-dimension bounds on sub-Finsler Heisenberg groups.
problem Investigating synthetic curvature-dimension bounds in sub-Finsler Heisenberg groups.
method Study of measure contraction property (MCP) and curvature-dimension condition (CD).
result Sub-Finsler Heisenberg groups do not satisfy MCP or CD for any parameters.
We consider the parametric learning problem, where the objective of the learner is determined by a parametric loss function. Employing empirical risk minimization with possibly regularization, the inferred parameter vector will be biased toward the training samples. Such bias is measured by the cross validation procedu…
Bayes factors and relative belief ratios are compared as measures of statistical evidence.
problem Which measure of evidence is more appropriate: Bayes factors or relative belief ratios?
method Comparison of Bayes factors and relative belief ratios, considering properties and restrictions.
result Relative belief ratio has better properties as a measure of evidence.
Bayesian framework forecasts financial tail risks using realized volatility and nonlinear thresholds.
problem Forecasting financial tail risks using realized volatility and nonlinear thresholds.
method Bayesian Markov Chain Monte Carlo method for model estimation; nonlinear threshold regression specification.
result The proposed framework produces competitive tail risk forecasts compared to GARCH and Realized-GARCH models.
New method corrects bias in feature importance measures of GBM.
problem Bias in feature importance measures of GBM.
method Cross-validated unbiased base learners.
result Significant improvement in feature importance measures with minimal computational cost.
Generative model learns object variability from MRI measurements.
problem Establishing stochastic object models from medical imaging data.
method Advanced AmbientGANs with multiresolution training.
result AmbientGANs reliably learn object distributions from incomplete or noisy data.
A new CVI called DSI evaluates clustering results without true labels.
problem No universal CVI for clustering without true labels.
method DSI based on data separability measure.
result DSI is an effective, unique, and competitive CVI.
Transformer models improve financial sentiment measurement.
problem Capturing nuanced sentiment from financial news articles.
method Transformer-based language models for sentiment classification and aggregation.
result Transformer models outperform traditional dictionary-based methods in sentiment classification.
Method measures weight similarity in neural networks using normalization and statistical inference.
problem Quantifying weight similarity in non-convex neural networks.
method Chain normalization rule and hypothesis-training-testing statistical inference.
result Weights of identical neural networks converge to similar local solutions.
Undetected overfitting can occur when there are significant redundancies between training and validation data. We describe AVE, a new measure of training-validation redundancy for ligand-based classification problems that accounts for the similarity amongst inactive molecules as well as active. We investigated seven wi…
The problem of robust hedging requires to solve the problem of superhedging under a nondominated family of singular measures. Recent progress was achieved by [9,11]. We show that the dual formulation of this problem is valid in a context suitable for martingale optimal transportation or, more generally, for optimal tra…
New measure captures differences across entire distributions of counterfactual outcomes.
problem Capturing differences across entire distributions of counterfactual outcomes.
method Entropic optimal transport measure, statistical functional, smooth transformation of embeddings.
result Established first-order and second-order pathwise differentiability.
The paper clarifies conditions for using benchmark scores in machine learning.
problem Using benchmark scores to draw scientific inferences about learning problems.
method Developing conditions of construct validity inspired by psychological measurement theory.
result Clarifies conditions under which benchmark scores support diverse scientific claims.
Improved statistical inference for expensive data using machine learning predictions.
problem Statistical inference under adaptive two-phase multiwave sampling with expensive measurements.
method Multiwave Predict-Then-Debias estimator combining proxy information and expensive measurements.
result Valid estimators and confidence intervals for M-estimation under adaptive sampling.
New risk measures assess cryptocurrency market vulnerabilities during financial distress.
problem Capturing systemic risk in cryptocurrency markets during financial distress.
method Introducing Vulnerability Conditional Risk Measures (VCoES) and related measures.
result Validated theoretical insights and demonstrated practical relevance in cryptocurrency market.
Evaluating data separation in a geometrical space is fundamental for pattern recognition. A plethora of dimensionality reduction (DR) algorithms have been developed in order to reveal the emergence of geometrical patterns in a low dimensional visible representation space, in which high-dimensional samples similarities …
Paper introduces statistical learning for point processes.
problem Statistical learning for point processes in general spaces.
method Combines bivariate innovations and point process cross-validation.
result Statistical learning approach outperforms state of the art.
The Bakry-Emery tensor gives an analog of the Ricci tensor for a Riemannian manifold with a smooth measure. We show that some of the topological consequences of having a positive or nonnegative Ricci tensor are also valid for the Bakry-Emery tensor. We show that the Bakry-Emery tensor is nondecreasing under a Riemannia…
A fast bootstrap method estimates cross-validation standard error.
problem Uncertainty quantification in cross-validation estimates.
method Random-effects model to estimate variance component.
result Valid confidence intervals for model performance.
Generalization of deep networks has been of great interest in recent years, resulting in a number of theoretically and empirically motivated complexity measures. However, most papers proposing such measures study only a small set of models, leaving open the question of whether the conclusion drawn from those experiment…
Improves test set performance and reduces out-of-sample disappointment for unstable models.
problem Ensuring strong test set performance via cross-validation for unstable models.
method Nested k-fold cross-validation with hyperparameter selection based on a weighted sum of cross-validation metric and model stability measure.
result Improves out-of-sample MSE for sparse ridge regression and CART by 4% and 2% respectively, compared to k-fold cross-validation.
The paper validates a method for recovering over-parameterized matrices and images from noisy measurements.
problem Recovering a low-rank matrix from noisy measurements when the rank is unknown.
method Using gradient descent with small random initialization on a nonconvex objective function built from a rank-overspecified factored representation of the matrix variable.
result Gradient descent iterations converge to the ground-truth matrix under certain conditions and can be stopped efficiently to detect a nearly optimal estimator.
Study on geometric Jensen-Shannon divergence for Gaussian measures in Hilbert space.
problem Computing divergence between Gaussian measures in infinite-dimensional Hilbert space.
method Closed form expression and regularization for divergence calculation.
result Closed form expression and regularization for Geometric Jensen-Shannon divergence.
Proposes Neural Complexity (NC) for predicting and explaining generalization in deep neural networks.
problem Challenges in specifying a suitable complexity measure for deep neural networks to predict and explain generalization.
method A meta-learning framework that learns a scalar complexity measure through interactions with many heterogeneous tasks.
result Trained NC model can be added to standard training loss to regularize any task learner.
We describe an abstract control-theoretic framework in which the validity of the dynamic programming principle can be established in continuous time by a verification of a small number of structural properties. As an application we treat several cases of interest, most notably the lower-hedging and utility-maximization…
Cross-validation estimates model performance on unseen data, not training data.
problem Understanding how cross-validation estimates prediction error and its limitations.
method Analyzing linear models and popular prediction error estimates, introducing nested cross-validation.
result Cross-validation estimates the average prediction error of models fit on other unseen training sets, not the model at hand.
The goal of the paper is to study the angle between two curves in the framework of metric (and metric measure) spaces. More precisely, we give a new notion of angle between two curves in a metric space. Such a notion has a natural interplay with optimal transportation and is particularly well suited for metric measure …