This paper reviews and proposes a new approach for evaluating internal cluster validation indices.
problem Selecting the best-performing unsupervised classification algorithm without external information.
method Examines and proposes a new evaluation approach for internal validation indices.
result Suggests a new evaluation approach for internal validation indices.
Cluster analysis is used to explore structure in unlabeled data sets in a wide range of applications. An important part of cluster analysis is validating the quality of computationally obtained clusters. A large number of different internal indices have been developed for validation in the offline setting. However, thi…
The paper challenges the validity of cluster validity measures in unsupervised learning.
problem The validity of cluster validity measures in selecting optimal clusterings.
method The authors investigate the use of cluster validity measures as objective functions in unsupervised learning and introduce a new variant of the Dunn index.
result Many cluster validity measures promote clusterings that do not match expert knowledge well.
Develops a new cluster validity index to find multiple optimal cluster numbers.
problem Finding the optimal number of clusters in real-world data with varying densities, sizes, and shapes.
method A new correlation-based cluster validity index that yields multiple local peaks.
result The new index finds multiple optimal cluster numbers in various scenarios.
Introduces BCVI, a Bayesian cluster validity index for better cluster selection.
problem Selecting the optimal number of clusters in clustering algorithms.
method Develops BCVI using Dirichlet or generalized Dirichlet priors, evaluated against existing indices.
result BCVI offers clear advantages in scenarios where user expertise is valuable.
New indices for determining cluster compactness and separability.
problem Challenges in identifying true clusters in data sets.
method Developed absolute cluster indices to measure compactness and separability.
result Demonstrated improved performance compared to existing indices.
The paper extends cluster validity indices for incremental analysis.
problem Providing incremental alternatives for cluster validation.
method Extending iCVI family to include 6 incremental indices and examining their behavior under under- and over-partitioning.
result Over-partitioning is more challenging to detect than under-partitioning.
This paper tackles deep clustering evaluation challenges in high-dimensional data.
problem Evaluation of deep clustering methods is problematic due to the curse of dimensionality and variations in embedding spaces.
method Develops a theoretical framework to highlight the ineffectiveness of internal validation measures and proposes a systematic approach to applying clustering validity indices in deep learning.
result The proposed framework reduces misguidance from improper use of clustering validity indices in deep learning.
OUI tool detects optimal Weight Decay for DNNs without validation data.
problem Optimal Weight Decay hyperparameter selection for DNNs.
method Overfitting-Underfitting Indicator (OUI) tool.
result OUI correlates with improved generalization and validation scores.
Enhances clustering quality evaluation in noisy data.
problem Reliable clustering quality assessment in noisy Gaussian mixtures.
method Feature Importance Rescaling (FIR) method.
result FIR improves correlation between cluster validity indices and ground truth.
iCVI-ARTMAP accelerates clustering with adaptive resonance theory and validity indices.
problem Improving clustering efficiency and accuracy using adaptive resonance theory.
method Integrates adaptive resonance theory (ARTMAP) with incremental cluster validity indices (iCVIs) for clustering.
result Significantly reduces clustering time and outperforms other methods on synthetic and real-world data.
New validity index for fuzzy-possibilistic c-means clustering.
problem Conflicting results in determining the optimal number of clusters due to noisy data points and outliers.
method Introducing a new validity index (FP index) for fuzzy-possibilistic c-means clustering.
result FP index works well in datasets with varying cluster shapes and densities.
Stacked conformal prediction simplifies model validation.
problem Validating stacked predictive models efficiently.
method Meta-learner at the top of a stacked ensemble for approximate marginal validity.
result The method achieves approximate marginal validity without a separate calibration sample.
New algorithm for fuzzy clustering improves clustering quality.
problem Seeding iterative fuzzy algorithms for high-quality clustering.
method MaxMin Linear initialization algorithm for fuzzy C-Means.
result Validation through experiments on various datasets.
Study uses topological signatures to quantify financial market complexity.
problem Capturing temporal organization beyond volatility measures.
method Null validated topological approach using L1 norm of persistence landscapes. result Persistence landscape norms reveal dynamical structure during market stress.
Study finds RVIs unreliable for SP selection in clustering.
problem Reliability of RVIs for selecting Similarity Paradigms (SPs) in clustering.
method Extensive experiments with 7 RVIs on synthetic and real-world datasets.
result RVIs are unreliable for SP selection.
Novel approach detects early warning indicators in complex systems.
problem Detecting abrupt transitions in complex systems.
method Directed anisotropic diffusion map and latent stochastic dynamical systems.
result Early warning indicators can detect tipping points in state transitions.
The paper analyzes indices based on counting object pairs for assessing partition agreement in unsupervised learning.
problem The difficulty in interpreting overall indices like Rand and adjusted Rand indices.
method Analysis of three families of indices based on counting object pairs, decomposing overall indices into cluster-level indices.
result Overall indices based on pair-counting approach are sensitive to cluster size imbalance and provide limited information on smaller clusters.
CARVE validates clustering results using resampling and stability analysis.
problem Inconsistent and unreliable clustering results due to algorithm, preprocessing, and k sensitivity. method CARVE uses resampling-based validation and stability analysis to evaluate multiple clustering algorithms and hyperparameters.
result CARVE consistently recovers near-optimal clusterings and finer biological structure.
Method identifies causal interactions between time series using extreme eigenvalue variability.
problem Detecting causal interactions between time series.
method Largest eigenvalue of lagged correlation matrices, measuring causal interactions through variability.
result The method outperforms traditional Granger causality tests in detecting structural changes.
Statistical physics of complex systems exploits network theory not only to model, but also to effectively extract information from many dynamical real-world systems. A pivotal case of study is given by financial systems: market prediction represents an unsolved scientific challenge yet with crucial implications for soc…
The paper evaluates integrals for fBm with various Hurst indices.
problem Evaluating integrals for stochastic processes with fractional Brownian motion for different Hurst indices.
method Analytic continuation from complex analysis to extend integral domain.
result Integral formulas for fBm with Hurst indices H∈(0,1) are derived. Paper uses stacking with neural networks to predict cryptocurrency price direction.
problem Predicting the direction of cryptocurrency prices.
method Generative and discriminative classifiers stacked over a one-layer neural network, using technical indicators and sentiment analysis.
result Stacking method outperformed individual models in accuracy.
Study finds macroeconomic indicators predict health workforce and infrastructure measures.
problem Evaluating the predictive value of macroeconomic indicators for public health targets.
method Examined multiple forecasting approaches including neural networks, generalized additive models, random forests, and time series models with exogenous indicators.
result Macroeconomic indicators provide consistent and reproducible predictive signals for health workforce and infrastructure measures, but less so for other targets.
TINs use neural networks to interpret technical indicators for trading.
problem Lack of interpretable neural architectures for technical indicators in trading.
method Introduced TINs, a neural architecture that reformulates technical indicators into trainable modules.
result Improved risk-adjusted performance compared to traditional indicator-based strategies.
Paper proposes a cost-sensitive conformal training method with provably controllable learning bounds.
problem Uncertainty quantification and learning bounds in conformal prediction.
method Cost-sensitive conformal training algorithm that minimizes the expected size of prediction sets using rank weighting.
result Theoretical analysis shows tightness between weighted objective and expected size of conformal prediction sets.
HD-BWDM improves clustering validation in high-dimensional data.
problem Determining the right number of clusters in high-dimensional data.
method HD-BWDM integrates random projection, PCA, trimmed clustering, and medoid-based distances.
result HD-BWDM remains stable and interpretable under high-dimensional projections and contamination.
Improved genetic algorithm optimizes SVR for robust long-term stock index forecasting.
problem Inaccurate long-term stock price predictions.
method Adaptive Weighted Genetic Algorithm-Optimized SVR (IGA-SVR).
result Reduction in MAPE by 19.87% compared to LSTM and 50.03% compared to OGA-SVR.
Recent studies have shown that tuning prediction models increases prediction accuracy and that Random Forest can be used to construct prediction intervals. However, to our best knowledge, no study has investigated the need to, and the manner in which one can, tune Random Forest for optimizing prediction intervals { thi…
Transformer models improve financial sentiment measurement.
problem Capturing nuanced sentiment from financial news articles.
method Transformer-based language models for sentiment classification and aggregation.
result Transformer models outperform traditional dictionary-based methods in sentiment classification.
A very simple interpretation of matrix completion problem is introduced based on statistical models. Combined with the well-known results from missing data analysis, such interpretation indicates that matrix completion is still a valid and principled estimation procedure even without the missing completely at random (M…
Develops CPL for optimal prediction set length and validity.
problem Balancing conditional validity and length efficiency in conformal prediction.
method Conformal Prediction with Length-Optimization (CPL).
result Achieves optimal prediction set length while maintaining conditional validity.
Predicting stock jumps using liquidity and technical indicators.
problem Predicting intraday stock jumps in finance.
method Divide trading day into 5-minute intervals, use liquidity measures and technical indicators, apply machine learning algorithms.
result Initial evidence of predictability of jump arrivals and directions using level-2 stock data.
Payments data and machine learning improve nowcasting accuracy for macroeconomic indicators.
problem Lagged indicators in linear models are insufficient during crisis periods.
method Non-traditional payments data, nonlinear machine learning, and tailored cross-validation.
result Improved macroeconomic nowcasting accuracy up to 40% during crises.
Paper presents a new validation method for simulation workflow.
problem Validation of simulation workflow is challenging and domain-dependent.
method Empirical learning-based validation procedure using AHP and machine learning.
result Validation procedure is semi-automated and efficient.
This paper optimizes cryptocurrency portfolios by integrating sentiment analysis with technical indicators.
problem Effective portfolio management in volatile cryptocurrency markets.
method Dynamic portfolio strategy using technical indicators and sentiment analysis.
result The integrated approach outperforms traditional benchmarks and achieves stronger risk-adjusted returns.
Following the financial crisis of the late 2000s, policy makers have shown considerable interest in monitoring financial stability. Several central banks now publish indices of financial stress, which are essentially based upon market related data. In this paper, we examine the potential for improving the indices by de…
Study on time-varying APT validity in Japanese stock market.
problem Validity of Arbitrage Pricing Theory (APT) in Japanese stock market over time.
method Rolling window method applied to Fama and MacBeth's two-step regression and Kamstra and Shi's generalized GRS test.
result APT validity is unstable over time in Japanese stock market, influenced by monetary policy and business cycle.
A new model validation framework for agentic AI systems based on POMDPs.
problem Model validation of agentic AI systems.
method A POMDP-based framework for belief-state, forecast, and policy validation.
result The framework decomposes autonomous decision making into information, beliefs, forecasts, actions, and utility.
New AI stock indices classify firms' AI engagement using 10-K filings.
problem Opaque AI selection criteria in existing ETFs.
method NLP analysis of 10-K filings to classify AI stocks.
result Companies with higher AI engagement have greater positive returns.
Paper proposes a new GPR-HS framework for accurate VCV estimation in global equity indices.
problem Accurate forecasting of Volatility-Covariance Matrix (VCV) for regulatory processes.
method Hybrid Gaussian Process Regression-Historical Simulation (GPR-HS) framework.
result GPR-HS framework achieves regulatory compliance and outperforms static VaR benchmarks.
This paper deals with the problem of describing the vector spaces of divergence-free, natural tensors on a pseudo-Riemannian manifold that are second-order; i.e., that are defined using only second derivatives of the metric. The main result establishes isomorphisms between these spaces and certain spaces of tensors (at…
A global agreement on how to reduce and cap human footprint, especially their GHG emissions, is very unlikely in near future. At the same time, bilateral agreements would be inefficient because of their neural and balanced nature. Therefore, unilateral actions would have attracted attention as a practical option. Howev…
Study detects Chinese stock market bubbles using LPPLS confidence indicator.
problem Early detection of stock market bubbles in China.
method LPPLS confidence indicator applied to CSI 300 index data.
result LPPLS detects positive and negative bubbles with high accuracy.
Deep asymmetric networks prioritize feature importance.
problem Learning feature importance in neural networks.
method Nodes with smaller indices are more sensitive to variant activation functions.
result Features are learned in order of importance, and nodes can be pruned.
Study finds short-term wage increases due to COVID-19, contrary to expectations.
problem Impact of COVID-19 on wages over time.
method Empirical analysis controlling for GDP as a demand proxy.
result Short-term positive wage effect, contrary to expectations.
This paper explores how train-validation splits help in NAS to prevent overfitting.
problem NAS overfits with train-validation splits and needs better generalization guarantees.
method Established refined properties of validation loss and risk for NAS.
result NAS with train-validation splits can select the most generalizable model.
New research investigates why influence functions are fragile and proposes new validation procedures.
problem Understanding and mitigating the fragility of influence functions in deep learning model explanations.
method Verification of influence functions using various conditions and procedures, including convexity and non-convexity.
result Validation procedures may cause the observed fragility of influence functions.