Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

222445667889 · Jun 202019922001200920172026
48 results for Dominant Set Clustering

Unified framework for efficient Frank-Wolfe optimization of Dominant Set Clustering.

problem Optimizing Dominant Set Clustering with various Frank-Wolfe algorithms.
method Unified framework for pairwise, standard, and away-steps Frank-Wolfe algorithms, with explicit convergence rates.
result Explicit convergence rates for Frank-Wolfe methods in Dominant Set Clustering.

This paper proposes a new clustering method based on Stochastic Dominance for asset allocation.

problem Traditional clustering methods fail to capture risk dominance relationships among assets.
method Integrates Stochastic Dominance theory with machine learning algorithms to construct a Stochastic Dominance Coefficient Matrix and modify clustering algorithms.
result The proposed method effectively facilitates customized asset allocation for investors.

The increasing needs of clustering massive datasets and the high cost of running clustering algorithms poses difficult problems for users. In this context it is important to determine if a data set is clusterable, that is, it may be partitioned efficiently into well-differentiated groups containing similar objects. We …

2019-08-28abs ↗pdf ↗

We propose a combination of cluster analysis and stochastic process analysis to characterize high-dimensional complex dynamical systems by few dominating variables. As an example, stock market data are analyzed for which the dynamical stability as well as transitions between different stable states are found. This comb…

2015-02-26abs ↗pdf ↗

We extend the theoretical analysis of a recently proposed single subspace learning algorithm, called Dual Principal Component Pursuit (DPCP), to the case where the data are drawn from of a union of hyperplanes. To gain insight into the properties of the 1\ell_1 non-convex problem associated with DPCP, we develop a geo…

2017-06-06abs ↗pdf ↗

Connected domination numbers found for plane triangulations up to 13 vertices.

problem Finding connected domination numbers for plane triangulations.
method Analyzing triangulations of up to 13 vertices and proving the difference between connected and regular domination numbers can be arbitrarily large.
result Connected domination numbers for triangulations up to 13 vertices and upper bound for larger triangulations.

Kernel methods are popular in clustering due to their generality and discriminating power. However, we show that many kernel clustering criteria have density biases theoretically explaining some practically significant artifacts empirically observed in the past. For example, we provide conditions and formally prove the…

2017-05-16abs ↗pdf ↗

FOSC-X: An extended framework for extracting multiple optimal flat clusterings from hierarchical cluster trees

problem Extracting multiple optimal flat clusterings from hierarchical cluster trees
method Dynamic programming with lower and upper feasibility bounds
result Guaranteed optimal rankings of top-M solutions with linear-time complexity

Paper proposes a new method to improve clustering ensemble performance.

problem Improving clustering ensemble performance by refining co-association matrix.
method Low-rank tensor approximation to derive coherent-link matrix and refine co-association matrix.
result The proposed method achieves breakthrough in clustering performance compared to state-of-the-art methods.

MAS scores cluster size consistency from points, robust to label changes.

problem Desired uniformity in cluster sizes, stability under label perturbations.
method Mass Agreement Score (MAS) measures point-centric cluster size consistency, robust to label changes.
result MAS yields similar scores for partitions with similar bulk structure, sensitive to genuine redistribution of cluster mass.

This paper presents a neural network-based end-to-end clustering framework. We design a novel strategy to utilize the contrastive criteria for pushing data-forming clusters directly from raw data, in addition to learning a feature embedding suitable for such clustering. The network is trained with weak labels, specific…

2015-11-19abs ↗pdf ↗

Complex network analysis reveals dominant stocks in financial stock returns correlations.

problem Inferring financial stock returns correlations from complex network analysis.
method Simulated geometric Brownian motion for stocks, complex network analysis, eigenvector centrality, clustering.
result Returns correlation matrix is dominated by stocks with high eigenvector centrality and clustering.

Entropy regularization improves interpretability of probabilistic clustering models.

problem Bayesian nonparametric mixture models often produce unbalanced cluster frequencies.
method Interpreting the posterior as penalized likelihood, entropy regularization reduces sparsely-populated clusters.
result The proposed entropy-regularized estimator enhances interpretability without sacrificing computational convenience.

New method separates market motion from stock correlations.

problem Understanding the dynamics of stock correlations relative to market motion.
method Cluster reduced-rank correlation matrices by subtracting the largest eigenvalue.
result Extracted market states are quasi-stationary over long periods.

Framework clusters noisy MTS with robust fuzzy clustering, improving accuracy over existing methods.

problem Challenges in clustering multivariate time series due to non-stationary dependencies, noise, and state boundaries.
method Spectral fuzzy clustering using Kendall's tau-based canonical coherence for frequency-specific monotonic relationships.
result Framework outperforms existing methods in clustering noisy, high-dimensional MTS.

This paper considers the problem of estimating a high-dimensional vector of parameters θRn\boldsymbolθ \in \mathbb{R}^n from a noisy observation. The noise vector is i.i.d. Gaussian with known variance. For a squared-error loss function, the James-Stein (JS) estimator is known to dominate the simple maximum-likelihood (…

2016-02-01abs ↗pdf ↗

Proves rigidity for specific initial data sets under the dominant energy condition.

problem Rigidity of initial data sets with boundary and convex polytopes.
method Solution of boundary value problems for Dirac operators and approximations by manifolds with smooth boundary.
result Proves rigidity for compact smooth spin manifolds and convex polytopes under the dominant energy condition.

The study uses unsupervised machine learning to identify top European football teams.

problem Selecting teams for the new European football Super League.
method Used Laplacian eigenmaps clustering on performance data.
result Successfully identified four clusters of teams based on performance metrics.

New method ranks multivariate distributions in SMOOP using q-dominance.

problem Lack of reliable methods to rank multivariate distributions in SMOOP.
method Introduces center-outward q-dominance and develops empirical test procedures.
result Proves q-dominance implies FSD and establishes a sample size threshold.

Kernel-based K-means clustering has gained popularity due to its simplicity and the power of its implicit non-linear representation of the data. A dominant concern is the memory requirement since memory scales as the square of the number of data points. We provide a new analysis of a class of approximate kernel methods…

2016-08-26abs ↗pdf ↗

Proves positive mass theorem for spin initial data sets with arbitrary ends and dominant energy shields.

problem Proving the positive mass theorem for spin initial data sets with various ends and energy shields.
method Modification of Witten's approach involving an additional independent timelike direction in the spinor bundle.
result Positive mass theorem for spin initial data sets with arbitrary ends and dominant energy shields.

Hierarchical clustering is a class of algorithms that seeks to build a hierarchy of clusters. It has been the dominant approach to constructing embedded classification schemes since it outputs dendrograms, which capture the hierarchical relationship among members at all levels of granularity, simultaneously. Being gree…

2018-06-09abs ↗pdf ↗

New insights into Bartnik mass from improvability of dominant energy scalar.

problem Characterizing Bartnik mass minimizing initial data sets.
method Introducing improvability concept, proving non-improvability consequences, and analyzing pp-wave counterexamples.
result Bartnik mass minimizing initial data sets are characterized, advancing conjectures.

New framework for ranking distributions using variable fractional parameters.

problem Ordering distributions with varying steepness and local non-concavities.
method Introducing a function γ:Ro[0,1]\boldsymbolγ: \mathbb{R} o [0,1] to replace the fixed parameter in fractional SD.
result Enables ranking of a broader range of distributions and incorporates dynamic greediness.

Automatic cover detection -- the task of finding in a audio dataset all covers of a query track -- has long been a challenging theoretical problem in MIR community. It also became a practical need for music composers societies requiring to detect automatically if an audio excerpt embeds musical content belonging to the…

2019-10-22abs ↗pdf ↗

Smooth dec initial data sets may not extend to smooth spacetimes.

problem Whether every dec initial data set can be extended to a smooth spacetime.
method Examined the converse of the dominant energy condition for initial data sets and spacelike hypersurfaces.
result Not all dec initial data sets can be extended to smooth spacetimes.

X-DC improves speech separation by making DNNs more interpretable.

problem Black-box nature of DNNs in speech separation tasks.
method Introduces X-DC, a DNN architecture that interprets as spectrogram template fitting followed by Wiener filtering.
result X-DC achieves comparable speech separation performance to DC but with enhanced interpretability.

New algorithm finds sparse matrices on Stiefel manifold for optimisation.

problem Finding sparse matrices on Stiefel manifold for optimisation.
method Modified Orthogonal Iteration algorithm for sparse global optimality.
result Proposed method finds globally optimal sparse Stiefel matrices.

SC-InfoNCE improves InfoNCE for feature clustering in contrastive learning.

problem Lack of theoretical understanding of InfoNCE's feature clustering mechanism.
method Introduced a transition probability matrix to model data augmentation dynamics and optimize feature similarity.
result SC-InfoNCE achieves strong performance across diverse domains, aligning feature similarity with downstream data.

Characterizes causal structure dominance for latent variables.

problem Determining dominance relations between causal structures with latent variables.
method Complete characterization for three visible variables, partial for four; uses nontrivial inequality constraints.
result Equivalence classes with nontrivial inequality constraints become ubiquitous as the number of visible variables increases.