Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

4285127169 · Jun 202019922001200920182026
48 results for distance threshold

Automatically tunes distance threshold in metric learning.

problem Manual tuning of distance threshold in ITML-based methods is sensitive and time-consuming.
method Optimized metric learning algorithm using Dykstra algorithm to solve nonlinear equation efficiently.
result The proposed metric learning algorithm automatically tunes the distance threshold and achieves comparable accuracy.

New algorithm identifies good arms with fewer samples when thresholds are close.

problem Good arm identification in bandit problems with small threshold gaps.
method Proposes lil'HDoC algorithm to improve GAI under small threshold gaps.
result Sample complexity of first λ output arm is nearly identical to HDoC algorithm when thresholds are close.

Revisits fuzzy neural networks with generalized Hamming distance, simplifying BN and ReLU.

problem Improving neural network techniques using fuzzy logic and generalized Hamming distance.
method Introducing generalized Hamming distance to reinterpret BN and ReLU, proposing GHN.
result Batch normalization and ReLU can be simplified or removed without loss of performance.

Study examines stock market connections before, during, and after the 2008 financial crisis.

problem Effects of the 2008 global financial crisis on stock market connectivity.
method Generated complex networks from cross-correlation matrices, using threshold networks and minimal spanning trees.
result During the crisis, countries in different zones had varying levels of connectivity.

Optimizes risk assessment tools using mixed-integer programming.

problem Challenges in healthcare risk assessment due to label scarcity and asymmetric misclassification costs.
method Jointly optimizes scoring weights and category thresholds via mixed-integer programming (MIP).
result Prevents label-scarce category collapse and achieves more accurate risk categorization.

The study bounds the stability of Gaussian mixtures under small perturbations.

problem Stability of Gaussian mixtures under small changes in distribution.
method Deriving an explicit bound on parameter stability of spherical Gaussian Mixture Models (sGMM) in a pre-defined model class.
result Upper bound on parameter distance of close sGMMs to the original sGMM, dependent only on the original model.

Paper introduces WWAggr for ensemble CPD, improving accuracy and decision threshold selection.

problem Challenges in detecting abrupt distribution shifts in high-dimensional data streams.
method Introduces WWAggr, a novel task-specific ensemble aggregation method based on Wasserstein distance.
result Demonstrates WWAggr outperforms standard aggregation techniques and decision threshold selection.

Study examines financial market structure changes during the COVID-19 crash using a novel MI approach.

problem Analyzing nonlinear dependencies among major stocks during market crashes.
method Conditional p-threshold mutual information (MI) and Minimum Spanning Tree (MST) framework.
result Financial networks become more integrated during crashes, with increased periphery vulnerability.

The paper identifies clusters in the World Trade Network using communicability distances.

problem Identifying clusters in the World Trade Network.
method Uses Estrada and vibrational communicability distances to find clusters maximizing a modularity function.
result Identifies specific distance thresholds that maximize modularity, revealing unique relationships between countries.

Transforms distance-based outlier scores into interpretable probabilistic estimates.

problem Difficult interpretation of distance-based outlier scores.
method Generic transformation of scores into probabilistic estimates using distance probability distributions.
result Probabilistic transformation improves interpretability without impacting detection performance.

Optimal threshold resetting reduces search time for multiple diffusive searchers.

problem Optimizing search time for multiple diffusive searchers in a one-dimensional space.
method Threshold resetting (TR) is introduced as an event-driven optimization strategy, coupling resetting to the internal dynamics of searchers.
result Optimal threshold distance uu significantly reduces mean first-passage time for N2N \geq 2 searchers, with a minimum at Nopt(u)N_{\mathrm{opt}}(u).

For a certain class of distributions, we prove that the linear programming relaxation of kk-medoids clustering---a variant of kk-means clustering where means are replaced by exemplars from within the dataset---distinguishes points drawn from nonoverlapping balls with high probability once the number of points drawn a…

2013-09-12abs ↗pdf ↗

Develops a hypothesis testing framework for generalized Thurstone models.

problem Determining whether pairwise comparison data fits a generalized Thurstone model.
method Introduces separation distance and derives upper and lower bounds for testing.
result Critical threshold for testing depends on observation graph topology and scales as Θ((nk)1/2)Θ((nk)^{-1/2}) for complete graphs.

New method learns local metrics for k-NN classification using sample similarity.

problem Improving k-NN classification accuracy through better distance metrics.
method Local distance metric learning based on sample similarity, using conical combinations of metric weight matrices.
result New metrics yield smaller distances for similar samples and larger distances for dissimilar ones.

We improve deep threshold networks' memorization capacity exponentially.

problem Memorizing datasets with randomized labels using deep neural networks.
method Using Gaussian random weights in the first layer and binary or integer weights in subsequent layers, we prove a new dependence on minimum distance.
result We show that O~(1δ+n)\widetilde{\mathcal{O}}(\frac{1}{\delta} + \sqrt{n}) neurons and O~(dδ+n)\widetilde{\mathcal{O}}(\frac{d}{\delta} + n) weights are sufficient.

Efficient method classifies locally stationary time series based on second-order characteristics.

problem Classifying locally stationary time series for various applications.
method Autoregressive approximation, ensemble aggregation, distance-based threshold.
result Zero misclassification error rate asymptotically for mildly differing second-order characteristics.

The study uses persistent homology to determine when Voronoi interpolation should stop.

problem Interpolating complex topological data sets accurately.
method Persistent homology is applied to the Voronoi tessellation to detect changes in the data's topology.
result The method effectively identifies when the interpolation has captured the data's topology changes.

New algorithm achieves strong consistency in binary non-uniform hypergraph classification.

problem Node classification on binary non-uniform hypergraphs with varying edge probabilities.
method Proposes a refinement algorithm using power iteration on weighted adjacency matrices.
result Proves optimality of the refinement algorithm, achieving strong consistency and IT lower bound.

This paper shows how to estimate distances in latent space of random graphs using entropic OT.

problem Estimating distances between groups of nodes in latent space of random graphs.
method Entropic Optimal Transport (OT) with stability results for perturbations of the cost matrix.
result Consistent estimation of entropic OT distances between groups of nodes in latent space.

Support Vector Data Description (SVDD) is a machine learning technique used for single class classification and outlier detection. SVDD based K-chart was first introduced by Sun and Tsung for monitoring multivariate processes when underlying distribution of process parameters or quality characteristics depart from Norm…

2016-07-25abs ↗pdf ↗

This paper finds the noise threshold for learning Gaussian mixture models equals channel capacity.

problem Learning Gaussian mixture models with noisy data.
method Bayesian formulation with uniformly distributed centers on a sphere, analyzing the large system limit.
result The maximal noise level σ2σ^2 for which GMM learning is as easy as labeled observations is the channel capacity.

Adaptive sampling theory has shown that, with proper assumptions on the signal class, algorithms exist to reconstruct a signal in Rd\mathbb{R}^{d} with an optimal number of samples. We generalize this problem to the case of spatial signals, where the sampling cost is a function of both the number of samples taken and t…

2015-09-28abs ↗pdf ↗

New method for optimal transport with missing data, debiased and efficient.

problem Solving optimal transport between two distributions with missing values.
method Debiasing Wasserstein distance for empirical Gaussian distributions, entropic regularized optimal transport using ISVT.
result Efficient and consistent estimation of entropic regularized optimal transport.

Subspace clustering refers to the problem of clustering high-dimensional data points into a union of low-dimensional linear subspaces, where the number of subspaces, their dimensions and orientations are all unknown. In this paper, we propose a variation of the recently introduced thresholding-based subspace clustering…

2014-03-13abs ↗pdf ↗

This work bounds the run-time of nonconvex optimization with early stopping.

problem Bounding the expected run-time of nonconvex optimization with early stopping.
method Derives conditions for well-defined early stopping based on validation function norms and bounds the expected number of iterations and gradient evaluations.
result Guarantees the validity of early stopping and provides bounds on the expected run-time for various optimization algorithms.

The problem of clustering noisy and incompletely observed high-dimensional data points into a union of low-dimensional subspaces and a set of outliers is considered. The number of subspaces, their dimensions, and their orientations are assumed unknown. We propose a simple low-complexity subspace clustering algorithm, w…

2013-07-18abs ↗pdf ↗

The paper improves conformal prediction by analyzing the beta law of conditional coverage.

problem Improving finite-sample marginal coverage guarantees for non-i.i.d. data.
method The method uses Wasserstein distances to quantify deviations from the beta law of conditional coverage.
result The framework provides direct bounds on marginal coverage gaps and bad-calibration probabilities.

Optimal iterative thresholding algorithms improve upon hard and soft thresholding.

problem Optimizing sparsity or rank constraints in optimization problems.
method Developed the notion of relative concavity for thresholding operators, finding a new class of operators that are optimal.
result A new class of thresholding operators, including q\ell_q thresholding and reciprocal thresholding, achieves the strongest convergence guarantee.

Proposes COLA, a communication-efficient algorithm for decentralized optimization.

problem Decentralized consensus optimization over a network.
method Linearization and communication-censoring strategy to reduce computation and communication costs.
result Proven convergence and established convergence rates for COLA.