Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

105210315420 · Jun 202019922001200920172026
48 results for statistical similarity

Method measures weight similarity in neural networks using normalization and statistical inference.

problem Quantifying weight similarity in non-convex neural networks.
method Chain normalization rule and hypothesis-training-testing statistical inference.
result Weights of identical neural networks converge to similar local solutions.

Defines a similarity measure for classification distributions.

problem Measuring similarity between classification distributions.
method Proposes task similarity, a novel measure quantifying performance of source distributions on target distributions.
result Empirical task similarity correlates with transfer efficiency and semantic similarity of source distributions.

Recent work has sought to understand the behavior of neural networks by comparing representations between layers and between different trained models. We examine methods for comparing neural network representations based on canonical correlation analysis (CCA). We show that CCA belongs to a family of statistics for mea…

2019-05-01abs ↗pdf ↗

The study improves the assessment of fairness in face recognition using ROC curves and statistical guarantees.

problem Improving the assessment of fairness in face recognition systems.
method Proves asymptotic guarantees for empirical ROC curves and fairness metrics, and introduces a recentering technique to avoid bootstrap pitfalls.
result Demonstrates the practical relevance of the methods for assessing fairness in face recognition systems.

Proposes a method to compare noisy high-dimensional datasets with low-dimensional manifolds.

problem Comparing distributions on manifolds in noisy high-dimensional datasets.
method Linking low-rank structure to manifold geometry, developing a scale-invariant distance measure.
result Superior robustness and statistical power compared to existing methods.

A new test assesses text similarity between two groups of documents.

problem Comparing similarity between two groups of documents.
method Neural network-based language models estimate entropy, and a test statistic derived from an estimation-and-inference framework is used.
result The proposed test maintains the nominal Type one error rate while offering greater power compared to existing methods.

The study analyzes the conflict between group fairness and individual fairness in machine learning.

problem The conflict between group fairness (optimal statistical parity) and individual fairness in machine learning.
method Established sufficient conditions for the compatibility between optimal statistical parity and individual fairness requirements.
result Identified regions along the Pareto frontier that satisfy individual fairness requirements.

For many analytical problems the challenge is to handle huge amounts of available data. However, there are data science application areas where collecting information is difficult and costly, e.g., in the study of geological phenomena, rare diseases, faults in complex systems, insurance frauds, etc. In many such cases,…

2019-09-12abs ↗pdf ↗

Study compares Web3 tokens to traditional assets, finding similar statistical properties.

problem Understanding statistical properties of Web3 tokens compared to traditional financial assets.
method Statistical analysis of various Web3 tokens across multiple time scales, comparing with traditional financial assets.
result Most Web3 tokens exhibit similar stylized facts to traditional financial assets, including heavy tails and volatility clustering.

Adaptive methods learn from multiple datasets, leveraging similarities and robust to outliers.

problem Simultaneously analyze multiple datasets with possible similarities and differences.
method Adaptive multi-task learning methods that automatically utilize similarities and handle differences.
result Sharp statistical guarantees and robustness against outlier tasks demonstrated.

This paper introduces individual fairness in clustering using ff-divergence.

problem Ensuring fair clustering by treating similar individuals similarly.
method Uses ff-divergence to measure statistical similarity and assigns individuals to probability distributions over cluster centers.
result Provides an algorithm with provable approximation guarantee for clustering with individual fairness constraints.

From a sequence of similarity networks, with edges representing certain similarity measures between nodes, we are interested in detecting a change-point which changes the statistical property of the networks. After the change, a subset of anomalous nodes which compares dissimilarly with the normal nodes. We study a sim…

2016-12-05abs ↗pdf ↗

Robust method estimates self-similarity for mammogram images, improving cancer detection.

problem Statistical assessment of self-similarity in real data with large mean level shifts.
method Theil-type weighted regression for wavelet-based estimation, compared to OLS and AV.
result Robust approach shows nearly 68% accuracy in cancer vs non-cancer classification.

Machine learning and statistical modeling complement each other in healthcare analytics.

problem Choosing between machine learning and statistical modeling for analytics challenges.
method Choosing based on problem, data, and desired outcomes.
result Machine learning and statistical modeling are complementary, using similar principles but different tools.

New framework to test neural network representation similarity measures.

problem Disagreements among dissimilarity measures in neural networks.
method Statistical testing framework to evaluate measures based on functional behavior.
result Current metrics have different weaknesses; a classical baseline performs surprisingly well.

Geometry of hypersurfaces defined by the relation which generalizes classical formula for free energy in terms of microstates is studied. Induced metric, Riemann curvature tensor, Gauss-Kronecker curvature and associated entropy are calculated. Special class of ideal statistical hypersurfaces is analyzed in details. No…

2016-02-25abs ↗pdf ↗

The study explores whether model selection guarantees apply to contextual bandits.

problem Applying model selection guarantees to contextual bandits.
method Investigates whether similar guarantees for model selection in statistical learning can be extended to contextual bandit learning.
result Initial findings suggest that model selection guarantees may not directly apply to contextual bandits.

We study the problem of discriminative sub-trajectory mining. Given two groups of trajectories, the goal of this problem is to extract moving patterns in the form of sub-trajectories which are more similar to sub-trajectories of one group and less similar to those of the other. We propose a new method called Statistica…

2019-05-06abs ↗pdf ↗

We investigate scaling and memory effects in return intervals between price volatilities above a certain threshold qq for the Japanese stock market using daily and intraday data sets. We find that the distribution of return intervals can be approximated by a scaling function that depends only on the ratio between the …

2007-09-11abs ↗pdf ↗

The paper examines how to test if two learning algorithms produce similar outcomes.

problem Testing if two learning algorithms produce similar outcomes when trained on different data sets.
method Using Total Variation (TV) distance to measure similarity of posterior distributions.
result TV indistinguishable learning rules are equivalent to existing stability notions and can be statistically amplified.

New statistics are introduced that maintain the Fisher metric structure closely, akin to sufficient statistics.

problem Maintaining the Fisher metric structure in statistical models.
method Characterizing statistics that maintain the Fisher metric structure bi-Lipschitz equivalently.
result Characterized statistics that preserve the Fisher metric structure closely.

Develops a framework for identifying mispriced assets through attention factors for statistical arbitrage.

problem Identifying mispriced assets in statistical arbitrage trading.
method Uses conditional latent factors learned from firm characteristic embeddings to identify time-series signals and form a trading strategy.
result Achieves an out-of-sample Sharpe ratio above 4 on the largest U.S. equities over a 24-year period.

We introduce a new discrepancy score between two distributions that gives an indication on their similarity. While much research has been done to determine if two samples come from exactly the same distribution, much less research considered the problem of determining if two finite samples come from similar distributio…

2012-10-15abs ↗pdf ↗

Unified understanding of neural representation similarity measures.

problem Fragmented research landscape of neural network similarity measures.
method Observation and exploration of connections between shape distances and normalized Bures similarity.
result Cosine of the Riemannian shape distance equals normalized Bures similarity.

Improves Gower's similarity for mixed-type variables with automatic weighting.

problem Handling missing values and unbalanced variable contributions in Gower's similarity for mixed-type data.
method Automatic weighting scheme minimizing differences in correlation between contributing dissimilarities and weighted Gower's dissimilarity.
result Improved performance in classification and imputation of missing values.

Estimates shared parameters across related learning problems using robust statistics and LASSO.

problem Simultaneously learning related but heterogeneous problems like store demand or patient risk.
method Two-stage multitask learning estimator combining robust statistics and LASSO regression.
result Improved sample complexity bounds for multitask learning, especially beneficial for 'data-poor' instances.

Recently, metric learning and similarity learning have attracted a large amount of interest. Many models and optimisation algorithms have been proposed. However, there is relatively little work on the generalization analysis of such methods. In this paper, we derive novel generalization bounds of metric and similarity …

2012-07-23abs ↗pdf ↗

Excessive reuse of test data has become commonplace in today's machine learning workflows. Popular benchmarks, competitions, industrial scale tuning, among other applications, all involve test data reuse beyond guidance by statistical confidence bounds. Nonetheless, recent replication studies give evidence that popular…

2019-05-29abs ↗pdf ↗

Solves a 60-year-old question on agreement measures in statistics.

problem The challenge of measuring agreement between two raters or measures.
method Developed a new algorithm to minimize diagonals in contingency tables, formulated the minimum feasible agreement, and studied the lower limit of maximum feasible agreement.
result Formulated the lower limit of Cohen's kappa and two statistics for agreement analysis.

Method embeds numeric tabular datasets into a shared vector space for similarity and retrieval.

problem Lack of meaningful representation for numeric tabular datasets in large language models.
method Structured exploratory data analysis descriptors, sentence transformer embedding, CCA for cross-dataset alignment.
result Total P@1 score of 0.9 across 15 datasets, robust nearest-neighbor retrieval and cluster structure.

Active seriation recovers item order from noisy pairwise similarity measurements.

problem Recovering an unknown item ordering from noisy pairwise similarity measurements.
method Proposes an active seriation algorithm that provably recovers the latent ordering with high probability.
result Establishes optimal performance guarantees for successful recovery under a uniform separation condition.

The paper reports the construction of artificial stock market that emerges the similar statistical facts with real data in Indonesian stock market. We use the individual but dominant data, i.e.: PT TELKOM in hourly interval. The artificial stock market shows standard statistical facts, e.g.: volatility clustering, the …

2004-08-16abs ↗pdf ↗

Following the work of Okuyama, Takayasu and Takayasu [Okuyama, Takayasu and Takayasu 1999] we analyze huge databases of Japanese companies' financial figures and confirm that the Zipf's law, a power law distribution with the exponent -1, has been maintained over 30 years in the income distribution of Japanese companies…

2003-08-19abs ↗pdf ↗

Paper proposes a supervised similarity framework for corporate bonds using RF proximities.

problem Challenges in measuring similarity for corporate bonds due to noisy data and lack of ground truth.
method Proposes a supervised similarity framework using Random Forest for corporate bonds, introducing a novel metric to evaluate similarities.
result Random Forest outperforms other methods in evaluating similarities for corporate bonds.

Similarity plays a fundamental role in many areas, including data mining, machine learning, statistics and various applied domains. Inspired by the success of ensemble methods and the flexibility of trees, we propose to learn a similarity kernel called rpf-kernel through random projection forests (rpForests). Our theor…

2019-08-28abs ↗pdf ↗