Wavelet scattering spectra model non-Gaussian time-series, proving scale invariance for self-similar processes.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
CAST improves spectral clustering for multi-scale data by integrating reachability similarity.
Inverse depth scaling found in LLMs due to similar layers averaging error.
We propose a method to visualize class similarity in large-scale classifiers.
The cognitive framework of conceptual spaces proposes to represent concepts as regions in psychological similarity spaces. These similarity spaces are typically obtained through multidimensional scaling (MDS), which converts human dissimilarity ratings for a fixed set of stimuli into a spatial representation. One can d…
Pyramid Attention Networks improve image restoration by leveraging self-similarities across scales.
In this paper we show similarities between turbulence and financial systems. Motivated by similarities between the two systems, we construct a multiscale model for hierarchical financial structures that exhibits a constant cascade of wealth from large financial entities to small financial entities. According to our mod…
Excessive reuse of test data has become commonplace in today's machine learning workflows. Popular benchmarks, competitions, industrial scale tuning, among other applications, all involve test data reuse beyond guidance by statistical confidence bounds. Nonetheless, recent replication studies give evidence that popular…
The presence of log-periodic structures before and after stock market crashes is considered to be an imprint of an intrinsic discrete scale invariance (DSI) in this complex system. The fractal framework of the theory leaves open the possibility of observing self-similar log-periodic structures at different time scales.…
Spectral analysis of neighborhood graphs is one of the most widely used techniques for exploratory data analysis, with applications ranging from machine learning to social sciences. In such applications, it is typical to first encode relationships between the data samples using an appropriate similarity function. Popul…
We investigate scaling and memory effects in return intervals between price volatilities above a certain threshold for the Japanese stock market using daily and intraday data sets. We find that the distribution of return intervals can be approximated by a scaling function that depends only on the ratio between the …
Mandelbrot unified diverse fields with scaling concept.
We present network embedding algorithms that capture information about a node from the local distribution over node attributes around it, as observed over random walks following an approach similar to Skip-gram. Observations from neighborhoods of different sizes are either pooled (AE) or encoded distinctly in a multi-s…
Quantum-assisted VAE improves similarity search in high-dimensional datasets.
Active Search has become an increasingly useful tool in information retrieval problems where the goal is to discover as many target elements as possible using only limited label queries. With the advent of big data, there is a growing emphasis on the scalability of such techniques to handle very large and very complex …
Framework for inferring latent structure from sparse, imperfectly detected bipartite networks.
We make use of wavelet transform to study the multi-scale, self similar behavior and deviations thereof, in the stock prices of large companies, belonging to different economic sectors. The stock market returns exhibit multi-fractal characteristics, with some of the companies showing deviations at small and large scale…
Despite the success of the popular kernelized support vector machines, they have two major limitations: they are restricted to Positive Semi-Definite (PSD) kernels, and their training complexity scales at least quadratically with the size of the data. Many natural measures of similarity between pairs of samples are not…
Under the framework of spectral clustering, the key of subspace clustering is building a similarity graph which describes the neighborhood relations among data points. Some recent works build the graph using sparse, low-rank, and -norm-based representation, and have achieved state-of-the-art performance. Howeve…
The paper studies entropy calibration in language models and finds that miscalibration improves slowly with scale.
Efficiently selects nearest neighbors for labeling to speed up active learning.
A new method matches similar regions in non-rigid shapes using spectra of differential operators.
Regionalization is the task of dividing up a landscape into homogeneous patches with similar properties. Although this task has a wide range of applications, it has two notable challenges. First, it is assumed that the resulting regions are both homogeneous and spatially contiguous. Second, it is well-recognized that l…
CoSimGNN improves graph similarity computation for large graphs.
This paper studies the matched network inference problem, where the goal is to determine if two networks, defined on a common set of nodes, exhibit a specific form of stochastic similarity. Two notions of similarity are considered: (i) equality, i.e., testing whether the networks arise from the same random graph model,…
Non-Markovian point process shows power-law scaling, similar to nonlinear Markovian process.
We develop a family of techniques to align word embeddings which are derived from different source datasets or created using different mechanisms (e.g., GloVe or word2vec). Our methods are simple and have a closed form to optimally rotate, translate, and scale to minimize root mean squared errors or maximize the averag…
We study distributions which have both fractal and non-fractal scale regions by introducing a typical scale into a scale invariant system. As one of models in which distributions follow power law in the large scale region and deviate further from the power law in the smaller scale region, we employ 2-dim quantum gravit…
We propose a new approach for analyzing price fluctuations in their strongly correlated regime ranging from minutes to months. This is done by employing a self-similarity assumption for the magnitude of coarse-grained price fluctuation or volatility. The existence of a Cramer function, the characteristic function for s…
We define regularity scales to study the behavior of the Calabi flow. Based on estimates of the regularity scales, we obtain convergence theorems of the Calabi flow on extremal Kahler surfaces, under the assumption of global existence of the Calabi flow solutions. Our results partially confirm Donaldson's conjectural p…
The determination of cluster centers generally depends on the scale that we use to analyze the data to be clustered. Inappropriate scale usually leads to unreasonable cluster centers and thus unreasonable results. In this study, we first consider the similarity of elements in the data as the connectivity of nodes in an…
We define the intrinsic scale at which a network begins to reveal its identity as the scale at which subgraphs in the network (created by a random walk) are distinguishable from similar sized subgraphs in a perturbed copy of the network. We conduct an extensive study of intrinsic scale for several networks, ranging fro…
A new clustering algorithm considers data smoothness for better performance.
Dissertation tackles zero-shot anomaly detection, focusing on consistent anomalies and proposing CoDeGraph framework.
A new topology design improves zero-shot classification performance in contrastive learning.
Study uses trajectory embedding to measure place function similarity at fine spatial granularity.
There are plenty of problems where the data available is scarce and expensive. We propose a generator of semi-artificial data with similar properties to the original data which enables development and testing of different data mining algorithms and optimization of their parameters. The generated data allow a large scal…
There is a large body of work, built on tools developed in mathematics and physics, demonstrating that financial market prices exhibit self-similarity at different scales. In this paper, we explore the use of analytical topology to characterize financial price series. While wavelet and Fourier transforms decompose a si…
SEMASIA provides a large dataset of latent representations for model comparison.
We study the problem of large scale, multi-label visual recognition with a large number of possible classes. We propose a method for augmenting a trained neural network classifier with auxiliary capacity in a manner designed to significantly improve upon an already well-performing model, while minimally impacting its c…
BiLRP explains deep similarity models by decomposing scores into feature contributions.
New clustering method using point-set kernel measures similarity.
We apply a novel spectral graph technique, that of locally-biased semi-supervised eigenvectors, to study the diversity of galaxies. This technique permits us to characterize empirically the natural variations in observed spectra data, and we illustrate how this approach can be used in an exploratory manner to highlight…
Solves a 60-year-old question on agreement measures in statistics.
We analyze the memory in volatility by studying volatility return intervals, defined as the time between two consecutive fluctuations larger than a given threshold, in time periods following stock market crashes. Such an aftercrash period is characterized by the Omori law, which describes the decay in the rate of after…
The concepts of scale invariance, self-similarity and scaling have been fruitfully applied to the study of price fluctuations in financial markets. After a brief review of the properties of stable Levy distributions and their applications to market data we indicate the shortcomings of such models and describe the trunc…
Study shows optimal model performance at critical level of feature learning.
Recently, randomly mapping vectorial data to strings of discrete symbols (i.e., sketches) for fast and space-efficient similarity searches has become popular. Such random mapping is called similarity-preserving hashing and approximates a similarity metric by using the Hamming distance. Although many efficient similarit…