Data analysis in high-dimensional spaces aims at obtaining a synthetic description of a data set, revealing its main structure and its salient features. We here introduce an approach providing this description in the form of a topography of the data, namely a human-readable chart of the probability density from which t…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new clustering algorithm reduces density peaks clustering's computational complexity.
As one type of efficient unsupervised learning methods, clustering algorithms have been widely used in data mining and knowledge discovery with noticeable advantages. However, clustering algorithms based on density peak have limited clustering effect on data with varying density distribution (VDD), equilibrium distribu…
Hierarchical nucleation patterns emerge in deep neural network layers.
Most density-based clustering methods largely rely on how well the underlying density is estimated. However, density estimation itself is also a challenging problem, especially the determination of the kernel bandwidth. A large bandwidth could lead to the over-smoothed density estimation in which the number of density …
This paper focuses on density-based clustering, particularly the Density Peak (DP) algorithm and the one based on density-connectivity DBSCAN; and proposes a new method which takes advantage of the individual strengths of these two methods to yield a density-based hierarchical clustering algorithm. Our investigation be…
New clustering algorithm for mixed data improves applicability and efficiency.
We propose a Fourier-based approach for optimization of several clustering algorithms. Mathematically, clusters data can be described by a density function represented by the Dirac mixture distribution. The density function can be smoothed by applying the Fourier transform and a Gaussian filter. The determination of th…
The heuristic identification of peaks from noisy complex spectra often leads to misunderstanding of the physical and chemical properties of matter. In this paper, we propose a framework based on Bayesian inference, which enables us to separate multipeak spectra into single peaks statistically and consists of two steps.…
Equity auctions show linear price impact up to a large volume, then non-linear.
New clustering method using point-set kernel measures similarity.
Power spectrum densities for the number of tick quotes per minute (market activity) on three currency markets (USD/JPY, EUR/USD, and JPY/EUR) for periods from January 1999 to December 2000 are analyzed. We find some peaks on the power spectrum densities at a few minutes. We develop the double-threshold agent model and …
Study examines local extrema and crossing statistics in financial markets.
Paper proposes bypassing implicit assumption in GM-based AD methods.
I report a new statistical distribution formulated to confront the infamous, long-standing, computational/modeling challenge presented by highly skewed and/or leptokurtic ("fat- or heavy-tailed") data. The distribution is straightforward, flexible and effective. Even when working with far fewer data points than are rou…
A recent proposal of data dependent similarity called Isolation Kernel/Similarity has enabled SVM to produce better classification accuracy. We identify shortcomings of using a tree method to implement Isolation Similarity; and propose a nearest neighbour method instead. We formally prove the characteristic of Isolatio…
This study investigates that a characteristic time scale on an exchange rate market (USD/JPY) is examined for the period of 1998 to 2000. Calculating power spectrum densities for the number of tick quotes per minute and averaging them over the year yield that the mean power spectrum density has a peak at high frequenci…
We analyze the data on personal income distribution from the Australian Bureau of Statistics. We compare fits of the data to the exponential, log-normal, and gamma distributions. The exponential function gives a good (albeit not perfect) description of 98% of the population in the lower part of the distribution. The lo…
A robust method for decomposing spectral peaks robust to distortion and interference.
SPADE improves demand forecasting accuracy by 4.5% for post-promotion periods.
We find empirically a characteristic sharp peak-flat trough pattern in a large set of commodity prices. We argue that the sharp peak structure reflects an endogenous inter-market organization, and that peaks may be seen as local ``singularities'' resulting from imitation and herding. These findings impose a novel strin…
New matrix ensembles better match deep neural network spectral densities.
New method for estimating lead-lag times between non-synchronously observed point processes.
Joint peak detection is a central problem when comparing samples in genomic data analysis, but current algorithms for this task are unsupervised and limited to at most 2 sample types. We propose PeakSegJoint, a new constrained maximum likelihood segmentation model for any number of sample types. To select the number of…
Mass spectrometry (MS) is an important technique for chemical profiling which calculates for a sample a high dimensional histogram-like spectrum. A crucial step of MS data processing is the peak picking which selects peaks containing information about molecules with high concentrations which are of interest in an MS in…
Study on-chain peak shaving to reduce Ethereum transaction costs.
The paper explains two distinct peaks in generalization error for neural networks and simpler models, each governed by different factors.
We study the statistical regularities of opening call auction using the ultra-high-frequency data of 22 liquid stocks traded on the Shenzhen Stock Exchange in 2003. The distribution of the relative price, defined as the relative difference between the order price in opening call auction and the closing price of last tr…
In this letter, we propose a method for period estimation in light curves from periodic variable stars using correntropy. Light curves are astronomical time series of stellar brightness over time, and are characterized as being noisy and unevenly sampled. We propose to use slotted time lags in order to estimate corrent…
Bayesian framework integrates spectral deconvolution with expert reasoning for robust peak estimation.
FLOPART solves peak detection by creating accurate train and test set predictions.
The paper shows how the generalization curve can have multiple peaks, influenced by data and learning algorithm biases.
PEAKS selects key training examples incrementally based on prediction error and kernel similarity.
PEAK tests means of multiple data streams with sequential betting.
LLMs learn peaked distributions slowly due to power-law losses.
Finite-time queue peaks in stochastic networks have logarithmic scaling after geometric thresholds.
Bayesian Quadrature improves ensembling for neural networks with dispersed likelihood peaks.
During a stock market peak the price of a given stock () jumps from an initial level to a peak level before falling back to a bottom level . The ratios and are referred to as the peak- and bottom-amplitude respectively. The paper show…
In this paper, the fractional order curvature equation in is considered. Assuming has two critical points satisfying certain local conditions, we prove the existence of two-peak solutions.
A nonparametric method for time series analysis extracts envelopes, detects peaks, and clusters data.
Populations of species in ecosystems are often constrained by availability of resources within their environment. In effect this means that a growth of one population, needs to be balanced by comparable reduction in populations of others. In neutral models of biodiversity all populations are assumed to change increment…
Paper proposes a network framework for prosumers to manage peak loads in Iran.
Traditionally in regression one minimizes the number of fitting parameters or uses smoothing/regularization to trade training (TE) and generalization error (GE). Driving TE to zero by increasing fitting degrees of freedom (dof) is expected to increase GE. However modern big-data approaches, including deep nets, seem to…
The paper analyzes Kernel Density Estimation in high dimensions with varying data and dimensionality.
New algorithm tackles constrained Markov decision processes with peak constraints.
We win EVA2025 by estimating extreme precipitation events using Peaks Over Thresholds and martingale testing.
This paper explains why double descent sometimes occurs weakly or not at all from an optimization perspective.
Develops a new cluster validity index to find multiple optimal cluster numbers.