LLMs learn peaked distributions slowly due to power-law losses.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper proposes a network framework for prosumers to manage peak loads in Iran.
Paper proposes combining GAM and DNN for accurate peak demand estimation from lower-resolution data.
Populations of species in ecosystems are often constrained by availability of resources within their environment. In effect this means that a growth of one population, needs to be balanced by comparable reduction in populations of others. In neutral models of biodiversity all populations are assumed to change increment…
This article considers a model for alternative processes for securities prices and compares this model with actual return data of several securities. The distributions of returns that appear in the model can be Gaussian as well as non-Gaussian; in particular they may have two peaks. We consider a discrete Markov chain …
As one type of efficient unsupervised learning methods, clustering algorithms have been widely used in data mining and knowledge discovery with noticeable advantages. However, clustering algorithms based on density peak have limited clustering effect on data with varying density distribution (VDD), equilibrium distribu…
We win EVA2025 by estimating extreme precipitation events using Peaks Over Thresholds and martingale testing.
Financial time series typically exhibit strong fluctuations that cannot be described by a Gaussian distribution. In recent empirical studies of stock market indices it was examined whether the distribution P(r) of returns r(tau) after some time tau can be described by a (truncated) Levy-stable distribution L_{alpha}(r)…
Paper presents a method for identifying isotope envelopes in MALDI-ToF data.
Paper proposes GAS-ALD model for financial risk prediction.
This paper develops the Jungle model in a credit portfolio framework. The Jungle model is able to model credit contagion, produce doubly-peaked probability distributions for the total default loss and endogenously generate quasi phase transitions, potentially leading to systemic credit events which happen unexpectedly …
A robust method for decomposing spectral peaks robust to distortion and interference.
SPADE improves demand forecasting accuracy by 4.5% for post-promotion periods.
A new clustering algorithm reduces density peaks clustering's computational complexity.
Distributed, controllable energy storage devices offer several benefits to electric power system operation. Three such benefits include reducing peak load, providing standby power, and enhancing power quality. These benefits, however, are only realized during peak load or during an outage, events that are infrequent. T…
We find empirically a characteristic sharp peak-flat trough pattern in a large set of commodity prices. We argue that the sharp peak structure reflects an endogenous inter-market organization, and that peaks may be seen as local ``singularities'' resulting from imitation and herding. These findings impose a novel strin…
We investigate intra-day foreign exchange (FX) time series using the inverse statistic analysis developed in [1,2]. Specifically, we study the time-averaged distributions of waiting times needed to obtain a certain increase (decrease) in the price of an investment. The analysis is performed for the Deutsch mark (DM…
Joint peak detection is a central problem when comparing samples in genomic data analysis, but current algorithms for this task are unsupervised and limited to at most 2 sample types. We propose PeakSegJoint, a new constrained maximum likelihood segmentation model for any number of sample types. To select the number of…
Mass spectrometry (MS) is an important technique for chemical profiling which calculates for a sample a high dimensional histogram-like spectrum. A crucial step of MS data processing is the peak picking which selects peaks containing information about molecules with high concentrations which are of interest in an MS in…
Study on-chain peak shaving to reduce Ethereum transaction costs.
Quantile gradient boosted trees outperform other models in predicting NO2 concentration distributions.
The paper explains two distinct peaks in generalization error for neural networks and simpler models, each governed by different factors.
The heuristic identification of peaks from noisy complex spectra often leads to misunderstanding of the physical and chemical properties of matter. In this paper, we propose a framework based on Bayesian inference, which enables us to separate multipeak spectra into single peaks statistically and consists of two steps.…
Dual ML approach predicts peak temperatures in AFSD, improving process optimization.
Paper proposes bypassing implicit assumption in GM-based AD methods.
Bayesian framework integrates spectral deconvolution with expert reasoning for robust peak estimation.
FLOPART solves peak detection by creating accurate train and test set predictions.
Study on eigenvalue distribution of correlated time series deforming the semi-circle law.
The paper shows how the generalization curve can have multiple peaks, influenced by data and learning algorithm biases.
PEAKS selects key training examples incrementally based on prediction error and kernel similarity.
PEAK tests means of multiple data streams with sequential betting.
Data analysis in high-dimensional spaces aims at obtaining a synthetic description of a data set, revealing its main structure and its salient features. We here introduce an approach providing this description in the form of a topography of the data, namely a human-readable chart of the probability density from which t…
Finite-time queue peaks in stochastic networks have logarithmic scaling after geometric thresholds.
Bayesian Quadrature improves ensembling for neural networks with dispersed likelihood peaks.
During a stock market peak the price of a given stock () jumps from an initial level to a peak level before falling back to a bottom level . The ratios and are referred to as the peak- and bottom-amplitude respectively. The paper show…
In this paper, the fractional order curvature equation in is considered. Assuming has two critical points satisfying certain local conditions, we prove the existence of two-peak solutions.
A nonparametric method for time series analysis extracts envelopes, detects peaks, and clusters data.
Traditionally in regression one minimizes the number of fitting parameters or uses smoothing/regularization to trade training (TE) and generalization error (GE). Driving TE to zero by increasing fitting degrees of freedom (dof) is expected to increase GE. However modern big-data approaches, including deep nets, seem to…
Optimizes sampling in continuous domains by adjusting search distribution.
New algorithm tackles constrained Markov decision processes with peak constraints.
We study the statistical regularities of opening call auction using the ultra-high-frequency data of 22 liquid stocks traded on the Shenzhen Stock Exchange in 2003. The distribution of the relative price, defined as the relative difference between the order price in opening call auction and the closing price of last tr…
This paper explains why double descent sometimes occurs weakly or not at all from an optimization perspective.
We propose a deep-learning approach based on generative adversarial networks (GANs) to reduce noise in weak lensing mass maps under realistic conditions. We apply image-to-image translation using conditional GANs to the mass map obtained from the first-year data of Subaru Hyper Suprime-Cam (HSC) survey. We train the co…
Support Vector Data Description (SVDD) provides a useful approach to construct a description of multivariate data for single-class classification and outlier detection with various practical applications. Gaussian kernel used in SVDD formulation allows flexible data description defined by observations designated as sup…
We propose a Fourier-based approach for optimization of several clustering algorithms. Mathematically, clusters data can be described by a density function represented by the Dirac mixture distribution. The density function can be smoothed by applying the Fourier transform and a Gaussian filter. The determination of th…
I report a new statistical distribution formulated to confront the infamous, long-standing, computational/modeling challenge presented by highly skewed and/or leptokurtic ("fat- or heavy-tailed") data. The distribution is straightforward, flexible and effective. Even when working with far fewer data points than are rou…
In this paper, we introduce a new sparsity-promoting prior, namely, the "normal product" prior, and develop an efficient algorithm for sparse signal recovery under the Bayesian framework. The normal product distribution is the distribution of a product of two normally distributed variables with zero means and possibly …
This paper focuses on density-based clustering, particularly the Density Peak (DP) algorithm and the one based on density-connectivity DBSCAN; and proposes a new method which takes advantage of the individual strengths of these two methods to yield a density-based hierarchical clustering algorithm. Our investigation be…