Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

11.6%23.2%34.8%46.3% · Jun 202019922001200920172026
48 results for log data correlation

Dividing deep learning models for consistent anomaly detection in changing log data.

problem Anomaly detection methods fail when log data types change, leading to false negatives.
method Divide deep learning models based on log data correlation and extract correlations.
result Continues anomaly detection accuracy even when log data changes.

Study shows Merton model limits to Poisson process with log-normal intensity, improving default portfolio prediction.

problem Improving prediction of default portfolios using complex models.
method Applying Merton model with log-normal intensity function to Poisson process, discussing temporal correlation effects.
result Power decay model provides better generalization for long-term default portfolio data.

The paper assesses dimensionality reduction for cryptocurrency link prediction.

problem Establishing a link between cryptocurrencies using dimensionality reduction techniques.
method Used canonical correlation analysis and principal component analysis on log returns and covariates of Bitcoin and Ethereum.
result Performance of dimensionality reduction techniques in forecasting Ethereum returns with Bitcoin features.

We simplify matrix computations for block matrices, especially useful for covariance and correlation matrices.

problem Complex computations for block matrices, especially for covariance and correlation matrices.
method Obtained a canonical representation for block matrices, facilitating computation of various matrix operations.
result Simplified computation of matrix operations for block matrices, particularly useful for covariance and correlation matrices.

Efficiently matches random graphs with inhomogeneous edge probabilities.

problem Matching latent vertex correspondence between two correlated random graphs with inhomogeneous edge probabilities.
method Inspired by Ding et al. (2021), an efficient matching algorithm is developed with conditions on minimal average degree and minimal correlation.
result An efficient matching algorithm is obtained as long as the minimal average degree is at least Ω(log2n)Ω(\log^{2} n) and the minimal correlation is at least 1O(log2n)1 - O(\log^{-2} n).

In this paper, we investigate a new compressive sensing model for multi-channel sparse data where each channel can be represented as a hierarchical tree and different channels are highly correlated. Therefore, the full data could follow the forest structure and we call this property as \emph{forest sparsity}. It exploi…

2012-11-20abs ↗pdf ↗

Botnet, a group of coordinated bots, is becoming the main platform of malicious Internet activities like DDOS, click fraud, web scraping, spam/rumor distribution, etc. This paper focuses on design and experiment of a new approach for botnet detection from streaming web server logs, motivated by its wide applicability, …

2018-12-19abs ↗pdf ↗

New research challenges the flatness-generalization link in deep neural networks.

problem The correlation between flatness of the loss landscape and generalization in deep neural networks is questioned.
method The study examines various flatness measures and popular SGD variants, finding some break the flatness-generalization link. It proposes using logP(f)\log P(f), a global quantity, as a predictor of generalization.
result The log of Bayesian prior upon initialization, logP(f)\log P(f), is a significantly more robust predictor of generalization than flatness measures.

Gradient flow of ReLU networks converges in low-correlation high-dimensional data.

problem Convergence of shallow ReLU networks trained on weakly interacting data.
method Gradient flow analysis with Polyak-Łojasiewicz viewpoint.
result Network width of order log(n) neurons suffices for global convergence with high probability.

Algorithm learns stock correlation matrix embedding using graph machine learning.

problem Understanding complex relationships among stocks based on their correlation matrix.
method Proposes a graph machine learning approach called Node2Vec to compress the correlation network into an embedding.
result The algorithm can learn an embedding from the correlation network of S&P 500 stock data.

The study examines volatility models and finds decoupling of short- and long-term correlation structures.

problem Understanding the dynamic of volatility at different time scales.
method Developed a composite likelihood estimation framework for parametric continuous-time stationary Gaussian processes.
result The short- and long-term correlation structures of stochastic volatility are decoupled.

The paper proposes new cross-correlators using Price's Theorem and piecewise-linear decomposition.

problem Optimal method for estimating cross-correlations using finite samples.
method General mathematical framework using Price's Theorem and piecewise-linear decomposition.
result Some cross-correlators based on Huber's loss functions, MP functions, and LSE functions have higher SNR.

We present sharp tail asymptotics for the density and the distribution function of linear combinations of correlated log-normal random variables, that is, exponentials of components of a correlated Gaussian vector. The asymptotic behavior turns out to depend on the correlation between the components, and the explicit s…

2013-09-12abs ↗pdf ↗

Algorithm recovers permutations of high-dimensional Gaussian vectors with constant correlation.

problem Recovering permutations of high-dimensional Gaussian vectors with constant correlation.
method Computing and comparing weighted counts of specially chosen wide trees.
result Polynomial-time algorithm for exact recovery at constant correlation.

Correlated anomaly detection (CAD) from streaming data is a type of group anomaly detection and an essential task in useful real-time data mining applications like botnet detection, financial event detection, industrial process monitor, etc. The primary approach for this type of detection in previous researches is base…

2018-12-19abs ↗pdf ↗

Estimates Hurst exponent of log-volatility using KS statistic, addressing serial correlation in financial data.

problem Estimating Hurst exponent of log-volatility in financial time series with serial correlation.
method Proposes a random permutation procedure to remove serial correlation, using the Kolmogorov-Smirnov statistic for distribution-based estimation.
result Establishes the asymptotic variance of the estimator and reveals statistically significant hierarchy of roughness in volatility measures.

This paper compares log-likelihood and BLEU scores for sequence generation tasks.

problem The discrepancy between density estimation and sequence generation performance.
method Comparing several density estimators on five machine translation tasks.
result The correlation between log-likelihood and BLEU varies depending on model families.

Temporal aggregation reveals latent default correlation from monthly data.

problem Understanding effective default correlation from monthly default data.
method Temporal coarse-graining of latent default-probability paths.
result Temporal coarse-graining improves identifiability and reduces over-allocation of long-horizon fluctuations.

ResNets approximate log-Gaussian at initialization, improving network performance.

problem Understanding the initialization behavior of deep neural networks like ResNets.
method Analyzing ReLU ResNets in the infinite-depth-and-width limit, showing log-Gaussian behavior.
result ResNets at initialization exhibit hypoactivation and interlayer correlations, which are not captured by Gaussian limits.

Proposes integrating random effects into deep neural networks for better predictive performance.

problem Correlated data in real-life applications are not handled well by traditional deep neural networks.
method Uses mixed models with random effects to handle correlations in deep neural networks, minimizing Gaussian negative log-likelihood with SGD.
result Improves predictive performance over natural competitors in various correlation scenarios.

We consider the problem of clustering a set of high-dimensional data points into sets of low-dimensional linear subspaces. The number of subspaces, their dimensions, and their orientations are unknown. We propose a simple and low-complexity clustering algorithm based on thresholding the correlations between the data po…

2013-03-15abs ↗pdf ↗

We conclude from an analysis of high resolution NYSE data that the distribution of the traded value fif_i (or volume) has a finite variance σiσ_i for the very large majority of stocks ii, and the distribution itself is non-universal across stocks. The Hurst exponent of the same time series displays a crossover from we…

2006-08-02abs ↗pdf ↗

We address the curse of dimensionality in dynamic covariance estimation by modeling the underlying co-volatility dynamics of a time series vector through latent time-varying stochastic factors. The use of a global-local shrinkage prior for the elements of the factor loadings matrix pulls loadings on superfluous factors…

2016-08-30abs ↗pdf ↗

The inference of correlated signal fields with unknown correlation structures is of high scientific and technological relevance, but poses significant conceptual and numerical challenges. To address these, we develop the correlated signal inference (CSI) algorithm within information field theory (IFT) and discuss its n…

2016-12-26abs ↗pdf ↗

Temporal coarse-graining of latent default paths explains effective correlation in corporate defaults.

problem Understanding effective default correlation in corporate defaults.
method Temporal coarse-graining of latent default-probability paths, applied to corporate default-count data.
result Temporal coarse-graining provides a scale-consistent baseline that improves identifiability and reduces over-allocation of long-horizon fluctuations.

With the recent popularity of graphical clustering methods, there has been an increased focus on the information between samples. We show how learning cluster structure using edge features naturally and simultaneously determines the most likely number of clusters and addresses data scale issues. These results are parti…

2016-05-05abs ↗pdf ↗

The paper presents a novel approach to multi-output regression using probabilistic circuits.

problem Capturing correlations between multiple output dimensions in large-scale regression problems.
method Employing a mixture of single-output Gaussian process experts encoded via a probabilistic circuit.
result The method can capture correlations between output dimensions and often outperforms other approaches.

We study the volatility of the MIB30-stock-index high-frequency data from November 28, 1994 through September 15, 1995. Our aim is to empirically characterize the volatility random walk in the framework of continuous-time finance. To this end, we compute the index volatility by means of the log-return standard deviatio…

1999-03-14abs ↗pdf ↗

We generalize the log Gaussian Cox process (LGCP) framework to model multiple correlated point data jointly. The observations are treated as realizations of multiple LGCPs, whose log intensities are given by linear combinations of latent functions drawn from Gaussian process priors. The combination coefficients are als…

2018-05-24abs ↗pdf ↗

We study the sample complexity of canonical correlation analysis (CCA), \ie, the number of samples needed to estimate the population canonical correlation and directions up to arbitrarily small error. With mild assumptions on the data distribution, we show that in order to achieve εε-suboptimality in a properly define…

2017-02-21abs ↗pdf ↗

We consider a multi-armed bandit framework where the rewards obtained by pulling different arms are correlated. We develop a unified approach to leverage these reward correlations and present fundamental generalizations of classic bandit algorithms to the correlated setting. We present a unified proof technique to anal…

2019-11-06abs ↗pdf ↗

Bayesian inference and superstatistics model financial volatility dynamics across different timescales.

problem Modeling correlated volatility in financial time series with heavy tails and long memory.
method Superstatistical dynamics, Bayesian Inference, Metropolis-Hasting sampling.
result The log-Normal model is reliable for short timescales, while inverse-Gamma is preferred for long timescales.

A Bayesian procedure is developed for multivariate stochastic volatility, using state space models. An autoregressive model for the log-returns is employed. We generalize the inverted Wishart distribution to allow for different correlation structure between the observation and state innovation vectors and we extend the…

2008-02-01abs ↗pdf ↗

The aim of our work is to propose a natural framework to account for all the empirically known properties of the multivariate distribution of stock returns. We define and study a "nested factor model", where the linear factors part is standard, but where the log-volatility of the linear factors and of the residuals are…

2013-09-12abs ↗pdf ↗

Lead-lag relationships among assets represent a useful tool for analyzing high frequency financial data. However, research on these relationships predominantly focuses on correlation analyses for the dynamics of stock prices, spots and futures on market indexes, whereas foreign exchange data have been less explored. To…

2019-06-25abs ↗pdf ↗

Recent applications of machine learning algorithms in the seismic domain have shown great potential in different areas such as seismic inversion and interpretation. However, such algorithms rarely enforce geophysical constraints - the lack of which might lead to undesirable results. To overcome this issue, we have deve…

2019-08-19abs ↗pdf ↗