Study of eigenvalues in nonlinear kernels for classification of separable data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper proves local laws for non-separable sample covariance matrices.
The nowadays massive amounts of generated and communicated data present major challenges in their processing. While capable of successfully classifying nonlinearly separable objects in various settings, subspace clustering (SC) methods incur prohibitively high computational complexity when processing large-scale data. …
The immense amount of daily generated and communicated data presents unique challenges in their processing. Clustering, the grouping of data without the presence of ground-truth labels, is an important tool for drawing inferences from data. Subspace clustering (SC) is a relatively recent method that is able to successf…
For many years, a combination of principal component analysis (PCA) and independent component analysis (ICA) has been used for blind source separation (BSS). However, it remains unclear why these linear methods work well with real-world data that involve nonlinear source mixtures. This work theoretically validates that…
Notwithstanding the popularity of conventional clustering algorithms such as K-means and probabilistic clustering, their clustering results are sensitive to the presence of outliers in the data. Even a few outliers can compromise the ability of these algorithms to identify meaningful hidden structures rendering their o…
In response to the need for learning tools tuned to big data analytics, the present paper introduces a framework for efficient clustering of huge sets of (possibly high-dimensional) data. Building on random sampling and consensus (RANSAC) ideas pursued earlier in a different (computer vision) context for robust regress…
Study left invariant spray geometry on Lie groups using parallel translations.
This paper focuses on scalability and robustness of spectral clustering for extremely large-scale datasets with limited resources. Two novel algorithms are proposed, namely, ultra-scalable spectral clustering (U-SPEC) and ultra-scalable ensemble clustering (U-SENC). In U-SPEC, a hybrid representative selection strategy…
New k-means method handles random data better than traditional techniques.
We find exact solutions describing Ricci flows of four dimensional pp-waves nonlinearly deformed by two/three dimensional solitons. Such solutions are parametrized by five dimensional metrics with generic off-diagonal terms and connections with nontrivial torsion which can be related, for instance, to antisymmetric ten…
Optimal transport learns Riemannian metrics for evolving probability measures.
Paper explores solving HJB equations using neural networks.
This work closes the gap between theory and practice for nICA identifiability.
TabPFN model shows strong robustness to noisy data.
CKA with Gaussian RBF kernels converges linearly as bandwidth increases.
The existing approaches to intrinsic dimension estimation usually are not reliable when the data are nonlinearly embedded in the high dimensional space. In this work, we show that the explicit accounting to geometric properties of unknown support leads to the polynomial correction to the standard maximum likelihood est…
We obtain a lower asymptotic bound on the decay rate of the probability of a portfolio's underperformance against a benchmark over a large time horizon. It is assumed that the prices of the securities are governed by geometric Brownian motions with the coefficients depending on an economic factor, possibly nonlinearly.…
Market illiquidity, feedback effects, presence of transaction costs, risk from unprotected portfolio and other nonlinear effects in PDE based option pricing models can be described by solutions to the generalized Black-Scholes parabolic equation with a diffusion term nonlinearly depending on the option price itself. Di…
We analyze analytic approximation formulae for pricing zero-coupon bonds in the case when the short-term interest rate is driven by a one-factor mean-reverting process with a volatility nonlinearly depending on the interest rate itself. We derive the order of accuracy of the analytical approximation due to Choi and Wir…
New method identifies latent components in PNL mixtures without strong assumptions.
Slow feature analysis (SFA) is a method for extracting slowly varying features from a quickly varying multidimensional signal. An open source Matlab-implementation sfa-tk makes SFA easily useable. We show here that under certain circumstances, namely when the covariance matrix of the nonlinearly expanded data does not …
SSINNs learn Hamiltonian systems from data with interpretable, low-memory models.
Canonical correlation analysis (CCA) is a technique to find statistical dependencies between a pair of multivariate data. However, its application to high dimensional data is limited due to the resulting time complexity. While the conventional CCA algorithm requires polynomial time, we have developed an algorithm that …
This paper presents a geometric description on Lie algebroids of Lagrangian systems subject to nonholonomic constraints. The Lie algebroid framework provides a natural generalization of classical tangent bundle geometry. We define the notion of nonholonomically constrained system, and characterize regularity conditions…
Develops CLDS models to model neural activity with nonlinear dynamics.
We develop estimates for the solutions and derive existence and uniqueness results of various local boundary value problems for Dirac equations that improve all relevant results known in the literature. With these estimates at hand, we derive a general existence, uniqueness and regularity theorem for solutions of Dirac…
Optimizes prediction error method for time-varying models.
Recently, deep learning becomes the main focus of machine learning research and has greatly impacted many important fields. However, deep learning is criticized for lack of interpretability. As a successful unsupervised model in deep learning, the autoencoder embraces a wide spectrum of applications, yet it suffers fro…
This paper examines the short-run relationships between oil prices and GCC stock markets. Since GCC countries are major world energy market players, their stock markets may be susceptible to oil price shocks. To account for the fact that stock markets may respond nonlinearly to oil price shocks, we have examined both l…
Neural networks can separate non-separable data using feature maps.
NCPF model improves traffic data imputation with neural and tensor methods.
New covariance estimator for financial portfolios.
DSI measures dataset separability for neural networks.
Can multilayer neural networks -- typically constructed as highly complex structures with many nonlinearly activated neurons across layers -- behave in a non-trivial way that yet simplifies away a major part of their complexities? In this work, we uncover a phenomenon in which the behavior of these complex networks -- …
Stable blowup profile identified for wave maps in all dimensions.
Expected centre of mass for random embeddings is constant.
This paper investigates how data augmentation improves linear separation of manifold data.
In this paper, we consider the problem of optimal investment by an insurer. The insurer invests in a market consisting of a bank account and risky assets. The mean returns and volatilities of the risky assets depend nonlinearly on economic factors that are formulated as the solutions of general stochastic different…
A new measure DCSI quantifies separability for density-based clustering.
In the last decades the estimation of the intrinsic dimensionality of a dataset has gained considerable importance. Despite the great deal of research work devoted to this task, most of the proposed solutions prove to be unreliable when the intrinsic dimensionality of the input dataset is high and the manifold where th…
Improved clustering of extra-financial data using NMF with data separation.
The paper solves optimal bounds for separating data points in high dimensions.
Scattering networks maximize separation on low-dimensional data.
Among econophysics investigations, studies of religious groups have been of interest. On one hand, the present paper concerns the Antoinist community financial reports, - a community which appeared at the end of the 19-th century in Belgium. Several growth-decay regimes have been previously found over different time sp…
This work addresses the problem of learning sparse representations of tensor data using structured dictionary learning. It proposes learning a mixture of separable dictionaries to better capture the structure of tensor data by generalizing the separable dictionary learning model. Two different approaches for learning m…
Bayesian approach uses generative models as priors for better source separation.
In this paper, the performance of three deep learning methods for predicting short-term evolution and for reproducing the long-term statistics of a multi-scale spatio-temporal Lorenz 96 system is examined. The methods are: echo state network (a type of reservoir computing, RC-ESN), deep feed-forward artificial neural n…