LLMs can memorize economic data and recall exact values before their training cutoff.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We determine the information-theoretic cutoff value on separation of cluster centers for exact recovery of cluster labels in a -component Gaussian mixture model with equal cluster sizes. Moreover, we show that a semidefinite programming (SDP) relaxation of the -means clustering method achieves such sharp threshol…
In this paper we study the common distance between points and the behavior of a constant length step discrete random walk on finite area hyperbolic surfaces. We show that if the second smallest eigenvalue of the Laplacian is at least 1/4, then the distances on the surface are highly concentrated around the minimal poss…
Develops a tool to identify abnormal blood smear results based on CBC tests.
New cutoff phenomenon found for geodesic paths on hyperbolic manifolds.
New transport method simplifies cutoff phenomenon for Markov processes.
In this paper we tackle the problem of estimating the power-law tail exponent of income distributions by using the Hill's estimator. A subsample semi-parametric bootstrap procedure minimising the mean squared error is used to choose the power-law cutoff value optimally. This technique is applied to personal income data…
Proposes a method to detect anomalies in financial time series using PCA and neural networks.
High-dimensional curved diffusions show abrupt convergence at a critical time.
Non-negative curvature affects Markov chains' mixing and expansion properties.
Some general features of kinetic multi-agent models are reviewed, with particular attention to the relation between the agent saving propensities and the form of the equilibrium wealth distribution. The effect of a finite cutoff of the saving propensity distribution on the corresponding wealth distribution is studied. …
A low rank matrix X has been contaminated by uniformly distributed noise, missing values, outliers and corrupt entries. Reconstruction of X from the singular values and singular vectors of the contaminated matrix Y is a key problem in machine learning, computer vision and data science. In this paper we show that common…
The so-called Pareto-Levy or power-law distribution has been successfully used as a model to describe probabilities associated to extreme variations of worldwide stock markets indexes data and it has the form from empirical d…
The dynamics of generalized Lotka-Volterra systems is studied by theoretical techniques and computer simulations. These systems describe the time evolution of the wealth distribution of individuals in a society, as well as of the market values of firms in the stock market. The individual wealths or market values are gi…
We show that Chern-Simons gauge theory with appropriate cutoffs is equivalent, term by term in perturbation theory, to a Fermionic theory with a nonlocal interaction term. When an additional cutoff is placed on the Fermi fields, this Fermionic theory gives rise to a convergent perturbation expansion. This leads us to c…
One of the main obstacles regarding Barky Emery curvature on graphs is that the results require a global uniform lower curvature bounds where no exception sets are allowed. We overcome this obstacle by introducing the perpetual cutoff method. As applications, we prove gradient estimates only requiring curvature bounds …
Paper generalizes paracomposition and change of variables for paradifferential operators.
DatedGPT prevents lookahead bias in financial forecasting models.
We introduce a new, high-throughput, synchronous, distributed, data-parallel, stochastic-gradient-descent learning algorithm. This algorithm uses amortized inference in a compute-cluster-specific, deep, generative, dynamical model to perform joint posterior predictive inference of the mini-batch gradient computation ti…
New models avoid lookahead bias by training on past data only.
A new calibration metric bridges testability and actionability.
Bayesian model estimates treatment effects near cutoffs in regression discontinuity designs.
Study challenges the necessity of data augmentation for improving predictions on imbalanced text datasets.
The paper diagnoses factor models using characteristic axes and zero-curve restrictions.
The paper diagnoses factor-model pricing errors using characteristic axes and bridge-alpha curves.
Blockchain markets with paid-priority trading can lead to biased prices and reduced liquidity.
We show that the Yang-Mills quantum field theory with momentum and spacetime cutoffs in four Euclidean dimensions is equivalent, term by term in an appropriately resummed perturbation theory, to a Fermionic theory with nonlocal interaction terms. When a further momentum cutoff is imposed, this Fermionic theory has a co…
We detect lookahead bias in LLM forecasts using a novel statistical method.
An important class of economic models involve agents whose wealth changes due to transactions with other agents. Several authors have pointed out an analogy with kinetic theory, which describes molecules whose momentum and energy changes due to interactions with other molecules. We pursue this analogy and derive a Bolt…
AI bias arises from human-defined goals, not algorithmic flaws.
The paper is aware of the importance of certain figures that are essential to an understanding of Credit Scoring models in credit acceptance process optimization, namely if the power of discrimination measured by Gini value is increased by 5% then the profit of the process can be increased monthly by about 1 500 kPLN (…
Optimal cutoff interval for risk scores improves binary classification accuracy.
Graph neural networks fail to distinguish certain 3D atom configurations.
New method calibrates uncertainty estimates for image classifiers without labeled data.
ChatGPT snapshots predict future stock returns.
The CGMY model's ATM call-price asymptotics are derived using characteristic function.
We prove precompactness in an orbifold Cheeger-Gromov sense of complete gradient Ricci shrinkers with a lower bound on their entropy and a local integral Riemann bound. We do not need any pointwise curvature assumptions, volume or diameter bounds. In dimension four, under a technical assumption, we can replace the loca…
In this paper, by employ the cutoff function and the maximum principle, some Hamilton-Souplet-Zhang type gradient estimates for porous medium type equation are deduced. As a special case, an Hamilton-Souplet-Zhang type gradient estimates of the heat equation is derived which is different from the result of Souplet-Zhan…
Develops RES metrics for stable rare-event forecasting evaluation.
This paper proposes BAT to balance accuracy and robustness in adversarial training.
We investigate Inverse Mean Curvature Flow (IMCF) of non-compact hypersurfaces in hyperbolic space. Specifically, we look at bounded graphs over horospheres in and show long time existence of the flow. Along the way many important local estimates as well as global estimates are obtained. In addition,…
Econophysics and econometrics agree that there is a correlation between volume and volatility in a time series. Using empirical data and their distributions, we further investigate this correlation and discover new ways that volatility and volume interact, particularly when the levels of both are high. We find that the…
Binary classification rules based on covariates typically depend on simple loss functions such as zero-one misclassification. Some cases may require more complex loss functions. For example, individual-level monitoring of HIV-infected individuals on antiretroviral therapy (ART) requires periodic assessment of treatment…
New methods train neural networks without changing weights, achieving similar or higher performance.
The paper decomposes unsupervised learning's generalization error into model, data, and variance components.
Digital personas improve survey results for stable attributes but fail for subjective responses.
Low-rank modeling has a lot of important applications in machine learning, computer vision and social network analysis. While the matrix rank is often approximated by the convex nuclear norm, the use of nonconvex low-rank regularizers has demonstrated better recovery performance. However, the resultant optimization pro…
Low-rank modeling has many important applications in computer vision and machine learning. While the matrix rank is often approximated by the convex nuclear norm, the use of nonconvex low-rank regularizers has demonstrated better empirical performance. However, the resulting optimization problem is much more challengin…