LLMs can memorize economic data and recall exact values before their training cutoff.
problem Evaluating the trustworthiness of LLMs' economic forecasts during their training period.
method Demonstrated through counterfactual forecasting and analysis of LLMs' recall ability.
result LLMs have memorized economic and financial data, leading to recall-level accuracy before their knowledge cutoff.
New models avoid lookahead bias by training on past data only.
problem Lookahead bias in language models.
method Chronologically consistent training on data before a knowledge-cutoff date.
result Elimination of lookahead bias in predictions.
DatedGPT prevents lookahead bias in financial forecasting models.
problem Lookahead bias in large language models trained on internet-scale data.
method Time-aware pretraining with annual data cutoffs and instruction fine-tuning.
result Models' knowledge is effectively bounded by their data cutoff year, improving forecasting validity.
New cutoff phenomenon found for geodesic paths on hyperbolic manifolds.
problem Understanding the cutoff phenomenon for geodesic paths on hyperbolic manifolds.
method Spectral strategy and detailed spectral analysis of the spherical mean operator.
result Geodesic paths on compact hyperbolic manifolds exhibit cutoff for spatially localized initial conditions.
New transport method simplifies cutoff phenomenon for Markov processes.
problem Understanding the cutoff phenomenon for Markov processes.
method A new W-TV transport inequality combined with a parabolic regularization estimate.
result Recovery and extension of previous results on cutoff phenomena.
We introduce a new, high-throughput, synchronous, distributed, data-parallel, stochastic-gradient-descent learning algorithm. This algorithm uses amortized inference in a compute-cluster-specific, deep, generative, dynamical model to perform joint posterior predictive inference of the mini-batch gradient computation ti…
In this paper we study the common distance between points and the behavior of a constant length step discrete random walk on finite area hyperbolic surfaces. We show that if the second smallest eigenvalue of the Laplacian is at least 1/4, then the distances on the surface are highly concentrated around the minimal poss…
High-dimensional curved diffusions show abrupt convergence at a critical time.
problem Understanding abrupt convergence in high-dimensional curved diffusions.
method Functional inequalities and spectral rigidity.
result Abrupt convergence (cutoff) occurs in high dimensions, linked to spectral rigidity.
This paper proposes BAT to balance accuracy and robustness in adversarial training.
problem Balancing accuracy and robustness in adversarial training models.
method Blind adversarial training (BAT) uses a cutoff-scale strategy to adaptively estimate a nonuniform budget for AEs.
result BAT improves the overall robustness of adversarial training models.
Non-negative curvature affects Markov chains' mixing and expansion properties.
problem Understanding the behavior of Markov chains with non-negative curvature.
method Analyzing conductance, displacement, and cutoff phenomenon in sparse Markov chains.
result Non-negatively curved Markov chains exhibit specific, non-standard behavior in terms of mixing and expansion.
We determine the information-theoretic cutoff value on separation of cluster centers for exact recovery of cluster labels in a K-component Gaussian mixture model with equal cluster sizes. Moreover, we show that a semidefinite programming (SDP) relaxation of the K-means clustering method achieves such sharp threshol…
We detect lookahead bias in LLM forecasts using a novel statistical method.
problem Detecting lookahead bias in LLM-generated economic forecasts.
method Developed a statistical procedure using date-only recall queries and estimated Lookahead Propensity (LAP).
result LLM forecasts are contaminated with lookahead bias, as indicated by a positive interaction between LAP and the forecast in accuracy regressions.
We show that Chern-Simons gauge theory with appropriate cutoffs is equivalent, term by term in perturbation theory, to a Fermionic theory with a nonlocal interaction term. When an additional cutoff is placed on the Fermi fields, this Fermionic theory gives rise to a convergent perturbation expansion. This leads us to c…
In this paper we revisit the hypothesis needed to define the "paracomposition" operator, an analogue to the classic pull-back operation in the low regularity setting, first introduced by S. Alinhac in [3]. More precisely we do so in two directions. First we drop the diffeomorphism hypothesis. Secondly we give estimates…
One of the main obstacles regarding Barky Emery curvature on graphs is that the results require a global uniform lower curvature bounds where no exception sets are allowed. We overcome this obstacle by introducing the perpetual cutoff method. As applications, we prove gradient estimates only requiring curvature bounds …
Develops a tool to identify abnormal blood smear results based on CBC tests.
problem Manual review of blood smears by technologists is time-consuming and inconsistent.
method Cost-sensitive Lasso-penalized additive logistic regression combined with stability selection.
result The tool correctly identifies true cutoff values for abnormal smear results.
ChatGPT predicts stock market reactions from news headlines without financial training.
problem Predicting stock price movements using non-financial data.
method Used post-knowledge-cutoff headlines to train ChatGPT-4, which forecasts stock market reactions.
result ChatGPT-4 can predict stock market reactions with high accuracy, especially for small stocks and negative news.
Large language models improve futures market factor models in China.
problem Designing effective factor models for Chinese futures markets.
method Used large language models (GPT) to generate 40 factors for single and multi-factor portfolios.
result GPT-generated factors outperform benchmarks with high Sharpe ratios and alphas.
A new calibration metric bridges testability and actionability.
problem Combining testability and actionable insights for forecast probabilities.
method Cutoff Calibration Error (CCE) that assesses calibration over intervals of forecasted probabilities.
result Cutoff Calibration Error is both testable and actionable.
Bayesian model estimates treatment effects near cutoffs in regression discontinuity designs.
problem Estimating conditional average treatment effects in regression discontinuity designs.
method Develops a Bayesian additive regression tree (BART) model with linear leaf-level regressions.
result Adapts to different slopes on the running variable near the cutoff, providing interpretable inference.
Smooth KANs improve model reliability in computational biomedicine.
problem Limited convergence of KANs in representing generic smooth functions.
method Introducing smooth, structurally informed KANs that can approximate MLPs in specific function classes.
result Smooth KANs can achieve equivalence to MLPs in specific function classes, enhancing model reliability and performance.
The paper decomposes unsupervised learning's generalization error into model, data, and variance components.
problem Understanding the components of unsupervised learning's generalization error.
method Information-geometric decomposition of the Kullback-Leibler generalization error.
result The optimal rank in ε-PCA is the noise floor, balancing model-error gain and data-bias cost. Study challenges the necessity of data augmentation for improving predictions on imbalanced text datasets.
problem Improving predictions on imbalanced text datasets.
method Comparing classifier cutoff adjustments to data augmentation techniques.
result Classifier cutoff adjustments can produce similar results to data augmentation without the need for additional data.
New methods train neural networks without changing weights, achieving similar or higher performance.
problem Training neural networks efficiently with randomly initialized weights.
method Switching connections on and off, flipping weights' signs, minimizing changed connections.
result Achieves similar or higher performance with less computational cost than training all weights.
Blockchain markets with paid-priority trading can lead to biased prices and reduced liquidity.
problem Discrete clearing and paid-priority in blockchain markets lead to biased prices and reduced liquidity.
method Developed a model to evaluate the viability of blockchain markets under discrete clearing and paid-priority.
result Paid-priority ordering induces endogenous selection, leading to biased prices and reduced liquidity.
We show that the Yang-Mills quantum field theory with momentum and spacetime cutoffs in four Euclidean dimensions is equivalent, term by term in an appropriately resummed perturbation theory, to a Fermionic theory with nonlocal interaction terms. When a further momentum cutoff is imposed, this Fermionic theory has a co…
The so-called Pareto-Levy or power-law distribution has been successfully used as a model to describe probabilities associated to extreme variations of worldwide stock markets indexes data and it has the form Pr(X>x) x∗∗(−alpha)forgamma<x<infinity.Theselectionofthethresholdparametergamma from empirical d…
In this paper we tackle the problem of estimating the power-law tail exponent of income distributions by using the Hill's estimator. A subsample semi-parametric bootstrap procedure minimising the mean squared error is used to choose the power-law cutoff value optimally. This technique is applied to personal income data…
A new method for virtual drug screening detects top treatments.
problem Understanding model performance in virtual drug screening tasks.
method Regression Enrichment Surfaces (RES) method.
result RES detects more top-performing treatments than existing methods.
AI bias arises from human-defined goals, not algorithmic flaws.
problem AI bias due to human-defined goals in LLMs.
method Purpose-conditioned cognition and revealing downstream use of LLM outputs.
result AI bias can be reduced by purpose-aware prompting but not fully by regularization.
Optimal cutoff interval for risk scores improves binary classification accuracy.
problem Improving binary classification accuracy with abstention.
method Determines optimal cutoff interval for risk scores, refraining from decisions outside this interval.
result Minimizes classification margin and maximizes accuracy within the interval.
Graph neural networks fail to distinguish certain 3D atom configurations.
problem Graph neural networks (GNN) fail to distinguish certain 3D atom configurations.
method Construction of degenerate 3D atom configurations that are indistinguishable by first-order GNNs.
result First-order GNNs are incomplete for 3D atom configurations.
New method calibrates uncertainty estimates for image classifiers without labeled data.
problem Uncertainty estimates for modern classifiers are unreliable without labeled calibration data.
method Calibrates uncertainty estimates using unlabeled examples for distribution shifts.
result Proposes a method that provides excellent uncertainty estimates under natural distribution shifts.
This work proposes an efficient autoregressive model for text generation.
problem The challenge of generating high-quality text with autoregressive models.
method Introduces a cascaded decoding approach using Markov transformers to achieve sub-linear parallel time generation.
result Shows competitive accuracy/speed tradeoff compared to existing methods on five machine translation datasets.
ChatGPT snapshots predict future stock returns.
problem Predicting future stock returns using pre-cutoff text.
method Extracted LLM outlook scores from OpenAI snapshots.
result Outlook scores positively correlate with future stock returns.
The CGMY model's ATM call-price asymptotics are derived using characteristic function.
problem Deriving short-time asymptotics for the CGMY model's ATM call prices.
method Using the characteristic function, derived short-time asymptotics for the CGMY model's ATM call prices. Extracted higher-order coefficients by dynamic cutoff partitioning.
result Higher-order coefficients are derived for the CGMY model's ATM call prices.
We prove precompactness in an orbifold Cheeger-Gromov sense of complete gradient Ricci shrinkers with a lower bound on their entropy and a local integral Riemann bound. We do not need any pointwise curvature assumptions, volume or diameter bounds. In dimension four, under a technical assumption, we can replace the loca…
In this paper, by employ the cutoff function and the maximum principle, some Hamilton-Souplet-Zhang type gradient estimates for porous medium type equation are deduced. As a special case, an Hamilton-Souplet-Zhang type gradient estimates of the heat equation is derived which is different from the result of Souplet-Zhan…
Proposes a method to detect anomalies in financial time series using PCA and neural networks.
problem Anomalies in financial time series lead to miscalibrated risk models.
method Extract features using PCA, define anomaly score with neural network, calibrate cutoff value.
result The proposed PCA NN approach outperforms other anomaly detection methods.
The trimming scheme with a prefixed cutoff portion is known as a method of improving the robustness of statistical models such as multivariate Gaussian mixture models (MG- MMs) in small scale tests by alleviating the impacts of outliers. However, when this method is applied to real- world data, such as noisy speech pro…
Gradient descent optimally trains RNNs without overparameterization.
problem Training recurrent neural networks (RNNs) with gradient descent.
method Nonasymptotic analysis of gradient descent for RNNs with diagonal weight matrices.
result Gradient descent can achieve optimality in RNNs with a network size scaling logarithmically with the number of samples.
We investigate Inverse Mean Curvature Flow (IMCF) of non-compact hypersurfaces in hyperbolic space. Specifically, we look at bounded graphs over horospheres in Hn+1 and show long time existence of the flow. Along the way many important local estimates as well as global estimates are obtained. In addition,…
New method corrects biased predictions and uncertainty estimates in classification with nuisance parameters.
problem Tackles biased predictions and invalid uncertainty estimates in classification with nuisance parameters.
method Proposes a method that estimates ROC across the entire nuisance parameter space to devise invariant cutoffs.
result Demonstrates effective domain adaptation and valid prediction sets with high power.
Some general features of kinetic multi-agent models are reviewed, with particular attention to the relation between the agent saving propensities and the form of the equilibrium wealth distribution. The effect of a finite cutoff of the saving propensity distribution on the corresponding wealth distribution is studied. …
The study examines machine learning classification algorithms and their generalizability using Framingham Heart Study data.
problem Addressing biases and generalizability issues in machine learning classification algorithms.
method Comparison of eight machine learning classification algorithms on Framingham Heart Study data.
result Double discriminant scoring of type I is the most generalizable algorithm.
Study shows consistency of shallow GCNNs on sampled point clouds under manifold assumption.
problem Consistency of shallow GCNNs on sampled point clouds under manifold assumption.
method Functional analysis perspective, weakly compact product of unit balls, Sobolev regularity, frequency cutoff.
result Proves Γ-convergence of regularized empirical risk minimization functionals and convergence of their global minimizers. The dynamics of generalized Lotka-Volterra systems is studied by theoretical techniques and computer simulations. These systems describe the time evolution of the wealth distribution of individuals in a society, as well as of the market values of firms in the stock market. The individual wealths or market values are gi…
Large vessel occlusion (LVO) plays an important role in the diagnosis of acute ischemic stroke. Identifying LVO of patients in the early stage on admission would significantly lower the probabilities of suffering from severe effects due to stroke or even save their lives. In this paper, we utilized both structural and …