Optimistic estimate predicts best fitting performance of nonlinear models.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Meta-learners improve causal effect estimation in small samples.
We present the use of the fitted Q iteration in algorithmic trading. We show that the fitted Q iteration helps alleviate the dimension problem that the basic Q-learning algorithm faces in application to trading. Furthermore, we introduce a procedure including model fitting and data simulation to enrich training data as…
We propose a data aggregation-based algorithm with monotonic convergence to a global optimum for a generalized version of the L1-norm error fitting model with an assumption of the fitting function. The proposed algorithm generalizes the recent algorithm in the literature, aggregate and iterative disaggregate (AID), whi…
In quantitative finance, we often fit a parametric semimartingale model to asset prices. To ensure our model is correct, we must then perform goodness-of-fit tests. In this paper, we give a new goodness-of-fit test for volatility-like processes, which is easily applied to a variety of semimartingale models. In each cas…
This work improves AutoML systems by dynamically evaluating fitness to reduce overfitting.
We study a resource utilization scenario characterized by intrinsic fitness. To describe the growth and organization of different cities, we consider a model for resource utilization where many restaurants compete, as in a game, to attract customers using an iterative learning process. Results for the case of restauran…
Computational models in fields such as computational neuroscience are often evaluated via stochastic simulation or numerical approximation. Fitting these models implies a difficult optimization problem over complex, possibly noisy parameter landscapes. Bayesian optimization (BO) has been successfully applied to solving…
Smoothed fitness landscape improves protein optimization.
Improved SSD for faster and more accurate goodness-of-fit tests and model learning.
To find efficient screening methods for high dimensional linear regression models, this paper studies the relationship between model fitting and screening performance. Under a sparsity assumption, we show that a subset that includes the true submodel always yields smaller residual sum of squares (i.e., has better model…
HAR model outperforms ML in stock forecasting with correct fitting schemes.
New spectral tests assess network model fits efficiently.
Examines performance metrics for ML models in engineering.
Study uses copulas and DCC-GARCH for multivariate risk analysis of VaR and CVaR.
New method extends fitted Q-evaluation for distributional off-policy reinforcement learning.
GRASP tests goodness-of-fit for binary classifiers without parametric assumptions.
A new kernel Stein test assesses fit for variable-length sequential data.
Interactive visualization helps understand complex machine learning models.
Paper explores how text generation quality and diversity metrics relate to distribution fitting.
Breakthroughs in machine learning are rapidly changing science and society, yet our fundamental understanding of this technology has lagged far behind. Indeed, one of the central tenets of the field, the bias-variance trade-off, appears to be at odds with the observed behavior of methods used in the modern machine lear…
The paper addresses over-fitting in deep learning models trained on imbalanced data.
Hawkes processes have seen a number of applications in finance, due to their ability to capture event clustering behaviour typically observed in financial systems. Given a calibrated Hawkes process, of concern is the statistical fit to empirical data, particularly for the accurate quantification of self- and mutual-exc…
New method uses neural networks to estimate parameters without needing detector simulations.
Theoretical study explains why federated optimization fails to achieve perfect fitting.
QLBS and RLOP methods improve option pricing and hedging performance.
The goal of chemmodlab is to streamline the fitting and assessment pipeline for many machine learning models in R, making it easy for researchers to compare the utility of new models. While focused on implementing methods for model fitting and assessment that have been accepted by experts in the cheminformatics field, …
The paper introduces a method for fitting complex models using simulation and optimization.
LLMs are vulnerable to task-irrelevant data changes, limiting their use for data fitting.
We present a new variable selection method based on model-based gradient boosting and randomly permuted variables. Model-based boosting is a tool to fit a statistical model while performing variable selection at the same time. A drawback of the fitting lies in the need of multiple model fits on slightly altered data (e…
The paper uses FRFT to fit GTS distribution to asset returns.
Recently we proposed a general, ensemble-based feature engineering wrapper (FEW) that was paired with a number of machine learning methods to solve regression problems. Here, we adapt FEW for supervised classification and perform a thorough analysis of fitness and survival methods within this framework. Our tests demon…
Many fits of Hawkes processes to financial data look rather good but most of them are not statistically significant. This raises the question of what part of market dynamics this model is able to account for exactly. We document the accuracy of such processes as one varies the time interval of calibration and compare t…
A new framework improves kernel Stein discrepancy tests for validating distributions.
In this paper, we apply machine learning to distributed private data owned by multiple data owners, entities with access to non-overlapping training datasets. We use noisy, differentially-private gradients to minimize the fitness cost of the machine learning model using stochastic gradient descent. We quantify the qual…
New model allocates features sublinearly, improving model fit and performance.
Two methods monitor high-dimensional processes via manifold fitting or learning.
We propose a novel method that makes use of deep neural networks and gradient decent to perform automated design on complex real world engineering tasks. Our approach works by training a neural network to mimic the fitness function of a design optimization task and then, using the differential nature of the neural netw…
WeakNAS uses a set of weaker predictors to find top architectures with fewer samples.
Proposes a method to quantify uncertainty in PFNs.
New NMF algorithms improve topic model fits.
Deep nonlinear models pose a challenge for fitting parameters due to lack of knowledge of the hidden layer and the potentially non-affine relation of the initial and observed layers. In the present work we investigate the use of information theoretic measures such as mutual information and Kullback-Leibler (KL) diverge…
GAMI-Tree uses model-based trees to fit low-order fANOVA models.
Develops a goodness-of-fit test for self-exciting processes.
In this paper, a statistical analysis of log-return fluctuations of the IPC, the Mexican Stock Market Index is presented. A sample of daily data covering the period from was analyzed, and fitted to different distributions. Tests of the goodness of fit were performed in order to quantitatively as…
Automates fitting semiconductor device models using approximate Bayesian computation.
We propose a novel adaptive test of goodness-of-fit, with computational cost linear in the number of samples. We learn the test features that best indicate the differences between observed samples and a reference model, by minimizing the false negative rate. These features are constructed via Stein's method, meaning th…
Proposes a method to forecast spatial-temporal data with limited training data.