CTEF fits ellipsoids to noisy data in any dimension.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The coefficient of determination, known as , is commonly used as a goodness-of-fit criterion for fitting linear models. is somewhat controversial when fitting nonlinear models, although it may be generalised on a case-by-case basis to deal with specific models such as the logistic model. Assume we are fittin…
We propose a data aggregation-based algorithm with monotonic convergence to a global optimum for a generalized version of the L1-norm error fitting model with an assumption of the fitting function. The proposed algorithm generalizes the recent algorithm in the literature, aggregate and iterative disaggregate (AID), whi…
New findings show a balance between data fit and complexity in kernel hyperparameters.
The Bezier simplex fitting is a novel data modeling technique which exploits geometric structures of data to approximate the Pareto front of multi-objective optimization problems. There are two fitting methods based on different sampling strategies. The inductive skeleton fitting employs a stratified subsampling from e…
CEDA improves understanding of data fit to models.
This paper illustrates a procedure for fitting financial data with -stable distributions. After using all the available methods to evaluate the distribution parameters, one can qualitatively select the best estimate and run some goodness-of-fit tests on this estimate, in order to quantitatively assess its quality. I…
Study presents MMC model for better fitting multiple choice data.
A new kernel Stein test assesses fit for variable-length sequential data.
The first step to realize automatic experimental data analysis for fusion plasma experiments is fitting noisy data of temperature and density spatial profiles, which are obtained routinely. However, it has been difficult to construct algorithms that fit all the data without over- and under-fitting. In this paper, we sh…
We present techniques for effective Gaussian process (GP) modelling of multiple short time series. These problems are common when applying GP models independently to each gene in a gene expression time series data set. Such sets typically contain very few time points. Naive application of common GP modelling techniques…
In regression modelling approach, the main step is to fit the regression line as close as possible to the target variable. In this process most algorithms try to fit all of the data in a single line and hence fitting all parts of target variable in one go. It was observed that the error between predicted and target var…
The paper fits a seven-parameter GTS distribution to financial data.
Fitting models for non-Poisson point processes is complicated by the lack of tractable models for much of the data. By using large samples of independent and identically distributed realizations and statistical learning, it is possible to identify absence of fit through finding a classification rule that can efficientl…
New spectral tests assess network model fits efficiently.
Neural networks fit fewer samples than their parameters suggest in practice.
Proposes a method to forecast spatial-temporal data with limited training data.
Missing values, irregularly collected samples, and multi-resolution signals commonly occur in multivariate time series data, making predictive tasks difficult. These challenges are especially prevalent in the healthcare domain, where patients' vital signs and electronic records are collected at different frequencies an…
This paper is a step-by-step tutorial for fitting a mixture distribution to data. It merely assumes the reader has the background of calculus and linear algebra. Other required background is briefly reviewed before explaining the main algorithm. In explaining the main algorithm, first, fitting a mixture of two distribu…
Analyzes premium data of Indian non-life insurers, finding GEV distribution best fits Lognormal and GEV extremes.
Kernel Multigrid accelerates Back-fitting for additive Gaussian Processes.
The paper fits cash management models to data using stochastic and linear programming.
A new method treats all variables equally in fitting data.
Hawkes processes have seen a number of applications in finance, due to their ability to capture event clustering behaviour typically observed in financial systems. Given a calibrated Hawkes process, of concern is the statistical fit to empirical data, particularly for the accurate quantification of self- and mutual-exc…
LLMs are vulnerable to task-irrelevant data changes, limiting their use for data fitting.
Two methods monitor high-dimensional processes via manifold fitting or learning.
We present the use of the fitted Q iteration in algorithmic trading. We show that the fitted Q iteration helps alleviate the dimension problem that the basic Q-learning algorithm faces in application to trading. Furthermore, we introduce a procedure including model fitting and data simulation to enrich training data as…
Graphical lasso may fail to fit models when data points are insufficient.
Spectrahedral regression fits convex functions via a non-convex optimization problem.
We show that univariate and symmetric multivariate Hawkes processes are only weakly causal: the true log-likelihoods of real and reversed event time vectors are almost equal, thus parameter estimation via maximum likelihood only weakly depends on the direction of the arrow of time. In ideal (synthetic) conditions, test…
Paper solves best approximation by exponential functions for economic data.
We present a new variable selection method based on model-based gradient boosting and randomly permuted variables. Model-based boosting is a tool to fit a statistical model while performing variable selection at the same time. A drawback of the fitting lies in the need of multiple model fits on slightly altered data (e…
Study evaluates different mathematical models for three case studies using statistical fitting.
Proposes a method to quantify uncertainty in PFNs.
OccamNet finds interpretable symbolic fits to data efficiently.
Typically, operational risk losses are reported above a threshold. Fitting data reported above a constant threshold is a well known and studied problem. However, in practice, the losses are scaled for business and other factors before the fitting and thus the threshold is varying across the scaled data sample. A report…
New deep learning methods improve estimation and GOF assessment for large-scale IFA.
New robustness test for kernel goodness-of-fit tests.
Smoothed fitness landscape improves protein optimization.
Software package assesses spherical data distributions and clusters.
NNEinFact fits any nonnegative tensor factorization quickly and accurately.
The aim of the present article is to treat the Greek public debt issue strictly as a curve fitting problem. Thus, based on Eurostat data and using the Mathematica technical computing software, an exponential function that best fits the data is determined modelling how the Greek public debt expands with time. Exploring …
Improved modeling of persistence diagrams for data analysis.
Cross validation residuals are well known for the ordinary least squares model. Here leave-M-out cross validation is extended to generalised least squares. The relationship between cross validation residuals and Cook's distance is demonstrated, in terms of an approximation to the difference in the generalised residual …
Improved nuclear cross section fitting with weighted Levenberg-Marquardt method.
Streaming tensor factorization is a powerful tool for processing high-volume and multi-way temporal data in Internet networks, recommender systems and image/video data analysis. Existing streaming tensor factorization algorithms rely on least-squares data fitting and they do not possess a mechanism for tensor rank dete…
Two algorithms improve fitting autoregressive models for big data.
New method improves causal structure discovery with Prior-Fitted Networks.