Study shows the corrected Akaike criterion is inadmissible for estimating Kullback-Leibler discrepancy.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Automates model selection for GLMs using optimization.
SplitWise enhances stepwise regression by adaptively encoding numeric predictors into binary features.
We introduce a new criterion to determine the order of an autoregressive model fitted to time series data. It has the benefits of the two well-known model selection techniques, the Akaike information criterion and the Bayesian information criterion. When the data is generated from a finite order autoregression, the Bay…
We propose to use nonparametric Bernstein copulas as bivariate pair-copulas in high-dimensional vine models. The resulting smooth and nonparametric vine copulas completely obviate the error-prone need for choosing the pair-copulas from parametric copula families. By means of a simulation study and an empirical analysis…
In this paper we introduce a new feature selection algorithm to remove the irrelevant or redundant features in the data sets. In this algorithm the importance of a feature is based on its fitting to the Catastrophe model. Akaike information crite- rion value is used for ranking the features in the data set. The propose…
GWRBoost improves GWR for better spatial relationship quantification.
Paper uses referenced thermodynamic integration for Bayesian model selection in a complex COVID-19 transmission model.
We investigate the prediction capability of the orthogonal greedy algorithm (OGA) in high-dimensional regression models with dependent observations. The rates of convergence of the prediction error of OGA are obtained under a variety of sparsity conditions. To prevent OGA from overfitting, we introduce a high-dimension…
Modeling air pollutants using data-driven techniques and sparse identification of nonlinear dynamics.
When the in-sample Sharpe ratio is obtained by optimizing over a k-dimensional parameter space, it is a biased estimator for what can be expected on unseen data (out-of-sample). We derive (1) an unbiased estimator adjusting for both sources of bias: noise fit and estimation error. We then show (2) how to use the adjust…
We study the dynamical behavior of high-frequency data from the Korean Stock Price Index (KOSPI) using the movement of returns in Korean financial markets. The dynamical behavior for a binarized series of our models is not completely random. The conditional probability is numerically estimated from a return series of K…
Study predicts stream turbidity using surrogate data and meta-model.
A new method for averaging model predictions using minimum divergence.
In the information-based paradigm of inference, model selection is performed by selecting the candidate model with the best estimated predictive performance. The success of this approach depends on the accuracy of the estimate of the predictive complexity. In the large-sample-size limit of a regular model, the predicti…
We test three common information criteria (IC) for selecting the order of a Hawkes process with an intensity kernel that can be expressed as a mixture of exponential terms. These processes find application in high-frequency financial data modelling. The information criteria are Akaike's information criterion (AIC), the…
We propose an estimator of prediction error using an approximate message passing (AMP) algorithm that can be applied to a broad range of sparse penalties. Following Stein's lemma, the estimator of the generalized degrees of freedom, which is a key quantity for the construction of the estimator of the prediction error, …
Birg{é} and Massart proposed in 2001 the slope heuristics as a way to choose optimally from data an unknown multiplicative constant in front of a penalty. It is built upon the notion of minimal penalty, and it has been generalized since to some "minimal-penalty algorithms". This paper reviews the theoretical results ob…
Efficiently estimates covariance for sparse functional data.
Emergent and unscheduled cardiology admissions from cardiac catheterization laboratory add complexity to the management of Cardiology and in-patient department. In this article, we sought to study the behavior of cardiology admissions from Catheterization laboratory using time series models. Our research involves retro…
There are three principle paradigms of statistical inference: (i) Bayesian, (ii) information-based and (iii) frequentist inference. We describe an objective prior (the weighting or -prior) which unifies objective Bayes and information-based inference. The -prior is chosen to make the marginal probability an unbia…
We study tick-by-tick financial returns belonging to the FTSE MIB index of the Italian Stock Exchange (Borsa Italiana). We can confirm previously detected non-stationarities. However, scaling properties reported in the previous literature for other high-frequency financial data are only approximately valid. As a conseq…
A challenging problem in estimating high-dimensional graphical models is to choose the regularization parameter in a data-dependent way. The standard techniques include -fold cross-validation (-CV), Akaike information criterion (AIC), and Bayesian information criterion (BIC). Though these methods work well for lo…
Bayesian BIC for multi-trial data improves VAR model order selection.
Bayesian methods improve OoD detection in deep networks.
The paper derives an equation linking WAIC and WBIC for singular models.
Principal component analysis (PCA) is a popular method for projecting data onto uncorrelated components in lower dimension, although the optimal number of components is not specified. Likewise, multiple signal classification (MUSIC) algorithm is a popular PCA-based method for estimating directions of arrival (DOAs) of …
Modeling high-frequency order book data with Hawkes-Markovian process.
Proposes SNML for selecting word2vec Skip-gram dimensionality.
SIC detects elbows in error curves automatically.
In the economic literature, geographic distances are considered fundamental factors to be included in any theoretical model whose aim is the quantification of the trade between countries. Quantitatively, distances enter into the so-called gravity models that successfully predict the weight of non-zero trade flows. Howe…
A new strategy selects k in k-NN regression without hold-out data.
New method speeds up model selection for complex scientific tasks.
Bayesian method discovers PDEs with variable coefficients robustly.
Paper introduces NICc for fast cluster-based validation of prediction models.
The extension of the classical Bayesian penalized spline method to inference on vector-valued functions is considered, with an emphasis on characterizing the suitability of the method for general application.We show that the standard quadratic penalty is exactly analogous to the energy of a stretched string, with the p…
The -1 norm based optimization is widely used in signal processing, especially in recent compressed sensing theory. This paper studies the solution path of the -1 norm penalized least-square problem, whose constrained form is known as Least Absolute Shrinkage and Selection Operator (LASSO). A solution path …
Paper proposes a novel method to accurately determine the number of experts in Gaussian-gated Gaussian MoE models.
Unified framework for SGMoE resolves estimation and selection issues.
The study finds that firm membership in flagship indices and TCFD endorsement are strong predictors of a wider Disclosure-Performance Gap.
Develops new Markov processes with switching rates and past dependence.