Characterizes statistical complexity of realizable regression in PAC and online learning.
problem Understanding the statistical complexity of realizable regression in both PAC and online learning settings.
method Introduces minimax instance optimal learners, novel and combinatorial dimensions to characterize learnability.
result Characterizes which classes of real-valued predictors are learnable and provides necessary conditions for learnability.
In this paper we formulate a regression problem to predict realized volatility by using option price data and enhance VIX-styled volatility indices' predictability and liquidity. We test algorithms including regularized regression and machine learning methods such as Feedforward Neural Networks (FNN) on S&P 500 Index a…
New bandit algorithm works without realizability assumption.
problem Contextual bandit problems without realizability assumption.
method Computes a constrained regression problem in every epoch, ensuring similar regret guarantees as realizability-based algorithms.
result Ensures similar regret guarantees as realizability-based algorithms, up to a misspecification term.
A new realized conditional autoregressive Value-at-Risk (VaR) framework is proposed, through incorporating a measurement equation into the original quantile regression model. The framework is further extended by employing various Expected Shortfall (ES) components, to jointly estimate and forecast VaR and ES. The measu…
Faster algorithm reduces contextual bandit regret with fewer offline regression calls.
problem Optimizing reward in contextual bandits with unknown functions.
method Designing a simple algorithm with O(logT) offline regression calls. result Achieves statistically optimal regret with minimal offline calls.
The joint Value at Risk (VaR) and expected shortfall (ES) quantile regression model of Taylor (2017) is extended via incorporating a realized measure, to drive the tail risk dynamics, as a potentially more efficient driver than daily returns. Both a maximum likelihood and an adaptive Bayesian Markov Chain Monte Carlo m…
Study bounds noise level in linear regression with dependent data.
problem Analyzing noise level in linear regression with dependent data.
method Derive upper bounds for random design linear regression with β-mixing data, without realizability assumptions. result Correctly recovers the noise level of the problem, exhibiting graceful degradation with misspecification.
Optimal algorithm for maximizing rewards in contextual bandits with resource constraints.
problem Maximizing rewards in contextual bandits with resource constraints.
method Proposed a universal and optimal algorithmic framework for CBwK by reducing it to online regression.
result Established the optimality of the proposed algorithm for various function classes.
We study distributions of realized variance (squared realized volatility) and squared implied volatility, as represented by VIX and VXO indices. We find that Generalized Beta distribution provide the best fits. These fits are much more accurate for realized variance than for squared VIX and VXO -- possibly another indi…
This paper investigates how the conditional quantiles of future returns and volatility of financial assets vary with various measures of ex-post variation in asset prices as well as option-implied volatility. We work in the flexible quantile regression framework and rely on recently developed model-free measures of int…
Bayesian framework forecasts financial tail risks using realized volatility and nonlinear thresholds.
problem Forecasting financial tail risks using realized volatility and nonlinear thresholds.
method Bayesian Markov Chain Monte Carlo method for model estimation; nonlinear threshold regression specification.
result The proposed framework produces competitive tail risk forecasts compared to GARCH and Realized-GARCH models.
Paper develops neural network for distribution regression.
problem Regression with probability measures.
method Develops a novel fully connected neural network (FNN) for distribution inputs.
result Almost optimal learning rates for distribution regression derived.
Paper tackles MLR prediction error without assuming realizable models.
problem Prediction error in mixture of linear regressions without realizable assumptions.
method Developed algorithms for list-decoding MLR predictions and minimized empirical risk.
result Alternating minimization algorithm finds best fit lines in non-realizable settings.
This study improves tail risk forecasting by integrating overnight information into semi-parametric models.
problem Improving tail risk forecasting in financial markets.
method Proposes RES-CAViaR-oc models combining overnight return and realized volatility, using Bayesian estimation.
result Realized volatility and overnight return significantly improve tail risk forecasting.
New active learning framework for multiclass classification beyond realizability assumption.
problem Active learning in non-realizable settings with convex model classes.
method Surrogate risk minimization, epoch-based fitting, aggregation of models.
result Achieves label and sample complexity comparable to prior work in non-realizable settings.
In this work, we highlight a connection between the incremental proximal method and stochastic filters. We begin by showing that the proximal operators coincide, and hence can be realized with, Bayes updates. We give the explicit form of the updates for the linear regression problem and show that there is a one-to-one …
New method tests independence with single nonstationary time series.
problem Testing independence in nonstationary nonlinear time series.
method Time-varying nonlinear regression, local long-run covariance estimation, strong Gaussian approximation.
result First framework for conditional independence testing with a single realization of a nonstationary nonlinear process.
A major challenge in contextual bandits is to design general-purpose algorithms that are both practically useful and theoretically well-founded. We present a new technique that has the empirical and computational advantages of realizability-based approaches combined with the flexibility of agnostic methods. Our algorit…
Characterizes the sample complexity of list regression tasks.
problem Understanding the sample complexity of list learning tasks in regression.
method Introducing two combinatorial dimensions: k-OIG dimension and k-fat-shattering dimension.
result These dimensions characterize realizable and agnostic k-list regression.
New ensemble SVM model reduces prediction error without choosing best kernel.
problem Reducing prediction error in regression problems.
method Bagged-weighted support vector regression model with random machines.
result Regression Random Machines achieve lower generalization error.
New bounds for agnostic learning with average smoothness.
problem Distribution-free nonparametric regression with average smoothness.
method Distribution-free uniform convergence bounds and agnostic learning algorithm.
result Distribution-free uniform convergence bounds for average-smoothness classes in the agnostic setting.
CAVI speeds up Bayesian MIDAS regression by 107x-1,772x with similar accuracy.
problem Efficiently estimating Bayesian MIDAS regression models with many predictors.
method Coordinate Ascent Variational Inference (CAVI) for linear MIDAS regression.
result CAVI produces posterior means nearly identical to Gibbs sampling with significant speedup.
Unified framework for realizable and agnostic learning.
problem Lack of a unified theory for realizable and agnostic learnability.
method Three-line blackbox reduction.
result Unified understanding across various learning settings.
This paper investigates how realized and option implied volatilities are related to the future quantiles of commodity returns. Whereas realized volatility measures ex-post uncertainty, volatility implied by option prices reveals the market's expectation and is often used as an ex-ante measure of the investor sentiment.…
Enhances Gaussian process regression with multi-fidelity models and active subspaces for high-dimensional problems.
problem Data scarcity and high-dimensional input spaces with low intrinsic dimensionality.
method Employ Gaussian processes in a Bayesian setting, augmenting with low-fidelity models, and exploiting active subspaces.
result Improves model accuracy through multi-fidelity Gaussian process regression with active subspaces.
New algorithm reduces contextual bandits to efficient regression.
problem Developing efficient algorithms for contextual bandits with general function classes.
method Reduction from contextual bandits to online regression with oracle.
result First universal and optimal reduction with no overhead.
Adaptive sparseness enhances robust regression using MCC and ARD.
problem Developing a robust regression method with adaptive sparseness.
method Integrating MCC with ARD in a Bayesian framework using variational Bayesian inference.
result MCC-ARD regression outperforms existing methods in prediction and feature selection.
New algorithms achieve near-optimal cumulative loss in nonparametric online learning and games.
problem Fast rates of convergence in nonparametric online regression and classification.
method Randomized proper learning algorithms, hierarchical aggregation, multi-scale extension, stability proof.
result Achieved near-optimal cumulative loss bounds for real-valued and binary games.
Bayesian framework improves robustness in nonlinear regression models.
problem Measurement error, model misspecification, and distributional misspecification in regression analyses.
method Joint Dirichlet process prior on latent covariate-response distribution, updating with posterior pseudo-samples.
result Improved stability and consistency in estimators under increasing measurement error.
Study on ReLU regression with Massart noise, achieving exact parameter recovery.
problem Efficiently fitting ReLUs to data in the presence of Massart noise.
method Developed an efficient algorithm for exact parameter recovery under mild assumptions.
result Achieved exact parameter recovery in ReLU regression with Massart noise.
This paper presents the nonparametric inference for nonlinear volatility functionals of general multivariate Itô semimartingales, in high-frequency and noisy setting. Pre-averaging and truncation enable simultaneous handling of noise and jumps. Second-order expansion reveals explicit biases and a pathway to bias correc…
This paper revisits the fractional cointegrating relationship between ex-ante implied volatility and ex-post realized volatility. We argue that the concept of corridor implied volatility (CIV) should be used instead of the popular model-free option-implied volatility (MFIV) when assessing the fractional cointegrating r…
Study examines how imputation accuracy affects prediction accuracy in regression problems with missing covariates.
problem Missing covariates in regression or classification problems.
method Simulation and empirical analysis using UCI datasets and statistical inference.
result Imputation accuracy impacts prediction accuracy, especially with Machine Learning methods.
Algorithm identifies bilinear dynamical systems from noisy data.
problem Learning a realization of a partially observed bilinear dynamical system.
method Regression of outputs to highly correlated covariates for Markov-like parameters.
result High probability error bounds on identification algorithm under uniform stability assumption.
The paper analyzes the performance of empirical risk minimization for p-norm linear regression.
problem Empirical risk minimization on p-norm linear regression. method Analyzes performance under various conditions and moment assumptions.
result High probability excess risk bounds for empirical risk minimizer, matching asymptotic rates.
Study highlights how model choice affects uncertainty estimation in neural network regression.
problem Uncertainty estimation under model misspecification in neural network regression.
method Analyzed the impact of model choice on uncertainty estimation in neural network regression, focusing on aleatoric and epistemic uncertainties.
result Model misspecification leads to unreliable uncertainty estimates, highlighting the importance of choosing appropriate models.
Personalized medicine seeks to identify the causal effect of treatment for a particular patient as opposed to a clinical population at large. Most investigators estimate such personalized treatment effects by regressing the outcome of a randomized clinical trial (RCT) on patient covariates. The realized value of the ou…
The study uses reproducing kernels to model bond discount curves.
problem Estimating bond discount curves under no-arbitrage conditions.
method Introduced reproducing kernels as a regression basis for estimating bond discount curves.
result Reproducing kernels provide a tractable solution for calibrating models to market data.
New method predicts y distributions from imperfect data.
problem Predicting y from imperfect data (discrete, truncated, censored).
method Optimal transformations to estimate p(y|x).
result Estimates location, scale, and shape of y distribution.
In this work, we propose a new Gaussian process regression (GPR) method: physics information aided Kriging (PhIK). In the standard data-driven Kriging, the unknown function of interest is usually treated as a Gaussian process with assumed stationary covariance with hyperparameters estimated from data. In PhIK, we compu…
New method calibrates probabilistic regression models without restrictive assumptions.
problem Ensuring predictive distributions accurately reflect true uncertainty.
method Nonparametric re-calibration algorithm based on conditional kernel mean embeddings.
result Consistently outperforms prior re-calibration approaches across various benchmarks.
Selecting important features in non-linear or kernel spaces is a difficult challenge in both classification and regression problems. When many of the features are irrelevant, kernel methods such as the support vector machine and kernel ridge regression can sometimes perform poorly. We propose weighting the features wit…
This paper proposed a new regression model called l1-regularized outlier isolation and regression (LOIRE) and a fast algorithm based on block coordinate descent to solve this model. Besides, assuming outliers are gross errors following a Bernoulli process, this paper also presented a Bernoulli estimate model which, …
The paper proposes a mixed-frequency quantile regression model for VaR and ES forecasting.
problem Forecasting VaR and ES with mixed-frequency data.
method Mixed-frequency quantile regression model to estimate VaR and ES.
result The proposed model outperforms other models in VaR and ES backtesting tests.
Paper analyzes agnostic learning of mixed linear regression without generative models.
problem Learning mixed linear regression without assuming stochastic generation.
method Expectation Maximization (EM) and Alternating Minimization (AM) algorithms.
result AM and EM algorithms converge to population loss minimizers under standard conditions.
Realized statistics based on high frequency returns have become very popular in financial economics. In recent years, different non-parametric estimators of the variation of a log-price process have appeared. These were developed by many authors and were motivated by the existence of complete records of price data. Amo…
We introduce a concept of autoregressive (AR)state-space realization that could be applied to all transfer functions T(L) with T(0) invertible. We show that a theorem of Kalman implies each Vector Autoregressive model (with exogenous variables) has a minimal AR-state-space realization …
New GLS estimator handles high-dimensional data with autocorrelated errors.
problem High-dimensional regressions with autocorrelated errors.
method LASSO regression, autoregressive model fitting, and whitening.
result The method outperforms unadjusted LASSO in estimating errors driven by autoregressive processes.