Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

3468101135 · May 202619922001200920182026
48 results for covariance interpolation

Geometric families of low-rank covariances improve flexibility and tractability in high dimensions.

problem Interpolating and identifying covariance matrices in high dimensions with limited data.
method Differential geometric construction of low-rank covariance families, interpolation on manifolds, and distance minimization for identification.
result Differential geometric covariance families offer significant flexibility and computational tractability.

New bounds for linear interpolators show how they generalize under covariate shifts.

problem Understanding how linear interpolators generalize under covariate shifts.
method Proved non-asymptotic excess risk bounds for benignly-overfit linear interpolators in transfer learning.
result Identified beneficial and malignant covariate shifts based on overparameterization degree.

Study on clustering in high dimensions with anisotropic Gaussian mixtures, showing interpolation can be optimal and robust.

problem Clustering in high-dimensional anisotropic Gaussian mixtures.
method Derive minimax bounds, analyze 2\ell_2-regularized classifiers, and investigate interpolation's robustness.
result Interpolating solutions can be optimal and robust under certain conditions.

The paper addresses frequency-dependent distortions in massive MIMO systems and proposes a method to recover covariance matrices.

problem Frequency-dependent distortions in the covariance matrix of massive MIMO systems.
method Proposes a novel UL-DL covariance interpolation technique under a mild reciprocity condition.
result The proposed method can recover the covariance matrix in the DL from an estimate in the UL, especially in FDD massive MIMO systems.

Inflating the minimum norm interpolator improves linear regression generalization error.

problem Highly anisotropic covariances and diverging d/nd/n in linear regression.
method Inflating the minimum 2\ell_2 norm interpolator by a constant greater than one.
result Inflating the minimum norm interpolator improves generalization error.

Paper investigates optimal interpolation methods in linear regression.

problem Understanding when interpolating methods generalize well in linear regression.
method Investigates optimal response-linear interpolators using functions linear in the response variable.
result Provides a closed-form expression for the optimal interpolator and shows it can be derived as the limit of gradient descent.

The paper analyzes the generalization error of min-norm interpolators in transfer learning with limited test samples.

problem Characterizing the generalization error of min-norm interpolators in transfer learning with limited test samples.
method Characterizes the bias and variance of pooled min-2\ell_2-norm interpolation under covariate shift and model shift.
result Shows that adding data can hurt when SNR is low and is beneficial at higher SNR levels under certain conditions.

The study evaluates different parameter selection methods for Gaussian process interpolation.

problem Choosing optimal parameters for Gaussian process interpolation.
method Empirical study using scoring rules and leave-one-out selection criteria.
result The choice of model family is often more important than the selection criterion.

Near-interpolating models grow norms quickly, affecting generalization.

problem Understanding the trade-off between interpolation and generalization in near-interpolating models.
method Random matrix theory and eigendecay analysis of data covariance matrix.
result Near-interpolating models exhibit rapid norm growth and worse generalization trade-offs.

We are interested in comparing probability distributions defined on Riemannian manifold. The traditional approach to study a distribution relies on locating its mean point and finding the dispersion about that point. On a general manifold however, even if two distributions are sufficiently concentrated and have unique …

2008-07-04abs ↗pdf ↗

Study shows unusual non-monotonic risk behavior in minimum-norm interpolants for various data scaling.

problem Understanding the risk behavior of minimum-norm interpolants in RKHS for different data scaling.
method Analysis of spectral properties of the random kernel matrix restricted to eigen-spaces of the population covariance operator.
result Minimum-norm interpolants in RKHS exhibit multiple descent in risk for d=nαd = n^α with α(0,1)α\in(0,1).

Estimates individual treatment effects using gradient interpolation and kernel smoothing.

problem Estimating individualized continuous treatment effects in observational data.
method Augment training data with independently sampled treatments and inferred counterfactual outcomes using gradient interpolation and kernel smoothing.
result Our method outperforms state-of-the-art methods on counterfactual estimation error.

CNNs predict spatial fields from sparse data.

problem Predicting complete spatial fields from limited observations.
method Convolutional Neural Networks (CNNs) trained on a single partially observed field.
result CNNs can flexibly capture local spatial patterns without explicit covariance modeling.

Our paper examines binary linear classification under Gaussian mixtures, revealing conditions for optimal performance.

problem Understanding the conditions for optimal performance of binary linear classifiers under Gaussian mixtures.
method We study max-margin SVM and min-norm interpolating classifiers, deriving bounds and conditions for optimal performance.
result Interpolating estimators achieve asymptotically optimal performance under certain conditions, emphasizing the role of SNR and covariance.

The study prevents model collapse in overparameterized linear regression by mixing real and synthetic labels.

problem Preventing model collapse in overparameterized linear regression.
method Iterative mixing of real and synthetic labels, deriving generalization error formulae.
result Optimal mixing ratio converges to the reciprocal of the golden ratio for isotropic features.

A new GP interpolation method for better predictive distributions in ranges of interest.

problem Improving predictive distributions in specific ranges of interest.
method Relaxed Gaussian process interpolation, relaxing interpolation constraints outside ranges of interest.
result Better predictive distributions in ranges of interest, especially in non-stationary cases.

Kernel Ridgeless Regression can generalize well without explicit regularization.

problem The challenge of achieving generalization in Kernel Ridgeless Regression without additional regularization.
method Minimum-norm interpolated solutions with high-dimensional data, curvature of kernel function, and favorable data geometry.
result Implicit regularization leads to good generalization in Kernel Ridgeless Regression.

Linear regression can overfit without harm when data is high-dimensional.

problem Understanding when a perfect fit to noisy training data in linear regression leads to accurate predictions.
method Characterization of linear regression problems with a minimum norm interpolating prediction rule that has near-optimal prediction accuracy.
result Overparameterization is essential for benign overfitting in this setting, and the number of unimportant prediction directions must exceed the sample size.

Gaussian Processes offer a flexible method for modeling and predicting outcomes with uncertainty estimates.

problem Capturing uncertainty in predictions at new data points, especially with poor overlap and extrapolation.
method Gaussian Processes model a posterior distribution over outcomes, reflecting the range of plausible models.
result GPs provide a principled approach to handling extrapolation and uncertainty in predictions.

New kernel HMK improves Gaussian process expressiveness and supports harmonizable covariances.

problem Improving the expressiveness of Gaussian processes with non-stationary kernels.
method Proposed harmonizable mixture kernel (HMK) and variational Fourier features.
result HMK interpolates between local patterns and offers robust kernel learning.

Study shows interpolating predictor's risk is optimal in low-dimensional factor regression models.

problem Understanding the risk of interpolating predictors in high-dimensional factor regression models.
method Detailed finite-sample analysis of minimum-norm interpolating predictor's risk in factor regression models.
result The risk of the minimum-norm interpolating predictor approaches optimal benchmarks in low-dimensional factor regression models.

The study analyzes robustness of estimators in linear models with adversarial errors.

problem Analyzing robustness of estimators in linear models with adversarial errors.
method Develops a general theory for minimum norm interpolating estimators and RERM in linear models without conditions on errors.
result Quantitative bound for the prediction error relating it to Rademacher complexity, norm of minimum norm interpolator of errors, and subdifferential size.

We analyze ridge interpolators in correlated factor regression models using RDT.

problem Performance analysis of ridge interpolators in correlated factor regression models.
method Utilizing Random Duality Theory (RDT), we obtain precise closed form characterizations of optimization problems.
result Ridge interpolators can smooth out the excess prediction risk and exhibit double-descent behavior.

Proposes bivariate DeepKriging for efficient wind field prediction.

problem Challenges in predicting large-scale bivariate wind fields with high spatial variability and heterogeneity.
method Spatially dependent deep neural network (DNN) with embedding layer using spatial radial basis functions.
result Outperforms traditional cokriging predictors and reduces computation time.

A scalable algorithm for sampling and fine-tuning models using Tilt Matching.

problem Efficient sampling and fine-tuning of generative models.
method Tilt Matching, arising from a dynamical equation, minimizes variance and inherits regularity from stochastic interpolants.
result Empirically verified to be efficient and highly scalable, providing state-of-the-art results.

New distributed algorithm for gradient descent converges linearly in the interpolation limit.

problem Analyzing convergence of stochastic gradient descent in the interpolation limit of high-dimensional data fitting.
method Introduced a distributed gradient descent algorithm for the penalized distributed loss function, showing linear convergence rate.
result Distributed SGD algorithm converges linearly with rate 1-η/nλ_min(H), where λ_min(H) is the smallest nonzero eigenvalue of the Hessian.

Transfer learning improves MNI's performance in high-dimensional linear regression.

problem Improving model performance in high-dimensional linear regression with diverse data.
method Proposes a Transfer MNI approach, analyzing its excess risk and conditions for outperformance.
result Identifies free-lunch covariate shift regimes where knowledge transfer benefits.

Gaussian processes are rich distributions over functions, which provide a Bayesian nonparametric approach to smoothing and interpolation. We introduce simple closed form kernels that can be used with Gaussian processes to discover patterns and enable extrapolation. These kernels are derived by modelling a spectral dens…

2013-02-18abs ↗pdf ↗

The paper examines how kernel approximations affect Gaussian process regression in large data applications.

problem Effect of kernel approximations on Gaussian process regression in large data applications.
method Unified framework to analyze Gaussian process regression under computational and epistemic misspecification.
result Theoretical analysis of Gaussian process regression under various misspecifications.

Neural networks can interpolate random data but still generalize well, studied in the NT regime.

problem Understanding how neural networks interpolate random labels and generalize well in the overparametrized regime.
method Characterization of the eigenstructure of the empirical NT kernel and generalization error of NT ridge regression.
result The generalization error is well approximated by polynomial ridge regression with an increased regularization parameter.

Task shift from classification to regression is possible in overparameterized linear models with limited additional data.

problem Transferability of latent knowledge from classification to regression in overparameterized linear models.
method Investigation of task shift in overparameterized linear regression, zero-shot and few-shot cases, with a focus on minimum-norm interpolation.
result Minimum-norm interpolators can transfer latent knowledge from classification to regression with limited additional data.

Study resolves conjecture on overparameterized linear models' generalization.

problem Asymptotic generalization of multiclass classification with overparameterized models.
method Gaussian covariates bi-level model, Hanson-Wright inequality variant.
result Min-norm interpolating classifier can be suboptimal compared to noninterpolating classifiers.

The paper analyzes boosting and minimum-1\ell_1-norm classifiers in high dimensions.

problem Understanding the generalization error and optimal Bayes error in boosting.
method High-dimensional asymptotic theory, Gaussian comparison techniques, uniform deviation argument.
result Precise characterizations of boosting test error and optimal Bayes error.

New method for causal inference with observed covariates improves learning rates.

problem Causal inference with observed covariates in nonparametric instrumental variable regression.
method Introduces novel Fourier measure for partial smoothing and adapts kernel lengthscales for anisotropic smoothness.
result Upper and lower learning rates for KIV-O show interpolation between NPIV and NPR rates.

A new kernel-based nonconformity score improves multivariate prediction regions.

problem Tackling the challenge of compressing multivariate residual vectors into scalars while preserving geometric structure.
method Introducing a Multivariate Kernel Score (MKS) that decomposes into an anisotropic MMD, providing finite-sample coverage guarantees and convergence rates.
result The MKS produces prediction regions that explicitly adapt to geometric structure, reducing volume compared to ellipsoidal baselines.

Kernelized PCovR reveals structure-property relations in chemistry and materials.

problem Understanding structure-property relations in complex systems.
method Kernel Principal Covariates Regression (kernel PCovR) with sparsification.
result Kernelized PCovR effectively reveals and predicts structure-property relations.

Principal Component Analysis (PCA) is the most common nonparametric method for estimating the volatility structure of Gaussian interest rate models. One major difficulty in the estimation of these models is the fact that forward rate curves are not directly observable from the market so that non-trivial observational e…

2014-08-26abs ↗pdf ↗