Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

183367550733 · Jun 202019922001200920182026
48 results for dependent measurement errors

The paper addresses errors-in-variables models with dependent measurement errors.

problem Estimating sparse parameters from errors-in-variables models with dependent measurement errors.
method Developed methods to recover sparse ββ^* from a single observation matrix XX and response vector yy under specific conditions.
result Consistent estimation of sparse ββ^* and obtained rates of convergence in the q\ell_q norm.

A new weighted dissimilarity measure reduces positioning errors in feature-based systems.

problem Reducing errors in feature-based positioning systems, especially in areas with high variability.
method Iterative scheme using location-dependent standard deviations as weights.
result Maximum radial positioning error reduced by 40% using the weighted dissimilarity measure.

Proposes a general framework for fairness-aware learning using f-divergences.

problem Ensuring fairness in classifier predictions without compromising accuracy.
method Introduces a general framework using f-divergences and provides a unified analysis of the upper bound of the estimation error.
result Guarantees low dependencies on unseen samples for any f-divergence.

The paper evaluates forecast accuracy of realized volatility measures in large cross-sections.

problem Forecast evaluation of realized volatility measures in large cross-sections of financial data.
method Equal predictive accuracy testing procedures, LASSO shrinkage, measurement error correction, cross-sectional jump component measures.
result The augmented HAR model outperforms the standard HAR model in forecasting realized volatility.

We study the total least squares (TLS) problem that generalizes least squares regression by allowing measurement errors in both dependent and independent variables. TLS is widely used in applied fields including computer vision, system identification and econometrics. The special case when all dependent and independent…

2014-06-01abs ↗pdf ↗

Study proposes method to estimate causal effects from noisy treatment data.

problem Estimating causal effects from noisy treatment data without side information.
method Deep latent variable model with neural network parameterization and amortized importance-weighted variational objective.
result Causal effect estimates are identifiable without side information and measurement error variance knowledge.

This work extends learning theory to complexly dependent data under Dobrushin's condition.

problem Learning from weakly dependent data sampled on networks or spatial domains.
method Developed complexity measures and learning bounds for hypothesis classes under Dobrushin's condition.
result Generalization and learnability bounds degrade by constant and log factors compared to i.i.d. settings.

New weighted Lasso estimates improve logistic regression performance with measurement error.

problem Improper Lasso estimates in sparse logistic regression with equal penalties.
method Proposed weighted Lasso estimates using McDiarmid inequality for non-asymptotic oracle inequalities.
result Finite sample behavior illustrated by non-asymptotic oracle inequalities for estimation and prediction errors.

Paper examines risk measure expansions under FGM dependence, improving accuracy at extreme levels.

problem Capturing higher-order tail behavior and dependence effects in risk measures.
method Second-order asymptotic expansions using extreme value theory and regular variation theory.
result Second-order approximations reduce approximation errors, especially at extreme confidence levels.

Paper proposes Gini distance statistics for estimating feature-label dependence.

problem Identifying statistical dependence between features and categorical labels.
method Generalized Gini distance in RKHS for feature-label dependence estimation.
result Gini distance statistics converge faster and have tighter error bounds than distance covariance.

Sparse feature selection improves batch RL efficiency.

problem High-dimensional batch RL with many features.
method Sparse linear function approximation, Lasso, group Lasso, fitted Q-evaluation, fitted Q-iteration.
result Sparse feature selection makes batch RL more sample efficient.

This paper improves the robustness of risk estimation for financial positions.

problem Ensuring robustness of risk measures in the presence of data noise.
method Proposes a quantitative approach using the Fortet-Mourier metric to quantify the variation of true probability measures.
result Derives explicit error bounds for discrepancies between laws of estimators based on true and perturbed data.

Paper introduces new bounds linking data compressibility to generalization error.

problem Establishing data-dependent generalization bounds.
method Variable-size compressibility framework linking generalization error to compression rate of input data.
result New bounds depend on empirical data measure, subsuming existing PAC-Bayes and intrinsic dimension bounds.

The banking systems that deal with risk management depend on underlying risk measures. Following the Basel II accord, there are two separate methods by which banks may determine their capital requirement. The Value at Risk measure plays an important role in computing the capital for both approaches. In this paper we an…

2011-11-18abs ↗pdf ↗

Suppose that we observe yRfy \in \mathbb{R}^f and XRf×mX \in \mathbb{R}^{f \times m} in the following errors-in-variables model: \begin{eqnarray*} y & = & X_0 β^* + ε\\ X & = & X_0 + W \end{eqnarray*} where X0X_0 is a f×mf \times m design matrix with independent subgaussian row vectors, εRfε\in \mathbb{R}^f is a noise vector…

2015-02-09abs ↗pdf ↗

ST-SAN predicts flow with spatial-temporal dependencies using self-attention.

problem Challenges in predicting flow due to spatial-temporal dependencies.
method Spatial-Temporal Self-Attention Network (ST-SAN) that addresses temporal and spatial dependencies.
result Significant improvement in flow prediction accuracy (9% in inflow, 4% in outflow) compared to state-of-the-art methods.

Paper proposes a new method for learning kernels that depend on both inputs and outputs.

problem Common kernels are limited in their ability to handle complex tasks.
method Developed a spectral kernel learning framework that uses non-stationary kernels and learns from data.
result Derived a data-dependent generalization error bound and suggested regularization terms.

Paper estimates EOT maps for non-compactly supported measures with subGaussian target.

problem Estimating EOT maps between non-compactly supported measures.
method Uses bias-variance decomposition, T1-transport inequalities, and concentration of measure results.
result Shows error decay rates for different cases of subGaussian measures.

We study prediction and estimation problems using empirical risk minimization, relative to a general convex loss function. We obtain sharp error rates even when concentration is false or is very restricted, for example, in heavy-tailed scenarios. Our results show that the error rate depends on two parameters: one captu…

2014-10-13abs ↗pdf ↗

New bounds derived for machine learning algorithms using convex functions.

problem Bounding generalization error in machine learning.
method Using strongly convex functions and subgaussian loss tails, derived new generalization bounds.
result Generalization bounds can be derived using any strongly convex function of the joint input-output distribution.

GANs learn distributions well from samples, with rates depending on intrinsic dimension.

problem Learning distributions from samples using GANs.
method Oracle inequality, Hölder functions approximation, neural network approximation, integral probability metrics.
result Convergence rates of GANs depend on intrinsic dimension, not ambient dimension.

Study of estimation errors in surrogate loss minimizers, providing stronger guarantees than existing methods.

problem Estimation errors in surrogate loss minimizers for various hypothesis sets.
method Detailed study of H\mathscr{H}-consistency estimation error bounds, proving general theorems for distribution-dependent and independent settings.
result Explicit bounds for zero-one and adversarial losses, showing enhancements under distributional assumptions.

This work analyzes discrete diffusion models using stochastic integrals, providing error bounds and insights.

problem Error analysis for discrete diffusion models remains less understood.
method Proposes a comprehensive framework based on Lévy-type stochastic integrals.
result Obtains the first error bound for the ττ-leaping scheme in KL divergence.

The paper learns dependency graphs from non-Gaussian data.

problem Capturing dependency graphs from real data that may not be Gaussian.
method Additive over-parametrization with shrinkage to incorporate variable dependencies; iterative Gaussian graph learning algorithm.
result The estimators achieve satisfactory accuracy in measuring dependency structures.

Paper advances sparse regularisation theory for measures with new kernel insights.

problem Estimating sparse measures from noisy observations using continuous sparse regularisation.
method Develops new continuous sparse regularisation theory on measures with Beurling-LASSO, introduces kernel switch analysis.
result Proves the ``sinc-4'' kernel satisfies a technical LPC assumption for error bounds.

Framework identifies discrepancies in physics models, improving sensor accuracy.

problem Model inaccuracies leading to poor control algorithms.
method Learning systematic state-space residuals and deterministic dynamical errors.
result Improved quantification of system dynamics and control algorithms.

MOB-dS uses permutation to correct for dependency in discrete survival data.

problem Identifying subgroups in discrete event time data with potential spurious results.
method Model-based recursive partitioning (MOB) with modified data matrix and permutation test.
result MOB-dS controls type I error rate better than standard MOB for discrete survival data.

Develops efficient inference for noise heterogeneity in machine learning models.

problem Downstream procedures based on residuals can be biased in additive noise models.
method Semiparametrically efficient inference using a novel Hilbert-valued one-step estimator.
result Constructs tests and confidence intervals for residual independence and goodness of fit.

MEBoost selects variables in regression with measured error, improving accuracy over naive methods.

problem Variable selection in regression models with covariates measured with error.
method Iterative algorithm that corrects for measurement error using estimating equations.
result MEBoost outperforms naive methods in selecting accurate covariates, especially under high measurement error.

New approach uses Gaussian processes to learn and track complex systems with guaranteed accuracy.

problem Inaccurate first principle models for complex systems due to data complexity.
method Bayesian prediction error bound for Gaussian process regression, derived from kernel-based data density.
result Achieves vanishing tracking error with increasing data density, providing time-varying accuracy guarantees.

Two theories improve SGLD's generalization bounds for non-convex learning.

problem Improving generalization error bounds for non-convex learning problems.
method Stochastic Gradient Langevin Dynamics (SGLD) with stability and PAC-Bayesian analysis.
result First algorithm-dependent bounds with reasonable aggregated step size dependence.

DIET tests conditional independence using marginal dependence measures of residual information.

problem Computational intractability of conditional randomization tests (CRTs).
method DIET avoids fitting large models by leveraging marginal independence statistics of information residuals.
result DIET achieves higher power than other tractable CRTs on synthetic and real benchmarks.

Study confirms link between unemployment and real GDP growth in developed countries.

problem Predicting unemployment based on real GDP growth in developed nations.
method Revised Okun's law using real GDP per capita and unemployment rate data from 2010-2019.
result Accurate prediction of unemployment rate changes using real GDP growth rate.

Robust variable selection for high-dimensional data with missing and measurement errors.

problem Missing data and measurement errors confound data distribution.
method Exponential loss function with inverse probability weighting and additive error models.
result The Atan punishment method improves robust variable selection.

Paper bounds generalization error for noisy gradient methods on non-convex learning.

problem Proving tight generalization error bounds for non-convex learning.
method Developed Bayes-Stability framework combining PAC-Bayesian theory and algorithmic stability.
result Obtained new data-dependent generalization bounds for SGLD and noisy gradient methods.

Score matching errors are not sufficient for measuring diffusion model quality.

problem The L2L^2 score matching error is not a reliable measure of diffusion model performance.
method Decomposed score errors into gradient and solenoidal components and analyzed their geometric properties.
result Only the gradient component of the score error affects the marginal distributional quality.

This paper analyzes stability of decision trees and logistic regression.

problem Stability of decision trees and logistic regression is analyzed to understand their performance and sensitivity.
method Two stability notions (hypothesis and pointwise hypothesis stability) are derived for decision trees and logistic regression. The stability of decision trees depends on the number of leaves, while for logistic regression, it depends on the smallest eigenvalue of the Hessian matrix. Upper bounds on generalization error are constructed.
result Logistic regression is not a stable learning algorithm.