Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

118237355473 · Jun 202019922001200920172026
48 results for HAC estimation

Study improves financial risk assessment using ARMA-APARCH-EVT models with HACs.

problem Improving risk assessment in financial portfolios.
method ARMA-APARCH-EVT-HAC model for volatility and extreme value forecasting.
result Empirical analysis shows the model's effectiveness in international stock market data.

Develops inequalities for high-dimensional linear processes with dependent innovations.

problem Estimating high-dimensional VAR(p) systems and HAC covariance estimation.
method Concentration inequalities for ll_\infty norm of vector linear processes with sub-Weibull, mixingale innovations.
result Obtained concentration bounds for the maximum entrywise norm of lag-hh autocovariance matrices.

Fair HAC algorithms ensure clustering fairness across protected groups.

problem Ensuring clustering fairness in HAC algorithms when datasets contain biases.
method Proposes fair algorithms for HAC that enforce fairness constraints regardless of distance linkage criteria.
result Our fair HAC algorithms find fairer clusterings compared to vanilla HAC and other fair clustering approaches.

Proposes a method to make statistical inferences robust in spatially dependent settings with missing at random labels.

problem Statistical inference challenges with missing at random labels and spatial dependence.
method Doubly robust estimator with cross-fit nuisances and jackknife spatial HAC variance correction.
result Asymptotically valid confidence intervals with improved finite-sample calibration.

Hierarchical clustering is a widely used approach for clustering datasets at multiple levels of granularity. Despite its popularity, existing algorithms such as hierarchical agglomerative clustering (HAC) are limited to the offline setting, and thus require the entire dataset to be available. This prohibits their use o…

2019-09-20abs ↗pdf ↗

As the size nn of datasets become massive, many commonly-used clustering algorithms (for example, kk-means or hierarchical agglomerative clustering (HAC) require prohibitive computational cost and memory. In this paper, we propose a solution to these clustering problems by extending threshold clustering (TC) to probl…

2019-07-05abs ↗pdf ↗

In many applications that involve processing high-dimensional data, it is important to identify a small set of entities that account for a significant fraction of detections. Rather than formalize this as a clustering problem, in which all detections must be grouped into hard or soft categories, we formalize it as an i…

2018-05-08abs ↗pdf ↗

Study analyzes Airbnb lead-time distributions for Nights Booked and Gross Booking Value, finding divergent shapes and tail behavior.

problem Analyzing lead-time distributions for Airbnb demand metrics.
method Compositional analysis of daily lead-time vectors, fitting Gamma, Weibull, and Lognormal distributions, using generalized Pareto for tail inference.
result Lead-time distributions for Nights Booked and Gross Booking Value diverge, with GBV concentrating more in mid-range horizons.

Study reduces emissions in portfolios with error-prone emissions data.

problem Portfolio optimization with firm-level emissions intensities measured inaccurately.
method Introduced a scope-specific penalty operator to rescale asset payoffs based on revenue-normalized emissions intensity.
result Reduces average Scope~1 emissions intensity by roughly 92% while maintaining similar Sharpe ratios.

Study uses RL to optimize global equity portfolios, finds mixed results.

problem Optimizing dynamic portfolio weights across diverse global markets.
method Deep reinforcement learning with Soft Actor-Critic, incorporating various constraints and reward formulations.
result RL strategies achieve competitive performance, but no strategy consistently outperforms Buy and Hold.

New clustering method improves climate data analysis in Lesser Antilles.

problem Inducing undesirable effects in clustering algorithms using Euclidean distance.
method Replacing Euclidean distance with Expert Deviation (ED) based on symmetrized Kullback-Leibler divergence.
result KMS-ED produces more interpretable clusters with better discrimination of daily situations.

RGRR allocates between QQQ and DIA based on relative states, improving Sharpe and CAGR.

problem Optimizing ETF allocation between QQQ and DIA for better risk-adjusted returns.
method Screened relative and macro states, globally screened interactions, fixed position mapping, walk-forward validation.
result RGRR improves Sharpe and CAGR compared to 100% QQQ and 50/50 QQQ-DIA allocations.

New estimators outperform maximum likelihood without hyper-parameter estimation.

problem Improving system identification performance without hyper-parameter estimation.
method Developed generalized Bayes and closed-form biased estimators using excess MSE.
result New estimators have comparable performance to empirical-Bayes-based regularized estimator.

New framework converts offline to online estimation using black-box offline estimators.

problem Convert offline estimation algorithms to online estimation algorithms.
method Oracle-Efficient Online Estimation (OEOE) framework.
result Achieves near-optimal online estimation error via black-box offline estimators.

New estimator reduces variance in discrete random variables.

problem Estimating gradients for discrete random variables with reduced variance.
method Sampling without replacement and Rao-Blackwellization.
result Our estimator is the most consistent gradient estimator across different entropy settings.

SCOPE estimator improves covariance and precision matrix estimation.

problem Estimating covariance and precision matrices accurately.
method Distributionally robust optimization with convex spectral divergence.
result SCOPE estimator reduces spectral bias and improves condition number.

We present a multi-task learning approach to jointly estimate the means of multiple independent data sets. The proposed multi-task averaging (MTA) algorithm results in a convex combination of the single-task maximum likelihood estimates. We derive the optimal minimum risk estimator and the minimax estimator, and show t…

2011-07-21abs ↗pdf ↗

Obtaining more accurate equity value estimates is the starting point for stock selection, value-based indexing in a noisy market, and beating benchmark indices through tactical style rotation. Unfortunately, discounted cash flow, method of comparables, and fundamental analysis typically yield discrepant valuation estim…

2007-07-24abs ↗pdf ↗

The maximum mean discrepancy (MMD) is a kernel-based distance between probability distributions useful in many applications (Gretton et al. 2012), bearing a simple estimator with pleasing computational and statistical properties. Being able to efficiently estimate the variance of this estimator is very helpful to vario…

2019-06-05abs ↗pdf ↗

Stochastic volatility modelling of financial processes has become increasingly popular. The proposed models usually contain a stationary volatility process. We will motivate and review several nonparametric methods for estimation of the density of the volatility process. Both models based on discretely sampled continuo…

2009-10-27abs ↗pdf ↗

This paper reviews SDR methods for multivariate response regression.

problem Handling sufficient dimension reduction for multivariate response regression.
method Characterizes SDR estimators as inverse or forward regression methods.
result Pooled marginal, projective resampling, distance-based, ordinary least squares, partial least squares, and semiparametric SDR estimators are discussed.

Density ratio estimation is a vital tool in both machine learning and statistical community. However, due to the unbounded nature of density ratio, the estimation procedure can be vulnerable to corrupted data points, which often pushes the estimated ratio toward infinity. In this paper, we present a robust estimator wh…

2017-03-09abs ↗pdf ↗

TAKDE optimizes kernel density estimation for real-time dynamic processes.

problem Real-time density estimation in applications like computer vision and signal processing.
method Derives asymptotic mean integrated squared error (AMISE) upper bound for 'sliding window' kernel density estimator and proposes TAKDE as a novel, theoretically optimal estimator.
result TAKDE outperforms other dynamic density estimators in terms of test log-likelihood and runtime.

We introduce two new estimators of the bivariate Hurst exponent in the power-law cross-correlations setting -- the cross-periodogram and local XX-Whittle estimators -- as generalizations of their univariate counterparts. As the spectrum-based estimators are dependent on a part of the spectrum taken into consideration …

2014-08-28abs ↗pdf ↗

New method for fast volatility estimation robust to change points.

problem Robust high-frequency volatility estimation with change points.
method ℓ1-regularized power variation estimators using LARS for sparse estimation and dynamic programming for change point refinement.
result Minimax rates achieved for volatility estimators, providing accurate and smooth forecasts.

ROME improves density estimation for multi-modal, non-normal data.

problem Robust multi-modal density estimation in non-normal, highly correlated distributions.
method ROME uses clustering to segment multi-modal data into uni-modal clusters, then combines KDE estimates for each cluster.
result ROME outperforms state-of-the-art methods and is more robust to various distributions.

Private estimation of many quantiles using differential privacy.

problem Estimating quantiles of a distribution privately.
method Two approaches: 1) Private estimation of empirical quantiles, 2) Uniform density estimation.
result There is a tradeoff between estimating quantiles at specific points and uniformly estimating the quantile function.

Paper proposes robust LAD estimators for 2D sinusoidal model, proving consistency and normality.

problem Estimation of parameters in 2D sinusoidal models with outliers or heavy-tailed noise.
method Least absolute deviation (LAD) estimators for robust parameter estimation.
result Strong consistency and asymptotic normality of LAD estimators for 2D sinusoidal model parameters.

Optimal and safe semi-supervised learning estimator for high-dimensional data.

problem Improving regression parameter estimation with unlabeled data in high-dimensional settings.
method Established minimax lower bound, proposed optimal and safe semi-supervised estimators.
result Optimal semi-supervised estimator achieves the minimax lower bound.