Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,236 papers · 148 categories

Trend · papers per month

0.3%0.5%0.8%0.9% · Jan 200419922001200920182026
48 results for Sub-Sampling

Active covariance estimation using random sub-sampling of variable subsets.

problem Estimating covariance matrices for partially observed random vectors.
method Unbiased covariance estimator under a model of partially observed variables and active learning framework.
result Derivation of error bounds revealing relations between sub-sampling probabilities and covariance matrix entries.

Paper analyzes Nyström regularization for time series forecasting with sequential sub-sampling.

problem Learning rate analysis of Nyström regularization for ττ-mixing time series.
method Banach-valued Bernstein inequality and integral operator approach for ττ-mixing sequences.
result Almost optimal learning rates for Nyström regularization with sequential sub-sampling.

LOUPE optimizes MRI sub-sampling patterns using machine learning.

problem Optimizing sub-sampling patterns for MRI scans to improve reconstruction accuracy.
method End-to-end learning strategy combining sub-sampling pattern optimization and reconstruction model training.
result LOUPE yields more accurate reconstructions compared to standard under-sampling schemes.

Efficient bandit exploration for various distributions without distribution-specific tuning.

problem Optimizing exploration in multi-armed bandit models for different distributions.
method Sub-sampling Duelling Algorithms (SDA) with Random Block sampling for efficient exploration.
result Achieves asymptotically optimal regret for Bernoulli, Gaussian, and Poisson distributions.

Many data-fitting applications require the solution of an optimization problem involving a sum of large number of functions of high dimensional parameter. Here, we consider the problem of minimizing a sum of nn functions over a convex constraint set XRp\mathcal{X} \subseteq \mathbb{R}^{p} where both nn and pp are lar…

2016-01-18abs ↗pdf ↗

Large scale optimization problems are ubiquitous in machine learning and data analysis and there is a plethora of algorithms for solving such problems. Many of these algorithms employ sub-sampling, as a way to either speed up the computations and/or to implicitly implement a form of statistical regularization. In this …

2016-01-18abs ↗pdf ↗

This study addresses the challenges of dynamic mini-batch sub-sampling in neural network training.

problem Challenges in training neural networks due to dynamic mini-batch sub-sampling.
method Distinguishes between static and dynamic sub-sampling, recasting optimization to find SNN-GPPs.
result SNN-GPPs are less susceptible to sub-sampling-induced discontinuities and better approximate true optima.

This paper shows using sub-sample estimates can improve optimization results in large-scale problems.

problem Large-scale optimization problems with uncertain parameters often lead to suboptimal solutions due to mis-specifications or extreme sample characteristics.
method The paper introduces the use of sub-sample estimates to reduce errors in stochastic optimization models, providing theoretical analysis and numerical examples.
result Sub-sample optimization can achieve improved results over full-sample solution estimates in large-scale problems.

New method uses PDMPs with sub-sampling for efficient sampling from posterior distributions.

problem Efficient sampling from posterior distributions with limited data access.
method Approximate simulation of PDMPs with sub-sampling and stochastic gradient estimation.
result Stochastic-gradient PDMPs are efficient and robust compared to Langevin dynamics.

Efficiently estimate risk of large portfolios using MLMC and sub-sampling.

problem Estimating risk of large portfolios with high computational cost.
method Apply Multilevel Monte Carlo (MLMC) with adaptive inner sampling and sub-sampling strategy.
result Sub-sampling strategy reduces computational complexity without portfolio size increase.

We consider the problem of minimizing a sum of nn functions over a convex parameter set CRp\mathcal{C} \subset \mathbb{R}^p where np1n\gg p\gg 1. In this regime, algorithms which utilize sub-sampling techniques are known to be effective. In this paper, we use sub-sampling techniques together with low-rank approximation …

2015-08-12abs ↗pdf ↗

Paper introduces efficient online sub-sampling for RL with function approximation, reducing policy updates.

problem Efficiently managing computation complexity in RL with general function approximation.
method Online sub-sampling framework that measures information gain and guides exploration.
result Policy updates reduced to polylog(K)\propto\operatorname{poly}\log(K) times for near-optimal regret.

The standard approach to compressive sampling considers recovering an unknown deterministic signal with certain known structure, and designing the sub-sampling pattern and recovery algorithm based on the known structure. This approach requires looking for a good representation that reveals the signal structure, and sol…

2016-02-01abs ↗pdf ↗

New methods for non-convex optimization using inexact Hessian approximations.

problem Optimization of non-convex functions with inexact Hessian information.
method Trust-region and cubic regularization methods with inexact Hessian approximations.
result Iteration complexity to achieve ε-approximate second-order optimality.

Efficiently approximates statistical leverage scores for faster KRR.

problem Accurately estimating statistical leverage scores for fast KRR.
method Analytic formula for statistical leverage scores, leveraging kernel spectral density.
result Linear time approximation with theoretical guarantees, significantly faster than existing methods.

Adapts to high dimensions for estimating conditional moments.

problem Estimation and inference in high-dimensional settings with unknown intrinsic dimension.
method Sub-sampled kk-NN ZZ-estimator, adaptive data-driven sub-sampling.
result Estimation error of n1/(d+2)n^{-1/(d+2)} and asymptotic normality with n1/(d+2)n^{1/(d+2)} rate.

Improved MCMC for rare events in hidden Markov models.

problem Slow inference and prediction for rare latent states in hidden Markov models.
method Targeted sub-sampling (TASS) over-samples rare latent states, reducing variance in gradient estimation.
result Substantial gains in predictive and inferential accuracy on real and synthetic examples.

Bayesian realized EGARCH models improve tail risk forecasting.

problem Forecasting tail risks in financial markets.
method Developed a Bayesian framework for realized EGARCH models, incorporating multiple realized volatility measures and using robust adaptive Metropolis algorithm for estimation.
result Standardized skewed Student-t distribution and sub-sampled realized range models outperform other models in tail risk forecasting.

Bayesian realized-GARCH models forecast financial tail risks using two-sided Weibull distribution.

problem Forecasting financial tail risks in volatile markets.
method Adaptive Bayesian Markov Chain Monte Carlo for estimation and forecasting, incorporating sub-sampled realized range and variance.
result Realized-GARCH models with two-sided Weibull distribution outperform other models in tail risk forecasting.

End-to-end analysis of SGD for STL with adaptive sub-sampling.

problem Designing SGD for STL with statistical guarantees without prior knowledge of source quality.
method Mixed-sample SGD procedure that alternates between source and target data, maintaining transfer guarantees.
result Mixed-sample SGD converges to a target-adaptive solution with 1/T1/\sqrt{T} rate.

Paper proposes an efficient bandit-based algorithm for hyperparameter optimization.

problem Efficiently evaluating hyperparameters in deep learning models with large search spaces.
method Sub-Sampling (SS) algorithm combined with Bayesian Optimization (BOSS).
result Theoretical proof of optimality and empirical validation of superior performance.

A novel distributed adaptive NN classifier for large data sets.

problem Handling large and distributed data for efficient classification.
method Distributed adaptive nearest neighbor classifier with stochastic tuning parameter selection and early stopping rule.
result Achieves nearly optimal convergence rate under large sub-sample sizes.

This paper enhances stability selection by evaluating overall results robustness and identifying optimal regularization values.

problem Improving the robustness and reliability of high-dimensional variable selection.
method Developed a stability estimator to evaluate stability of stability selection results, calibrating key parameters.
result Identified optimal regularization value and improved stability of variable selection.

This paper compares and analyzes random projections and column sub-sampling for dimension reduction in regression.

problem Computational efficiency in dimension reduction for large datasets.
method Analysis of random projections and column sub-sampling methods for regression.
result Random projections and column sub-sampling can achieve similar prediction error to Principal Components Regression (PCR) but with less computational cost.

Under-bagging kk-NN improves performance on imbalanced classification.

problem Imbalanced classification problems where one class is significantly underrepresented.
method Proposes an under-bagging kk-NN ensemble learning algorithm, analyzing convergence rates and efficiency.
result Achieves optimal convergence rates under mild assumptions and reduces sub-sample size and kk for highly imbalanced data.

Improved tail risk forecasting using realized measures and Bayesian methods.

problem Forecasting tail risk in financial markets.
method Extended Taylor's model with realized measures, using maximum likelihood and Bayesian MCMC methods.
result Bayesian approach outperforms maximum likelihood in forecasting accuracy.

Novel Newton method for large-scale kernel methods using random features.

problem Efficiently solving large-scale finite-sum minimization problems in RKHS.
method Randomized feature-based Newton method for empirical risk minimization.
result Local superlinear and global linear convergence of the method.

A new method reduces computational cost for gene expression inference in large microarray data sets.

problem Efficiently predicting gene expression in large datasets with limited resources.
method Adaptive Lipschitz constant inspired learning rate, random sub-sampling, and A-ReLU activation function.
result Remarkable improvement in saving computational cost while maintaining prediction accuracy.

This paper tackles efficient and scalable estimation of a complex model involving stochastic linear combinations of non-linear regressions.

problem Estimating a model involving stochastic linear combinations of non-linear regressions efficiently and scalably.
method The paper provides algorithms for estimating the model under specific assumptions about the variate vector and sample size, using techniques like zero-bias transformation and sub-sampling.
result The paper provides theoretical guarantees for the estimation of the model, showing that the estimation errors are of the order O(pn)O(\sqrt{\frac{p}{n}}) and O(1p+pn)O(\frac{1}{\sqrt{p}}+\sqrt{\frac{p}{n}}) with high probability.

Proposes α\alpha-integration pooling for CNNs to improve performance.

problem Finding optimal pooling method for CNNs is challenging.
method Introduces α\alpha-integration pooling with a trainable parameter α\alpha.
result Demonstrates α\alpha-integration pooling outperforms other pooling methods in image recognition.

This paper proposes an improved active learning method using classification trees.

problem Reducing the size of training sets while maintaining high accuracy in supervised learning.
method A wrapper active learning method using a classification tree to sub-sample from low-entropy regions.
result The proposed method constructs accurate classification models even with severely restricted labeled data.

Study examines line search approximations for neural networks using MBSS.

problem Reducing computational cost in training large-scale neural networks.
method Empirical study of quadratic line search approximations for dynamic MBSS loss functions, enforcing different types of function and derivative information.
result Selectively enforcing information in approximations reduces the variance of predicted step sizes.