Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

0111 · Jun 201219922001200920182026
11 results for out-of-core

Improved Bayesian network classifiers using HDPs for better parameter estimation.

problem Inaccurate parameter estimation in Bayesian network classifiers limits their performance.
method Hierarchical Dirichlet Processes (HDPs) for accurate parameter estimation.
result HDPs improve BNCs' performance, matching or outperforming Random Forest on categorical datasets.

biglasso solves memory and computation issues for lasso models on large data.

problem Memory and computational limitations in fitting lasso models to large datasets.
method Memory-mapped files, out-of-core computation, efficient feature screening rules.
result biglasso efficiently handles ultrahigh-dimensional, multi-gigabyte data sets.

We present RandomizedCCA, a randomized algorithm for computing canonical analysis, suitable for large datasets stored either out of core or on a distributed file system. Accurate results can be obtained in as few as two data passes, which is relevant for distributed processing frameworks in which iteration is expensive…

2014-11-13abs ↗pdf ↗

Efficient kernel methods for large datasets using GPU acceleration.

problem Handling large-scale nonparametric learning problems efficiently.
method Preconditioned gradient solver, GPU acceleration, parallelization, out-of-core linear algebra, numerical precision optimization.
result Dramatic speedups on datasets with billions of points, maintaining state-of-the-art performance.

We propose a new analytical approximation to the χ2χ^2 kernel that converges geometrically. The analytical approximation is derived with elementary methods and adapts to the input distribution for optimal convergence rate. Experiments show the new approximation leads to improved performance in image classification and …

2012-06-18abs ↗pdf ↗

Residual Networks are shown to be equivalent to boosting feature representation.

problem Improving feature representation in deep learning models.
method Proved ResNet's equivalence to Online Gradient Boosting and proposed decision tree residual modules.
result ResNet can achieve Online Gradient Boosting regret bounds through architectural changes.

Stabilizes online learning by using weighted reservoir sampling.

problem Real-world deployment sensitivity to outliers causes low accuracy in final solutions.
method Weighted reservoir sampling to stabilize ensemble model without additional data passes.
result Risk of ensemble classifier is bounded with respect to the underlying online learning method's regret.