Divide-and-conquer method splits large data sets for efficient analysis.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper analyzes divide-and-conquer estimators for functional linear regression without assuming target function in the RKHS.
In this article, we advance divide-and-conquer strategies for solving the community detection problem in networks. We propose two algorithms which perform clustering on a number of small subgraphs and finally patches the results into a single clustering. The main advantage of these algorithms is that they bring down si…
We propose a novel class of Sequential Monte Carlo (SMC) algorithms, appropriate for inference in probabilistic graphical models. This class of algorithms adopts a divide-and-conquer approach based upon an auxiliary tree-structured decomposition of the model of interest, turning the overall inferential task into a coll…
DiCoLa recursively decomposes causal structure learning for latent variables.
SwISS improves scalability of Bayesian inference for large datasets.
Tuning parameter selection is of critical importance for kernel ridge regression. To this date, data driven tuning method for divide-and-conquer kernel ridge regression (d-KRR) has been lacking in the literature, which limits the applicability of d-KRR for large data sets. In this paper, by modifying the Generalized Cr…
Study risk bounds for distributed ERM with general loss functions and hypothesis spaces.
Paper connects two portfolio methods, HRP and Minimum Variance, revealing their underlying similarity.
We improve kernel ridge regression for skewed responses using oversampling and adaptive partitioning.
We present a parallelized bijective graph matching algorithm that leverages seeds and is designed to match very large graphs. Our algorithm combines spectral graph embedding with existing state-of-the-art seeded graph matching procedures. We justify our approach by proving that modestly correlated, large stochastic blo…
New method for efficient inference in large datasets.
Twinning splits data into fast, statistically similar sets.
Divide-and-conquer framework speeds up black-box inference for large data.
New method merges MCMC samples without distributional assumptions.
This paper presents a class of new algorithms for distributed statistical estimation that exploit divide-and-conquer approach. We show that one of the key benefits of the divide-and-conquer strategy is robustness, an important characteristic for large distributed systems. We establish connections between performance of…
Divide-and-conquer method speeds sparse factorization for large matrices.
A new Fusion method combines multiple distributions efficiently.
We consider the learning of algorithmic tasks by mere observation of input-output pairs. Rather than studying this as a black-box discrete regression problem with no assumption whatsoever on the input-output mapping, we concentrate on tasks that are amenable to the principle of divide and conquer, and study what are it…
Study analyzes crypto asset risk exposures using a divide-and-conquer approach.
For Bayesian computation in big data contexts, the divide-and-conquer MCMC concept splits the whole data set into batches, runs MCMC algorithms separately over each batch to produce samples of parameters, and combines them to produce an approximation of the target distribution. In this article, we embed random forests …
New algorithm groups variables by ancestral relationships to improve causal graph estimation accuracy.
DC-NAS improves neural architecture search by clustering and evaluating sub-networks.
Develops a new method to recover large latent tree models efficiently.
SDSR reconstructs species trees from genetic markers efficiently.
Active learning (AL) repeatedly trains the classifier with the minimum labeling budget to improve the current classification model. The training process is usually supervised by an uncertainty evaluation strategy. However, the uncertainty evaluation always suffers from performance degeneration when the initial labeled …
The deep layers of modern neural networks extract a rather rich set of features as an input propagates through the network. This paper sets out to harvest these rich intermediate representations for quantization with minimal accuracy loss while significantly reducing the memory footprint and compute intensity of the DN…
Paper speeds up Gaussian process inference using Matérn kernels.
New method quantifies uncertainty in distributed regression.
In this paper, a new method is proposed for sparse PCA based on the recursive divide-and-conquer methodology. The main idea is to separate the original sparse PCA problem into a series of much simpler sub-problems, each having a closed-form solution. By recursively solving these sub-problems in an analytical way, an ef…
We show how to use Bar-Natan's `divide and conquer' approach to computations to efficiently compute the universal sl(2) dotted foam cohomology groups, even for big knots and links. We also describe a purely topological version of the sl(2) foam theory, in the sense that no dots are needed on foams.
A new method for sampling complex posterior distributions in DDMs.
Two new scalable K-means initialization methods proposed for large-scale clustering.
This paper introduces a divide-and-conquer inspired adversarial learning (DACAL) approach for photo enhancement. The key idea is to decompose the photo enhancement process into hierarchically multiple sub-problems, which can be better conquered from bottom to up. On the top level, we propose a perception-based division…
A new predictive coding algorithm improves machine learning performance.
New algorithm ranks players from partial comparisons with optimal rate.
Dual-T method improves transition matrix estimation in noisy label learning.
We reduce a broad class of machine learning problems, usually addressed by EM or sampling, to the problem of finding the extremal rays spanning the conical hull of a data point set. These "anchors" lead to a global solution and a more interpretable model that can even outperform EM and sampling on generalizatio…
Novel hybrid method for Bayesian network structure learning reduces computational time without sacrificing accuracy.
uMoE trains NNs with uncertain data by embedding uncertainty into training.
We present a novel cross-view classification algorithm where the gallery and probe data come from different views. A popular approach to tackle this problem is the multi-view subspace learning (MvSL) that aims to learn a latent subspace shared by multi-view data. Despite promising results obtained on some applications,…
Paper introduces PTL-SI for statistical inference in TL-HDR, controlling FPR.
We use the divide-and-conquer and scanning algorithms for calculating Khovanov cohomology directly on the Lee- or Bar-Natan deformations of the Khovanov complex to give an alternative way to compute Rasmussen -invariants of knots. By disregarding generators away from homological degree 0 we can considerably improve …
Efficient algorithm for Bayesian networks reduces marginal probability distribution computation.
Study metric learning from limited preference comparisons, showing how low-dimensional structure can still reveal metric information.
The paper develops efficient estimators for semi-parametric binary models in distributed computing.
Faced with the growing research towards crude oil price fluctuations influential factors following the accelerated development of Internet technology, accessible data such as Google search volume index are increasingly quantified and incorporated into forecasting approaches. In this paper, we apply multi-scale data tha…
Divide-and-conquer is a general strategy to deal with large scale problems. It is typically applied to generate ensemble instances, which potentially limits the problem size it can handle. Additionally, the data are often divided by random sampling which may be suboptimal. To address these concerns, we propose the $DC^…