Divide-and-conquer method splits large data sets for efficient analysis.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper analyzes divide-and-conquer estimators for functional linear regression without assuming target function in the RKHS.
In this article, we advance divide-and-conquer strategies for solving the community detection problem in networks. We propose two algorithms which perform clustering on a number of small subgraphs and finally patches the results into a single clustering. The main advantage of these algorithms is that they bring down si…
We propose a novel class of Sequential Monte Carlo (SMC) algorithms, appropriate for inference in probabilistic graphical models. This class of algorithms adopts a divide-and-conquer approach based upon an auxiliary tree-structured decomposition of the model of interest, turning the overall inferential task into a coll…
DiCoLa recursively decomposes causal structure learning for latent variables.
Divide-and-conquer framework speeds up black-box inference for large data.
SwISS improves scalability of Bayesian inference for large datasets.
Tuning parameter selection is of critical importance for kernel ridge regression. To this date, data driven tuning method for divide-and-conquer kernel ridge regression (d-KRR) has been lacking in the literature, which limits the applicability of d-KRR for large data sets. In this paper, by modifying the Generalized Cr…
Study risk bounds for distributed ERM with general loss functions and hypothesis spaces.
Paper connects two portfolio methods, HRP and Minimum Variance, revealing their underlying similarity.
We improve kernel ridge regression for skewed responses using oversampling and adaptive partitioning.
We present a parallelized bijective graph matching algorithm that leverages seeds and is designed to match very large graphs. Our algorithm combines spectral graph embedding with existing state-of-the-art seeded graph matching procedures. We justify our approach by proving that modestly correlated, large stochastic blo…
New method for efficient inference in large datasets.
Twinning splits data into fast, statistically similar sets.
New method merges MCMC samples without distributional assumptions.
Divide-and-conquer is a general strategy to deal with large scale problems. It is typically applied to generate ensemble instances, which potentially limits the problem size it can handle. Additionally, the data are often divided by random sampling which may be suboptimal. To address these concerns, we propose the $DC^…
This paper presents a class of new algorithms for distributed statistical estimation that exploit divide-and-conquer approach. We show that one of the key benefits of the divide-and-conquer strategy is robustness, an important characteristic for large distributed systems. We establish connections between performance of…
Divide-and-conquer method speeds sparse factorization for large matrices.
A new Fusion method combines multiple distributions efficiently.
We consider the learning of algorithmic tasks by mere observation of input-output pairs. Rather than studying this as a black-box discrete regression problem with no assumption whatsoever on the input-output mapping, we concentrate on tasks that are amenable to the principle of divide and conquer, and study what are it…
Study analyzes crypto asset risk exposures using a divide-and-conquer approach.
For Bayesian computation in big data contexts, the divide-and-conquer MCMC concept splits the whole data set into batches, runs MCMC algorithms separately over each batch to produce samples of parameters, and combines them to produce an approximation of the target distribution. In this article, we embed random forests …
Active learning (AL) repeatedly trains the classifier with the minimum labeling budget to improve the current classification model. The training process is usually supervised by an uncertainty evaluation strategy. However, the uncertainty evaluation always suffers from performance degeneration when the initial labeled …
New algorithm groups variables by ancestral relationships to improve causal graph estimation accuracy.
This paper introduces a divide-and-conquer inspired adversarial learning (DACAL) approach for photo enhancement. The key idea is to decompose the photo enhancement process into hierarchically multiple sub-problems, which can be better conquered from bottom to up. On the top level, we propose a perception-based division…
DC-NAS improves neural architecture search by clustering and evaluating sub-networks.
The deep layers of modern neural networks extract a rather rich set of features as an input propagates through the network. This paper sets out to harvest these rich intermediate representations for quantization with minimal accuracy loss while significantly reducing the memory footprint and compute intensity of the DN…
If learning methods are to scale to the massive sizes of modern datasets, it is essential for the field of machine learning to embrace parallel and distributed computing. Inspired by the recent development of matrix factorization methods with rich theory but poor computational complexity and by the relative ease of map…
New algorithm ranks players from partial comparisons with optimal rate.
Develops a new method to recover large latent tree models efficiently.
Universal probabilistic programming systems (PPSs) provide a powerful framework for specifying rich probabilistic models. They further attempt to automate the process of drawing inferences from these models, but doing this successfully is severely hampered by the wide range of non--standard models they can express. As …
We present a novel cross-view classification algorithm where the gallery and probe data come from different views. A popular approach to tackle this problem is the multi-view subspace learning (MvSL) that aims to learn a latent subspace shared by multi-view data. Despite promising results obtained on some applications,…
SDSR reconstructs species trees from genetic markers efficiently.
New method quantifies uncertainty in distributed regression.
Paper speeds up Gaussian process inference using Matérn kernels.
In this paper, a new method is proposed for sparse PCA based on the recursive divide-and-conquer methodology. The main idea is to separate the original sparse PCA problem into a series of much simpler sub-problems, each having a closed-form solution. By recursively solving these sub-problems in an analytical way, an ef…
We show how to use Bar-Natan's `divide and conquer' approach to computations to efficiently compute the universal sl(2) dotted foam cohomology groups, even for big knots and links. We also describe a purely topological version of the sl(2) foam theory, in the sense that no dots are needed on foams.
A new method for sampling complex posterior distributions in DDMs.
Two new scalable K-means initialization methods proposed for large-scale clustering.
We study the risk performance of distributed learning for the regularization empirical risk minimization with fast convergence rate, substantially improving the error analysis of the existing divide-and-conquer based distributed learning. An interesting theoretical finding is that the larger the diversity of each local…
A new predictive coding algorithm improves machine learning performance.
Dual-T method improves transition matrix estimation in noisy label learning.
We reduce a broad class of machine learning problems, usually addressed by EM or sampling, to the problem of finding the extremal rays spanning the conical hull of a data point set. These "anchors" lead to a global solution and a more interpretable model that can even outperform EM and sampling on generalizatio…
SURF simplifies distribution estimation with simple, robust, and fast algorithms.
Novel hybrid method for Bayesian network structure learning reduces computational time without sacrificing accuracy.
uMoE trains NNs with uncertain data by embedding uncertainty into training.
Hybrid approach improves crude oil price forecasting using multi-scale data.
Paper introduces PTL-SI for statistical inference in TL-HDR, controlling FPR.