New method unifies and formalizes data partitioning using a single vector.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Differentially private method for synthetic data generation from vertically partitioned data.
Efficiently calculates PL model likelihood for partitioned preference data.
Asynchronous federated learning for vertically partitioned data improves efficiency and privacy.
This research designs a data-driven partition to test independence between continuous variables.
Stochastic partition models divide a multi-dimensional space into a number of rectangular regions, such that the data within each region exhibit certain types of homogeneity. Due to the nature of their partition strategy, existing partition models may create many unnecessary divisions in sparse regions when trying to d…
The excessively increased volume of data in modern data management systems demands an improved system performance, frequently provided by data distribution, system scalability and performance optimization techniques. Optimized horizontal data partitioning has a significant influence of distributed data management syste…
This paper presents Sparse Partitioning, a Bayesian method for identifying predictors that either individually or in combination with others affect a response variable. The method is designed for regression problems involving binary or tertiary predictors and allows the number of predictors to exceed the size of the sa…
This work proposes a method to optimize hyperparameters without validation data.
Extends partitioned local depth concept with probabilistic considerations.
Space partitions of underlie a vast and important class of fast nearest neighbor search (NNS) algorithms. Inspired by recent theoretical work on NNS for general metric spaces [Andoni, Naor, Nikolov, Razenshteyn, Waingarten STOC 2018, FOCS 2018], we develop a new framework for building space partitions re…
SplitNN-driven Vertical Partitioning enables distributed learning from diverse data sources.
The fundamental aim of clustering algorithms is to partition data points. We consider tasks where the discovered partition is allowed to vary with some covariate such as space or time. One approach would be to use fragmentation-coagulation processes, but these, being Markov processes, are restricted to linear or tree s…
Online BSP-Forest improves space partitioning for large-scale classification and regression.
FDSKL algorithm trains vertically partitioned data with kernels securely and efficiently.
Bayesian approach for multifile record linkage and duplicate detection.
Distributed sparse learning with a cluster of multiple machines has attracted much attention in machine learning, especially for large-scale applications with high-dimensional data. One popular way to implement sparse learning is to use regularization. In this paper, we propose a novel method, called proximal \mb…
Bayesian nonparametric method partitions shapes using curves.
Improved supervised EM learning for shared kernel models with feature space partitioning.
Improved neural network robustness certification through tighter convex relaxations.
We explore the geometrical interpretation of the PCA based clustering algorithm Principal Direction Divisive Partitioning (PDDP). We give several examples where this algorithm breaks down, and suggest a new method, gap partitioning, which takes into account natural gaps in the data between clusters. Geometric features …
Partition Tree estimates conditional densities for mixed continuous and categorical variables.
The beta-negative binomial process (BNBP), an integer-valued stochastic process, is employed to partition a count vector into a latent random count matrix. As the marginal probability distribution of the BNBP that governs the exchangeable random partitions of grouped data has not yet been developed, current inference f…
Unsupervised space partitioning improves ANNS performance without pre-processing.
New classifiers converge under large data, simplifying complex models.
In this paper, we investigate a divide and conquer approach to Kernel Ridge Regression (KRR). Given n samples, the division step involves separating the points based on some underlying disjoint partition of the input space (possibly via clustering), and then computing a KRR estimate for each partition. The conquering s…
We construct new multivariate copulas on the basis of a generalized infinite partition-of-unity approach. This approach allows - in contrast to finite partition-of-unity copulas - for tail-dependence as well as for asymmetry. A possibility of fitting such copulas to real data from quantitative risk management is also p…
Comparing and aligning large datasets is a pervasive problem occurring across many different knowledge domains. We introduce and study MREC, a recursive decomposition algorithm for computing matchings between data sets. The basic idea is to partition the data, match the partitions, and then recursively match the points…
Study on estimating Gaussian mean from coarse data, resolving identifiability and computational efficiency questions.
Paper recovers lattice signal partitions efficiently.
This paper addresses clustering with missing data using Rubin's rules.
We present a new approach for learning compact and intuitive distributed representations with binary encoding. Rather than summing up expert votes as in products of experts, we employ for each variable the opinion of the most reliable expert. Data points are hence explained through a partitioning of the variables into …
We present a constructive and self-contained approach to data driven general partition-of-unity copulas that were recently introduced in the literature. In particular, we consider Bernstein-, negative binomial and Poisson copulas and present a solution to the problem of fitting such copulas to highly asymmetric data.
Backdoor attacks are possible in feature-partitioned collaborative learning, even without labels.
In this paper we propose a novel Bayesian methodology for Value-at-Risk computation based on parametric Product Partition Models. Value-at-Risk is a standard tool to measure and control the market risk of an asset or a portfolio, and it is also required for regulatory purposes. Its popularity is partly due to the fact …
We present a new way of constructing an ensemble classifier, named the Guided Random Forest (GRAF) in the sequel. GRAF extends the idea of building oblique decision trees with localized partitioning to obtain a global partitioning. We show that global partitioning bridges the gap between decision trees and boosting alg…
Mixed datasets consist of both numeric and categorical attributes. Various k-means-based clustering algorithms have been developed for these datasets. Generally, these algorithms use random partition as a starting point, which tends to produce different clustering results for different runs. In this paper, we propose, …
This study proposes a graph partitioning method to improve spatial prediction models.
Space partitioning methods such as random forests and the Mondrian process are powerful machine learning methods for multi-dimensional and relational data, and are based on recursively cutting a domain. The flexibility of these methods is often limited by the requirement that the cuts be axis aligned. The Ostomachion p…
We present Random Partition Kernels, a new class of kernels derived by demonstrating a natural connection between random partitions of objects and kernels between those objects. We show how the construction can be used to create kernels from methods that would not normally be viewed as random partitions, such as Random…
Clustering is an essential technique for discovering patterns in data. The steady increase in amount and complexity of data over the years led to improvements and development of new clustering algorithms. However, algorithms that can cluster data with mixed variable types (continuous and categorical) remain limited, de…
Fitting statistical models is computationally challenging when the sample size or the dimension of the dataset is huge. An attractive approach for down-scaling the problem size is to first partition the dataset into subsets and then fit using distributed algorithms. The dataset can be partitioned either horizontally (i…
Survey of Bayesian nonparametric space partition models and their applications.
We study the problem of learning to choose from m discrete treatment options (e.g., news item or medical drug) the one with best causal effect for a particular instance (e.g., user or patient) where the training data consists of passive observations of covariates, treatment, and the outcome of the treatment. The standa…
The method learns to partition event time space for better prediction.
A new learning rule consistently reduces error over data samples.
Chern-Simons theory on a closed contact three-manifold is studied when the Lie group for gauge transformations is compact, connected and abelian. A rigorous definition of an abelian Chern-Simons partition function is derived using the Faddeev-Popov gauge fixing method. A symplectic abelian Chern-Simons partition functi…
Recently, there has been an increasing interest in designing distributed convex optimization algorithms under the setting where the data matrix is partitioned on features. Algorithms under this setting sometimes have many advantages over those under the setting where data is partitioned on samples, especially when the …