Random Tessellation Process improves multi-dimensional data analysis.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Online BSP-Forest improves space partitioning for large-scale classification and regression.
A new tree-based method for adaptive dictionary learning.
The Mondrian process represents an elegant and powerful approach for space partition modelling. However, as it restricts the partitions to be axis-aligned, its modelling flexibility is limited. In this work, we propose a self-consistent Binary Space Partitioning (BSP)-Tree process to generalize the Mondrian process. Th…
Partition Tree estimates conditional densities for mixed continuous and categorical variables.
This paper introduces the Partition Tree Weighting technique, an efficient meta-algorithm for piecewise stationary sources. The technique works by performing Bayesian model averaging over a large class of possible partitions of the data into locally stationary segments. It uses a prior, closely related to the Context T…
We introduce a comprehensive and statistical framework in a model free setting for a complete treatment of localized data corruptions due to severe noise sources, e.g., an occluder in the case of a visual recording. Within this framework, we propose i) a novel algorithm to efficiently separate, i.e., detect and localiz…
Recent theory work has found that a special type of spatial partition tree - called a random projection tree - is adaptive to the intrinsic dimension of the data from which it is built. Here we examine this same question, with a combination of theory and experiments, for a broader class of trees that includes k-d trees…
Efficient algorithm for matching graphs with community structure.
The modern scale of data has brought new challenges to Bayesian inference. In particular, conventional MCMC algorithms are computationally very expensive for large data sets. A promising approach to solve this problem is embarrassingly parallel MCMC (EP-MCMC), which first partitions the data into multiple subsets and r…
Approximate nearest neighbor algorithms are used to speed up nearest neighbor search in a wide array of applications. However, current indexing methods feature several hyperparameters that need to be tuned to reach an acceptable accuracy--speed trade-off. A grid search in the parameter space is often impractically slow…
Paper solves graph matching for correlated Erdős--Rényi graphs.
Undirected graphical models encode in a graph the dependency structure of a random vector . In many applications, it is of interest to model given another random vector as input. We refer to the problem of estimating the graph of conditioned on as ``graph-valued regression.'' In this pap…
We propose a multiresolution Gaussian process to capture long-range, non-Markovian dependencies while allowing for abrupt changes. The multiresolution GP hierarchically couples a collection of smooth GPs, each defined over an element of a random nested partition. Long-range dependencies are captured by the top-level GP…
Efficient index structures for fast approximate nearest neighbor queries are required in many applications such as recommendation systems. In high-dimensional spaces, many conventional methods suffer from excessive usage of memory and slow response times. We propose a method where multiple random projection trees are c…
Supervised Machine Learning (SML) algorithms such as Gradient Boosting, Random Forest, and Neural Networks have become popular in recent years due to their increased predictive performance over traditional statistical methods. This is especially true with large data sets (millions or more observations and hundreds to t…
Study calculates Floer homology for binary polyhedral spaces.
DiFF-RF detects point-wise and collective anomalies using random partitioning trees.
Proposes HypCSE for enhanced hierarchical clustering.
The moduli space of smooth real binary octics has five connected components. They parametrize the real binary octics whose defining equations have 0, 1, ..., 4 complex-conjugate pairs of roots respectively. We show that the GIT-stable completion of each of these five components admits the structure of an arithmetic rea…
FSPT identifies training space to prevent ML model extrapolation.
Bayesian optimization for high-dimensional combinatorial spaces using embeddings.
New BDEs reveal singular surfaces from line congruences.
Study improves learning algorithms for convex polyhedra in Hilbert spaces.
Proposes MELODIC family for simultaneous binary logistic regression.
An attractive approach for fast search in image databases is binary hashing, where each high-dimensional, real-valued image is mapped onto a low-dimensional, binary vector and the search is done in this binary space. Finding the optimal hash function is difficult because it involves binary constraints, and most approac…
Study links K-stability of certain surfaces to binary forms, proving stability and non-stability conditions.
Quantum machine learning model for binary classification.
Observational data hints at a finite universe, with spherical manifolds such as the Poincare dodecahedral space tentatively providing the best fit. Simulating the physics of a model universe requires knowing the eigenmodes of the Laplace operator on the space. The present article provides explicit polynomial eigenmodes…
New method uses binary quadratic forms to classify Seifert surfaces in 4-ball.
We introduce a multiscale supervised dimension reduction method for SPatial Interaction Network (SPIN) data, which consist of a collection of spatially coordinated interactions. This type of predictor arises when the sampling unit of data is composed of a collection of primitive variables, each of them being essentiall…
Binary representation is desirable for its memory efficiency, computation speed and robustness. In this paper, we propose adjustable bounded rectifiers to learn binary representations for deep neural networks. While hard constraining representations across layers to be binary makes training unreasonably difficult, we s…
Binary embedding of high-dimensional data requires long codes to preserve the discriminative power of the input space. Traditional binary coding methods often suffer from very high computation and storage costs in such a scenario. To address this problem, we propose Circulant Binary Embedding (CBE) which generates bina…
Two binary Sine Cosine Algorithms improve feature selection in medical datasets.
Paper studies binary random projections with controllable sparsity patterns for computational and accuracy advantages.
Classifies singularities in quiver varieties for specific Dynkin quivers.
Symmetric binary matrices representing relations among entities are commonly collected in many areas. Our focus is on dynamically evolving binary relational matrices, with interest being in inference on the relationship structure and prediction. We propose a nonparametric Bayesian dynamic model, which reduces dimension…
In distributed machine learning, data is dispatched to multiple machines for processing. Motivated by the fact that similar data points often belong to the same or similar classes, and more generally, classification rules of high accuracy tend to be "locally simple but globally complex" (Vapnik & Bottou 1993), we propo…
We generalize Conway's approach to integral binary quadratic forms on Q to study integral binary hermitian forms on quadratic imaginary extensions of Q. In Conway's case, an indefinite form that doesn't represent 0 determines a line ("river") in the spine T associated with SL(2,Z) in the hyperbolic plane. In our genera…
A new algebraic structure emerges from reductive homogeneous spaces.
We give a graphical theory of integral indefinite binary Hamiltonian forms analogous to the one by Conway for binary quadratic forms and the one of Bestvina-Savin for binary Hermitian forms. Given a maximal order in a definite quaternion algebra over , we define the waterworld of , analog…
Lower bound for VAE training objective for binary data.
GMBL uses graph embedding to learn binary codes from multiple views for clustering.
Privacy-preserving binary classification using locally differential private data.
Lectures explore how differential methods improve understanding of algebraic group orbit spaces.
This work extends score-based methods to binary data on the Boolean hypercube.
Defines computable learning for binary classification over metric spaces.
A new method for binary ICA using non-stationary sources.