A new method for multi-task learning improves performance without weakening inductive bias.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
DEEPLY improves cloud service partitioning across multiple datasets and goals.
Many popular random partition models, such as the Chinese restaurant process and its two-parameter extension, fall in the class of exchangeable random partitions, and have found wide applicability in model-based clustering, population genetics, ecology or network analysis. While the exchangeability assumption is sensib…
We propose an algorithm, HPREF (Hierarchical Partitioning by Repeated Features), that produces a hierarchical partition of a set of clusterings of a fixed dataset, such as sets of clusterings produced by running a clustering algorithm with a range of parameters. This gives geometric structure to such sets of clustering…
Fitting statistical models is computationally challenging when the sample size or the dimension of the dataset is huge. An attractive approach for down-scaling the problem size is to first partition the dataset into subsets and then fit using distributed algorithms. The dataset can be partitioned either horizontally (i…
MPF method improves parameter estimation in probabilistic models.
Graph partitioning is the problem of dividing the nodes of a graph into balanced partitions while minimizing the edge cut across the partitions. Due to its combinatorial nature, many approximate solutions have been developed, including variants of multi-level methods and spectral clustering. We propose GAP, a Generaliz…
Homology of partition algebras matches symmetric group homology under certain conditions.
Study optimal partitions on spheres using fractional Q-curvature and variational methods.
Spectral clustering is sensitive to how graphs are constructed from data particularly when proximal and imbalanced clusters are present. We show that Ratio-Cut (RCut) or normalized cut (NCut) objectives are not tailored to imbalanced data since they tend to emphasize cut sizes over cut values. We propose a graph partit…
Novel method recursively partitions sample space for density estimation.
The thermodynamics of markets is analyzed, revealing parallels with thermodynamic laws.
Study on optimal partitions and nodal solutions for the Yamabe equation.
New examples show non-rotational annuli in a ball, solving a uniqueness problem.
This paper tackles multi-modal label disentanglement in partition-based XMC.
Efficiently resolves entities via scaled Ewens--Pitman model.
Spectral clustering methods which are frequently used in clustering and community detection applications are sensitive to the specific graph constructions particularly when imbalanced clusters are present. We show that ratio cut (RCut) or normalized cut (NCut) objectives are not tailored to imbalanced cluster sizes sin…
We propose moment-based variational inference as a flexible framework for approximate smoothing of latent Markov jump processes. The main ingredient of our approach is to partition the set of all transitions of the latent process into classes. This allows to express the Kullback-Leibler divergence between the approxima…
Backdoor attacks are possible in feature-partitioned collaborative learning, even without labels.
This research designs a data-driven partition to test independence between continuous variables.
Traditionally, neural networks are parameterized using optimization procedures such as stochastic gradient descent, RMSProp and ADAM. These procedures tend to drive the parameters of the network toward a local minimum. In this article, we employ alternative "sampling" algorithms (referred to here as "thermodynamic para…
Lectures detail field theory dynamics and exact WKB analysis.
This work proposes a method to optimize hyperparameters without validation data.
Nonparametric regression for massive numbers of samples (n) and features (p) is an increasingly important problem. In big n settings, a common strategy is to partition the feature space, and then separately apply simple models to each partition set. We propose an alternative approach, which avoids such partitioning and…
Proposes a new method for subgroup analysis using optimal trees with parameter fusion.
Homologies of Jones and partition algebras match cyclic and symmetric groups.
Bayesian classifiers converge under certain exchangeability conditions with more data.
Mathematical construction of Chern-Simons partition function using reflection positivity.
Unsupervised space partitioning improves ANNS performance without pre-processing.
Researchers compute dimensions of GLN-skein modules for genus-one mapping tori.
Capital distribution curve is defined as log-log plot of normalized stock capitalizations ranked in descending order. The curve displays remarkable stability over periods of time. Theory of exchangeable distributions on set partitions, developed for purposes of mathematical genetics and recently applied in non-parametr…
Comparing and aligning large datasets is a pervasive problem occurring across many different knowledge domains. We introduce and study MREC, a recursive decomposition algorithm for computing matchings between data sets. The basic idea is to partition the data, match the partitions, and then recursively match the points…
An autonomous variational inference algorithm for arbitrary graphical models requires the ability to optimize variational approximations over the space of model parameters as well as over the choice of tractable families used for the variational approximation. In this paper, we present a novel combination of graph part…
In this paper we propose a new parameter-free method for trajectory classification which finds the best trajectory partition and dimension combination for robust trajectory classification. Preliminary experiments show that our approach is very promising.
We consider log-supermodular models on binary variables, which are probabilistic models with negative log-densities which are submodular. These models provide probabilistic interpretations of common combinatorial optimization tasks such as image segmentation. In this paper, we focus primarily on parameter estimation in…
Novel hybrid method for Bayesian network structure learning reduces computational time without sacrificing accuracy.
An ant colony optimization approach for partitioning a set of objects is proposed. In order to minimize the intra-variance, or within sum-of-squares, of the partitioned classes, we construct ant-like solutions by a constructive approach that selects objects to be put in a class with a probability that depends on the di…
ConvNets improve nonstationary covariance estimation for large-scale spatial data.
Improved supervised EM learning for shared kernel models with feature space partitioning.
We consider the community detection problem in sparse random hypergraphs. Angelini et al. (2015) conjectured the existence of a sharp threshold on model parameters for community detection in sparse hypergraphs generated by a hypergraph stochastic block model. We solve the positive part of the conjecture for the case of…
We investigate a class of feature allocation models that generalize the Indian buffet process and are parameterized by Gibbs-type random measures. Two existing classes are contained as special cases: the original two-parameter Indian buffet process, corresponding to the Dirichlet process, and the stable (or three-param…
Given a vector of probability distributions, or arms, each of which can be sampled independently, we consider the problem of identifying the partition to which this vector belongs from a finitely partitioned universe of such vector of distributions. We study this as a pure exploration problem in multi armed bandit sett…
MEI model improves knowledge graph completion by efficiently modeling interactions between embeddings.
Divide-and-conquer framework speeds up black-box inference for large data.
Efficient Bayesian LMM framework for high-dimensional longitudinal data.
SPAQL improves RL by adaptively partitioning state-action space and learning a time-invariant policy.
K-fold cross validation (CV) is a popular method for estimating the true performance of machine learning models, allowing model selection and parameter tuning. However, the very process of CV requires random partitioning of the data and so our performance estimates are in fact stochastic, with variability that can be s…
Study of 5D SYM theory on toric surfaces yields refined Vafa-Witten invariants.