The fundamental aim of clustering algorithms is to partition data points. We consider tasks where the discovered partition is allowed to vary with some covariate such as space or time. One approach would be to use fragmentation-coagulation processes, but these, being Markov processes, are restricted to linear or tree s…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
ConvNets improve nonstationary covariance estimation for large-scale spatial data.
We show that the space of algebraic covariant derivative curvature tensors R' is generated by Young symmetrized tensor products W*U or U*W, where W and U are covariant tensors of order 2 and 3 whose symmetry classes are irreducible and characterized by the following pairs of partitions: {(2),(3)}, {(2),(2 1)} or {(1 1)…
New method splits unknown covariance Gaussians into independent parts.
Causal trees struggle with accuracy in estimating treatment effects.
We consider the sparse inverse covariance regularization problem or graphical lasso with regularization parameter . Suppose the co- variance graph formed by thresholding the entries of the sample covariance matrix at is decomposed into connected components. We show that the vertex-partition induced by the thresh…
Localized transfer learning improves nonparametric regression performance.
WHOMP optimizes randomized controlled trials by minimizing subgroup bias.
Study on estimating Gaussian mean from coarse data, resolving identifiability and computational efficiency questions.
CDL index improves clustering validation for non-convex data.
Linear regression models depend directly on the design matrix and its properties. Techniques that efficiently estimate model coefficients by partitioning rows of the design matrix are increasingly popular for large-scale problems because they fit well with modern parallel computing architectures. We propose a simple me…
We study the problem of learning to choose from m discrete treatment options (e.g., news item or medical drug) the one with best causal effect for a particular instance (e.g., user or patient) where the training data consists of passive observations of covariates, treatment, and the outcome of the treatment. The standa…
This paper considers the problem of estimating multiple related Gaussian graphical models from a -dimensional dataset consisting of different classes. Our work is based upon the formulation of this problem as group graphical lasso. This paper proposes a novel hybrid covariance thresholding algorithm that can effecti…
Proposes a method to improve learning when training data is not representative.
Bayesian framework for model uncertainty identifies complex heterogeneity without strong assumptions.
Framework for domain adaptation using pseudo-labels from unlabeled data.
GPU-accelerated particle methods outperform neural samplers in LFT benchmarks.
We consider a multi-armed bandit problem in a setting where each arm produces a noisy reward realization which depends on an observable random covariate. As opposed to the traditional static multi-armed bandit problem, this setting allows for dynamically changing rewards that better describe applications where side inf…
Bayesian model detects communities in networks with covariates.
New tree and forest methods use oblique splits for better risk bounds.
Study addresses covariate mismatch in federated learning, improving model accuracy.
A new method estimates conditional canonical correlations using random forests.
DRUM transfers cardiac arrest models across registries with missing covariates.
Probabilistic learning for binary classification with categorical variables.
To construct flexible nonlinear predictive distributions, the paper introduces a family of softplus function based regression models that convolve, stack, or combine both operations by convolving countably infinite stacked gamma distributions, whose scales depend on the covariates. Generalizing logistic regression that…
This paper presents a new approach for Gaussian process (GP) regression for large datasets. The approach involves partitioning the regression input domain into multiple local regions with a different local GP model fitted in each region. Unlike existing local partitioned GP approaches, we introduce a technique for patc…
In the classical Gaussian SVM classification we use the feature space projection transforming points to normal distributions with fixed covariance matrices (identity in the standard RBF and the covariance of the whole dataset in Mahalanobis RBF). In this paper we add additional information to Gaussian SVM by considerin…
The paper challenges the use of decision trees for pointwise inference due to slow convergence rates.
Abstract M5 branes on ADE singularities yields BPS spectrum and partition functions.
We consider a blind identification problem in which we aim to recover a statistical model of a network without knowledge of the network's edges, but based solely on nodal observations of a certain process. More concretely, we focus on observations that consist of single snapshots taken from multiple trajectories of a d…
We construct a topological Chern-Simons sigma model on a Riemannian three-manifold M with gauge group G whose hyperkahler target space X is equipped with a G-action. Via a perturbative computation of its partition function, we obtain new topological invariants of M that define new weight systems which are characterized…
We consider generators of algebraic covariant derivative curvature tensors R' which can be constructed by a Young symmetrization of product tensors W*U or U*W, where W and U are covariant tensors of order 2 and 3. W is a symmetric or alternating tensor whereas U belongs to a class of the infinite set S of irreducible s…
A method for non-parametric conditional distribution estimation using CRPS-optimal binning.
The paper improves clustering algorithms for noisy data samples.
Efficient algorithms for -means clustering frequently converge to suboptimal partitions, and given a partition, it is difficult to detect -means optimality. In this paper, we develop an a posteriori certifier of approximate optimality for -means clustering. The certifier is a sub-linear Monte Carlo algorithm b…
Generalizes underlap coefficient for multivariate group separation.
New method targets relative risk heterogeneity in clinical trials.
Binscatter is a popular method for visualizing bivariate relationships and conducting informal specification testing. We study the properties of this method formally and develop enhanced visualization and econometric binscatter tools. These include estimating conditional means with optimal binning and quantifying uncer…
We study large-scale spatial systems that contain exogenous variables, e.g. environmental factors that are significant predictors in spatial processes. Building predictive models for such processes is challenging because the large numbers of observations present makes it inefficient to apply full Kriging. In order to r…
The paper develops a theory for random forests, separating variance components and providing methods for estimating prediction intervals.
The Random Projection Tree structures proposed in [Freund-Dasgupta STOC08] are space partitioning data structures that automatically adapt to various notions of intrinsic dimensionality of data. We prove new results for both the RPTreeMax and the RPTreeMean data structures. Our result for RPTreeMax gives a near-optimal…
Generalizes underlap coefficient for multivariate group separation.
Recent results in coupled or temporal graphical models offer schemes for estimating the relationship structure between features when the data come from related (but distinct) longitudinal sources. A novel application of these ideas is for analyzing group-level differences, i.e., in identifying if trends of estimated ob…
While studying response trajectory, often the population of interest may be diverse enough to exist distinct subgroups within it and the longitudinal change in response may not be uniform in these subgroups. That is, the timeslope and/or influence of covariates in longitudinal profile may vary among these different sub…
BPR matches NN accuracy in crop classification while being more transparent.
The study limits how many parts regular simplicial partitions can overlap.
Hypergraph partitioning lies at the heart of a number of problems in machine learning and network sciences. Many algorithms for hypergraph partitioning have been proposed that extend standard approaches for graph partitioning to the case of hypergraphs. However, theoretical aspects of such methods have seldom received …
Recent papers have formulated the problem of learning graphs from data as an inverse covariance estimation with graph Laplacian constraints. While such problems are convex, existing methods cannot guarantee that solutions will have specific graph topology properties (e.g., being -partite), which are desirable for so…