STICC clusters geographic objects considering both spatial contiguity and attributes.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We prove lognormal distribution for symmetric perceptron model, solving key conjectures.
Regionalization is the task of dividing up a landscape into homogeneous patches with similar properties. Although this task has a wide range of applications, it has two notable challenges. First, it is assumed that the resulting regions are both homogeneous and spatially contiguous. Second, it is well-recognized that l…
Hyperspectral imaging is a powerful technology that is plagued by large dimensionality. Herein, we explore a way to combat that hindrance via non-contiguous and contiguous (simpler to realize sensor) band grouping for dimensionality reduction. Our approach is different in the respect that it is flexible and it follows …
Identification of regions of interest (ROI) associated with certain disease has a great impact on public health. Imposing sparsity of pixel values and extracting active regions simultaneously greatly complicate the image analysis. We address these challenges by introducing a novel region-selection penalty in the framew…
Sharp thresholds and contiguity for community detection in contextual SBM.
Paper proves computational hardness for graph matching and detection problems.
New method explains computational barriers in high-dimensional statistical models.
The paper develops distribution-free methods for ordinal classification.
The classic double bubble theorem says that the least-perimeter way to enclose and separate two prescribed volumes in is the standard double bubble. We seek the optimal double bubble in with density, which we assume to be strictly log-convex. For we show that the solution is sometime…
We consider the problems of detection and localization of a contiguous block of weak activation in a large matrix, from a small number of noisy, possibly adaptive, compressive (linear) measurements. This is closely related to the problem of compressed sensing, where the task is to estimate a sparse vector using a small…
Using ideas from shape theory we embed the coarse category of metric spaces into the category of direct sequences of simplicial complexes with bonding maps being simplicial. Two direct sequences of simplicial complexes are equivalent if one of them can be transformed to the other by contiguous factorizations of bonding…
Density-based clustering is the task of discovering high-density regions of entities (clusters) that are separated from each other by contiguous regions of low-density. DBSCAN is, arguably, the most popular density-based clustering algorithm. However, its cluster recovery capabilities depend on the combination of the t…
We present two identities (contiguity relation and variation formula) concerning the volume of a spherically faced simplex in the Euclidean space. These identities are described in terms of Cayley-Menger determinants and their differentials involved with hypersphere arrangements. They are derived as a limit of fundamen…
This paper proposes new search algorithms for counterfactual explanations based upon mixed integer programming. We are concerned with complex data in which variables may take any value from a contiguous range or an additional set of discrete states. We propose a novel set of constraints that we refer to as a "mixed pol…
Functional data analysis involves data described by regular functions rather than by a finite number of real valued variables. While some robust data analysis methods can be applied directly to the very high dimensional vectors obtained from a fine grid sampling of functional data, all methods benefit from a prior simp…
Interpretable semantic textual similarity (iSTS) task adds a crucial explanatory layer to pairwise sentence similarity. We address various components of this task: chunk level semantic alignment along with assignment of similarity type and score for aligned chunks with a novel system presented in this paper. We propose…
Hidden Markov models (HMMs) are one of the most widely used statistical methods for analyzing sequence data. However, the reporting of output from HMMs has largely been restricted to the presentation of the most-probable (MAP) hidden state sequence, found via the Viterbi algorithm, or the sequence of most probable marg…
Complex non-linear interactions between banks and assets we model by two time-dependent Erdős Renyi network models where each node, representing bank, can invest either to a single asset (model I) or multiple assets (model II). We use dynamical network approach to evaluate the collective financial failure---systemic ri…
We study a generalized framework for structured sparsity. It extends the well-known methods of Lasso and Group Lasso by incorporating additional constraints on the variables as part of a convex optimization problem. This framework provides a straightforward way of favouring prescribed sparsity patterns, such as orderin…
Scalable method for regionalizing and extracting temporal patterns from time series data.
PatchUp improves CNN robustness with mixed feature blocks.
New insights into bias-variance tradeoff for data-driven optimization under local misspecification.
We present a novel statistical inference framework for convex empirical risk minimization, using approximate stochastic Newton steps. The proposed algorithm is based on the notion of finite differences and allows the approximation of a Hessian-vector product from first-order information. In theory, our method efficient…
Improves multi-objective learning by adapting to local subintervals.
In this paper, we propose a Ward-like hierarchical clustering algorithm including spatial/geographical constraints. Two dissimilarity matrices and are inputted, along with a mixing parameter . The dissimilarities can be non-Euclidean and the weights of the observations can be non-uniform. The fi…
FDS tackles long horizon hyperparameter optimization issues.
SEEK algorithm selects minimal state in reinforcement learning for better policy learning.
In this paper, we work to construct mosaic representations of knots on the torus, rather than in the plane. This consists of a particular choice of the ambient group, as well as different definitions of contiguous and suitably connected. We present conditions under which mosaic numbers might decrease by this projection…
This paper sharpens privacy guarantees for high-dimensional PCA under differential privacy.
Autonomous systems can be used to search for sparse signals in a large space; e.g., aerial robots can be deployed to localize threats, detect gas leaks, or respond to distress calls. Intuitively, search algorithms may increase efficiency by collecting aggregate measurements summarizing large contiguous regions. However…
Nowadays, online learning is an appealing learning paradigm, which is of great interest in practice due to the recent emergence of large scale applications such as online advertising placement and online web ranking. Standard online learning assumes a finite number of samples while in practice data is streamed infinite…
Multi-SpaCE generates valid counterfactual explanations for multivariate time series data.
New algorithms allow multiple robots to search efficiently without central coordination.
There is a growing interest in joint multi-subject fMRI analysis. The challenge of such analysis comes from inherent anatomical and functional variability across subjects. One approach to resolving this is a shared response factor model. This assumes a shared and time synchronized stimulus across subjects. Such a model…
Maximum A posteriori Probability (MAP) inference in graphical models amounts to solving a graph-structured combinatorial optimization problem. Popular inference algorithms such as belief propagation (BP) and generalized belief propagation (GBP) are intimately related to linear programming (LP) relaxation within the She…
This paper deals with the notion of a large financial market and the concepts of asymptotic arbitrage and strong asymptotic arbitrage (both of the first kind), introduced by Yu.M. Kabanov and D.O. Kramkov. We show that the arbitrage properties of a large market are completely determined by the asymptotic behavior of th…
We compute exact values respectively bounds of "distances" - in the sense of (transforms of) power divergences and relative entropy - between two discrete-time Galton-Watson branching processes with immigration GWI for which the offspring as well as the immigration is arbitrarily Poisson-distributed (leading to arbitra…
In this paper, we prove than given two cubic knots , in , they are isotopic if and only if one can pass from one to the other by a finite sequence of cubulated moves. These moves are analogous to the Reidemeister moves for classical tame knots. We use the fact that a cubic knot is determined by…
Markov regime switching models have been used in numerous empirical studies in economics and finance. However, the asymptotic distribution of the likelihood ratio test statistic for testing the number of regimes in Markov regime switching models has been an unresolved problem. This paper derives the asymptotic distribu…
We give characterizations of asymptotic arbitrage of the first and second kind and of strong asymptotic arbitrage for large financial markets with small proportional transaction costs $\la_n$ on market in terms of contiguity properties of sequences of equivalent probability measures induced by $\la_n$--consistent p…
Biclustering techniques have been widely used to identify homogeneous subgroups within large data matrices, such as subsets of genes similarly expressed across subsets of patients. Mining a max-sum sub-matrix is a related but distinct problem for which one looks for a (non-necessarily contiguous) rectangular sub-matrix…
Rényi divergence is related to Rényi entropy much like Kullback-Leibler divergence is related to Shannon's entropy, and comes up in many settings. It was introduced by Rényi as a measure of information that satisfies almost the same axioms as Kullback-Leibler divergence, and depends on a parameter that is called its or…
Spotlight method finds hidden errors in deep learning models.
DRL agents learn to trade Intel stock with stable positive returns.
Associating distinct groups of objects (clusters) with contiguous regions of high probability density (high-density clusters), is central to many statistical and machine learning approaches to the classification of unlabelled data. We propose a novel hyperplane classifier for clustering and semi-supervised classificati…
New algorithms improve uncertainty estimation in satellite precipitation predictions.
A new method for faster spatial modeling on exascale computers.