STICC clusters geographic objects considering both spatial contiguity and attributes.
problem Discovering repeated geographic patterns with spatial contiguity.
method Spatial Toeplitz Inverse Covariance-Based Clustering (STICC) method.
result STICC significantly outperforms baseline methods in adjusted rand index and macro-F1 score.
Spatially constrained clustering divides landscapes into homogeneous regions with spatial contiguity.
problem Dividing landscapes into homogeneous patches with spatial contiguity and hierarchy.
method Developed a spatially constrained spectral clustering framework using a flexible kernel and recursive bisection.
result The proposed framework outperforms baseline methods in balancing region contiguity and homogeneity.
Sharp thresholds and contiguity for community detection in contextual SBM.
problem Community detection in graphs with high-dimensional node-covariates.
method Contextual Stochastic Block Model, non-rigorous cavity method, information theory.
result Established the sharp threshold for detection and weak recovery in the contextual SBM.
We prove lognormal distribution for symmetric perceptron model, solving key conjectures.
problem Understanding the performance of learning algorithms in neural networks.
method Lognormal distribution characterization and small graph conditioning method.
result Established lognormal distribution and several conjectures for the symmetric perceptron model.
Flexible band grouping and kernel fusion for hyperspectral image processing.
problem Large dimensionality in hyperspectral imaging.
method Non-contiguous and contiguous band grouping for dimensionality reduction; improved visual clustering; unsupervised clustering algorithms; diverse features via different proximity metrics and kernel functions; l∞-norm multiple kernel learning. result Heterogeneous features and kernels lead to performance gain.
In 1D, optimal double bubbles are intervals or spheres.
problem Finding the least-perimeter way to enclose two volumes with a log-convex density.
method Analyzing the density function's log-convexity to determine the optimal configuration.
result In 1D, the optimal configuration can be intervals or spheres.
Identification of regions of interest (ROI) associated with certain disease has a great impact on public health. Imposing sparsity of pixel values and extracting active regions simultaneously greatly complicate the image analysis. We address these challenges by introducing a novel region-selection penalty in the framew…
We consider the problems of detection and localization of a contiguous block of weak activation in a large matrix, from a small number of noisy, possibly adaptive, compressive (linear) measurements. This is closely related to the problem of compressed sensing, where the task is to estimate a sparse vector using a small…
Using ideas from shape theory we embed the coarse category of metric spaces into the category of direct sequences of simplicial complexes with bonding maps being simplicial. Two direct sequences of simplicial complexes are equivalent if one of them can be transformed to the other by contiguous factorizations of bonding…
The paper develops distribution-free methods for ordinal classification.
problem Constructing valid prediction sets for ordinal classification problems.
method Leveraging conformal prediction and multiple testing with FWER control.
result The proposed methods achieve satisfactory levels of marginal and class-specific conditional coverages.
New method explains computational barriers in high-dimensional statistical models.
problem Understanding detection-recovery gaps in high-dimensional inference.
method Combining algorithmic contiguity and cross-validation reduction to obtain conditional computational lower bounds.
result Mild control of low-degree advantage is sufficient to explain computational barriers for recovery.
Paper proves computational hardness for graph matching and detection problems.
problem Computational hardness for graph matching and detection problems in correlated random graphs.
method Algorithmic contiguity and low-degree advantage bounds.
result No efficient algorithms exist for certain graph matching and detection problems.
Complex non-linear interactions between banks and assets we model by two time-dependent Erdős Renyi network models where each node, representing bank, can invest either to a single asset (model I) or multiple assets (model II). We use dynamical network approach to evaluate the collective financial failure---systemic ri…
New identities for simplex volume in Euclidean space.
problem Calculating the volume of a spherically faced simplex.
method Using Cayley-Menger determinants and hypergeometric integrals.
result Derivation of two identities describing simplex volume.
Scalable method for regionalizing and extracting temporal patterns from time series data.
problem Static spatial snapshots and ad hoc regularization limit effective spatial analysis and resource management.
method Minimum description length principle for fully nonparametric spatial partitioning and time series archetypes.
result Accurately recovers planted regional structure and drivers in synthetic and empirical data.
PatchUp improves CNN robustness with mixed feature blocks.
problem High generalization gap in deep learning models with limited labeled data.
method Block-level regularization of hidden feature maps from mixed samples.
result PatchUp improves model robustness and generalization.
New insights into bias-variance tradeoff for data-driven optimization under local misspecification.
problem Understanding the relative performance of SAA, IEO, and ETO under local misspecification.
method Developed a local misspecification perspective using contiguity theory in statistics.
result Explicit expressions for decision bias and geometric understanding of variance.
Proposes a test for stochastic block models with bounded degrees.
problem Testing Erdös-Rényi model versus bisection stochastic block model with bounded degrees.
method Likelihood-ratio (LR) type procedure based on regularization.
result Limit distributions as power Poisson laws under null and alternative hypotheses.
In this paper, we work to construct mosaic representations of knots on the torus, rather than in the plane. This consists of a particular choice of the ambient group, as well as different definitions of contiguous and suitably connected. We present conditions under which mosaic numbers might decrease by this projection…
This paper sharpens privacy guarantees for high-dimensional PCA under differential privacy.
problem Understanding the exact privacy loss in high-dimensional PCA with differential privacy.
method Analyzes the exponential mechanism in a model-free setting for high-dimensional PCA.
result Sharp utility and privacy characterizations in high dimensions show the difficulty of detecting a target individual's presence.
PCA can detect a low-rank signal in spiked random matrix models, but not always optimally.
problem Understanding when PCA can detect a low-rank signal in the presence of noise.
method Le Cam's notion of contiguity, analysis of spiked Wishart ensemble, and non-spectral tests.
result PCA is sub-optimal for detection in non-Gaussian Wigner ensembles and certain negative spikes in Gaussian Wishart ensemble.
There is a growing interest in joint multi-subject fMRI analysis. The challenge of such analysis comes from inherent anatomical and functional variability across subjects. One approach to resolving this is a shared response factor model. This assumes a shared and time synchronized stimulus across subjects. Such a model…
Functional data analysis involves data described by regular functions rather than by a finite number of real valued variables. While some robust data analysis methods can be applied directly to the very high dimensional vectors obtained from a fine grid sampling of functional data, all methods benefit from a prior simp…
This paper deals with the notion of a large financial market and the concepts of asymptotic arbitrage and strong asymptotic arbitrage (both of the first kind), introduced by Yu.M. Kabanov and D.O. Kramkov. We show that the arbitrage properties of a large market are completely determined by the asymptotic behavior of th…
We compute exact values respectively bounds of "distances" - in the sense of (transforms of) power divergences and relative entropy - between two discrete-time Galton-Watson branching processes with immigration GWI for which the offspring as well as the immigration is arbitrarily Poisson-distributed (leading to arbitra…
In this paper, we prove than given two cubic knots K1, K2 in R3, they are isotopic if and only if one can pass from one to the other by a finite sequence of cubulated moves. These moves are analogous to the Reidemeister moves for classical tame knots. We use the fact that a cubic knot is determined by…
Efficiently finds diverse coherent counterfactual explanations.
problem Finding coherent counterfactual explanations for complex data.
method Mixed integer programming with mixed polytope constraints.
result Efficiently generates diverse coherent counterfactual explanations.
Study on detecting and estimating a rank-one tensor in noisy data.
problem Detecting and estimating a rank-one deformation in symmetric random Gaussian tensors.
method Established upper and lower bounds on critical signal-to-noise ratios for various priors.
result Upper and lower bounds match up to a 1+o(1) factor for large tensor order, and are asymptotically tight for sparse signals.
We give characterizations of asymptotic arbitrage of the first and second kind and of strong asymptotic arbitrage for large financial markets with small proportional transaction costs $\la_n$ on market n in terms of contiguity properties of sequences of equivalent probability measures induced by $\la_n$--consistent p…
Rényi divergence is related to Rényi entropy much like Kullback-Leibler divergence is related to Shannon's entropy, and comes up in many settings. It was introduced by Rényi as a measure of information that satisfies almost the same axioms as Kullback-Leibler divergence, and depends on a parameter that is called its or…
Interpretable semantic textual similarity (iSTS) task adds a crucial explanatory layer to pairwise sentence similarity. We address various components of this task: chunk level semantic alignment along with assignment of similarity type and score for aligned chunks with a novel system presented in this paper. We propose…
Spotlight method finds hidden errors in deep learning models.
problem Systematic errors in deep learning models on rare data subsets.
method Shining a spotlight on hidden layer representations to find poor performance areas.
result Identifies semantically meaningful areas of weakness in various models.
This paper resolves the test for Markov regime switching models' regime number.
problem Testing the number of regimes in Markov regime switching models.
method Derives the asymptotic distribution of the likelihood ratio test statistic.
result Establishes the asymptotic validity of the parametric bootstrap.
Associating distinct groups of objects (clusters) with contiguous regions of high probability density (high-density clusters), is central to many statistical and machine learning approaches to the classification of unlabelled data. We propose a novel hyperplane classifier for clustering and semi-supervised classificati…
We study a generalized framework for structured sparsity. It extends the well-known methods of Lasso and Group Lasso by incorporating additional constraints on the variables as part of a convex optimization problem. This framework provides a straightforward way of favouring prescribed sparsity patterns, such as orderin…
Hidden Markov models (HMMs) are one of the most widely used statistical methods for analyzing sequence data. However, the reporting of output from HMMs has largely been restricted to the presentation of the most-probable (MAP) hidden state sequence, found via the Viterbi algorithm, or the sequence of most probable marg…
New clustering algorithm uses reverse nearest neighbour for better density-based clustering.
problem Density-based clustering of separated high-density regions.
method Uses reverse nearest neighbour (RNN) queries to estimate densities and recover clusters.
result Outperforms DBSCAN and ISDBSCAN on synthetic and real-world data.
A new method for statistical inference using approximate Newton steps from stochastic gradients.
problem Efficient statistical inference for convex and non-convex learning problems.
method Approximate stochastic Newton steps based on finite differences.
result Efficient computation of statistical error covariance without exact second-order information.
Training-free looped transformers improve model performance without additional training.
problem Improving model performance without additional training or fine-tuning.
method A lightweight inference-time wrapper loops a contiguous mid-stack block of layers of a frozen checkpoint without additional fine-tuning.
result Our method improves model performance across various model families.
Multi-SpaCE generates valid counterfactual explanations for multivariate time series data.
problem Lack of transparency in deep learning models for multivariate time series data.
method Multi-objective counterfactual explanation method using NSGA-II for multivariate time series data.
result Ensures perfect validity and superior performance compared to existing methods.
SliceOut speeds up deep learning training without sacrificing accuracy.
problem Frequent model re-training and large model training workloads in deep learning.
method SliceOut uses dropout-inspired scheme to drop contiguous sets of units at random, leveraging GPU memory layout.
result 10-40% speedups and memory reduction with minimal accuracy loss.
ClustGeo uses Ward-like clustering with spatial constraints in R.
problem Hierarchical clustering with spatial/geographical constraints.
method Ward-like hierarchical clustering algorithm with two dissimilarity matrices and a mixing parameter.
result Determines optimal spatial contiguity without sacrificing variable quality.
Generative adversarial approach for satellite image time series land cover classification.
problem Enhance interpretability of land cover classification models.
method Generative adversarial counterfactual approach for multi-class land cover classification.
result Discovery of interesting information on land cover class relationships and sparser, interpretable solutions.
Prototype for adaptive electron microscopy scans reduces dose and time.
problem Reduce electron microscopy scan time and dose with minimal loss.
method Adaptive partial scanning with reinforcement learning.
result Reinforcement learning trained neural network optimizes scan paths.
FDS tackles long horizon hyperparameter optimization issues.
problem Memory scaling and gradient degradation in long horizon tasks.
method Forward-mode differentiation with sharing (FDS).
result Significantly outperforms greedy gradient-based alternatives.
Study robustness of split conformal prediction under adversarial attacks.
problem Ensuring distribution-free coverage guarantees in CP under adversarial conditions.
method Theoretical analysis and extensive experiments on split conformal prediction robustness.
result Prediction coverage varies with calibration-time attack strength, enabling control over coverage under adversarial tests.
Improves multi-objective learning by adapting to local subintervals.
problem Learning a predictor satisfying multiple objectives in an online, changing data setting.
method Adapting an existing multi-objective learning method with an adaptive online algorithm.
result Improves predictions over subgroups and remains robust under distribution shift.
ABC method improves subseasonal weather forecasting by 60-90%.
problem Improving subseasonal temperature and precipitation forecasting accuracy.
method Combines dynamical forecasts with machine learning-based bias correction.
result Significant improvement in temperature and precipitation forecasting skills.