While Multiple Instance (MI) data are point patterns -- sets or multi-sets of unordered points -- appropriate statistical point pattern models have not been used in MI learning. This article proposes a framework for model-based MI learning using point process theory. Likelihood functions for point pattern data derived …
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
PoPPy is a Point Process toolbox based on PyTorch, which achieves flexible designing and efficient learning of point process models. It can be used for interpretable sequential data modeling and analysis, e.g., Granger causality analysis of multi-variate point processes, point process-based simulation and prediction of…
The paper classifies circle actions on 6D manifolds with isolated fixed points.
In this paper, we study a circle action on a compact oriented manifold with a discrete fixed point set. The fixed point data consists of the weights of the -representations at the fixed points. We prove various results and properties of the action, in terms of the fixed point data. We show that the manifold can be…
High-dimensional data often lie in low-dimensional subspaces corresponding to different classes they belong to. Finding sparse representations of data points in a dictionary built using the collection of data helps to uncover low-dimensional subspaces and address problems such as clustering, classification, subset sele…
This paper proposes a centroid-based clustering algorithm which is capable of clustering data-points with n-features, without having to specify the number of clusters to be formed. The core logic behind the algorithm is a similarity measure, which collectively decides whether to assign an incoming data-point to a pre-e…
Sparse subspace clustering (SSC) is one of the current state-of-the-art methods for partitioning data points into the union of subspaces, with strong theoretical guarantees. However, it is not practical for large data sets as it requires solving a LASSO problem for each data point, where the number of variables in each…
Method infers dynamics from incomplete time series data.
The paper develops methods to create private synthetic spatial point patterns.
We consider the problem of learning a manifold from a teacher's demonstration. Extending existing approaches of learning from randomly sampled data points, we consider contexts where data may be chosen by a teacher. We analyze learning from teachers who can provide structured data such as individual examples (isolated …
Evaluates change point detection algorithms on real-world data.
In a typical online learning scenario, a learner is required to process a large data stream using a small memory buffer. Such a requirement is usually in conflict with a learner's primary pursuit of prediction accuracy. To address this dilemma, we introduce a novel Bayesian online classi cation algorithm, called the Vi…
New method upsamples sparse, non-uniform point clouds more accurately.
Point patterns are sets or multi-sets of unordered elements that can be found in numerous data sources. However, in data analysis tasks such as classification and novelty detection, appropriate statistical models for point pattern data have not received much attention. This paper proposes the modelling of point pattern…
Paper finds new realizable data for maps with three branch points.
The paper develops a neural network-based method for detecting change points in large-scale time-evolving data.
Clustering methods group a set of data points into a few coherent groups or clusters of similar data points. As an example, consider clustering pixels in an image (or video) if they belong to the same object. Different clustering methods are obtained by using different notions of similarity and different representation…
Paper introduces a neural network-based non-stationary influence kernel for complex event data.
A new method clusters rows of a matrix of point processes.
NN-CUSUM detects changes in high-dimensional data using neural networks.
DALES offers a large annotated aerial LiDAR dataset for 3D deep learning.
A new method detects change points in time series with conceptors.
Paper introduces a novel point process model for graph data using GNNs.
Two methods using low-discrepancy points improve data compression for neural networks.
Proves critical points of ADM mass correspond to specific initial data sets.
New algorithm detects changes in high-dimensional data with mean and variance.
One key use of k-means clustering is to identify cluster prototypes which can serve as representative points for a dataset. However, a drawback of using k-means cluster centers as representative points is that such points distort the distribution of the underlying data. This can be highly disadvantageous in problems wh…
AUCRSS detects change points in partially observed multivariate autocorrelated data.
Shapley value is a classic notion from game theory, historically used to quantify the contributions of individuals within groups, and more recently applied to assign values to data points when training machine learning models. Despite its foundational role, a key limitation of the data Shapley framework is that it only…
Online detection of abrupt changes in high-dimensional data streams.
The paper proves -convergence of discrete tangent-point energies to continuous energies and ropelength, with applications to biarc curves.
We introduce a novel geometry-oriented methodology, based on the emerging tools of topological data analysis, into the change point detection framework. The key rationale is that change points are likely to be associated with changes in geometry behind the data generating process. While the applications of topological …
A new model predicts spatio-temporal data using adaptive decision trees and point processes.
Unified framework detects changes in complex system models.
We propose the Autoencoding Binary Classifiers (ABC), a novel supervised anomaly detector based on the Autoencoder (AE). There are two main approaches in anomaly detection: supervised and unsupervised. The supervised approach accurately detects the known anomalies included in training data, but it cannot detect the unk…
DSNE visualizes data velocity in lower dimensions.
Kernel methods on discrete domains have shown great promise for many challenging data types, for instance, biological sequence data and molecular structure data. Scalable kernel methods like Support Vector Machines may offer good predictive performances but do not intrinsically provide uncertainty estimates. In contras…
Develops a deep non-stationary kernel for non-stationary spatio-temporal point processes.
A conjugate Bayesian method detects change points in Hawkes processes efficiently.
Local GP approach improves simulation efficiency for large datasets.
A given set of data-points in some feature space may be associated with a Schrodinger equation whose potential is determined by the data. This is known to lead to good clustering solutions. Here we extend this approach into a full-fledged dynamical scheme using a time-dependent Schrodinger equation. Moreover, we approx…
In this paper we present a loss-based approach to change point analysis. In particular, we look at the problem from two perspectives. The first focuses on the definition of a prior when the number of change points is known a priori. The second contribution aims to estimate the number of change points by using a loss-ba…
Data sets are often modeled as point clouds in , for large. It is often assumed that the data has some interesting low-dimensional structure, for example that of a -dimensional manifold , with much smaller than . When is simply a linear subspace, one may exploit this assumption for encoding ef…
Semi-supervised clustering is the task of clustering data points into clusters where only a fraction of the points are labelled. The true number of clusters in the data is often unknown and most models require this parameter as an input. Dirichlet process mixture models are appealing as they can infer the number of clu…
We present a numerical algorithm for nonnegative matrix factorization (NMF) problems under noisy separability. An NMF problem under separability can be stated as one of finding all vertices of the convex hull of data points. The research interest of this paper is to find the vectors as close to the vertices as possible…
GOCPD detects change points by maximizing the probability of two independent models.
ADS filters data points for efficient batch active learning.
IDK improves anomaly detection for points and groups without explicit learning.