This paper quantifies and mitigates a bias in the Hayashi-Yoshida estimator causing data loss.
problem Formulaic bias in the Hayashi-Yoshida estimator leading to data loss.
method Formalizes and quantifies the data loss, introduces (a,b)-asynchronous adversary, and provides algorithms.
result Proves that for equal rates, the minimal average cumulative data loss is 25%.
While Multiple Instance (MI) data are point patterns -- sets or multi-sets of unordered points -- appropriate statistical point pattern models have not been used in MI learning. This article proposes a framework for model-based MI learning using point process theory. Likelihood functions for point pattern data derived …
PoPPy is a Point Process toolbox based on PyTorch, which achieves flexible designing and efficient learning of point process models. It can be used for interpretable sequential data modeling and analysis, e.g., Granger causality analysis of multi-variate point processes, point process-based simulation and prediction of…
Clustering groups similar data points into clusters.
problem Grouping similar data points into coherent clusters.
method Different clustering methods based on similarity and data representations.
result Various clustering methods exist.
The paper classifies circle actions on 6D manifolds with isolated fixed points.
problem Classifying circle actions on 6D manifolds with isolated fixed points.
method Performing equivariant connected sums at fixed points with specific manifolds.
result A sequence of operations can reduce the fixed point data to the empty collection.
In this paper, we study a circle action on a compact oriented manifold with a discrete fixed point set. The fixed point data consists of the weights of the S1-representations at the fixed points. We prove various results and properties of the action, in terms of the fixed point data. We show that the manifold can be…
High-dimensional data often lie in low-dimensional subspaces corresponding to different classes they belong to. Finding sparse representations of data points in a dictionary built using the collection of data helps to uncover low-dimensional subspaces and address problems such as clustering, classification, subset sele…
This paper proposes a centroid-based clustering algorithm which is capable of clustering data-points with n-features, without having to specify the number of clusters to be formed. The core logic behind the algorithm is a similarity measure, which collectively decides whether to assign an incoming data-point to a pre-e…
New method uses topological data analysis for better change point detection.
problem Detecting change points in time series data.
method Integrates topological data analysis with existing nonparametric change point detection methods.
result Enhanced detection accuracy of change point locations.
Sparse subspace clustering (SSC) is one of the current state-of-the-art methods for partitioning data points into the union of subspaces, with strong theoretical guarantees. However, it is not practical for large data sets as it requires solving a LASSO problem for each data point, where the number of variables in each…
Method infers dynamics from incomplete time series data.
problem Challenges in inferring stochastic dynamics from time series with missing data.
method Expectation Maximization (EM) algorithm that iterates between E-step and M-step.
result The EM algorithm effectively recovers missing data points and infers underlying network models from real neuronal activities.
The paper develops methods to create private synthetic spatial point patterns.
problem Generating private synthetic spatial point patterns.
method Developed differentially private Poisson and Cox point synthesizers.
result The synthesizers effectively maintain privacy and utility of synthetic data.
We consider the problem of learning a manifold from a teacher's demonstration. Extending existing approaches of learning from randomly sampled data points, we consider contexts where data may be chosen by a teacher. We analyze learning from teachers who can provide structured data such as individual examples (isolated …
Evaluates change point detection algorithms on real-world data.
problem Insufficient evaluation of change point detection algorithms on real-world time series.
method Developed a data set of 37 time series from various domains, annotated by human experts, and evaluated 14 algorithms using consistency metrics.
result Demonstrates the need for better evaluation methods in change point detection.
In a typical online learning scenario, a learner is required to process a large data stream using a small memory buffer. Such a requirement is usually in conflict with a learner's primary pursuit of prediction accuracy. To address this dilemma, we introduce a novel Bayesian online classi cation algorithm, called the Vi…
New method upsamples sparse, non-uniform point clouds more accurately.
problem Suboptimal results from existing point cloud upsampling methods.
method Imposes manifold distribution constraints using Gaussian functions.
result Generates higher-quality, more uniformly distributed dense point clouds.
Point patterns are sets or multi-sets of unordered elements that can be found in numerous data sources. However, in data analysis tasks such as classification and novelty detection, appropriate statistical models for point pattern data have not received much attention. This paper proposes the modelling of point pattern…
A new framework assigns values to data points considering their distribution.
problem Limited applicability of data Shapley to points outside the fixed data set.
method Proposes distributional Shapley, defining point value in context of data distribution.
result Distributional Shapley values are stable under data point and distribution perturbations.
Paper finds new realizable data for maps with three branch points.
problem Existence of rational maps with specific branch points.
method New families of branch data identified through football decomposition method.
result Identifies new realizable branch data and exceptional data.
The paper develops a neural network-based method for detecting change points in large-scale time-evolving data.
problem Detecting and locating change points in multivariate time-evolving data.
method Two-step procedure involving neural network training and test error function calibration over moving windows.
result Consistent estimates for the number and locations of change points under temporal dependence.
Paper introduces a neural network-based non-stationary influence kernel for complex event data.
problem Modeling complex, non-stationary, and dependent discrete event data.
method Neural Spectral Marked Point Processes (NSMPP) with a versatile non-stationary influence kernel.
result NSMPP outperforms state-of-the-art models on synthetic and real data.
A new method clusters rows of a matrix of point processes.
problem Challenges in analyzing structured point process data.
method Mixture model of multi-level marked point processes, combined with ES algorithm and FPCA.
result An efficient method for clustering rows of a matrix of point processes.
NN-CUSUM detects changes in high-dimensional data using neural networks.
problem Detecting abrupt changes in high-dimensional data.
method Neural network-based CUSUM for online change-point detection.
result NN-CUSUM performs well in detecting changes in high-dimensional data.
DALES offers a large annotated aerial LiDAR dataset for 3D deep learning.
problem Lack of large-scale annotated aerial LiDAR datasets for deep learning.
method Collection and annotation of over half a billion hand-labeled points from an ALS scanner.
result DALES is the most extensive publicly available ALS data set with improved resolution and coverage.
A new method detects change points in time series with conceptors.
problem Detecting change points in time series with nonlinear temporal dependence.
method Use of conceptor matrix to learn baseline dynamics and identify change points.
result The method provides a consistent estimate of the true change point.
Paper introduces a novel point process model for graph data using GNNs.
problem Modeling discrete event data over graphs with influence kernel.
method Combines Hawkes kernel and Graph Neural Networks (GNN) for event prediction.
result Achieves superior predictive performance compared to state-of-the-art.
Two methods using low-discrepancy points improve data compression for neural networks.
problem Efficiently compress large datasets for neural network training.
method Two methods based on low-discrepancy points: digital nets with averaging and clustering.
result Second method outperforms supercompress in compression error and neural network accuracy.
A new clustering method preserves data distribution.
problem Distorted cluster centers in k-means clustering.
method Distributional Clustering method ensuring cluster centers mimic data distribution.
result Cluster centers converge to data generating distribution.
Proves critical points of ADM mass correspond to specific initial data sets.
problem Finding initial data sets with fixed Bartnik boundary data.
method Proves existence of critical points on a Banach manifold.
result Critical points of ADM mass correspond to initial data sets with generalized Killing vector fields.
New algorithm detects changes in high-dimensional data with mean and variance.
problem Challenges in detecting changes in high-dimensional data with mean and variance.
method Complete graph-based approach to detect changes of mean and variance from low to high-dimensional online data.
result The proposed method outperforms existing methods in terms of detection power.
AUCRSS detects change points in partially observed multivariate autocorrelated data.
problem Detecting change points in multivariate autocorrelated data with limited sensing resources.
method Adaptive Upper Confidence Region (AUCRSS) with state space model (SSM), adaptive sampling policy, and generalized likelihood ratio test.
result The method outperforms existing approaches in detecting change points efficiently.
Online detection of abrupt changes in high-dimensional data streams.
problem Detecting abrupt changes in high-dimensional, streaming data with multiple subspaces.
method Dynamic sparse subspace learning approach with multiple structural change-point model, Bayesian information criterion for penalty coefficients selection, and Pruned Exact Linear Time algorithm.
result Effectiveness demonstrated through simulation and real gesture data studies.
The paper proves Γ-convergence of discrete tangent-point energies to continuous energies and ropelength, with applications to biarc curves.
problem Proving convergence of discrete tangent-point energies to continuous energies and ropelength.
method Using biarc curves and interpolation, the paper proves Γ-convergence of discretized tangent-point energies to the continuous tangent-point energies and ropelength functional. result Discrete almost minimizing biarc curves converge to ropelength minimizers and minimizers of continuous tangent-point energies.
A new model predicts spatio-temporal data using adaptive decision trees and point processes.
problem Predicting spatio-temporal data with real-life applications.
method Hawkes process, adaptive decision tree, joint optimization algorithm.
result Significant improvement in predictions compared to standard methods.
Unified framework detects changes in complex system models.
problem Accurate identification of dynamic changes in simulation models.
method Combines machine learning and process-driven simulation modeling.
result Significantly improves change point detection accuracy.
We propose the Autoencoding Binary Classifiers (ABC), a novel supervised anomaly detector based on the Autoencoder (AE). There are two main approaches in anomaly detection: supervised and unsupervised. The supervised approach accurately detects the known anomalies included in training data, but it cannot detect the unk…
DSNE visualizes data velocity in lower dimensions.
problem Understanding movement patterns in high-dimensional data.
method DSNE is a variation of Stochastic Neighbor Embedding that learns velocity embeddings using Euclidean distances on a unit sphere.
result DSNE enables visualization of data movement in lower dimensions.
Develops a deep non-stationary kernel for non-stationary spatio-temporal point processes.
problem Capturing non-stationary dependencies in point process data.
method Approximates the influence kernel with a novel low-rank decomposition and introduces a log-barrier penalty to maintain non-negativity.
result Demonstrates superior performance and computational efficiency compared to state-of-the-art methods.
Kernel methods on discrete domains have shown great promise for many challenging data types, for instance, biological sequence data and molecular structure data. Scalable kernel methods like Support Vector Machines may offer good predictive performances but do not intrinsically provide uncertainty estimates. In contras…
A conjugate Bayesian method detects change points in Hawkes processes efficiently.
problem Non-conjugacy between Hawkes process likelihood and prior causes inefficiency in change point detection.
method Data augmentation to propose a conjugate Bayesian two-step change point detection method.
result The conjugate method is more accurate and efficient than non-conjugate methods.
Local GP approach improves simulation efficiency for large datasets.
problem High computational cost of traditional Gaussian processes for large-scale simulations.
method Hybridizes global and local GP approximations with strategic placement of inducing points.
result Local inducing points enhance accuracy and computational efficiency.
A given set of data-points in some feature space may be associated with a Schrodinger equation whose potential is determined by the data. This is known to lead to good clustering solutions. Here we extend this approach into a full-fledged dynamical scheme using a time-dependent Schrodinger equation. Moreover, we approx…
In this paper we present a loss-based approach to change point analysis. In particular, we look at the problem from two perspectives. The first focuses on the definition of a prior when the number of change points is known a priori. The second contribution aims to estimate the number of change points by using a loss-ba…
Data sets are often modeled as point clouds in RD, for D large. It is often assumed that the data has some interesting low-dimensional structure, for example that of a d-dimensional manifold M, with d much smaller than D. When M is simply a linear subspace, one may exploit this assumption for encoding ef…
Semi-supervised clustering is the task of clustering data points into clusters where only a fraction of the points are labelled. The true number of clusters in the data is often unknown and most models require this parameter as an input. Dirichlet process mixture models are appealing as they can infer the number of clu…
We present a numerical algorithm for nonnegative matrix factorization (NMF) problems under noisy separability. An NMF problem under separability can be stated as one of finding all vertices of the convex hull of data points. The research interest of this paper is to find the vectors as close to the vertices as possible…
GOCPD detects change points by maximizing the probability of two independent models.
problem Large false discovery rates in online change point detection methods.
method GOCPD uses ternary search to find change points by maximizing the probability of two independent models.
result GOCPD accelerates CPD with logarithmic complexity for single change point detection.
ADS filters data points for efficient batch active learning.
problem Efficiently selecting data points for annotation in parallel settings.
method Active Data Shapley (ADS) using the Shapley value of data.
result Significantly increases efficiency of active learning by 6x.