New method clusters hydrological and sediment data for storm event analysis.
problem Analyzing storm events for water quality constituents like turbidity.
method Multivariate time series clustering of river discharge and sediment data.
result Clusters differ from 2-D hysteresis loop classifications.
We present a novel hierarchical distance-dependent Bayesian model for event coreference resolution. While existing generative models for event coreference resolution are completely unsupervised, our model allows for the incorporation of pairwise distances between event mentions -- information that is widely used in sup…
New clustering technique improves RNN event log predictions.
problem Leveraging event attributes for better RNN predictions.
method A novel clustering technique for event attributes.
result Improved prediction accuracy with reduced training time.
Study examines tech stocks' reactions to Facebook data leak scandal.
problem Impact of Facebook data leak scandal on U.S. tech stocks.
method Clustering method to identify related companies, CAR to measure impact.
result Overall tech sector showed no adverse impact, but Facebook's performance was negatively affected.
New method corrects bias in event study analysis using machine learning.
problem Bias in conventional event study analysis mechanisms.
method Topological machine-learning approach with self-organizing map (SOM).
result Identifies factors of abnormal stock returns and depicts event clusters.
We propose an effective method to solve the event sequence clustering problems based on a novel Dirichlet mixture model of a special but significant type of point processes --- Hawkes process. In this model, each event sequence belonging to a cluster is generated via the same Hawkes process with specific parameters, an…
We suggest a novel method of clustering and exploratory analysis of temporal event sequences data (also known as categorical time series) based on three-dimensional data grid models. A data set of temporal event sequences can be represented as a data set of three-dimensional points, each point is defined by three varia…
Bayesian approach clusters survival data for better risk prediction.
problem Identifying subpopulations with distinct risk profiles in survival analysis.
method Bayesian nonparametric approach in a clustered latent space.
result Consistent improvements in predictive performance and interpretability.
ClusterLOB clusters market events to identify different trading behaviors.
problem Understanding market microstructure and participant behavior in financial markets.
method ClusterLOB uses K-means++ algorithm to cluster market events based on six time-dependent features.
result ClusterLOB identifies three distinct trading behaviors: directional, opportunistic, and market-making participants.
Paper clusters event sequences using a reinforcement learning approach with policy mixture model.
problem Clustering event sequences with varying temporal patterns.
method Reinforcement learning with a policy mixture model, decomposing sequences into states and actions.
result Effective clustering of event sequences into underlying policies, outperforming existing methods.
New risk models use chaotic attractors to predict extreme events.
problem Predicting Black Swan events in financial markets.
method Combining heavy-tailed priors with chaotic dynamics (Lorenz and Rossler systems).
result Models generate volatility clustering, fat tails, and extreme events.
Sketch-based approach detects community events in evolving networks.
problem Community detection in time-varying networks.
method Maintains a small sketch graph to capture essential community structure.
result Efficiently identifies six key community events during network evolution.
Modeling financial crises and cryptocurrency shocks using copulae clustering.
problem Detecting financial crises and shock events in stock and cryptocurrency markets.
method Copulae clustering based on probability distribution distances.
result Successfully detected all past crises and shock events in stock and cryptocurrency markets.
System detects controversial events on social media and impacts markets.
problem Lack of systematic data on company social consciousness and sustainability.
method Uses Twitter data to identify and validate controversial events.
result Validated controversial events impact market volatility.
Novel connections between Neyman-Scott processes and Bayesian nonparametric mixture models enable scalable inference.
problem Efficiently modeling and detecting clusters in spatiotemporal data.
method Adapting collapsed Gibbs sampling for Neyman-Scott processes via connections to mixture of finite mixture models.
result Demonstrated scalability and effectiveness on neural spike trains and document streams.
Proposes a deep neural network for predicting clustered time-to-event data.
problem Predicting clustered time-to-event data with subject-specific frailties.
method Deep neural network based gamma frailty model (DNN-FM) trained using negative profiled h-likelihood.
result Enhances prediction performance compared to existing methods.
Finding the optimal k-means clustering is NP-hard in general and many heuristics have been designed for minimizing monotonically the k-means objective. We first show how to extend Lloyd's batched relocation heuristic and Hartigan's single-point relocation heuristic to take into account empty-cluster and single-poin…
Method detects and locates eavesdropping in optical links.
problem Detect and locate eavesdropping in optical links with small power losses.
method Cluster-based approach using OPM data at receiver and in-line OPM data for localization.
result Subtle eavesdropping losses can be detected and localized using OPM data.
New model for clustering dependent community Hawkes processes in temporal networks.
problem Modeling strong dependence and community structure in temporal networks.
method Dependent Community Hawkes (DCH) models combining stochastic block models and Hawkes processes.
result Spectral clustering error bound derived for DCH models.
Paper proposes PP-GCN for fine-grained social event categorization.
problem Challenges in mining social events due to heterogeneous event elements and social network structures.
method Design an event meta-schema, build an HIN, propose PP-GCN, and use KIES.
result PP-GCN outperforms other techniques in social event detection and clustering.
Clusters of financial market states identified over 2006-2019.
problem Understanding the statistical properties of financial markets.
method Clustering analysis of correlation matrices constructed from sliding epochs.
result Financial markets can be classified into distinct states with transitions indicating precursors to catastrophic events.
Paper tackles event stream outliers using a novel weight function.
problem Learning event streams with unexpected absences or occurrences.
method Temporal point process framework with a novel weight function.
result The proposed method effectively handles both commission and omission outliers.
A system is presented that segments, clusters and predicts musical audio in an unsupervised manner, adjusting the number of (timbre) clusters instantaneously to the audio input. A sequence learning algorithm adapts its structure to a dynamically changing clustering tree. The flow of the system is as follows: 1) segment…
We study large deviations and rare default clustering events in a dynamic large heterogeneous portfolio of interconnected components. Defaults come as Poisson events and the default intensities of the different components in the system interact through the empirical default rate and via systematic effects that are comm…
The team predicts foreign exchange rates using clustering and attention models.
problem Complexity and unexpected events in foreign exchange markets.
method Clustering and attention models applied to historical data.
result Improved event-driven price prediction for oversold scenarios.
Data of the form of event times arise in various applications. A simple model for such data is a non-homogeneous Poisson process (NHPP) which is specified by a rate function that depends on time. We consider the problem of having access to multiple independent observations of event time data, observed on a common inter…
Modeling cascading behavior in complex systems using CTBNs.
problem Understanding which states trigger cascading events in complex systems.
method Continuous-time Bayesian networks (CTBNs) for modeling and identifying likely sentry states.
result Identification of likely sentry states that may lead to cascading behavior.
Proposes online learning for Hawkes processes with network structure and event interaction.
problem Modeling complex interactions and latent structures in network events.
method Online learning approach for mixture of multivariate Hawkes processes.
result Efficacy demonstrated on synthetic and real-world data.
SurvMixClust clusters survival data and predicts individual survival curves.
problem Integrating clustering into survival analysis for precision medicine.
method SurvMixClust learns latent representations for clustering and predicts survival functions using a mixture of non-parametric experts.
result SurvMixClust creates balanced clusters with distinct survival curves, outperforming clustering baselines and competing with non-clustering models in predictive accuracy.
Paper models treatment effects by clustering patients with distinct survival characteristics.
problem Estimating treatment efficacy in clinical settings with censored outcomes.
method Latent variable approach to model heterogeneous treatment effects.
result The latent structure can mediate base survival rates and reveal actionable phenotypes.
The method integrates survival constraints into NMF for identifying survival-associated gene clusters.
problem Understanding and interpreting high-dimensional biological data for disease markers.
method Cox proportional hazards regression integrated with NMF via proportional hazards non-negative matrix factorization.
result The method can uncover survival-associated gene clusters in cancer gene expression data.
BCV method helps estimate clusters and hyper-parameters in large data sets.
problem Determining the number of clusters in large-scale data.
method Bi-cross validation (BCV) for spectral clustering.
result BCV directly applies to spectral clustering for estimating clusters and hyper-parameters.
Two machine learning methods detect insider trading from investor activity data.
problem Detecting insider trading from trading activity data is challenging.
method Two unsupervised machine learning methods: clustering and group identification.
result Identifies potential insider trading rings around price sensitive events.
Recent progress in applying machine learning for jet physics has been built upon an analogy between calorimeters and images. In this work, we present a novel class of recursive neural networks built instead upon an analogy between QCD and natural languages. In the analogy, four-momenta are like words and the clustering…
Learning from electronic medical records (EMR) is challenging due to their relational nature and the uncertain dependence between a patient's past and future health status. Statistical relational learning is a natural fit for analyzing EMRs but is less adept at handling their inherent latent structure, such as connecti…
Event detection has been one of the most important research topics in social media analysis. Most of the traditional approaches detect events based on fixed temporal and spatial resolutions, while in reality events of different scales usually occur simultaneously, namely, they span different intervals in time and space…
R package stagedtrees learns staged tree structures from data.
problem Learning the structure of staged trees from data.
method Score-based and clustering-based algorithms implemented.
result Illustrated capabilities using two datasets.
Analyzed banking crises across 66 countries, showing interconnections and clustering patterns.
problem Characterizing banking crises and their interconnections across different countries.
method Used dichotomous banking crises time series data from 1800 to 2014, analyzed via heatmap matrices and clustering.
result Countries exhibit pairwise correlation in banking crises, and crises tend to affect countries with financial links.
Cluster analysis aims at separating patients into phenotypically heterogenous groups and defining therapeutically homogeneous patient subclasses. It is an important approach in data-driven disease classification and subtyping. Acute coronary syndrome (ACS) is a syndrome due to sudden decrease of coronary artery blood f…
Deep learning clusters patient time-series data for better prognosis.
problem Clustering time-series data for patient phenotyping and prognosis.
method Deep predictive clustering with novel loss functions for future outcome distribution.
result Model achieves superior clustering performance and identifies meaningful patient subgroups.
Paper detects important sub-events in disaster tweets.
problem Identifying useful information during large-scale disasters.
method Extract noun-verb pairs, learn semantic embeddings, rank and cluster.
result Effective unsupervised learning framework for sub-event detection.
New method uses cluster shapes to improve track finding in particle collisions.
problem Combining timing and additional detector information for efficient track finding.
method Neural networks to analyze cluster shapes for track seeding.
result Cluster shapes reduce fake combinatorial backgrounds while maintaining high track efficiency.
Bayesian approach models earthquake clustering with spatial mainshocks and aftershocks.
problem Estimating uncertainty in earthquake clustering models due to complex likelihood functions.
method Nonparametric Dirichlet process mixture prior for spatial mainshocks and an auxiliary latent variable routine for efficient inference.
result Efficient Bayesian forecasting of spatial earthquake occurrences with uncertainty quantification.
Study shows pre-event L2 liquidity state predicts crypto futures liquidity better than event labels.
problem Understanding how crypto futures liquidity changes over time.
method Combining L2 order book data, trade-flow records, and macro-event windows to define discrete liquidity-state transitions and evaluate models.
result Pre-event L2 liquidity state predicts post-event liquidity regimes better than event labels, and order flow adds value only when layered on top of the state model.
We study the structure of locational marginal prices in day-ahead and real-time wholesale electricity markets. In particular, we consider the case of two North American markets and show that the price correlations contain information on the locational structure of the grid. We study various clustering methods and intro…
Efficiently clusters survival curves without computationally intensive resampling.
problem Identifying clusters of survival curves efficiently and scalably.
method Log-rank test combined with k-means clustering.
result Achieves comparable results to bootstrap-based methods but with improved efficiency.
This study examines non-retail trading on Polymarket, revealing unique behavior patterns and structural limitations.
problem Lack of address-level quote-lifecycle data in Polymarket prediction markets.
method Empirical analysis of 13 million order-filled events using DBSCAN clustering on a six-feature fill-side vector.
result Non-retail behavior is uni-modal, contradicting previous archetypal hypotheses.
We quantify the amount of information filtered by different hierarchical clustering methods on correlations between stock returns comparing it with the underlying industrial activity structure. Specifically, we apply, for the first time to financial data, a novel hierarchical clustering approach, the Directed Bubble Hi…