An unsupervised neural network learns event truths from social network data.
problem Estimating event truths from conflicting opinions in social networks.
method Autoencoder learns relationships, Bayesian network models agent reliability and social relationships, variational inference estimates hidden variables and parameters.
result The approach outperforms state-of-the-art methods on real datasets.
We investigate the problem of truth discovery based on opinions from multiple agents who may be unreliable or biased. We consider the case where agents' reliabilities or biases are correlated if they belong to the same community, which defines a group of agents with similar opinions regarding a particular event. An age…
New method falsifies causal discovery results without ground truth.
problem Evaluation of causal discovery algorithms without ground truth data.
method Detects incompatibilities between causal graphs learned on different subsets of variables.
result Detection of incompatibilities can falsify wrongly inferred causal relations.
The aim of process discovery, originating from the area of process mining, is to discover a process model based on business process execution data. A majority of process discovery techniques relies on an event log as an input. An event log is a static source of historical data capturing the execution of a business proc…
Automated process discovery is a class of process mining methods that allow analysts to extract business process models from event logs. Traditional process discovery methods extract process models from a snapshot of an event log stored in its entirety. In some scenarios, however, events keep coming with a high arrival…
Novel approach models life events using causal discovery and survival analysis.
problem Modeling life event choices and occurrence from a probabilistic perspective.
method Bi-level problem formulation: causal discovery for life events graph, survival analysis for time-to-event modeling.
result Identification of causal relationships and factors influencing transition rates between life events.
New causal distances improve evaluation of causal discovery algorithms.
problem Evaluating causal discovery algorithms using graphical distances is limited.
method Defined causal distances based on causal distributions rather than graphical structure.
result Improved evaluation of causal discovery algorithms on synthetic and real-world datasets.
We address the problem of latent truth discovery, LTD for short, where the goal is to discover the underlying true values of entity attributes in the presence of noisy, conflicting or incomplete information. Despite a multitude of algorithms to address the LTD problem that can be found in literature, only little is kno…
CausalRivers benchmarks causal discovery methods on real-world river discharge data.
problem Lack of in-the-wild evaluation of causal discovery methods on complex, real-world data.
method Introduces CausalRivers, a large-scale dataset of river discharge data for benchmarking.
result Demonstrates the utility of CausalRivers in evaluating causal discovery methods.
SurvSurf predicts first hitting times for intermittent events without monotonic violations.
problem Predicting first hitting times for intermittent events with monotonicity guarantees.
method Partially monotonic neural network for sequential events, incorporating unobserved events.
result SurvSurf outperforms existing models in MSE and IBS metrics.
New test uncovers causal links in rare event dynamics.
problem Causal discovery for rare event phenomena in dynamic systems.
method Nonparametric conditional independence test on time-invariant data.
result Validated across simulated and real-world datasets.
Generates synthetic manufacturing data for causal discovery benchmarking.
problem Lack of suitable real data for validating causal discovery algorithms.
method Distributional random forests for estimating conditional distributions.
result Semisynthetic manufacturing data adheres to a causal model.
We consider the problem of separating error messages generated in large distributed data center networks into error events. In such networks, each error event leads to a stream of messages generated by hardware and software components affected by the event. These messages are stored in a giant message log. We consider …
Latent truth discovery, LTD for short, refers to the problem of aggregating ltiple claims from various sources in order to estimate the plausibility of atements about entities. In the absence of a ground truth, this problem is highly challenging, when some sources provide conflicting claims and others no claims at all.…
Improved method for unbiased causal discovery in presence of unobserved confounding.
problem Unbiased data synthesis for causal discovery algorithms in the presence of unobserved confounding.
method Explicit block-hierarchical ancestral sampling to address limitations of implicit parameterization.
result Our approach fully covers the space of causal models, including those generated by implicit parameterization.
Process mining is a research field focused on the analysis of event data with the aim of extracting insights in processes. Applying process mining techniques on data from smart home environments has the potential to provide valuable insights in (un)healthy habits and to contribute to ambient assisted living solutions. …
New bounds on majority voting's accuracy for multi-class classification problems.
problem Determining the accuracy of majority voting for multi-class classification.
method Analyzing the majority voting function under different voter conditions and distributions.
result The error rate of majority voting exponentially decays or grows with the number of voters under certain conditions.
DisCoveR efficiently discovers declarative process models from event logs.
problem Mining declarative process models from event logs efficiently and accurately.
method DisCoveR precisely formalizes an algorithm, uses a bit vector implementation, and rigorously evaluates performance.
result DisCoveR outperforms other declarative miners in accuracy and runtime.
Media tone around earnings announcements predicts stock returns.
problem Determining if media tone around earnings announcements provides useful information for stock prices.
method Conducted an event study on media tone around earnings announcements for nonfinancial S&P 500 firms.
result Media tone around earnings announcements predicts abnormal stock returns.
As part of the 2016 public evaluation challenge on Detection and Classification of Acoustic Scenes and Events (DCASE 2016), the second task focused on evaluating sound event detection systems using synthetic mixtures of office sounds. This task, which follows the `Event Detection - Office Synthetic' task of DCASE 2013,…
Neural causal discovery methods fail to accurately uncover causal structures due to the faithfulness property.
problem Accuracy in neural causal discovery is limited, especially when distinguishing between existing and non-existing causal relationships.
method Systematic evaluation of neural causal discovery methods, focusing on their performance in finite sample regimes and their ability to recover ground-truth graphs.
result Neural networks lack the precision to reliably recover ground-truth causal graphs, even for small graphs and large sample sizes.
Impute missing events in continuous-time sequences using particle smoothing.
problem Missing events in continuous-time sequences.
method Particle smoothing with trainable bidirectional LSTM proposals.
result Imputed sequences have low Bayes risk compared to ground truth.
The paper examines how timing of observations affects causal discovery methods.
problem The sensitivity of causal discovery methods to mismatched observation timing.
method Empirical and theoretical analysis of classical and recent causal discovery methods.
result Causal discovery methods are sensitive to sampling rate and window length.
Interpretable framework evaluates structure learning methods for causal discovery from observational data.
problem Evaluation of structure learning methods under assumption violations in causal discovery.
method Six-dimensional evaluation metric (DOS) tailored for causal discovery.
result Amortized causal discovery delivers results with high proximity to the optimal solution.
New approach detects and ranks novel and developing cyber threats in Twitter.
problem Detecting and ranking novel and developing cyber threats in Twitter streams.
method Unsupervised machine learning approach focusing on novelty and trendiness.
result Ranking of cyber threat events based on importance score using extracted terms.
Simulates neuropathic pain to evaluate causal discovery algorithms.
problem Lack of benchmark datasets for evaluating causal discovery algorithms.
method Developed a neuropathic pain diagnosis simulator.
result Simulator produces data similar to real-world data.
System detects controversial events on social media and impacts markets.
problem Lack of systematic data on company social consciousness and sustainability.
method Uses Twitter data to identify and validate controversial events.
result Validated controversial events impact market volatility.
ELUQuant quantifies uncertainties in DIS events using BNNs and MNFs.
problem Uncertainty quantification in Deep Inelastic Scattering (DIS) events.
method Physics-informed Bayesian Neural Network with flow approximated posteriors.
result Effective extraction of kinematic variables x x x , Q 2 Q^2 Q 2 , and y y y with detailed event-level uncertainty. CausalTime generates realistic time-series for TSCD evaluation.
problem Lack of realistic synthetic datasets for TSCD performance evaluation.
method Harnessing deep neural networks and normalizing flow for dynamics, extracting causal graphs, and deriving ground truth causal graphs.
result Generated datasets accurately reflect real data and ground truth causal graphs.
Hybrid method uses LLM to filter lead-lag relationships in prediction markets.
problem Challenges in discovering robust lead-lag relationships in prediction markets due to spurious correlations.
method Two-stage approach: statistical Granger causality followed by LLM semantic re-ranking.
result LLM-based method outperforms statistical baseline, increasing win rate and reducing average loss magnitude.
A new method detects epileptic events in EEG signals by integrating labeler categories.
problem Human oversight of brief epileptic events in EEG signals leads to inaccurate diagnoses.
method Integrates EEG signal features with one-hot encoded labeler categories for improved detection.
result The method outperforms consensus-trained detectors and maintains confidence bounds.
Framework isolates causal effects from time series data, improving accuracy under non-stationarity and autocorrelation.
problem Causal inference in non-stationary, autocorrelated time series data.
method Decomposes time series into trend, seasonal, and residual components; performs component-specific causal analysis.
result Framework more accurately recovers ground-truth causal structure than state-of-the-art baselines, especially under strong non-stationarity and temporal autocorrelation.
DECI combines causal discovery and inference in a single model for diverse data types.
problem Combining causal discovery and inference methods for diverse data types.
method Develops a single flow-based non-linear additive noise model (DECI) for causal discovery and inference.
result DECI can recover ground truth causal graphs and perform (C)ATE estimation.
New methods identify concepts in trained embeddings reliably without human labels.
problem Identifying interpretable concepts in trained embedding spaces without human labels.
method Explicitly connecting concept discovery to PCA and ICA, proposing novel approaches for dependent concepts.
result Proven methods outperform competitors on a variety of experiments, achieving up to 29% better alignment with ground truth.
Calibration without labels in multiple testing
problem Interpretable error probabilities in large-scale hypothesis testing
method Constructing pseudo-labels from spacings of ordered p p p -values result Finding that q q q -value can be severely miscalibrated TimeGraph creates synthetic datasets for robust time-series causal discovery.
problem Lack of reliable synthetic benchmark datasets for robust time-series causal discovery.
method Developed comprehensive synthetic datasets with temporal properties, including trends, seasonality, and noise.
result Demonstrated significant variations in algorithm performance under realistic temporal conditions.
BAMS uses Bayesian sampling to discover AV failures more efficiently and accurately.
problem Discovering potential failure cases in autonomous vehicles efficiently and accurately.
method Bayesian adaptive multifidelity sampling (BAMS) prioritizes exploration of low performance regions.
result BAMS discovers 10 times more issues than traditional methods with narrower rate estimates.
Estimates classifier errors without ground truth using algebraic geometry.
problem Lack of ground truth in real-world production systems.
method Non-parametric estimation using algebraic geometry to solve the self-assessment problem.
result Accuracy estimators are better than one part in a hundred.
Proposes a new framework to evaluate causal discovery methods for time series data.
problem Lack of ground truth for causal discovery in time series data.
method Flexible framework for generating synthetic time series data.
result Demonstrates degradation in performance when assumptions are violated.
A deep learning framework discovers causal relationships from incomplete data.
problem Discovering causal knowledge from incomplete observational data.
method Imputated Causal Learning (ICL) framework for iterative missing data imputation and causal structure discovery.
result ICL outperforms state-of-the-art methods in various missing data scenarios.
The study analyzes and mitigates errors in PC-based causal discovery methods.
problem Errors in PC-based causal discovery methods can lead to incorrect graphs.
method The study introduces coherency scores to detect assumption violations and small sample errors in PC-based methods.
result The coherency scores can detect errors that other methods cannot, bridging between global and local error detection.
We release a large ECG dataset for arrhythmia subtype discovery.
problem Discovering unknown subtypes of arrhythmia from continuous raw signals.
method Unsupervised representation learning task using semi-supervised evaluation.
result Qualitative evaluations show potential for representation learning in arrhythmia sub-type discovery.
We have discovered 12 independent new empirical scaling laws in foreign exchange data-series that hold for close to three orders of magnitude and across 13 currency exchange rates. Our statistical analysis crucially depends on an event-based approach that measures the relationship between different types of events. The…
A novel method detects multiple mitosis events and mitigates annotation gaps in phase-contrast microscopy.
problem Detecting multiple mitosis events and handling annotation gaps in closely placed cells.
method Estimating a spatiotemporal likelihood map via 3DCNN to detect multiple mitosis events and mitigate annotation gaps.
result Our method outperformed compared methods in terms of F1-score using a challenging dataset.
Grinch efficiently clusters large datasets with complex structures.
problem Large-scale hierarchical clustering with complex linkage functions.
method Rotate and graft subroutines for efficient reconfiguration.
result Grinch guarantees accurate cluster trees for consistent models.
New algorithm detects unique events in time series data.
problem Detecting anomalous events in unknown properties.
method Model-free, unsupervised detection using Temporal Outlier Factor (TOF).
result TOF outperforms traditional outlier detection methods.
A new mechanism reduces expert belief regret in online forecasting.
problem Minimizing expert belief regret in strategic forecasting.
method Developed a no-regret mechanism for non-myopic experts using online I-ELF.
result Achieved i l d e O ( T N ) ilde{O}(\sqrt{T N}) i l d e O ( T N ) regret for full-information setting. RED-2400 is a public benchmark of trading events from a Solana exchange, labeled by algorithmic rejection.
problem Analyzing algorithmically-rejected trading events for insights into market dynamics.
method Public dataset of 6,660 algorithmically-rejected trading events, linked to post-rejection price and liquidity trajectories.
result First window of a planned series of datasets extending the time horizon and enabling regime-stratified analysis.