Proposes a new model for online anomaly detection in multivariate time series.
problem Inaccurate anomaly detection in multivariate time series due to spurious correlations and lack of temporal causality.
method Clusters channels based on correlations, embeds each cluster, and integrates information through a causal mixer while maintaining temporal causality.
result Consistently superior performance across six public benchmark datasets.
MLP-Mixer achieves better performance through sparsity and wider architecture.
problem Understanding why MLP-Mixer outperforms conventional MLPs.
method Revealed sparseness as a key mechanism, showed effective expression as wider MLP, demonstrated quantitative similarities, and applied a guiding principle.
result MLP-Mixer's performance improvement through sparseness and wider architecture.
This paper proposes a new approach to Transformers by integrating hierarchical associative memory with MetaFormers.
problem Theoretical framework for Transformers and MLP-Mixers remains underdeveloped.
method Integrating hierarchical associative memory with MetaFormers to create a parallelized MLP-Mixer.
result Symmetry-breaking effects improve the performance of the MLP-Mixer in image recognition tasks.
MixerFlow combines MLP-Mixer with normalizing flows for efficient image modeling.
problem Efficiently modeling complex image densities using generative models.
method Proposes MixerFlow, a novel architecture based on MLP-Mixer for normalizing flows.
result Demonstrates improved density estimation and better scaling with higher image resolutions.
Hybrid QAOA approach optimizes portfolios with strict constraints, outperforming classical methods.
problem Combinatorial optimization under strict cardinality constraints in portfolio management.
method Constraint-preserving QAOA with XY-mixers and Trotterized initialization.
result QAOA achieves a Sharpe Ratio of 1.81, significantly outperforming classical methods.
Interpolated-MLPs control inductive bias for better performance in low-compute tasks.
problem Low-compute performance gap between MLPs and CNNs.
method Introduced Interpolated MLP (I-MLP) approach to control inductive bias incrementally.
result Continuous logarithmic relationship between inductive bias and performance in low-compute tasks.
FEM improves attention mechanisms by applying value-driven log-linear tilts.
problem Standard attention mechanisms read via convex average, limiting channel-wise selection.
method Free Energy Mixer (FEM) applies a value-driven, per-channel log-linear tilt to a fast prior over indices.
result FEM outperforms strong baselines on NLP, vision, and time-series tasks.
Generative compression technique reduces neural network size and improves performance on microcontrollers.
problem Memory constraints on microcontrollers limit the size of neural networks, especially for 1x1 pointwise (PW) mixers.
method HYPER-TINYPW uses a shared micro-MLP to generate PW kernels from tiny per-layer codes, reducing memory usage.
result HYPER-TINYPW achieves comparable performance to larger models while being significantly smaller (225 kB vs 1.4 MB).
StrTransformer recovers sources without labels by optimizing latent matrices and enforcing structural constraints.
problem Unsupervised blind source recovery in signal processing.
method Source-wise structured Transformer framework with latent source matrix optimization, structural regularization, and branch-specific weights.
result StrTransformer learns distinct temporal-scale structures and recovers source-aligned latent trajectories.
A quantum framework optimizes collateral allocation for derivatives.
problem Legal constraints and operational rules in collateral allocation for derivatives.
method Certified higher-order quantum framework that normalizes margin requirements and builds a bounded neighborhood of actions.
result Quantum framework improves certified sample quality compared to classical methods.
Standard neural networks are often overconfident when presented with data outside the training distribution. We introduce HyperGAN, a new generative model for learning a distribution of neural network parameters. HyperGAN does not require restrictive assumptions on priors, and networks sampled from it can be used to qu…
New framework optimizes classification trees with logistic loss and ℓ1 regularization.
problem Improving interpretability and generalization of classification trees.
method Developed a generalized framework for CTs, incorporating logistic loss and ℓ1 regularization. result Optimal Logistic Tree model outperforms state-of-the-art MIP-based approaches in terms of interpretability and generalization.
There are many industrial situations where rods are used to stir a fluid, or where rods repeatedly stretch a material such as bread dough or taffy. The goal in these applications is to stretch either material lines (in a fluid) or the material itself (for dough or taffy) as rapidly as possible. The growth rate of mater…
NIST CTS Superset offers a large dataset for telephony speaker recognition.
problem Lack of a large-scale, uniform dataset for telephony speaker recognition.
method Compilation of speech segments from multiple corpora, including Greybeard, Switchboard, and Mixer series.
result Results on the NIST 2020 CTS Speaker Recognition Challenge serve as a reference baseline.
EMIX minimizes surprise in multi-agent reinforcement learning.
problem Surprise and approximation bias in multi-agent reinforcement learning.
method Energy-based MIXer (EMIX) for minimizing surprise across multiple agents.
result EMIX demonstrates consistent stable performance in challenging StarCraft II scenarios.
New framework learns disentangled causal representations from observed labels.
problem Learning meaningful disentangled causal representations from observed data.
method ICM-VAE framework using flow-based diffeomorphic functions and causal disentanglement prior.
result Induces highly disentangled causal factors and improves robustness.
A power-law fit to the empirical inference-compute frontier in LOB prediction suggests a scaling-law-style frontier.
problem Limit order book prediction
method Using a suite of models ranging from small decision trees to neural LOB architectures
result A power-law fit to the low- and mid-compute non-MLPLOB frontier extrapolates across multiple orders of magnitude and attains R2=0.941 on the excluded high-compute MLPLOB target frontier. Vision Transformers show different internal representations compared to CNNs.
problem Understanding how Vision Transformers solve image classification tasks.
method Comparative analysis of ViT and CNN architectures on image classification benchmarks.
result ViT has more uniform representations across all layers, while CNNs have more varied representations.
Reinterprets Granger causality with causal Bayesian networks and Reichenbach's principles.
problem Lack of a rigorous causal foundation in Granger causality.
method Reinterpreting Granger causality through Reichenbach's principles and causal Bayesian networks, implementing as c-GC.
result c-GC provides a more principled framework for causal discovery in observational datasets.
New measures for causal entropy and information gain studied.
problem Quantifying causal relationships in machine learning.
method Formal study of causal entropy and information gain.
result Established fundamental properties and relationships.
We quantify causal bias in continuous treatment settings.
problem Identifying and quantifying causal bias in continuous treatment scenarios.
method Developed a novel characterization of causal bias in structural causal models, proving conditions for zero bias and efficient estimation.
result Causal bias can be estimated efficiently under certain structural equation restrictions, allowing for causal regularization of predictive models.
Improved Granger causality method for dynamic time series data.
problem Traditional Granger causality method assumes constant causalities, failing to model dynamic causalities.
method Dynamic window-level Granger causality (DWGC) method with causality indexing.
result Improved DWGC method better detects window-level causalities.
Study of generalized Csiszár divergences and their application to Cramér-Rao bounds.
problem Deriving lower bounds for estimator variance using generalized divergences.
method Applied Eguchi's theory to derive Fisher information metric and dual affine connections.
result More widely applicable Cramér-Rao inequality for escort distributions.
The study examines causal razors and their logical relations, highlighting a dilemma in causal discovery.
problem Selecting a reasonable scoring criterion for causal discovery algorithms.
method Review and logical comparison of numerous causal razors, focusing on parameter minimality in multinomial models.
result Parameter minimality poses a dilemma in selecting a reasonable scoring criterion for causal discovery algorithms.
CIB compresses variables causally, preserving key causal interactions.
problem Constructing causal variable abstractions in complex systems.
method Causal Information Bottleneck (CIB) method, extending IB to include causal structures.
result CIB produces causally interpretable abstractions that accurately capture causal relations.
Paper characterizes and represents pairwise causal background knowledge for improved causal inference.
problem Improving causal inference by handling pairwise causal constraints.
method Graphical characterization, direct causal clause (DCC), unified representation, MPDAG, polynomial-time algorithms.
result Pairwise causal background knowledge uniquely decomposes into MPDAG and DCCs, improving causal effect identification.
A new method clusters heterogeneous subgroups for accurate causal learning.
problem Diverse causal relationships across different time spans, regions, or strategies.
method Nonlinear Causal Kernel Clustering
result Reduction in prediction error through enhanced causal learning.
Framework for Granger causality in extreme events.
problem Identifying causal links from extreme events in time series.
method Causal tail coefficient and novel inference method.
result Framework outperforms state-of-the-art methods in detecting Granger causality in extremes.
Proposes DCNAR for dynamic causal inference from neural time series.
problem Uncertainty and evolution of causal structure in real-world domains.
method Two-stage neural causal modeling integrating discovery and inference.
result Dynamic causal inferences are more stable and meaningful than alternatives.
New algorithms for causal bandits without knowing the graph structure.
problem Causal bandit problems with unknown graph structure.
method Developed novel causal bandit algorithms for causal trees, forests, and general graphs without prior knowledge of the causal graph.
result Regret guarantees significantly improved over standard MAB algorithms under mild conditions.
This review explores causal decision-making to improve decision quality.
problem Effective decision-making requires understanding causal relationships.
method Causal structure learning, causal effect learning, and causal policy learning.
result Challenges in causal decision-making are identified and recent advances are discussed.
New model improves data augmentation for causal tasks.
problem Optimizing causal models robustly under Wasserstein distances.
method Proposes a new G-Causal Normalizing Flow architecture.
result Empirically outperforms standard generative models.
DoWhy-GCM extends causal inference in graphical models for diverse queries.
problem Addressing diverse causal queries in graphical causal models.
method Specify cause-effect relations via a causal graph, fit causal mechanisms, pose causal queries.
result Identification of root causes, attribution of causal influences, diagnosis of causal structures.
ABCI infers causal models and queries simultaneously using Bayesian active learning.
problem Inference of causal models and effects in a two-stage process is inefficient and unnatural.
method Active Bayesian Causal Inference (ABCI) using Gaussian processes for sequentially designing experiments.
result ABCI is more data-efficient and accurate in learning causal queries from fewer samples.
Optimizes causal effects on unknown graphs using Causal Entropy Optimization.
problem Optimizing causal effects in unknown causal graphs.
method Causal Entropy Optimization (CEO) framework that generalizes Causal Bayesian Optimization (CBO). Incorporates causal structure uncertainty in surrogate models and intervention selection.
result CEO achieves faster convergence to global optimum compared to CBO and improves upon sequential structure learning.
Proposes Causal Loss to improve machine learning models' causal inference.
problem Machine learning algorithms often fail to capture causal relationships when data is inconsistent.
method Introduces Causal Loss, a model-agnostic loss function that enhances interventional capabilities.
result Causal Loss improves non-causal associative models to have interventional capabilities.
iCITRIS learns causal variables from interactive systems with instantaneous effects.
problem Identifying causal variables from temporal sequences with instantaneous effects.
method iCITRIS method for causal representation learning that handles instantaneous effects in intervened temporal sequences.
result iCITRIS accurately identifies causal variables and their causal graph from three interactive system datasets.
Paper constructs unfaithful probability distributions in binary causal graphs.
problem Unfaithful probability distributions in binary causal graphs.
method Constructs unfaithful probability distributions in binary causal graphs.
result Examples of unfaithful probability distributions in binary causal graphs.
New framework for dynamic causal graph modeling and effect estimation.
problem Dynamic changes in causal relationships over time.
method Score-based causal discovery with autoregressive model structure.
result Dynamic causal graph with time-varying causal relations.
Causal normalizing flows recover causal models from observational data.
problem Recovering causal models from observational data.
method Use autoregressive normalizing flows and analyze design choices.
result Causal normalizing flows can capture causal data-generating processes.
Meta-causal states group equivalent qualitative causal dynamics, useful for analyzing system changes.
problem Qualitative changes in causal relationships due to agent actions or environmental tipping points.
method Propose meta-causal states to group causal models based on equivalent qualitative behavior and parameterize specific mechanisms.
result Meta-causal states can be inferred from observed agent behavior and disentangled from unlabeled data.
BGNNs model particle-boundary interactions efficiently.
problem Efficiently modeling geometric boundaries in 3D simulations.
method Introduce Boundary Graph Neural Networks (BGNNs) to dynamically modify graph structures.
result BGNNs accurately reproduce 3D granular flows without handcrafted conditions.
New method prevents invalid inference after causal discovery.
problem Invalid inference after causal discovery.
method Developed tools for valid post-causal-discovery inference.
result Our method provides reliable coverage while achieving more accurate causal discovery.
RealCause provides a realistic benchmark for causal inference.
problem Lack of a reliable benchmark for comparing causal effect estimators.
method Flexible generative models to create a benchmark that is both ground-truth and realistic.
result Evaluation of over 1500 causal estimators provides evidence for choosing hyperparameters using predictive metrics.
GC-KAN uses KANs to detect Granger causality in time series data.
problem Detecting causal relationships in nonlinear time series data.
method Developed GC-KAN framework using Kolmogorov-Arnold networks for Granger causality detection.
result KANs outperform MLPs in identifying sparse Granger causal relationships.
New method identifies causal structure in exchangeable data.
problem Existing causal discovery methods struggle with i.i.d. data.
method Exchangeable data provides richer conditional independence structure.
result Exchangeable data allows for unique causal structure identification.
New causal distances improve evaluation of causal discovery algorithms.
problem Evaluating causal discovery algorithms using graphical distances is limited.
method Defined causal distances based on causal distributions rather than graphical structure.
result Improved evaluation of causal discovery algorithms on synthetic and real-world datasets.
Amortized Causal Discovery learns to infer causal graphs from time-series data, improving performance.
problem Inference of causal graphs from time-series data is inefficient due to fitting new models for each sample.
method Proposes Amortized Causal Discovery, a variational model that leverages shared dynamics across samples with different causal graphs.
result Significant improvements in causal discovery performance demonstrated experimentally.