Detects BGP routing anomalies using bursty announcement patterns.
problem Detecting disruptive routing updates in the Internet.
method Burstiness measure of inter-arrival times around disruptive update dates and times.
result High recall and better precision in detecting anomalous incidents.
Proposes a new model to accurately describe random series of events.
problem Accurately and parsimoniously characterize random series of events (RSEs).
method Burstiness Scale (BuSca) model, which views RSEs as a mix of Poissonian and self-exciting processes.
result BuSca accurately describes RSEs in diverse systems, even with only two parameters.
NBFA addresses burstiness in count data using negative binomial likelihood.
problem Limitation of Poisson factorization in capturing burstiness.
method Constructs NBFA under negative binomial likelihood, proposes Gibbs samplers.
result NBFA provides clear advantages over Poisson models in burstiness.
We investigate large changes, bursts, of the continuous stochastic signals, when the exponent of multiplicativity is higher than one. Earlier we have proposed a general nonlinear stochastic model which can be transformed into Bessel process with known first hitting (first passage) time statistics. Using these results w…
PRGDS models count tensors with sparsity and burstiness.
problem Modeling sequential count data with sparsity and burstiness.
method Poisson-randomized gamma dynamical system with alternating Poisson and gamma latent states.
result Sparse PRGDS often outperforms other models in predicting count data.
Algorithm uncovers two main patterns of online content popularity: bursty and steady.
problem Understanding how online content gains popularity over time.
method Multi-faceted temporal analysis using dipm-SC algorithm.
result Two main patterns of popularity: bursty and steady temporal behaviors.
Paper proposes using CNN for stock trading with data normalization.
problem Improving stock trading accuracy in volatile markets.
method Developed CNN-based trading framework with novel data normalization.
result CNN-based framework outperforms other methods on 29 stocks.
Bayesian method speeds up demand forecasting for e-commerce.
problem Demand forecasting for fast and bursty items at scale.
method Approximate Bayesian inference using Newton-Raphson algorithm and Kalman smoothing.
result Significantly outperforms competing approaches on large datasets.
Combustion reaction kinetics models are used for the description of a special class of bursty Financial Time Series. The small number of parameters they depend upon enable financial analysts to predict the time as well as the magnitude of the jump of the value of the portfolio. Several Financial Time Series are analyse…
We present a new, efficient method for automatically detecting severe conflicts `edit wars' in Wikipedia and evaluate this method on six different language WPs. We discuss how the number of edits, reverts, the length of discussions, the burstiness of edits and reverts deviate in such pages from those following the gene…
Improved error correction using neural networks and belief propagation.
problem Inference in factor graphs with loops or poor approximations.
method Hybrid model combining FG-GNN and belief propagation.
result Hybrid model outperforms belief propagation in error correction tasks.
This paper uses cPF to build recommender systems from raw count data.
problem Sparse, over-dispersed and bursty count data make direct use in recommender systems challenging.
method Compound Poisson Factorization (cPF) with a unified framework (dcPF) and adaptive algorithm.
result dcPF achieves better recommendation scores than Poisson Factorization on raw or binarized data.
This study reveals statistical patterns in ERC20 token transactions on Ethereum blockchain.
problem Understanding transactional dynamics in decentralized systems.
method Examined over 44 million ERC20 token transfers, categorized by address type (EOA or SC), and analyzed using scaling laws.
result EOA-driven transactions exhibit consistent statistical behavior, while SC-driven activity displays sublinear scaling and bursty activity.
Paper proposes a new Markov model for efficient PLC system design.
problem Efficient estimation of Markov model parameters for bursty error channels.
method Introduced a Block Diagonal Markov model and a modified Baum-Welch algorithm.
result Efficient estimation of state transition matrix Λ for PLC system design. The episodic, irregular and asynchronous nature of medical data render them difficult substrates for standard machine learning algorithms. We would like to abstract away this difficulty for the class of time-stamped categorical variables (or events) by modeling them as a renewal process and inferring a probability dens…
Hybrid model improves geopolitical conflict forecasting.
problem Forecasting geopolitical events from sparse, bursty data.
method Sparse Temporal Fusion Transformer (TFT) + Variational Nearest Neighbor Gaussian Process (VNNGP).
result Consistently outperforms standalone TFT in long-range horizons.
Paper proposes a new method to handle missing data in medical records using sequential variational autoencoders.
problem Missing data in medical records due to sensor off-times and uneven data collection.
method Sequential variational autoencoders (VAEs) with a new methodology called Shi-VAE.
result Shi-VAE achieves the best performance in terms of both metrics compared to state-of-the-art methods.
Poseidon optimizes deep learning training on GPU clusters by reducing network communication.
problem Substantial parameter synchronization over the network in distributed DL implementations.
method Overlap communication and computation, use a hybrid communication scheme.
result Achieves significant speed-ups in DL training on GPU clusters.
New method uses reinforcement learning to accurately estimate available network bandwidth.
problem Accurate and fast estimation of available bandwidth in networks with varying cross-traffic.
method Employed reinforcement learning, specifically the ε-greedy algorithm in a multi-armed bandit approach. result Proposed method identifies available bandwidth with high precision and converges under various challenging conditions.
Robust Bayesian models are appealing alternatives to standard models, providing protection from data that contains outliers or other departures from the model assumptions. Historically, robust models were mostly developed on a case-by-case basis; examples include robust linear regression, robust mixture models, and bur…
Paper proves EM algorithm convergence for mixtures of discrete and continuous parameters.
problem Nontrivial convergence analysis for EM algorithms with mixed-integer parameters.
method Introduces conditions for EM convergence in mixed-integer optimization.
result Proves convergence of EM-based sparse Bayesian learning algorithm.
Improves text clustering by incorporating sequential features and word embeddings.
problem Lack of sequential information and synonym handling in current text clustering methods.
method SiDPMM model that models documents as joint of bags of words, sequential features, and word embeddings.
result Significant improvement in performance and accurate inference of cluster numbers.
Order cancellation process plays a crucial role in the dynamics of price formation in order-driven stock markets and is important in the construction and validation of computational finance models. Based on the order flow data of 18 liquid stocks traded on the Shenzhen Stock Exchange in 2003, we investigate the empiric…
Proposes a new model to better handle overdispersed count time series.
problem Heterogeneous overdispersed count time series.
method Negative-Binomial Randomized Gamma Markov Process.
result Significantly improves predictive performance and fast convergence of inference algorithm.
A new metric, Weighted Regret, unifies FDR and power evaluation in online multiple testing.
problem The asymmetric costs of false positives and false negatives in automated pipelines.
method Introducing Weighted Regret and Decoupled-OMT (DOMT) to unify FDR and power evaluation.
result DOMT achieves an order-optimal sublinear mitigation of threshold depletion in bursty environments.
In online social media systems users are not only posting, consuming, and resharing content, but also creating new and destroying existing connections in the underlying social network. While each of these two types of dynamics has individually been studied in the past, much less is known about the connection between th…
Study develops curvature for contact-sequence networks, revealing temporal dynamics.
problem Lack of geometric analysis for temporal network sequences.
method Develops Forman--Ricci curvature on spatiotemporal prism complexes.
result Two curvature variants disagree on 56-67% of temporal edges.
ByteGen models LOB dynamics without tokenization, achieving realistic market metrics.
problem Modeling high-frequency LOB dynamics in finance.
method Autoregressive next-byte prediction on packed binary data, using H-Net architecture.
result Successfully reproduces stylized facts of financial markets.
Exact asymptotic solutions found for nonlinear Hawkes processes.
problem Analytical solutions for nonlinear Hawkes processes with positive and negative feedbacks.
method Field master equation approach to classify steady-state solutions.
result Explicit power law formulas for steady-state intensity distributions Pss(λ)∝λ−1−a, with a as a function of parameters. Introduces bounded scale measure and generalizes property A.
problem Defining property A for large scale spaces with bounded geometry.
method Introduces bounded scale measure, shows its coarse invariance, and generalizes property A.
result Definition of property A for large scale spaces with bounded scale measure is a coarse invariant.
Transformers can scale both context and task, but MLPs can only scale task.
problem Understanding and scaling In-Context Learning in transformers.
method Simplified transformer architecture, feature map, and MLP combination.
result Simplified transformer can perform ICL and context-scaling but not task-scaling.
New scaling framework for MoE architectures ensures stability and optimal performance at scale.
problem Lack of principled understanding of how hyperparameters should scale in MoE architectures.
method Developed a novel Dynamical Mean Field Theory (DMFT) for three scaling regimes of MoE architectures.
result Derived Maximally Scale-Stable Parameterization (MSSP) for SGD and Adam, providing robust learning rate transfer and monotonic improvement with scale.
New principles needed for scaling large language models, challenging traditional regularization methods.
problem The shift from generalization to scaling in machine learning requires new guiding principles.
method Examining the effectiveness of traditional regularization methods in the scaling-centric era.
result Traditional principles of regularization may not generalize to larger scales, highlighting new phenomena like scaling law crossover.
Improves U-Net for scale equivariance in semantic segmentation.
problem Improving generalization in semantic segmentation tasks with varying scales.
method Introduces Scale Equivariant U-Net (SEU-Net) with carefully applied subsampling and upsampling layers and scale-equivariant layers.
result Significantly improved generalization to different scales compared to U-Net and scale-equivariant architecture without upsampling.
DSS networks use scale-equivariant cross-correlations to improve image recognition.
problem Improving image recognition by exploiting scale invariance.
method Constructing scale-equivariant cross-correlations based on scale-spaces and semigroups.
result Demonstrated utility on Patch Camelyon and Cityscapes datasets.
Scale-equivariant CNNs handle scale changes for improved performance.
problem Translation equivariance is not sufficient for handling scale changes in CNNs.
method Developed scale-equivariant convolutional networks with steerable filters.
result Demonstrated state-of-the-art results on MNIST-scale and STL-10 datasets.
Introduces resemblance structure for large scale geometry.
problem Defining similarity in large scale geometry.
method Axiomatizing the concept of resemblance for subsets of a set.
result Large scale resemblance structures can induce nearness and generalize large scale properties.
New scaling laws optimize model size, training, and inference for better performance.
problem Trade-off between model size and inference cost in modern LLMs.
method Train-to-Test (T2) scaling laws that jointly optimize model size, training tokens, and inference samples. result Optimal pretraining decisions shift into overtraining regime, leading to stronger performance.
Derives a family of hyperparameter scaling strategies for neural networks.
problem Optimizing hyperparameters for wide and deep neural networks.
method Introduces a one-parameter family of hyperparameter scaling strategies.
result Reveals proper scaling of depth with width for large-scale models.
Small intrinsic scale reveals network structure.
problem Understanding the scale at which network identity is revealed.
method Defined intrinsic scale as distinguishability of subgraphs in random walks.
result Intrinsic scale is surprisingly small (7-20 vertices) across various networks.
To elucidate allometric scaling in complex systems, we investigated the underlying scaling relationships between typical three-scale indicators for approximately 500,000 Japanese firms; namely, annual sales, number of employees, and number of business partners. First, new scaling relations including the distributions o…
Proposes a new two-stage scaling method for data preprocessing.
problem Data scaling in statistical learning models.
method Two-stage scaling method: linear regression fitting followed by data scaling.
result Advantages of the new scaling method demonstrated through simulations and real data analysis.
New methods maintain scale to speed up deep learning.
problem Improper scaling between layers causes exploding gradients in deep neural networks.
method Two methods of maintaining isometry (exact and stochastic) are proposed.
result Maintaining scale speeds up learning, especially in the early stages.
Study calculates tail risk for various mixture distributions.
problem Estimating tail risk for complex distribution mixtures.
method Analyzes tail conditional expectation for location-scale mixtures of elliptical distributions.
result Developed methods for calculating tail risk in various distributions.
A new robust scaling approach improves downstream metabolomics analysis.
problem Challenges in choosing scaling techniques for metabolomics data.
method Introduces a weighted scaling approach robust to outliers.
result The proposed method outperforms traditional scaling techniques in both outlier-free and outlier-present datasets.
Novel neural network solves PDEs with multi-scale resolution.
problem Solving time-dependent PDEs with varying spatial and temporal scales.
method Multi-scale message passing neural network with temporal and spatial gating modules.
result Outperforms baselines on PDEs with diverse scales.
This study examines how reward scaling impacts non-saturating ReLU networks in reinforcement learning.
problem The impact of reward scaling on non-saturating ReLU networks in reinforcement learning.
method Proposes an Adaptive Network Scaling framework to find a suitable reward scale during learning.
result Empirical studies justify the effectiveness of the Adaptive Network Scaling framework.
New pruning method breaks power law scaling, potentially reducing error to exponential.
problem Improving neural network performance through scaling alone is costly.
method Developed a new data pruning metric to break power law scaling.
result Pruned datasets show better than power law scaling on various image datasets.