This work compresses sequences by treating them as continuous-time processes, enabling efficient discretization.
problem Efficient compression of sequences, especially with deep learning models that scale with sequence length.
method Treat sequences as continuous-time processes, learn efficient discretization, and decode at different time intervals.
result Automatic bit rate reductions in video and motion capture sequences using learned discretization.
Study on continuous sequence classification with distribution uncertainty.
problem Classifying continuous sequences with varying distribution uncertainty.
method Proposes distribution-free tests for three test designs: fixed-length, sequential, and two-phase tests.
result Error probabilities decay exponentially fast for all test designs.
Neural SDEs model continuous sequences using neural networks.
problem Modeling continuous-time dynamics in sequence data.
method Interprets time-series as samples from a continuous dynamical system, parameterized by Neural SDE.
result Demonstrates superior performance in diverse sequence modeling tasks.
Proposes a continuous relaxation for discrete Bayesian optimization.
problem Efficiently optimizing discrete data with limited target observations.
method Continuous relaxation of objective function, incorporating prior knowledge.
result Optimization can be computationally tractable with few observations.
Continuous-time event sequences represent discrete events occurring in continuous time. Such sequences arise frequently in real-life. Usually we expect the sequences to follow some regular pattern over time. However, sometimes these patterns may be interrupted by unexpected absence or occurrences of events. Identificat…
A new method learns from multi-modal sequences with external memory.
problem Learning new modes in a dynamic environment without prior knowledge.
method Maintains a neural episodic memory with a Dirichlet Process prior to store mode descriptors and transfers knowledge through retrieval.
result Performs continual learning favorably compared to mainstream approaches.
We study the relationship between catastrophic forgetting and properties of task sequences. In particular, given a sequence of tasks, we would like to understand which properties of this sequence influence the error rates of continual learning algorithms trained on the sequence. To this end, we propose a new procedure …
Study continuity of limit sets in symmetric spaces.
problem Continuity of limit sets for geometrically finite subgroups in symmetric spaces.
method Extended geometrically finite representations theory.
result Limit sets vary continuously with respect to Hausdorff distance under strong convergence.
Proves solution uniqueness for biomembrane shape prediction.
problem Proving solution uniqueness for the genus one Canham variational problem.
method Combining numeric analytic continuation and singularity analysis to prove non-negativity of a sequence.
result Proves positivity of the sequence, leading to solution uniqueness.
Events in the world may be caused by other, unobserved events. We consider sequences of events in continuous time. Given a probability model of complete sequences, we propose particle smoothing---a form of sequential importance sampling---to impute the missing events in an incomplete sequence. We develop a trainable fa…
Hybrid framework prevents forgetting in continual learning.
problem Avoiding forgetting in learning new tasks without forgetting old ones.
method Hybrid continual learning framework combining architecture growth and experience replay.
result Hybrid approach effectively avoids forgetting across multiple tasks.
The paper tackles finding optimal treatment sequences in continuous state spaces.
problem Finding counterfactually optimal action sequences in continuous state spaces.
method Formalizes the problem using finite horizon Markov decision processes and structural causal models. Develops a search method based on the A* algorithm.
result The method can find optimal action sequences in polynomial time under certain conditions.
Paper introduces efficient methods for probabilistic querying of event sequences.
problem Hard queries about future events in continuous-time sequences.
method Importance sampling framework for addressing query types.
result Method is more efficient than naive simulation, often 1,000 times.
Continuous time analysis of bubble formation in harmonic maps.
problem Understanding bubble formation in harmonic map heat flow.
method Continuous time approach to analyze bubbling sequences.
result Solutions approach multi-bubble configurations in continuous time.
Continuous MDS embeds sequences of dissimilarities in Euclidean space.
problem Embedding sequences of dissimilarities as n increases. method Continuous MDS reformulates MDS for sequences of dissimilarity matrices.
result Uniform convergence of interpolated embeddings.
Despite the widespread adoption of Transformer models for NLP tasks, the expressive power of these models is not well-understood. In this paper, we establish that Transformer models are universal approximators of continuous permutation equivariant sequence-to-sequence functions with compact support, which is quite surp…
DLM-One speeds up language generation by 500x with continuous models.
problem Efficiently generating text sequences in natural language processing.
method Score-distillation of continuous diffusion language models.
result Achieves up to 500x speedup in inference time with competitive performance.
The Softmax function is used in the final layer of nearly all existing sequence-to-sequence models for language generation. However, it is usually the slowest layer to compute which limits the vocabulary size to a subset of most frequent types; and it has a large memory footprint. We propose a general technique for rep…
Develops a framework for modeling set-valued data in continuous-time.
problem Handling sequences where each event is associated with a set of items.
method General framework for modeling set-valued data, developed inference methods, and importance sampling techniques.
result Orders-of-magnitude improvements in efficiency for probabilistic queries over direct sampling.
The definition of the grafting operation for quasifuchsian groups is extended by Bromberg to all b-groups. Although the grafting maps are not necessarily continuous at boundary groups, in this paper, we show that the grafting maps take every "standard" convergent sequence to a convergent sequence. As a consequence of…
ADD-THIN improves TPP forecasting by handling long-term data sequences.
problem Sequential limitations in autoregressive models for TPPs.
method Diffusion model for TPPs that operates on entire sequences.
result ADD-THIN outperforms state-of-the-art models in forecasting.
This paper introduces a new neural ODE model for continuous-time sequence generation.
problem Representing and predicting continuous-time sequences with high accuracy.
method A neural emission model and neural ODE define the latent state evolution, with an Energy-based model for prior distribution.
result The model outperforms existing methods in various tasks, including long-horizon predictions.
Study shows memory needs grow with task sequence length in continual learning.
problem Challenges in retaining aptitude for multiple learning tasks sequentially.
method Complexity-theoretic study using communication complexity and multiplicative weights update.
result Memory needs grow linearly with task sequence length, suggesting intractability.
New methods tackle adversarial attacks on categorical sequences, improving model security.
problem Adversarial attacks on categorical sequence models, especially for money transactions and medical fraud.
method Two black-box adversarial attacks: Monte-Carlo and continuous relaxation methods.
result Generated adversarial sequences fool machine learning models but remain close to original ones.
In this paper we investigate continuous speech recognition using electroencephalography (EEG) features using recently introduced end-to-end transformer based automatic speech recognition (ASR) model. Our results demonstrate that transformer based model demonstrate faster training compared to recurrent neural network (R…
Study shows volumes of knot complements are bounded by linear functions of geodesic periods.
problem Volume calculation of knot complements associated with geodesics on modular surfaces.
method Analyzes geodesics on modular surfaces, their associated knots, and their complements' volumes.
result Volumes of knot complements are bounded linearly by the period of geodesic continued fractions.
The study finds new infinite dilogarithm identities related to number sequences and continued fractions.
problem Finding new infinite dilogarithm identities.
method Demonstrating families of identities associated with specific number sequences and continued fractions.
result New infinite dilogarithm identities related to Fibonacci, Lucas numbers, convergents of even period continued fractions, and recurrence relations.
VAIOM models financial returns using continuous input and categorical output.
problem Modeling continuous, noisy, and heterogeneous financial data.
method VAIOM is a decoder-only Transformer that separates input representation from output likelihood.
result VAIOM models outperform fixed single-bar LightGBM baseline in both Test halves.
By a fixed continuous map from a 3-space to itself, a knot in the 3-space may be mapped to another knot in the 3-space. We analyze possible knot types of them. Then we map a knot repeatedly by a fixed continuous map and analyze possible infinite sequences of knot types.
New flexible confidence sequences for robust statistical inference.
problem Creating robust statistical inference methods that work under mild assumptions.
method Proposed a new class of asymptotic time-uniform confidence sequences.
result Sharp asymptotic time-uniform confidence sequences achieved under mild assumptions.
ARCADe detects anomalies in a sequence of tasks with limited data.
problem Learning a sequence of anomaly detection tasks with only normal class examples.
method Formulated as a meta-learning problem, ARCADe addresses catastrophic forgetting and overfitting.
result ARCADe outperforms baselines on three datasets.
Using techniques from the theories of convex polytopes, lattice paths, and indirect influences on directed manifolds, we construct continuous analogues for the binomial coefficients and the Catalan numbers. Our approach for constructing these analogues can be applied to a wide variety of combinatorial sequences. As an …
New method for learning evolving tasks with performance guarantees.
problem Learning tasks in a sequence with evolving similarity.
method Adaptable learning methodology with performance guarantees.
result Improved performance in multiple scenarios with reliable guarantees.
New analysis shows RPE-based Transformers can't approximate all functions.
problem Understanding the limitations of RPE-based Transformers in approximating continuous functions.
method Mathematical analysis and development of a novel attention module (URPE) to overcome limitations.
result RPE-based Transformers can't approximate all continuous sequence-to-sequence functions, even with depth and width.
DFMs enable flow-based models for multimodal discrete and continuous data.
problem Combining discrete and continuous data for generative models.
method Discrete Flow Models (DFMs) using Continuous Time Markov Chains.
result DFMs achieve state-of-the-art co-design performance for protein structure and sequence generation.
S2P2 model improves predictive likelihoods for MTPPs.
problem Modeling irregular time intervals in event sequences.
method State-space point process model using deep state-space techniques.
result Empirically, S2P2 achieves state-of-the-art predictive likelihoods.
Rough Transformers improve time series modeling with lower costs and better performance.
problem Inefficient modeling of irregularly sampled time series data.
method Signature patching for continuous-time representations, reducing computational costs.
result Rough Transformers outperform vanilla Transformers and Neural ODE models.
We present a homogenization theorem for isotropically-distributed point defects, by considering a sequence of manifolds with increasingly dense point defects. The loci of the defects are chosen randomly according to a weighted Poisson point process, making it a continuous version of the first passage percolation model.…
Study on neural networks forgetting in continual learning.
problem Forgetting in neural networks during continual learning.
method Gradient descent analysis on XOR-cluster datasets.
result Explicit bounds on forgetting rate and generalization gap.
We study the infimum of the renormalized volume for convex-cocompact hyperbolic manifolds, as well as describing how a sequence converging to such values behaves. In particular, we show that the renormalized volume is continuous under the appropriate notion of limit. This result generalizes previous work in the subject…
We establish the weak continuity of the Gauss-Coddazi-Ricci system for isometric embedding with respect to the uniform Lp-bounded solution sequence for p>2, which implies that the weak limit of the isometric embeddings of the manifold is still an isometric embedding. More generally, we establish a compensated comp…
The paper proves continuity of Morse index for Ricci shrinkers.
problem Lower and upper semi-continuity of the Morse index for gradient Ricci shrinkers.
method Adapting and refining recent arguments on CMC hypersurfaces and polynomially weighted Sobolev spaces, with techniques for non-compact shrinkers.
result Identifies a condition ensuring the Morse index of asymptotically conical shrinkers is bounded below by the f-index of their asymptotic cone.
Improved neural models for diverse user event sequences.
problem Challenges in modeling diverse user event sequences.
method Mixtures of latent embeddings with amortized variational inference.
result Systematic improvements over existing work for various predictive metrics.
The paper tackles online learning problems with monotone arm sequences, achieving optimal or near-optimal regret bounds.
problem Online learning problems with ordinal and monotone arm sequences, such as dynamic pricing and clinical trials.
method Proposes algorithms for continuum-armed bandit problems with monotone arm sequences, achieving optimal or near-optimal regret bounds.
result Achieves optimal or near-optimal regret bounds for monotone arm sequences, differing from the continuous-armed bandit literature.
This primer explains diffusion models in general state spaces.
problem Diffusion models in general state spaces are not well-introduced.
method Develops discrete-time and continuous-time views of diffusion models, deriving Fokker-Planck and master equations.
result Unified understanding of diffusion models across continuous and discrete domains.
Diffusion models generate music sequences without autoregressive loops.
problem Generating music sequences from symbolic data using diffusion models.
method Parameterize discrete symbolic data in continuous latent space, train diffusion model, generate sequences through reverse process.
result Strong unconditional generation and post-hoc conditional infilling compared to autoregressive models.
We consider a binary sequence generated by thresholding a hidden continuous sequence. The hidden variables are assumed to have a compound symmetry covariance structure with a single parameter characterizing the common correlation. We study the parameter estimation problem under such one-parameter models. We demonstrate…
Paper tackles missing data in irregularly-sampled time series.
problem Modeling irregularly-sampled time series data.
method Encoder-decoder framework based on variational autoencoders and generative adversarial networks.
result Models achieve competitive or better classification results on irregularly-sampled multivariate time series.