Consider a general machine learning setting where the output is a set of labels or sequences. This output set is unordered and its size varies with the input. Whereas multi-label classification methods seem a natural first resort, they are not readily applicable to set-valued outputs because of the growth rate of the o…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Seq-SetNet processes sequence sets directly, improving protein structure prediction.
Set-Sequence model learns cross-sectional dynamics directly from time series data.
Study Thompson Sampling in adversarial bit prediction, finding regret bounds and optimal sequences.
A novel transformer model improves classification of partially ordered sequences.
Study continuity of limit sets in symmetric spaces.
Data of practical interest - such as personal records, transaction logs, and medical histories - are sequential collections of events relevant to a particular source entity. Recent studies have attempted to link sequences that represent a common entity across data sets to allow more comprehensive statistical analyses a…
Sequences have become first class citizens in supervised learning thanks to the resurgence of recurrent neural networks. Many complex tasks that require mapping from or to a sequence of observations can now be formulated with the sequence-to-sequence (seq2seq) framework which employs the chain rule to efficiently repre…
We obtain an index of the complexity of a random sequence by allowing the role of the measure in classical probability theory to be played by a function we call the generating mechanism. Typically, this generating mechanism will be a finite automata. We generate a set of biased sequences by applying a finite state auto…
Predicts node sequences in graphs using multi-order network models.
Study concordance of decompositions from defining sequences in 3-sphere.
New algorithm minimizes expert selection regret in partial bandit feedback.
Using sequence to sequence algorithms for query expansion has not been explored yet in Information Retrieval literature nor in Question-Answering's. We tried to fill this gap in the literature with a custom Query Expansion engine trained and tested on open datasets. Starting from open datasets, we built a Query Expansi…
In this paper we consider monopoles on an asymptotically conical, oriented, Riemannian -manifold with one end. The connected components of the moduli space of monopoles in this setting are labeled by an integer called the charge. We analyse the limiting behavior of sequences of monopoles with fixed charg…
Rapid progress in deep learning has spurred its application to bioinformatics problems including protein structure prediction and design. In classic machine learning problems like computer vision, progress has been driven by standardized data sets that facilitate fair assessment of new methods and lower the barrier to …
Constructs Serre spectral sequence for bounded cohomology.
Machine learning and data mining techniques have been used extensively in order to detect credit card frauds. However, most studies consider credit card transactions as isolated events and not as a sequence of transactions. In this article, we model a sequence of credit card transactions from three different perspectiv…
Sequence to sequence learning has recently emerged as a new paradigm in supervised learning. To date, most of its applications focused on only one task and not much work explored this framework for multiple tasks. This paper examines three multi-task learning (MTL) settings for sequence to sequence models: (a) the onet…
Amino acid sequence portrays most intrinsic form of a protein and expresses primary structure of protein. The order of amino acids in a sequence enables a protein to acquire a particular stable conformation that is responsible for the functions of the protein. This relationship between a sequence and its function motiv…
Deep RL optimizes compiler passes for better performance.
Transformers can interpolate finite input sequences exactly.
The problem of universal outlying sequence detection is studied, where the goal is to detect outlying sequences among sequences of samples. A sequence is considered as outlying if the observations therein are generated by a distribution different from those generating the observations in the majority of the sequenc…
Paper tackles embedding attributed sequences in unsupervised learning.
An adversarial detector identifies anomalous sequences in sequential data.
We suggest a novel method of clustering and exploratory analysis of temporal event sequences data (also known as categorical time series) based on three-dimensional data grid models. A data set of temporal event sequences can be represented as a data set of three-dimensional points, each point is defined by three varia…
ProGen models protein sequences for synthetic biology.
This paper proposes a generative model, the latent Dirichlet hidden Markov models (LDHMM), for characterizing a database of sequential behaviors (sequences). LDHMMs posit that each sequence is generated by an underlying Markov chain process, which are controlled by the corresponding parameters (i.e., the initial state …
The paper solves the isoperimetric problem on noncompact metric spaces with lower Ricci bounds.
Adversarial learning for mixture Hawkes processes improves performance.
Develops methods to create manifolds with positive scalar curvature.
Differentiable losses for combinatorial optimization problems in sequence modeling.
The paper extends confidence sequences for infinite variance data.
Continuous MDS embeds sequences of dissimilarities in Euclidean space.
Attention is an operation that selects some largest element from some set, where the notion of largest is defined elsewhere. Applying this operation to sequence to sequence mapping results in significant improvements to the task at hand. In this paper we provide the mathematical definition of attention and examine its …
For germs of subanalytic sets, we define two finite sequences of new numerical invariants. The first one is obtained by localizing the classical Lipschitz-Killing curvatures, the second one is the real analogue of the evanescent characteristics introduced by M. Kashiwara. We show that each invariant of one sequence is …
Optimal Farey sequence for with upper bound .
The paper develops a convex parameterization for robust RNNs ensuring stability and robustness.
We present methods for online linear optimization that take advantage of benign (as opposed to worst-case) sequences. Specifically if the sequence encountered by the learner is described well by a known "predictable process", the algorithms presented enjoy tighter bounds as compared to the typical worst case bounds. Ad…
We show that for a strongly convergent sequence of purely loxodromic finitely generated Kleinian groups with incompressible ends, Cannon-Thurston maps, viewed as maps from a fixed base limit set to the Riemann sphere, converge uniformly. For algebraically convergent sequences we show that there exist examples where eve…
Paper proposes using LSTM for LSH-based sequence alignment.
When analyzing the genome, researchers have discovered that proteins bind to DNA based on certain patterns of the DNA sequence known as "motifs". However, it is difficult to manually construct motifs due to their complexity. Recently, externally learned memory models have proven to be effective methods for reasoning ov…
Given a sequence of curves on a surface, we provide conditions which ensure that (1) the sequence is an infinite quasi-geodesic in the curve complex, (2) the limit in the Gromov boundary is represented by a nonuniquely ergodic ending lamination, and (3) the sequence divides into a finite set of subsequences, each of wh…
Transformers can simulate MLE for Bayesian network sequences.
Learned factor graphs improve inference from time sequences using neural networks.
Improved algorithms for stochastic linear bandits using tighter confidence sequences.
Study SGD dynamics in sequence models, revealing training phases and influence of sequence length.
Microbial clades modeling is a challenging problem in biology based on microarray genome sequences, especially in new species gene isolates discovery and category. Marker family genome sequences play important roles in describing specific microbial clades within species, a framework of support vector machine (SVM) base…
P3BO optimizes biological sequence design by combining multiple methods.