XL-Editor improves sentence post-editing using XLNet's variable-length insertion probability.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Transformer improves sequence generation with insertion and deletion phases.
We present the Insertion Transformer, an iterative, partially autoregressive model for sequence generation based on insertion operations. Unlike typical autoregressive models which rely on a fixed, often left-to-right ordering of the output, our approach accommodates arbitrary orderings by allowing for tokens to be ins…
Paper tackles variable-length, incomplete wearable sensor data to improve personalized insights.
ID-ExpO fine-tunes neural networks for more faithful explanations.
In this work we explore the use of metric index structures, which accelerate nearest neighbor queries, in the scenario where we need to interleave insertions and queries during deployment. This use-case is inspired by a real-life need in malware analysis triage, and is surprisingly understudied. Existing literature ten…
New framework for consistent submodular maximization with insertions and deletions.
SummerTime summarizes variable-length time series for machine learning applications.
New attacks improve privacy audits by analyzing model updates.
Time series constitute a challenging data type for machine learning algorithms, due to their highly variable lengths and sparse labeling in practice. In this paper, we tackle this challenge by proposing an unsupervised method to learn universal embeddings of time series. Unlike previous works, it is scalable with respe…
New coding theorem shows achievable rate matches theoretical limit.
The paper compares inserting and stretching points for grid refinement near critical points.
ALT transforms time series data for better classification.
Shared Keyboard design improves phase I clinical trials by borrowing information across doses.
One of the ubiquitous representation of long DNA sequence is dividing it into shorter k-mer components. Unfortunately, the straightforward vector encoding of k-mer as a one-hot vector is vulnerable to the curse of dimensionality. Worse yet, the distance between any pair of one-hot vectors is equidistant. This is partic…
New formulas for feature importance tests in regression models.
IFH models graph generation with adjustable sequentiality.
Study on elastic curves with variable stiffness, derived from bending energy.
The task of clustering unlabeled time series and sequences entails a particular set of challenges, namely to adequately model temporal relations and variable sequence lengths. If these challenges are not properly handled, the resulting clusters might be of suboptimal quality. As a key solution, we present a joint clust…
A new kernel Stein test assesses fit for variable-length sequential data.
Universal perturbations misclassify text with high accuracy.
A well-known problem in data science and machine learning is {\em linear regression}, which is recently extended to dynamic graphs. Existing exact algorithms for updating the solution of dynamic graph regression require at least a linear time (in terms of : the size of the graph). However, this time complexity might…
We study topology of configuration spaces of planar linkages having one leg of variable length. Such telescopic legs are common in modern robotics where they are used for shock absorbtion and serve a variety of other purposes. Using a Morse theoretic technique, we compute explicitly, in terms of the metric data, the Be…
Current end-to-end deep Reinforcement Learning (RL) approaches require jointly learning perception, decision-making and low-level control from very sparse reward signals and high-dimensional inputs, with little capability of incorporating prior knowledge. This results in prohibitively long training times for use on rea…
SentenceMIM learns rich latent representations for variable-length language data.
Study curvature and torsion from cross-ratios in discrete curves.
We introduce backdrop, a flexible and simple-to-implement method, intuitively described as dropout acting only along the backpropagation pipeline. Backdrop is implemented via one or more masking layers which are inserted at specific points along the network. Each backdrop masking layer acts as the identity in the forwa…
While machine learning (ML) models are being increasingly trusted to make decisions in different and varying areas, the safety of systems using such models has become an increasing concern. In particular, ML models are often trained on data from potentially untrustworthy sources, providing adversaries with the opportun…
We design and study a Contextual Memory Tree (CMT), a learning memory controller that inserts new memories into an experience store of unbounded size. It is designed to efficiently query for memories from that store, supporting logarithmic time insertion and retrieval operations. Hence CMT can be integrated into existi…
We give a proof of Ilmanen's lemma, which asserts that between a locally semi-convex and a locally semi-concave function it is possible to find a C function.
New linear models improve time series classification efficiency and interpretability.
Inserts proximal mapping into deep networks for better regularization.
We describe our first-place solution to the Animal Behavior Challenge (ABC 2018) on predicting gender of bird from its GPS trajectory. The task consisted in predicting the gender of shearwater based on how they navigate themselves across a big ocean. The trajectories are collected from GPS loggers attached on shearwate…
Large-scale graph data in real-world applications is often not static but dynamic, i. e., new nodes and edges appear over time. Current graph convolution approaches are promising, especially, when all the graph's nodes and edges are available during training. When unseen nodes and edges are inserted after training, it …
Inserting label noise can improve model accuracy and fairness.
We prove that the property of admitting no cosmetic crossing changes is preserved under the operation of forming certain satellites of winding number zero. We also define strongly cosmetic crossing changes and we discuss their behavior under the operation of inserting full twists in the strings of closed braids.
We present the Latent Sequence Decompositions (LSD) framework. LSD decomposes sequences with variable lengthed output units as a function of both the input sequence and the output sequence. We present a training algorithm which samples valid extensions and an approximate decoding algorithm. We experiment with the Wall …
Tree-based LSTM improves sequential regression with missing data.
This is an English translation of the following paper, published several years ago: Nikonorov Yu.G. On the geodesic diameter of surfaces with involutive isometry (Russian), Tr. Rubtsovsk. Ind. Inst., 2001, V. 9, 62-65, Zbl. 1015.53041. All inserted footnotes provide additional information related to the mentioned probl…
We investigate anomaly detection in an unsupervised framework and introduce Long Short Term Memory (LSTM) neural network based algorithms. In particular, given variable length data sequences, we first pass these sequences through our LSTM based structure and obtain fixed length sequences. We then find a decision functi…
Bayesian optimization reduces RNN architecture search time.
Ensemble method detects time series anomalies without preselecting parameter values.
The notion of a pseudoknot is defined as an equivalence class of knot diagrams that may be missing some crossing information. We provide here a topological invariant schema for pseudoknots and their relatives, 4-valent rigid vertex spatial graphs and singular knots, that is obtained by replacing unknown crossings or ve…
When environmental interaction is expensive, model-based reinforcement learning offers a solution by planning ahead and avoiding costly mistakes. Model-based agents typically learn a single-step transition model. In this paper, we propose a multi-step model that predicts the outcome of an action sequence with variable …
This paper explores using a Long short-term memory (LSTM) based sequence autoencoder to learn interesting features for detecting surveillance aircraft using ADS-B flight data. An aircraft periodically broadcasts ADS-B (Automatic Dependent Surveillance - Broadcast) data to ground receivers. The ability of LSTM networks …
TimeAutoML learns effective representations for irregularly sampled MTS data without manual tuning.
We propose learning flexible but interpretable functions that aggregate a variable-length set of permutation-invariant feature vectors to predict a label. We use a deep lattice network model so we can architect the model structure to enhance interpretability, and add monotonicity constraints between inputs-and-outputs.…
This study provides benchmarks for different implementations of LSTM units between the deep learning frameworks PyTorch, TensorFlow, Lasagne and Keras. The comparison includes cuDNN LSTMs, fused LSTM variants and less optimized, but more flexible LSTM implementations. The benchmarks reflect two typical scenarios for au…