Sequence-to-sequence models with soft attention have been successfully applied to a wide variety of problems, but their decoding process incurs a quadratic time and space cost and is inapplicable to real-time sequence transduction. To address these issues, we propose Monotonic Chunkwise Attention (MoChA), which adaptiv…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper tackles online learning problems with monotone arm sequences, achieving optimal or near-optimal regret bounds.
Holonomy groups of metric connections converge in a monotonic way.
Proves existence of planar curves with specific curvature.
We define an integer graded symplectic Floer cohomology and a spectral sequence which are new invariants for monotone Lagrangian sub-manifolds and exact isotopies. Such an integer graded Floer cohomology is an integral lifting of the usual Floer-Oh cohomology with $Z_{\Si (L)}$ grading. As one of applications of the sp…
Unimodal sequences of moves connect 3-manifold triangulations.
New method learns low-dimensional representations of nonlinear time series without supervision.
Single-timescale analysis improves convergence in multi-sequence stochastic approximation.
Let C_n(M) be the configuration space of n distinct ordered points in M. We prove that if M is any connected orientable manifold (closed or open), the homology groups H_i(C_n(M); Q) are representation stable in the sense of [Church-Farb]. Applying this to the trivial representation, we obtain as a corollary that the un…
We define an integer graded symplectic Floer cohomology and a Fintushel-Stern type spectral sequence which are new invariants for monotone Lagrangian sub-manifolds and exact isotopes. The Z-graded symplectic Floer cohomology is an integral lifting of the usual Z_Sigma(L)-graded Floer-Oh cohomology. We prove the Kunneth…
Floer homotopy theory applies to Lagrangians, overcoming curvature issues.
This paper analyzes and improves monotonic accelerated algorithms like M-NAG and M-FISTA.
In this paper, we establish a general monotonicity formula of the following elliptic system $$ Δu_i+f_i(u_1,...,u_m)=0 \quad {\rm in} Ω, \label{0.1} $$ where is a bounded domain, , and is a given smooth function of …
Framework for online resource allocation using social welfare functions.
Study online monotone density estimation with expert aggregation and log-optimal calibration.
New algorithms avoid non-monotonic risk curves in statistical learning.
Investigates probability of error in structured thresholding bandit problems.
In this paper, we consider first-order convergence theory and algorithms for solving a class of non-convex non-concave min-max saddle-point problems, whose objective function is weakly convex in the variables of minimization and weakly concave in the variables of maximization. It has many important applications in mach…
This note has an experimental nature and contains no new theorems. We introduce certain moves for classical knot diagrams that for all the very many examples we have tested them on give a monotonic complete simplification. A complete simplification of a knot diagram D is a sequence of moves that transform D into a diag…
Traditional pairwise sequence alignment is based on matching individual samples from two sequences, under time monotonicity constraints. However, in many application settings matching subsequences (segments) instead of individual samples may bring in additional robustness to noise or local non-causal perturbations. Thi…
The paper studies convergence of cosmological spacetimes using null distance.
Undirected neural sequence models such as BERT (Devlin et al., 2019) have received renewed interest due to their success on discriminative natural language understanding tasks such as question-answering and natural language inference. The problem of generating sequences directly from these models has received relativel…
Study de Rham homomorphism for Lipschitz cohomologies on metric simplicial complexes.
In this paper, we present Neural Phrase-based Machine Translation (NPMT). Our method explicitly models the phrase structures in output sequences using Sleep-WAke Networks (SWAN), a recently proposed segmentation-based sequence modeling method. To mitigate the monotonic alignment requirement of SWAN, we introduce a new …
We analyze the signature type of a cascade of periodic orbits associated to period doubling renormalizable maps of the two dimensional disk. The signature is a sequence of rational numbers which describes how periodic orbits turn each other and is invariant by topological conjugacies that preserve orientation. We prove…
This paper studies rapidly forming singularities in the Yang-Mills flow. It is shown that a sequence of blow-ups near the singular point converges, modulo the gauge group, to a homothetically shrinking soliton with non-zero curvature. The proof uses Hamilton's monotonicity formula. Examples of homothetically shrinking …
New analysis of annealing paths in sampling and estimation.
Adaptive learning rate improves FTRL's performance in online learning.
We propose a new framework for how to use sequential Monte Carlo (SMC) algorithms for inference in probabilistic graphical models (PGM). Via a sequential decomposition of the PGM we find a sequence of auxiliary distributions defined on a monotonically increasing sequence of probability spaces. By targeting these auxili…
In this paper we consider a general matrix factorization model which covers a large class of existing models with many applications in areas such as machine learning and imaging sciences. To solve this possibly nonconvex, nonsmooth and non-Lipschitz problem, we develop a non-monotone alternating updating method based o…
This paper presents some partial answers to the following question. QUESTION. If a normal space X is the union of an increasing sequence of open sets U(1), U(2), U(3) ... such that each U(n) contracts to a point in X, must X be contractible? The main results of the paper are: THEOREM 1. If a normal space X is the union…
Meta-learning improves event prediction from short sequences.
The paper examines sequences of metric spaces converging to compact limits with specific properties.
Graph Shift (GS) algorithms are recently focused as a promising approach for discovering dense subgraphs in noisy data. However, there are no theoretical foundations for proving the convergence of the GS Algorithm. In this paper, we propose a generic theoretical framework consisting of three key GS components: simplex …
A key problem in reinforcement learning for control with general function approximators (such as deep neural networks and other nonlinear functions) is that, for many algorithms employed in practice, updates to the policy or -function may fail to improve performance---or worse, actually cause the policy performance …
Optimizes profit in targeted marketing across multiple markets with varying marketing expenditures.
Conformer encoder reverses sequence in time dimension, affecting decoder training.
The paper addresses monotonicity in machine learning models for fairness and accountability.
Bayesian approach improves online prediction accuracy without distributional assumptions.
We investigate a class of quadratic-exponential growth BSDEs with jumps. The quadratic structure introduced by Barrieu & El Karoui (2013) yields the universal bounds on the possible solutions. With local Lipschitz continuity and the so-called A_gamma-condition for the comparison principle to hold, we prove the existenc…
Paper develops an online covariance estimator for nonsmooth stochastic approximation problems.
Probit Monotone BART estimates binary outcomes using monotonic functions.
Study uses instanton Floer theory to obstruct knot unknotting operations.
Monotone neural networks can approximate and interpolate functions efficiently.
The Schwarz lemmas are well-known characterizations for holomorphic maps and we exhibit two examples of their applications. For a sequence family of biholomorphisms , it is useful to determine the location of for a fixed point in source manifolds (see Proposition \ref{2.5}). With it, we extend the For…
Probabilistic models are a critical part of the modern deep learning toolbox - ranging from generative models (VAEs, GANs), sequence to sequence models used in machine translation and speech processing to models over functional spaces (conditional neural processes, neural processes). Given the size and complexity of th…
Improves k-NN for monotonic data with robustness against noise.
Study examines explainable machine learning for monotonic models, finding Integrated gradients better for strong monotonicity.