Enhances relational reasoning with multi-layer architecture.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Many recent papers address reading comprehension, where examples consist of (question, passage, answer) tuples. Presumably, a model must combine information from both questions and passages to predict corresponding answers. However, despite intense interest in the topic, with hundreds of published papers vying for lead…
Machine reading using differentiable reasoning models has recently shown remarkable progress. In this context, End-to-End trainable Memory Networks, MemN2N, have demonstrated promising performance on simple natural language based reasoning tasks such as factual reasoning and basic deduction. However, other tasks, namel…
We further study the incidence relations that arise from the various subtowers, known as Baby Monster, which exist within the -Monster Tower. This allows us to complete the class spelling rules. We also present a method of calculating the various Baby Monster that appear within the Monster Tower.
MAC Net improves natural language question answering with data-driven reasoning.
Population growth (or decay) in a country can be due to various f socio-economic constraints, as demonstrated in this paper. For example, sexual intercourse is banned in various religions, during Nativity and Lent fasting periods. Data consisting of registered daily birth records for very long (35,429 points) time seri…
The Monster tower, also known as the Semple tower, is a sequence of manifolds with distributions of interest to both differential and algebraic geometers. Each manifold is a projective bundle over the previous. Moreover, each level is a fiber compactified jet bundle equipped with an action of finite jets of the diffeom…
Paper improves preterm birth prediction using neural networks with noisy labels.
New algorithms minimize dynamic regret for strongly convex losses.
Paper proposes a method to recover accurate labels from partially valid data in multi-label learning.
We highlight a pitfall when applying stochastic variational inference to general Bayesian networks. For global random variables approximated by an exponential family distribution, natural gradient steps, commonly starting from a unit length step size, are averaged to convergence. This useful insight into the scaling of…
The 20/60/20 rule improves risk management and portfolio optimization in finance.
For any positive integer r, we exhibit a knot Kr with (20 2 r--1 + 1) crossings whose Jones polynomial V (Kr) is equal to 1 mod-ulo 2 r. Our construction rests on a certain 20-crossing tangle T 20 which is undetectable by the Kauffman bracket polynomial pair mod 2.
Memory-augmented neural networks (MANNs) are designed for question-answering tasks. It is difficult to run a MANN effectively on accelerators designed for other neural networks (NNs), in particular on mobile devices, because MANNs require recurrent data paths and various types of operations related to external memory a…
Graph-structured data appears frequently in domains including chemistry, natural language semantics, social networks, and knowledge bases. In this work, we study feature learning techniques for graph-structured inputs. Our starting point is previous work on Graph Neural Networks (Scarselli et al., 2009), which we modif…
Pareto's 80/20 rule follows a Gaussian distribution with twice the mean standard deviation.
The Knowledge Base (KB) used for real-world applications, such as booking a movie or restaurant reservation, keeps changing over time. End-to-end neural networks trained for these task-oriented dialogs are expected to be immune to any changes in the KB. However, existing approaches breakdown when asked to handle such c…
We present a novel recurrent neural network (RNN) based model that combines the remembering ability of unitary RNNs with the ability of gated RNNs to effectively forget redundant/irrelevant information in its memory. We achieve this by extending unitary RNNs with a gating mechanism. Our model is able to outperform LSTM…
Let be the real form of a complex simple Jordan algebra such that the automorphism group is . By using some orbit types of on , for , explicitly, we give the Iwasawa decomposition, the Oshima--Sekiguchi's Iwasawa decomp…
Synthetic learning improves neonatal brain MRI segmentation robustness.
Oeljeklaus-Toma (OT) manifolds are certain compact complex manifolds built from number fields. Conversely, we show that the fundamental group often pins down the number field uniquely. We relate the first homology to some interesting ideal. OT manifolds are never Kähler, but carry an LCK metric (locally conformally Käh…
In this study, we investigate the limits of the current state of the art AI system for detecting buffer overflows and compare it with current static analysis tools. To do so, we developed a code generator, s-bAbI, capable of producing an arbitrarily large number of code samples of controlled complexity. We found that t…
Proves Montesinos-Nakanishi 3-move conjecture for links up to 20 crossings.
Every year, thousands of people receive consumer product related injuries. Research indicates that online customer reviews can be processed to autonomously identify product safety issues. Early identification of safety issues can lead to earlier recalls, and thus fewer injuries and deaths. A dataset of product reviews …
Fair quantile regression adjusts estimators to balance subpopulation quantiles.
Neural nets solve braid untangling up to length 20.
Proposes VILMAP for finding motifs and segmenting words in time series.
Deep Learning can significantly benefit cancer proteomics and genomics. In this study, we attempt to determine a set of critical proteins that are associated with the FLT3-ITD mutation in newly-diagnosed acute myeloid leukemia patients. A Deep Learning network consisting of autoencoders forming a hierarchical model fro…
The paper studies hyperkähler structures and adapted complex structures using the Monge-Ampère equation.
The divergence theorem in its usual form applies only to suitably smooth vector fields. For vector fields which are merely piecewise smooth, as is natural at a boundary between regions with different physical properties, one must patch together the divergence theorem applied separately in each region. We give an elegan…
The celebrated Sequence to Sequence learning (Seq2Seq) technique and its numerous variants achieve excellent performance on many tasks. However, many machine learning tasks have inputs naturally represented as graphs; existing Seq2Seq models face a significant challenge in achieving accurate conversion from graph form …
New algorithms minimize dynamic regret in non-stationary online learning.
In Peña (2007), MCMC sampling is applied to approximately calculate the ratio of essential graphs (EGs) to directed acyclic graphs (DAGs) for up to 20 nodes. In the present paper, we extend that work from 20 to 31 nodes. We also extend that work by computing the approximate ratio of connected EGs to connected DAGs, of …
This paper is not ready for public consumption, as the last step (Figure 20) is incorrect.
For any discrete, torsion-free subgroup of (resp.\ ) with no parabolic elements, we prove that (resp.\ for ) for any --module . The main technical advance is a new bound on the --Jacobian of the barycenter map of Besson--Cour…
The early layers of a deep neural net have the fewest parameters, but take up the most computation. In this extended abstract, we propose to only train the hidden layers for a set portion of the training run, freezing them out one-by-one and excluding them from the backward pass. Through experiments on CIFAR, we empiri…
We can overcome uncertainty with uncertainty. Using randomness in our choices and in what we control, and hence in the decision making process, could potentially offset the uncertainty inherent in the environment and yield better outcomes. The example we develop in greater detail is the news-vendor inventory management…
A graph is 2-apex if it is planar after the deletion of at most two vertices. Such graphs are not intrinsically knotted, IK. We investigate the converse, does not IK imply 2-apex? We determine the simplest possible counterexample, a graph on nine vertices and 21 edges that is neither IK nor 2-apex. In the process, we s…
This is lecture notes of a talk I gave at the Morningside Center of Mathematics on June 20, 2006. In this talk, I survey on Poincare and geometrization conjecture.
We prove the following: there are infinitely many finite-covolume (resp. cocompact) Coxeter groups acting on hyperbolic space H^n for every n < 20 (resp. n < 7). When n=7 or 8, they may be taken to be nonarithmetic. Furthermore, for 1 < n < 20, with the possible exceptions n=16 and 17, the number of essentially distinc…
The number of closed billiard trajectories in a rational-angled polygon grows quadratically in the length. This paper gives an analogue on K3 surfaces, by considering special Lagrangian tori. The analogue of the angle of a billiard trajectory is a point on a twistor sphere, and the number of directions admitting a spec…
We consider gradient estimates to positive solutions of porous medium equations and fast diffusion equations: associated with the Witten Laplacian on Riemannian manifolds. Under the assumption that the -dimensional Bakry-Emery Ricci curvature is bounded from below, we obtain gradient estimates which…
EMR learns to read and remember from streaming data for QA.
Study estimates risks of nuclear waste storage projects.
Germany's tax admin costs likely exceed 20% of total revenue, requiring system improvement.
The study classifies complex parallelisable nilmanifolds with unobstructed deformations.
We address the question of the growth of firm size. To this end, we analyze the Compustat data base comprising all publicly-traded United States manufacturing firms within the years 1974-1993. We find that the distribution of firm sizes remains stable for the 20 years we study, i.e., the mean value and standard deviati…
Maximal knotless graphs have at least 74% of their vertices' edges.