A method for non-projective dependency parsing without fixed edge order.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Expanding spoken language understanding to handle complex entities and intents.
This paper explores unsupervised learning of parsing models along two directions. First, which models are identifiable from infinite data? We use a general technique for numerically checking identifiability based on the rank of a Jacobian matrix, and apply it to several standard constituency and dependency parsing mode…
Paper proposes a reinforcement learning framework for distant supervision of question parsing.
Improves semantic parsing with human feedback in a counterfactual setup.
Flexible log file parsing using HMM adapts to evolving content.
DocParser parses document structures from renderings like PDFs and scans.
Domain-general semantic parsing is a long-standing goal in natural language processing, where the semantic parser is capable of robustly parsing sentences from domains outside of which it was trained. Current approaches largely rely on additional supervision from new domains in order to generalize to those domains. We …
Graph-to-Tree Neural Networks improve structured input-output translation in tasks like semantic parsing and math word problems.
Deep models learn to parse complex language structures from local data patterns.
Synthetic reference strings are as effective as real ones for training citation parsing models.
Tasks like code generation and semantic parsing require mapping unstructured (or partially structured) inputs to well-formed, executable outputs. We introduce abstract syntax networks, a modeling framework for these problems. The outputs are represented as abstract syntax trees (ASTs) and constructed by a decoder with …
Co-PLNet combines point and line predictions to improve wireframe parsing accuracy and efficiency.
We explore the problem of learning to decompose spatial tasks into segments, as exemplified by the problem of a painting robot covering a large object. Inspired by the ability of classical decision tree algorithms to construct structured partitions of their input spaces, we formulate the problem of decomposing objects …
Proposes a method to interpret linguistic data models using parse trees and least-squares scores.
In this paper, we propose a probabilistic parsing model, which defines a proper conditional probability distribution over non-projective dependency trees for a given sentence, using neural representations as inputs. The neural network architecture is based on bi-directional LSTM-CNNs which benefits from both word- and …
PATOIS synthesizes code from natural language using learned code idioms.
Survey on automating geometry problem solving with large models.
Unsupervised algorithm parses CSG images into CFG without pretraining.
Privacy-preserving syntactic parsing using obfuscation.
Automatically mined rules from dependency parsing help neural models learn from less labeled data.
While neural networks have shown impressive performance on large datasets, applying these models to tasks where little data is available remains a challenging problem. In this paper we propose to use feature transfer in a zero-shot experimental setting on the task of semantic parsing. We first introduce a new method fo…
Novel ramp loss method improves weakly supervised machine translation and parsing.
Scene parsing is an important and challenging prob- lem in computer vision. It requires labeling each pixel in an image with the category it belongs to. Tradition- ally, it has been approached with hand-engineered features from color information in images. Recently convolutional neural networks (CNNs), which automatica…
Sequence to sequence learning has recently emerged as a new paradigm in supervised learning. To date, most of its applications focused on only one task and not much work explored this framework for multiple tasks. This paper examines three multi-task learning (MTL) settings for sequence to sequence models: (a) the onet…
Counterfactual learning from human bandit feedback describes a scenario where user feedback on the quality of outputs of a historic system is logged and used to improve a target system. We show how to apply this learning framework to neural semantic parsing. From a machine learning perspective, the key challenge lies i…
Framework for universal graph function approximators outperforms existing methods.
Syntactic constituency parsing is a fundamental problem in natural language processing and has been the subject of intensive research and engineering for decades. As a result, the most accurate parsers are domain specific, complex, and inefficient. In this paper we show that the domain agnostic attention-enhanced seque…
New grammar model learns sentence structure with latent variables.
Deep generative models have been wildly successful at learning coherent latent representations for continuous data such as video and audio. However, generative modeling of discrete data such as arithmetic expressions and molecular structures still poses significant challenges. Crucially, state-of-the-art methods often …
Neurally-Guided Structure Inference combines search and data-driven methods for efficient, robust structure inference.
Unsupervised RNNGs perform similarly to supervised ones in language modeling and grammar induction.
The paper sets a lower bound on the crossing number of 2-bridge knots and answers a question about their epimorphism number.
We speed up marginal inference by ignoring factors that do not significantly contribute to overall accuracy. In order to pick a suitable subset of factors to ignore, we propose three schemes: minimizing the number of model factors under a bound on the KL divergence between pruned and full models; minimizing the KL dive…
Despite enormous progress in object detection and classification, the problem of incorporating expected contextual relationships among object instances into modern recognition systems remains a key challenge. In this work we propose Information Pursuit, a Bayesian framework for scene parsing that combines prior models …
Simple framework decouples word alignment and multilingual embedding mapping.
West Frisian lemmatizer, POS tagger, and parser created.
A new CNN-based code generator outperforms RNNs by 5%.
Novel unsupervised relation extraction framework using BERT.
Patients with epilepsy can manifest short, sub-clinical epileptic "bursts" in addition to full-blown clinical seizures. We believe the relationship between these two classes of events---something not previously studied quantitatively---could yield important insights into the nature and intrinsic dynamics of seizures. A…
We provide complete source code for building a fundamental industry classification based on publically available and freely downloadable data. We compare various fundamental industry classifications by running a horserace of short-horizon trading signals (alphas) utilizing open source heterotic risk models (https://ssr…
Malicious web content is a serious problem on the Internet today. In this paper we propose a deep learning approach to detecting malevolent web pages. While past work on web content detection has relied on syntactic parsing or on emulation of HTML and Javascript to extract features, our approach operates directly on a …
We present Memory Augmented Policy Optimization (MAPO), a simple and novel way to leverage a memory buffer of promising trajectories to reduce the variance of policy gradient estimate. MAPO is applicable to deterministic environments with discrete actions, such as structured prediction and combinatorial optimization ta…
HapNet predicts marketing campaign effects using a hierarchical structure.
Deep learning has improved performance on many natural language processing (NLP) tasks individually. However, general NLP models cannot emerge within a paradigm that focuses on the particularities of a single metric, dataset, and task. We introduce the Natural Language Decathlon (decaNLP), a challenge that spans ten ta…
A new method for spotting symbols in CAD images reduces annotation costs and improves accuracy.
MeRL learns from sparse, underspecified rewards by discounting spurious trajectories.
Automatically computes reference ranges for UK Biobank cardiac data.