Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

35810 · Nov 201919922001200920172026
48 results for POS tagging

Paper proposes set-valued prediction for historical POS tagging.

problem Difficult POS tagging in historical corpora due to lack of native speakers and sparse data.
method Set-valued prediction approach to allow uncertainty in tagging.
result Set-valued prediction improves POS tagging precision and robustness.

Part-of-speech (POS) tagging is a fundamental component for performing natural language tasks such as parsing, information extraction, and question answering. When POS taggers are trained in one domain and applied in significantly different domains, their performance can degrade dramatically. We present a methodology f…

2014-10-31abs ↗pdf ↗

West Frisian lemmatizer, POS tagger, and parser created.

problem Creating accurate lemmatization, POS tagging, and dependency parsing for West Frisian.
method Using a corpus of 44,714 words annotated according to Universal Dependency version 2. Applying Dutch POS tags and morphological/syntactic annotations to create Frisian translations.
result Significant improvement in lemma accuracy compared to default parameters.

ICP improves text infilling and POS tagging with valid confidence sets.

problem Statistical reliability of machine learning predictions.
method Inductive conformal prediction algorithms for text infilling and POS tagging.
result Valid set-valued predictions with small size for real-world applications.

Labeling of sequential data is a prevalent meta-problem for a wide range of real world applications. While the first-order Hidden Markov Models (HMM) provides a fundamental approach for unsupervised sequential labeling, the basic model does not show satisfying performance when it is directly applied to real world probl…

2019-04-05abs ↗pdf ↗

State-of-the-art sequence labeling systems traditionally require large amounts of task-specific knowledge in the form of hand-crafted features and data pre-processing. In this paper, we introduce a novel neutral network architecture that benefits from both word- and character-level representations automatically, by usi…

2016-03-04abs ↗pdf ↗

Recent studies have shown that neural models can achieve high performance on several sequence labelling/tagging problems without the explicit use of linguistic features such as part-of-speech (POS) tags. These models are trained only using the character-level and the word embedding vectors as inputs. Others have shown …

2019-09-29abs ↗pdf ↗

Conditional random fields (CRFs) have been shown to be one of the most successful approaches to sequence labeling. Various linear-chain neural CRFs (NCRFs) are developed to implement the non-linear node potentials in CRFs, but still keeping the linear-chain hidden structure. In this paper, we propose NCRF transducers, …

2018-11-04abs ↗pdf ↗

This new research explores the effects of various training methods on a Polish to English Statistical Machine Translation system for medical texts. Various elements of the EMEA parallel text corpora from the OPUS project were used as the basis for training of phrase tables and language models and for development, tunin…

2015-09-29abs ↗pdf ↗

Positive representations of surface groups in PO(p,q) form connected components of character varieties.

problem Characterizing representations of surface groups in special orthogonal groups PO(p,q).
method Using Anosov representations and root versus weight collar lemmas.
result Connected components of character varieties are formed by ΘΘ-positive Anosov representations.

RANDomized-exploration policy Optimization via Multiple Importance Sampling with Truncation (RANDOMIST) for PO with mediator feedback.

problem Policy Optimization in continuous control tasks.
method RANDomized-exploration policy Optimization via Multiple Importance Sampling with Truncation (RANDOMIST) for regret minimization in PO.
result Achieving constant regret under certain circumstances in PO with mediator feedback.

Network embedding is a method to learn low-dimensional representation vectors for nodes in complex networks. In real networks, nodes may have multiple tags but existing methods ignore the abundant semantic and hierarchical information of tags. This information is useful to many network applications and usually very sta…

2019-04-19abs ↗pdf ↗

Proposes PO-QA framework to optimize portfolios using quantum algorithms.

problem Optimizing investment portfolios with reduced risk and increased gains.
method Develops a scalable quantum framework (PO-QA) to investigate quantum algorithm parameters.
result Identifies efficient quantum circuit configurations for portfolio optimization.

PO-Flow models potential and counterfactual outcomes for personalized treatment decisions.

problem Predicting individualized treatment effects from observational data.
method Continuous normalizing flow (CNF) framework for causal inference.
result Unified approach to potential outcome prediction, treatment effect estimation, and counterfactual prediction.

We show that the non-arithmetic lattices in PO(n,1) of Belolipetsky and Thomson (2011), obtained as fundamental groups of closed hyperbolic manifolds with short systole, are quasi-arithmetic in the sense of Vinberg, and, by contrast, the well-known non-arithmetic lattices of Gromov and Piatetski-Shapiro are not quasi-a…

2014-12-16abs ↗pdf ↗

Paper eliminates warm-up phase for PO in linear MDPs, achieving optimal regret.

problem Costly warm-up phase in PO algorithms for linear MDPs.
method Simple contraction mechanism replaces warm-up phase.
result Achieves rate-optimal regret with improved dependence on problem parameters.

The paper extends influence functions to sequence tagging tasks for better model interpretability.

problem Lack of interpretability methods for sequence tagging models.
method Define and compute influence of training instance segments on test segment predictions.
result The segment influence method tracks with true influence and identifies annotation errors.

A new framework for performative prediction robust to distributional misspecification.

problem Performative prediction models can be influenced by their own predictions, leading to suboptimal outcomes.
method Introduces distributionally robust performative prediction (DRPO) to approximate the true performative optimum (PO) robustly.
result DRPO provides provable guarantees as a robust approximation to the true PO when the nominal distribution map is misspecified.

Stochastic particle-optimization sampling (SPOS) is a recently-developed scalable Bayesian sampling framework that unifies stochastic gradient MCMC (SG-MCMC) and Stein variational gradient descent (SVGD) algorithms based on Wasserstein gradient flows. With a rigorous non-asymptotic convergence theory developed recently…

2018-11-20abs ↗pdf ↗

Improved method reduces projection calls for nonsmooth convex optimization.

problem Optimizing nonsmooth convex functions with convex constraints.
method MOPES and MOLES methods combining Moreau-Yosida smoothing and accelerated first-order schemes.
result Achieves εε-suboptimality with significantly fewer projection calls.

New Teichmüller spaces found for higher-dimensional groups.

problem Finding new Teichmüller spaces for higher-dimensional groups.
method Proving representations of groups in pseudo-Riemannian hyperbolic spaces are convex cocompact.
result Set of representations forms connected components of Hom spaces.

Limit sets of AdS\mathrm{AdS}-quasi-Fuchsian groups of PO(n,2)\mathrm{PO}(n,2) are always Lipschitz submanifolds. The aim of this article is to show that they are never C1\mathcal{C}^1, except for the case of Fuchsian groups. As a byproduct we show that AdS\mathrm{AdS}-quasi-Fuchsian groups that are not Fuchsian are Zariski d…

2018-09-27abs ↗pdf ↗

Study extends Hausdorff dimension Hessian results to new hyperconvex representations.

problem Extending classical results on Hausdorff dimension Hessian.
method Analyzes (1,1,2)-hyperconvex representations and small complex deformations.
result Positive definiteness of Hessian of Hausdorff dimension for co-compact Γ in PO(n,1).

To date, there have been massive Semi-Structured Documents (SSDs) during the evolution of the Internet. These SSDs contain both unstructured features (e.g., plain text) and metadata (e.g., tags). Most previous works focused on modeling the unstructured text, and recently, some other methods have been proposed to model …

2015-07-30abs ↗pdf ↗

We study the cluster categories arising from marked surfaces (with punctures and non-empty boundaries). By constructing skewed-gentle algebras, we show that there is a bijection between tagged curves and string objects. Applications include interpreting dimensions of Ext1\operatorname{Ext}^1 as intersection numbers of ta…

2013-10-31abs ↗pdf ↗

MuLan links music audio to natural language tags.

problem Traditional music tagging systems use rigid attributes; MuLan aims to link audio directly to natural language.
method Joint audio-text embedding model trained on 44 million music recordings and text annotations.
result MuLan's embeddings enable zero-shot functionalities and transfer learning.

Identifying the flavour of neutral BB mesons production is one of the most important components needed in the study of time-dependent CPCP violation. The harsh environment of the Large Hadron Collider makes it particularly hard to succeed in this task. We present an inclusive flavour-tagging algorithm as an upgrade of…

2017-05-24abs ↗pdf ↗