Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

35810 · Nov 201919922001200920172026
48 results for POS tagging

Paper proposes set-valued prediction for historical POS tagging.

problem Difficult POS tagging in historical corpora due to lack of native speakers and sparse data.
method Set-valued prediction approach to allow uncertainty in tagging.
result Set-valued prediction improves POS tagging precision and robustness.

Part-of-speech (POS) tagging is a fundamental component for performing natural language tasks such as parsing, information extraction, and question answering. When POS taggers are trained in one domain and applied in significantly different domains, their performance can degrade dramatically. We present a methodology f…

2014-10-31abs ↗pdf ↗

West Frisian lemmatizer, POS tagger, and parser created.

problem Creating accurate lemmatization, POS tagging, and dependency parsing for West Frisian.
method Using a corpus of 44,714 words annotated according to Universal Dependency version 2. Applying Dutch POS tags and morphological/syntactic annotations to create Frisian translations.
result Significant improvement in lemma accuracy compared to default parameters.

ICP improves text infilling and POS tagging with valid confidence sets.

problem Statistical reliability of machine learning predictions.
method Inductive conformal prediction algorithms for text infilling and POS tagging.
result Valid set-valued predictions with small size for real-world applications.

Greek POS Tagger and Entity Recognizer built using spaCy.

problem Developing a machine learning model for Greek language POS tagging and named entity recognition.
method Machine learning approach with spaCy platform, focusing on morphological features and token classification.
result The Greek spaCy platform improved performance in part-of-speech tagging and named entity recognition.

Labeling of sequential data is a prevalent meta-problem for a wide range of real world applications. While the first-order Hidden Markov Models (HMM) provides a fundamental approach for unsupervised sequential labeling, the basic model does not show satisfying performance when it is directly applied to real world probl…

2019-04-05abs ↗pdf ↗

State-of-the-art sequence labeling systems traditionally require large amounts of task-specific knowledge in the form of hand-crafted features and data pre-processing. In this paper, we introduce a novel neutral network architecture that benefits from both word- and character-level representations automatically, by usi…

2016-03-04abs ↗pdf ↗

Recent studies have shown that neural models can achieve high performance on several sequence labelling/tagging problems without the explicit use of linguistic features such as part-of-speech (POS) tags. These models are trained only using the character-level and the word embedding vectors as inputs. Others have shown …

2019-09-29abs ↗pdf ↗

Conditional random fields (CRFs) have been shown to be one of the most successful approaches to sequence labeling. Various linear-chain neural CRFs (NCRFs) are developed to implement the non-linear node potentials in CRFs, but still keeping the linear-chain hidden structure. In this paper, we propose NCRF transducers, …

2018-11-04abs ↗pdf ↗

A deep learning approach generates math word problems in multiple languages.

problem Template-based mechanisms for generating mathematical word problems lack customizability and creativity.
method Character Level Long Short Term Memory Network (LSTM) and POS tags are used to generate and resolve constraints in generated problems.
result The approach generates accurate math word problems in English and Sinhala with over 90% accuracy.

This new research explores the effects of various training methods on a Polish to English Statistical Machine Translation system for medical texts. Various elements of the EMEA parallel text corpora from the OPUS project were used as the basis for training of phrase tables and language models and for development, tunin…

2015-09-29abs ↗pdf ↗

Positive representations of surface groups in PO(p,q) form connected components of character varieties.

problem Characterizing representations of surface groups in special orthogonal groups PO(p,q).
method Using Anosov representations and root versus weight collar lemmas.
result Connected components of character varieties are formed by ΘΘ-positive Anosov representations.

RANDomized-exploration policy Optimization via Multiple Importance Sampling with Truncation (RANDOMIST) for PO with mediator feedback.

problem Policy Optimization in continuous control tasks.
method RANDomized-exploration policy Optimization via Multiple Importance Sampling with Truncation (RANDOMIST) for regret minimization in PO.
result Achieving constant regret under certain circumstances in PO with mediator feedback.

Network embedding is a method to learn low-dimensional representation vectors for nodes in complex networks. In real networks, nodes may have multiple tags but existing methods ignore the abundant semantic and hierarchical information of tags. This information is useful to many network applications and usually very sta…

2019-04-19abs ↗pdf ↗

Proposes PO-QA framework to optimize portfolios using quantum algorithms.

problem Optimizing investment portfolios with reduced risk and increased gains.
method Develops a scalable quantum framework (PO-QA) to investigate quantum algorithm parameters.
result Identifies efficient quantum circuit configurations for portfolio optimization.

PO-Flow models potential and counterfactual outcomes for personalized treatment decisions.

problem Predicting individualized treatment effects from observational data.
method Continuous normalizing flow (CNF) framework for causal inference.
result Unified approach to potential outcome prediction, treatment effect estimation, and counterfactual prediction.

We show that the non-arithmetic lattices in PO(n,1) of Belolipetsky and Thomson (2011), obtained as fundamental groups of closed hyperbolic manifolds with short systole, are quasi-arithmetic in the sense of Vinberg, and, by contrast, the well-known non-arithmetic lattices of Gromov and Piatetski-Shapiro are not quasi-a…

2014-12-16abs ↗pdf ↗

Paper eliminates warm-up phase for PO in linear MDPs, achieving optimal regret.

problem Costly warm-up phase in PO algorithms for linear MDPs.
method Simple contraction mechanism replaces warm-up phase.
result Achieves rate-optimal regret with improved dependence on problem parameters.

The paper extends influence functions to sequence tagging tasks for better model interpretability.

problem Lack of interpretability methods for sequence tagging models.
method Define and compute influence of training instance segments on test segment predictions.
result The segment influence method tracks with true influence and identifies annotation errors.

Improved method reduces projection calls for nonsmooth convex optimization.

problem Optimizing nonsmooth convex functions with convex constraints.
method MOPES and MOLES methods combining Moreau-Yosida smoothing and accelerated first-order schemes.
result Achieves εε-suboptimality with significantly fewer projection calls.

A new framework for performative prediction robust to distributional misspecification.

problem Performative prediction models can be influenced by their own predictions, leading to suboptimal outcomes.
method Introduces distributionally robust performative prediction (DRPO) to approximate the true performative optimum (PO) robustly.
result DRPO provides provable guarantees as a robust approximation to the true PO when the nominal distribution map is misspecified.

Stochastic particle-optimization sampling (SPOS) is a recently-developed scalable Bayesian sampling framework that unifies stochastic gradient MCMC (SG-MCMC) and Stein variational gradient descent (SVGD) algorithms based on Wasserstein gradient flows. With a rigorous non-asymptotic convergence theory developed recently…

2018-11-20abs ↗pdf ↗

A new RL algorithm tackles PO tasks by modeling the environment and improving the policy.

problem Tackling unsatisfactory performance in RL agents in PO environments.
method Proposes a VRM for modeling the environment and an RL controller that uses both the environment and VRM.
result The proposed algorithm achieved better data efficiency and/or learned more optimal policies.

New Teichmüller spaces found for higher-dimensional groups.

problem Finding new Teichmüller spaces for higher-dimensional groups.
method Proving representations of groups in pseudo-Riemannian hyperbolic spaces are convex cocompact.
result Set of representations forms connected components of Hom spaces.

Study extends Hausdorff dimension Hessian results to new hyperconvex representations.

problem Extending classical results on Hausdorff dimension Hessian.
method Analyzes (1,1,2)-hyperconvex representations and small complex deformations.
result Positive definiteness of Hessian of Hausdorff dimension for co-compact Γ in PO(n,1).

Limit sets of AdS\mathrm{AdS}-quasi-Fuchsian groups of PO(n,2)\mathrm{PO}(n,2) are always Lipschitz submanifolds. The aim of this article is to show that they are never C1\mathcal{C}^1, except for the case of Fuchsian groups. As a byproduct we show that AdS\mathrm{AdS}-quasi-Fuchsian groups that are not Fuchsian are Zariski d…

2018-09-27abs ↗pdf ↗

To date, there have been massive Semi-Structured Documents (SSDs) during the evolution of the Internet. These SSDs contain both unstructured features (e.g., plain text) and metadata (e.g., tags). Most previous works focused on modeling the unstructured text, and recently, some other methods have been proposed to model …

2015-07-30abs ↗pdf ↗

We study the cluster categories arising from marked surfaces (with punctures and non-empty boundaries). By constructing skewed-gentle algebras, we show that there is a bijection between tagged curves and string objects. Applications include interpreting dimensions of Ext1\operatorname{Ext}^1 as intersection numbers of ta…

2013-10-31abs ↗pdf ↗

Staking and on-chain lending can reduce PoS network security if rewards are not calibrated properly.

problem Rational actors can reduce PoS network security if block rewards are not calibrated appropriately above on-chain lending yields.
method Simple stochastic model and agent-based simulations to validate the phase transition between staking and lending.
result Rational actors can reduce PoS network security if block rewards are not calibrated appropriately above on-chain lending yields.