Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,878 papers · 148 categories

Trend · papers per month

25.0%50.0%75.0%100.0% · Feb 199419922001200920172026
48 results for Question Entailment

We reformulate Lehmer's question from 1933 and a question due to Schinzel and Zassenhaus from 1965 in terms of a comparison of the Mahler measures and the houses, respectively, of monic integer reciprocal and skew-reciprocal polynomials of the same degree. This entails that understanding the difference between orientat…

2018-12-12abs ↗pdf ↗

System filters and ranks medical answers using pre-trained models.

problem Challenges in ranking and classifying medical answers due to input size and dataset limitations.
method Multi-task learning with pre-trained models as feature extractors.
result Achieved high performance on medical QA task (Spearman's Rho 0.338, MRR 0.9622).

End-to-end ASR error detection using audio-transcript entailment.

problem Detecting transcription errors in ASR systems to prevent error propagation.
method Proposes a novel end-to-end approach using audio-transcript entailment, with acoustic and linguistic encoders.
result Achieves CER of 26.2% on all transcription errors and 23% on medical errors specifically, improving by 12% and 15.4% respectively over a strong baseline.

The paper proposes methods to control errors in language generation models using textual entailment.

problem The lack of a correctness metric hinders applying principled methods to language generation tasks.
method The paper leverages textual entailment to evaluate correctness and proposes two selective generation algorithms: SGen^Sup and SGen^Semi.
result The proposed algorithms control the false discovery rate with respect to textual entailment and achieve comparable selection efficiency to baselines.

Word embeddings provide point representations of words containing useful semantic information. We introduce multimodal word distributions formed from Gaussian mixtures, for multiple word meanings, entailment, and rich uncertainty information. To learn these distributions, we propose an energy-based max-margin objective…

2017-04-27abs ↗pdf ↗

New framework for consistent submodular maximization with insertions and deletions.

problem Maintaining near-optimal solutions in a dynamic setting with insertions and deletions.
method Developed a general framework for fully dynamic submodular maximization, instantiated for cardinality and rank-k matroid constraints.
result First constant-factor approximations with sublinear consistency for both cardinality and rank-k matroid constraints.

Approaches KL divergence for learning multi-sense word distributions.

problem Capturing the polysemy and uncertainty of words in word embeddings.
method Modeling words as multi-sense Gaussian mixtures and using KL divergence for learning.
result The proposed approach effectively captures word entailment and distribution similarity.

Training modern deep learning models requires large amounts of computation, often provided by GPUs. Scaling computation from one GPU to many can enable much faster training and research progress but entails two complications. First, the training library must support inter-GPU communication. Depending on the particular …

2018-02-15abs ↗pdf ↗

Neurosymbolic predictors fail to model uncertainty under independence assumption.

problem Neurosymbolic predictors' reliance on independence assumption limits their ability to model uncertainty.
method Formal analysis of NeSy predictors under independence assumption.
result Assuming independence among symbolic concepts prevents NeSy predictors from representing uncertainty.

Learning graph representations via low-dimensional embeddings that preserve relevant network properties is an important class of problems in machine learning. We here present a novel method to embed directed acyclic graphs. Following prior work, we first advocate for using hyperbolic spaces which provably model tree-li…

2018-04-03abs ↗pdf ↗

American options are financial instruments that can be exercised at any time before expiration. In this paper we study the problem of pricing this kind of derivatives within a framework in which some of the properties --volatility and dividend policy-- of the underlaying stock can change at a random instant of time, bu…

2006-10-09abs ↗pdf ↗

In this article we extend cutting and blowing up to the nonrational symplectic toric setting. This entails the possibility of cutting and blowing up for symplectic toric manifolds and orbifolds in nonrational directions.

2016-06-02abs ↗pdf ↗

By representing words with probability densities rather than point vectors, probabilistic word embeddings can capture rich and interpretable semantic information and uncertainty. The uncertainty information can be particularly meaningful in capturing entailment relationships -- whereby general words such as "entity" co…

2018-04-26abs ↗pdf ↗

By a theorem of Kirchhoff if the six sphere admits an almost complex structure then the seven sphere is parallelizable, more crucial, he exhibited an explicit global frame constructed out of the given almost complex structure. This result implicitly equips the seven sphere with a definite H-space multiplication. We pro…

2018-04-16abs ↗pdf ↗

Recurrent neural networks have become ubiquitous in computing representations of sequential data, especially textual data in natural language processing. In particular, Bidirectional LSTMs are at the heart of several neural models achieving state-of-the-art performance in a wide variety of tasks in NLP. However, BiLSTM…

2018-05-18abs ↗pdf ↗

The paper introduces Relative Bias to quantify LLM bias systematically.

problem Quantifying bias in LLMs is challenging due to ambiguity and rapid model emergence.
method Relative Bias framework using Embedding Transformation and LLM-as-a-Judge methodologies.
result The two scoring methods show strong alignment, providing a systematic approach.

The last few years have seen a staggering number of empirical studies of the robustness of neural networks in a model of adversarial perturbations of their inputs. Most rely on an adversary which carries out local modifications within prescribed balls. None however has so far questioned the broader picture: how to fram…

2018-06-08abs ↗pdf ↗

New method uses LLMs to generate detailed scientific hypotheses.

problem Generating detailed, actionable scientific hypotheses from coarse initial directions.
method Hierarchical search method that incrementally adds details to hypotheses.
result Hierarchical search method consistently outperforms strong baselines on expert-annotated hypotheses.

This paper is an attempt at understanding the quantum-like dynamics of financial markets in terms of non-differentiable price-time continuum having fractal properties. The main steps of this development are the statistical scaling, the non-differentiability hypothesis, and the equations of motion entailed by this hypot…

2013-12-11abs ↗pdf ↗

We prove that a wide class of correlated stochastic volatility models exactly measure an empirical fact in which past returns are anticorrelated with future volatilities: the so-called ``leverage effect''. This quantitative measure allows us to fully estimate all parameters involved and it will entail a deeper study on…

2002-02-12abs ↗pdf ↗

A geometric analysis of protein folding, which complements many of the models in the literature, is presented. We examine the process from unfolded strand to the point where the strand becomes self-interacting. A central question is how it is possible that so many initial configurations proceed to fold to a unique fina…

2008-09-11abs ↗pdf ↗

Modern neural networks are often augmented with an attention mechanism, which tells the network where to focus within the input. We propose in this paper a new framework for sparse and structured attention, building upon a smoothed max operator. We show that the gradient of this operator defines a mapping from real val…

2017-05-22abs ↗pdf ↗

Paper proposes an ensemble approach to improve fairness in classifier decisions.

problem Improving fairness in classifier decisions to prevent bias.
method Inspired by dropout techniques, feature drop-out is used to reduce classifier dependence on sensitive features while maintaining accuracy.
result An ensemble of classifiers with reduced sensitivity to sensitive features and improved accuracy.

The Killing operator on a Riemannian manifold is a linear differential operator on vector fields whose kernel provides the infinitesimal Riemannian symmetries. The Killing operator is best understood in terms of its prolongation, which entails some simple tensor identities. These simple identities can be viewed as aris…

2010-06-08abs ↗pdf ↗

Embedding methods which enforce a partial order or lattice structure over the concept space, such as Order Embeddings (OE) (Vendrov et al., 2016), are a natural way to model transitive relational data (e.g. entailment graphs). However, OE learns a deterministic knowledge base, limiting expressiveness of queries and the…

2018-05-17abs ↗pdf ↗

Whatever information a deep neural network has gleaned from training data is encoded in its weights. How this information affects the response of the network to future data remains largely an open question. Indeed, even defining and measuring information entails some subtleties, since a trained network is a determinist…

2019-05-29abs ↗pdf ↗

We discuss the construction of Sp(2)Sp(1)-structures whose fundamental form is closed. In particular, we find 10 new examples of 8-dimensional nilmanifolds that admit an invariant closed 4-form with stabiliser Sp(2)Sp(1). Our constructions entail the notion of SO(4)-structures on 7-manifolds. We present a thorough inve…

2013-08-19abs ↗pdf ↗

This work explores the trade-offs between stability and accuracy in statistical estimation.

problem Understanding the statistical cost of algorithmic stability.
method Statistical decision-theoretic perspective, focusing on worst-case and average-case stability.
result Optimal stable estimators for mean estimation and regression settings are developed, revealing trade-offs between stability and accuracy.

Paper characterizes causal graphs from hard interventions and proposes a learning algorithm.

problem Discovering causal structure from hard interventions and observational data.
method Proposes graphical constraints and a learning algorithm based on do-calculus.
result Characterizes interventional equivalence classes of causal graphs with latent variables.

Multi-hop inference is necessary for machine learning systems to successfully solve tasks such as Recognising Textual Entailment and Machine Reading. In this work, we demonstrate the effectiveness of adaptive computation for learning the number of inference steps required for examples of different complexity and that l…

2016-10-24abs ↗pdf ↗

Deep convolutional neural networks (CNNs) used in practice employ potentially hundreds of layers and 1010,000000s of nodes. Such network sizes entail significant computational complexity due to the large number of convolutions that need to be carried out; in addition, a large number of parameters needs to be learned and…

2017-07-10abs ↗pdf ↗

Jacobian regularization boosts neural network robustness without degrading generalization.

problem Ensuring robustness of machine learning models against input perturbations.
method Developed a computationally efficient Jacobian regularization technique.
result Significant improvements in robustness measured against random and adversarial perturbations.