The paper optimizes LLM accuracy by stopping early based on consistent answers.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Prefix consistency improves model reliability by weighting answers based on their reproducibility.
CITE algorithm provides anytime-valid certification of model outputs.
We propose a new class of probabilistic neural-symbolic models, that have symbolic functional programs as a latent, stochastic variable. Instantiated in the context of visual question answering, our probabilistic formulation offers two key conceptual advantages over prior neural-symbolic models for VQA. Firstly, the pr…
We introduce the new task of Acoustic Question Answering (AQA) to promote research in acoustic reasoning. The AQA task consists of analyzing an acoustic scene composed by a combination of elementary sounds and answering questions that relate the position and properties of these sounds. The kind of relational questions …
Improved selection of best outputs from LLMs for better accuracy.
REALFIN benchmarks financial reasoning by removing implicit assumptions, revealing model weaknesses.
We propose a method to efficiently learn diverse strategies in reinforcement learning for query reformulation in the tasks of document retrieval and question answering. In the proposed framework an agent consists of multiple specialized sub-agents and a meta-agent that learns to aggregate the answers from sub-agents to…
Unified framework for certifying LLM reliability without extra supervision.
The paper certifies AI reliability via sampling and calibration, providing exact guarantees.
Proposes CRA framework for certifying fair predictive models.
Deep learning techniques are rapidly advanced recently, and becoming a necessity component for widespread systems. However, the inference process of deep learning is black-box, and not very suitable to safety-critical systems which must exhibit high transparency. In this paper, to address this black-box limitation, we …
Recently, several studies have explored methods for using KG embedding to answer logical queries. These approaches either treat embedding learning and query answering as two separated learning tasks, or fail to deal with the variability of contributions from different query paths. We proposed to leverage a graph attent…
The nearest neighbor rule is proven consistent in a broad setting.
Consistency of the kernel density estimator requires that the kernel bandwidth tends to zero as the sample size grows. In this paper we investigate the question of whether consistency is possible when the bandwidth is fixed, if we consider a more general class of weighted KDEs. To answer this question in the affirmativ…
In this note, we give a proof of the famous theorem of M. Morse dealing with the cancellation of a pair of non-degenerate critical points of a smooth function. Our proof consists of a reduction to the one-dimensional case where the question becomes easy to answer.
This paper presents the current state of a work in progress, whose objective is to better understand the effects of factors that significantly influence the performance of Latent Semantic Analysis (LSA). A difficult task, which consists in answering (French) biology Multiple Choice Questions, is used to test the semant…
We investigate the geometry of word metrics on fundamental groups of manifolds associated with the generating sets consisting of elements represented by closed geodesics. We ask whether the diameter of such a metric is finite or infinite. The first answer we interpret as an abundance of closed geodesics, while the seco…
Sequence probability predicts correctness in LLMs, but not for repeated prompts
ForecastQA creates a new QA task for event forecasting from text data.
Arriving at the complete probabilistic knowledge of a domain, i.e., learning how all variables interact, is indeed a demanding task. In reality, settings often arise for which an individual merely possesses partial knowledge of the domain, and yet, is expected to give adequate answers to a variety of posed queries. Tha…
In this paper we investigate the following existence problem for rational functions: for a given collection of partitions of a number to define whether there exists a rational function of degree for which is the branch datum. An important particular case when the answer to this problem is known is t…
Multi-domain dialogue state tracking (DST) is a critical component for conversational AI systems. The domain ontology (i.e., specification of domains, slots, and values) of a conversational AI system is generally incomplete, making the capability for DST models to generalize to new slots, values, and domains during inf…
Explicit answer is given for the HOMFLY polynomial of the figure eight knot in arbitrary symmetric representation R=[p]. It generalizes the old answers for p=1 and 2 and the recently derived results for p=3,4, which are fully consistent with the Ooguri-Vafa conjecture. The answer can be considered as a quantizati…
Motivated by the study in Morse theory and Smale's work in dynamics, the following questions are studied and answered: (1) When does a 3-manifold admit an automorphism having a knotted Smale solenoid as an attractor? (2) When does a 3-manifold admit an automorphism whose non-wandering set consists of Smale solenoids? T…
Develops PLL methods that are provably consistent and compatible with any deep network.
Many recent papers address reading comprehension, where examples consist of (question, passage, answer) tuples. Presumably, a model must combine information from both questions and passages to predict corresponding answers. However, despite intense interest in the topic, with hundreds of published papers vying for lead…
A new Q&A labeling method for assigning labels in machine learning.
Spectral clustering achieves strong consistency in the stochastic block model under certain conditions.
This work explores test-time scaling strategies for LLMs, improving sample efficiency and expressiveness.
KodeXv0.1 improves financial question answering over GPT-4.
The paper defines and characterizes conditional nonlinear expectations.
We construct two infinite families of algebraic minimal cones in . The first family consists of minimal cubics given explicitly in terms of the Clifford systems. We show that the classes of congruent minimal cubics are in one to one correspondence with those of geometrically equivalent Clifford systems. As a byp…
LLMs generate answers under incomplete context, and their uncertainty should scale with missing information.
Let S^3_i be a 3-sphere embedded in the 5-sphere S^5 (i=1,2). Let S^3_1 and S^3_2 intersect transversely. Then the intersection C of S^3_1 and S^3_2 is a disjoint collection of circles. Thus we obtain a pair of 1-links, C in S^3_i (i=1,2), and a pair of 3-knots, S^3_i in S^5 (i=1,2). Conversely let (L_1,L_2) be a pair …
A subset of a group is characteristic if it is invariant under every automorphism of the group. We study word length in fundamental groups of closed hyperbolic surfaces with respect to characteristic generating sets consisting of a finite union of orbits of the automorphism group, and show that the translation length o…
A conformal procedure improves CoT reasoning by aggregating reasoning paths and calibrating abstention rules.
Blend-ASC improves self-consistency efficiency by dynamically allocating samples, reducing costs.
A new framework evaluates LLM calibration in open-ended QA.
Researchers have used from 30 days to several years of daily returns as source data for clustering financial time series based on their correlations. This paper sets up a statistical framework to study the validity of such practices. We first show that clustering correlated random variables from their observed values i…
Unified framework for ranking-and-selection with multiple correct answers and non-answerable estimates
Nonconvex optimization algorithms with random initialization have attracted increasing attention recently. It has been showed that many first-order methods always avoid saddle points with random starting points. In this paper, we answer a question: can the nonconvex heavy-ball algorithms with random initialization avoi…
This paper uses LLMs and cycle consistency for better machine translation evaluation.
The paper explores how LLMs with CoT improve performance on complex tasks.
Paper introduces SDM for detecting LLM hallucinations, improving on entropy tests.
A new algorithm identifies one of several nearly optimal arms in linear bandits.
There is an increasing body of evidence suggesting that exact nearest neighbour search in high-dimensional spaces is affected by the curse of dimensionality at a fundamental level. Does it necessarily mean that the same is true for k nearest neighbours based learning algorithms such as the k-NN classifier? We analyse t…
Deep networks can handle noisy labels up to a certain threshold.