A new method for math reasoning that allows for iterative correction.
problem Standard reasoning models commit to each token and cannot recover from early errors.
method Generative framework with latent thought vectors for iterative self-correction.
result 30 rethinking iterations surpass baselines with 15 times more parameters.
Early stopping methods reduce unnecessary reasoning steps in LLMs by monitoring uncertainty signals.
problem LLMs sometimes generate unnecessary reasoning steps, especially under uncertainty.
method Statistically principled early stopping methods that monitor uncertainty signals during generation.
result Uncertainty-aware early stopping improves efficiency and reliability in LLM reasoning, especially in math reasoning.
Enhances math problem-solving models with multi-turn preference learning.
problem Improving mathematical problem-solving capabilities of large language models.
method Introduces a multi-turn direct preference learning framework for tool-integrated mathematical reasoning tasks.
result Significant performance improvements in model accuracy on math datasets.
Extends reinforcement learning alignment to scalar rewards, improving math reasoning.
problem Designing reinforcement learning algorithms for general LLM alignment.
method Introduces f-GRPO and f-HAL, estimating f-divergences between reward-aligned and unaligned distributions.
result Improves math reasoning RLVR tasks and mitigates reward hacking.
Estimates model performance based on compute budget and evaluates stability over time.
problem Understanding how model performance evolves with compute budget and stability over time.
method Large-scale observational evaluations, prescriptive scaling laws, smoothed quantile regression, I-optimal sampling algorithm.
result Estimates attainable accuracies and stability of model performance over time.
The paper optimizes LLM accuracy by stopping early based on consistent answers.
problem Improving LLM accuracy in math and reasoning problems.
method Bayesian stopping policy to save on sampling costs, tracking only the L-1 most frequent answer counts.
result The L=3 stopping policy is sufficient for asymptotic optimality and significantly reduces inference costs.
Benchmark for math reasoning models from human proofs.
problem Measuring and accelerating machine learning models in high-level mathematical reasoning.
method Built a non-synthetic dataset from theorem prover proofs, defined a task for model to fill in missing propositions, used hierarchical transformer to improve performance.
result Neural models can capture non-trivial mathematical reasoning, hierarchical transformer outperforms baseline.
This short note contains an explicit proof of the Jacobi identity for variational Schouten bracket in Z2-graded commutative setup. For the reasoning to be rigorous, it refers to the product bundle geometry of iterated variations (see arXiv:1312.1262 [math-ph]); no ad hoc regularizations occur anywhere in this theory…
RACER optimizes LLM-as-judge accuracy with dynamic reasoning selection.
problem Balancing reasoning accuracy with computational cost in LLM-as-judge settings.
method Formulates routing as a constrained distributionally robust optimization problem, accounting for distribution shift via KL-divergence uncertainty set.
result RACER achieves superior accuracy-cost trade-offs under distribution shift.
Generalizes soft noncommutative schemes to flag varieties.
problem Applying soft noncommutative schemes to flag varieties.
method Generalization via toric geometry and distinguished affine charts.
result Soft noncommutative schemes can be applied to flag varieties.
Enhances large language models' reasoning through simpler off-policy reinforcement learning.
problem Improving large language models' ability to reason and solve problems.
method EM Policy Gradient, optimizing expected return over reasoning trajectories using Expectation-Maximization (EM) optimization.
result Achieves comparable or slightly superior performance to state-of-the-art methods on reasoning datasets, with additional cognitive behaviors.
LLMs can fail to maximize aligned values even after training, due to irrational reasoning.
problem Value misalignment in LLMs' reasoning despite training alignment.
method Formalized rational value risk and decomposed estimation error.
result Rational value risk is widespread and cannot be fully eliminated.
Method guarantees coherent factuality for language model outputs in reasoning tasks.
problem Ensuring correctness of language model outputs in reasoning tasks.
method Developed a conformal-prediction-based method applied to subgraphs within a deducibility graph.
result Achieved coherent factuality across target coverage levels, 90% on stricter definition.
EORM boosts LLM accuracy with a lightweight, energy-based verifier.
problem Efficiently verifying mathematical reasoning in large language models.
method Energy-based framework to rank Chain-of-Thought solutions using simple outcome labels.
result EORM boosts LLM accuracy to 90.7% on GSM8k and 63.7% on MATH with only 55M parameters.
SFPO optimizes LLM reasoning by repositioning before updating, improving stability and efficiency.
problem Noisy gradients from low-quality rollouts cause instability and inefficient exploration in on-policy RL algorithms.
method Decomposes each step into three stages: a short fast trajectory, repositioning, and slow correction, preserving the objective and rollout process unchanged.
result SFPO consistently improves stability, reduces rollouts, and accelerates convergence, outperforming GRPO on math reasoning benchmarks.
ORCA calibrates LLMs for efficient, generalizable reasoning.
problem Miscalibration of large language models leading to inefficiencies.
method Online Reasoning Calibration (ORCA) using conformal prediction and test-time training.
result ORCA provides higher efficiency and generalization across different reasoning tasks.
Added examples of S^1-manifolds with finite 2nd homotopy group and non-zero A-genus.
problem Constructing examples of S^1-manifolds with finite 2nd homotopy group and non-zero A-genus.
method Explicit equivariant surgeries to construct examples.
result Construction of new examples with finite 2nd homotopy group and non-zero A-genus.
Inference-Time Scaling can be extended to domains prone to systematic failure using intrinsic statistics.
problem Scaling inference time in domains prone to systematic failure
method Intrinsic Selection (iS), Intrinsic Particle Filtering (iPF), and Particle Distillation (dPF)
result Intrinsic Selection improves engineering design selection by 20% and pass@1 by 6.1 points on average.
Model learns brevity by exposing to easy problems, improving efficiency without explicit length penalties.
problem Excessive verbosity in step-by-step reasoning models trained with RLVR.
method Retaining and up-weighting moderately easy problems as implicit length regularizers.
result Model generates solutions that are, on average, nearly twice as short without explicit length penalties.
Prefix consistency improves model reliability by weighting answers based on their reproducibility.
problem Improving the reliability of large language models' reasoning traces.
method Use prefix consistency to weight candidate answers based on their reproducibility during regeneration.
result Prefix consistency is the best correctness predictor, reaching Standard MV plateau accuracy with up to 21x fewer tokens.
wd1 improves reasoning in dLLMs by optimizing policies without policy ratios.
problem Improving reasoning in diffusion-based large language models through RL.
method wd1: ratio-free policy optimization using weighted log-likelihood.
result wd1 outperforms diffusion-based GRPO while requiring lower computational cost.
LiveTradeBench evaluates LLMs in live trading environments.
problem Static benchmarks fail to assess real-world trading ability.
method Live data streaming, portfolio management abstraction, multi-market evaluation.
result LLMs show distinct portfolio styles and adapt to live signals.
This is no longer available.
The papers math.QA/0403527 and math.QA/0409414 v.1 are now merged together. The final version is available at math.QA/0409414 v.2. To avoid duplication of papers, math.QA/0403527 is now removed.
Galactica learns from scientific literature to help researchers.
problem Information overload in scientific literature makes it hard to find useful insights.
method Trained on a large corpus of scientific papers, reference material, and knowledge bases.
result Outperforms existing models on various scientific tasks, including LaTeX equations and mathematical reasoning.
New RL algorithm GDPO improves DLM reasoning efficiency.
problem Adapting RL to DLMs for efficient, unbiased likelihood estimation.
method Group Diffusion Policy Optimization (GDPO) using semi-deterministic Monte Carlo.
result GDPO outperforms existing methods on math, reasoning, and coding benchmarks.
We survey what is known about singularities of special Lagrangian submanifolds (SL m-folds) in (almost) Calabi-Yau manifolds. The bulk of the paper summarizes the author's five papers math.DG/0211294, math.DG/0211295, math.DG/0302355, math.DG/0302356, math.DG/0303272 on SL m-folds X with isolated conical singularities.…
INFUSER improves reasoning by co-evolving a generator and solver with adaptive curriculum.
problem Improving reasoning in language models with minimal external supervision.
method INFUSER uses a Generator and Solver to co-evolve iteratively, rewarding the Generator with an influence score and the Solver with correctness rewards.
result INFUSER outperforms self-evolution baselines by over 20% on Olympiad and SuperGPQA benchmarks.
INFUSER improves reasoning by self-evolving with a generator and solver that co-learn from unstructured documents.
problem Improving reasoning through self-evolution
method INFUSER uses a generator and solver co-evolving in a document pool to improve reasoning.
result INFUSER outperforms strong self-evolution baselines on Olympiad and SuperGPQA benchmarks.
MathChat uses LLM agents to solve challenging math problems through conversational problem-solving.
problem Solving math problems expressed in natural language.
method MathChat is a conversational framework combining an LLM agent and a user proxy agent for collaborative problem-solving.
result MathChat improves tool-using prompting methods by 6% on difficult math problems.
The purpose of this note is to reconcile two different results concerning the model-free upper bound on the price of an American option, given a set of European option prices. Neuberger (2007, `Bounds on the American option') and Hobson and Neuberger (2016, `On the value of being American') argue that the cost of the c…
We give a topological interpretation of the core group invariant of a surface embedded in S^4. We show that the group is isomorphic to the free product of the fundamental group of the double branch cover of S^4 with the surface as a branched set, and the infinite cyclic group. We present a generalization for unoriented…
ePF improves PF for ITS by balancing exploration and exploitation, outperforming baselines.
problem Premature exploitation in PF leads to suboptimal solutions under constrained budgets.
method Integrates Entropic Annealing and Look-ahead Modulation to preserve diversity and evaluate potential.
result Significant improvement in task reward (up to 50% relative) on math benchmarks.
This is the second in a series of five papers math.DG/0211294, math.DG/0302355, math.DG/0302356, math.DG/0303272 studying special Lagrangian submanifolds (SL m-folds) X in (almost) Calabi-Yau m-folds M with singularities x_1,...,x_n locally modelled on special Lagrangian cones C_1,...,C_n in C^m with isolated singulari…
This is the first in a series of five papers math.DG/0211295, math.DG/0302355, math.DG/0302356, math.DG/0303272 studying special Lagrangian submanifolds (SL m-folds) X in (almost) Calabi-Yau m-folds M with singularities x_1,...,x_n locally modelled on special Lagrangian cones C_1,...,C_n in C^m with isolated singularit…
In this survey paper, we outline the proofs of the rigidity results for simple, thick, hyperbolic P-manifolds found in our three earlier papers math.GR/0506518, math.GT/0410476, and math.GR/0409586. We discuss how the arguments change in the two, three, and higher dimensional settings. This paper was written for the 22…
Math and dance blend in choreographer's research.
problem Exploring intersections between dance and mathematics.
method Choreographic practice and mathematical concepts.
result Examples of fractals, braids in choreography.
This is the fourth in a series of five papers math.DG/0211294, math.DG/0211295, math.DG/0302355, math.DG/0303272 studying compact special Lagrangian submanifolds (SL m-folds) X in (almost) Calabi-Yau m-folds M with singularities x_1,...,x_n locally modelled on special Lagrangian cones C_1,...,C_n in C^m with isolated s…
This is the third in a series of five papers math.DG/0211294, math.DG/0211295, math.DG/0302356, math.DG/0303272 studying compact special Lagrangian submanifolds (SL m-folds) X in (almost) Calabi-Yau m-folds M with singularities x_1,...,x_n locally modelled on special Lagrangian cones C_1,...,C_n in C^m with isolated si…
In Theorem 1.2 of the paper math.GT/0002110 the author claimed to have proved that all transversal knots whose topological knot type is that of an iterated torus knot (we call them cable knots) are transversally simple. That theorem is false, and the Erratum math.GT/0610565 identifies the gap. The purpose of this paper…
We prove that the leaves of the rescaled curvature flow considered in arXiv:math/0403485 [math.DG] converge to the graph of a constant function.
Continuing the work started in Part I and II of this series (see q-alg/9706004 and math.QA/9801049), we prove the relationship between the Aarhus integral and the invariant Ω (henceforth called LMO) defined by T.Q.T. Le, J. Murakami and T. Ohtsuki in q-alg/9512002. The basic reason for the relationship is that both c…
This is the last in a series of five papers math.DG/0211294, math.DG/0211295, math.DG/0302355, math.DG/0302356 studying compact special Lagrangian submanifolds (SL m-folds) X in (almost) Calabi-Yau m-folds M with singularities x_1,...,x_n locally modelled on special Lagrangian cones C_1,...,C_n in C^m with isolated sin…
Given a family of Dirac operators with vanishing spectral flow we construct a thin-invariant rank-one field theory in the sense of Turner and Willerton arXiv:math.AT/0201116. Our construction of the field theory generalizes the one of the index gerbe by Lott, arXiv:math.DG/0106177, and it also complements the relation …
Novel method diagnoses large language models' reasoning abilities.
problem Fine-grained evaluation of large language models' reasoning abilities.
method Adapting cognitive diagnosis models to LLMs, estimating mastery profiles and Q-matrix, incorporating textual information.
result Accurate parameter recovery and insights into LLMs' capabilities.
Transformer models can solve complex math problems with less data.
problem Solving complex symbolic mathematics problems with limited data.
method Pretrain transformer models on language translation tasks and fine-tune for symbolic math.
result Pretrained transformer models achieve comparable accuracy to state-of-the-art models with less data.
This is the author's PhD thesis, as submitted to the Princeton University. The results of this paper have already appeared in arXiv:math/0607777v4, arXiv:math/0607691 and arXiv:0901.2156.
This article is the third part of the series of articles where the theory of valuations on manifolds is constructed. In math.MG/0503399 the notion of a smooth valuation on a manifold was introduced. The goal of this article is to put a canonical multiplicative structure on the space of smooth valuations on general mani…