Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

36912 · May 202619922001200920172026
48 results for prefix notation

Improved math problem solvers using Transformer networks and diverse notations.

problem Challenges in constructing accurate and automatic solvers for math word problems.
method Transformer networks trained to translate math word problems to arithmetic expressions in infix, prefix, and postfix notations. Pre-training on general text corpus to improve performance.
result Significant improvements in accuracy, up to 10% over previous state of the art.

A new robust prefix-tuning framework improves model robustness against adversarial attacks.

problem Lack of robustness in prefix-tuning for adversarial attacks.
method Leveraging layerwise activations of pretrained models for additional prefix finetuning during the test phase.
result Framework substantially improves robustness over strong baselines while maintaining comparable accuracy on clean texts.

CROP verifies clean prefixes in reasoning traces, improving downstream repair accuracy.

problem Uncertainty in reasoning traces prevents full certification of entire responses.
method CROP selects a calibrated threshold to certify the longest prefix with low risk proxies.
result CROP improves downstream repair accuracy by preserving valid reasoning and discarding misleading suffixes.

Prefix consistency improves model reliability by weighting answers based on their reproducibility.

problem Improving the reliability of large language models' reasoning traces.
method Use prefix consistency to weight candidate answers based on their reproducibility during regeneration.
result Prefix consistency is the best correctness predictor, reaching Standard MV plateau accuracy with up to 21x fewer tokens.

ReOPD uses pre-collected teacher trajectories to distill knowledge from multi-turn interactions.

problem The cost of fully online on-policy distillation for multi-turn interactions.
method ReOPD, an off-environment alternative that reuses pre-collected teacher trajectories as replayed prefixes, addressing the prefix trap and distribution shift.
result ReOPD preserves or improves OPD-level accuracy, uses zero tool calls, and is at least 4imes imes faster per training step.

Unified notation simplifies information-theoretic concepts in machine learning.

problem Opaque notation for information-theoretic quantities in machine learning.
method Proposed a practical and unified notation for information-theoretic quantities.
result Unified notation facilitates new intuitions and rederivations in machine learning.

Graphical notation simplifies complex polynomial constraints in linear models.

problem Complex polynomial constraints in linear structural equation models are impractical.
method Developed a graphical notation to represent these constraints.
result The graphical notation simplifies the representation of many polynomial constraints.

KVCOMM optimizes multi-agent LLM systems by reusing KV-caches, reducing redundant processing.

problem Substantial overhead from reprocessing overlapping contexts across multi-agent systems.
method KVCOMM reuses KV-caches and adjusts offsets for shared content using a pool of cached examples (anchors).
result Achieves over 70% reuse rate across diverse multi-agent tasks, up to 7.8x speedup.

We describe a method of encoding various types of link diagrams, including those with classical, flat, rigid, welded, and virtual crossings. We show that this method may be used to encode link diagrams, up to equivalence, in a notation whose length is a cubic function of the number of 'riser marks'. For classical knots…

2012-08-01abs ↗pdf ↗

Autostackability for finitely generated groups is defined via a topological property of the associated Cayley graph which can be encoded in a finite state automaton. Autostackable groups have solvable word problem and an effective inductive procedure for constructing van Kampen diagrams with respect to a canonical fini…

2013-07-18abs ↗pdf ↗

Study risk-constrained Kelly optimization for mutually exclusive outcomes, proving support invariance and developing a structured algorithm.

problem Risk-constrained Kelly optimization for mutually exclusive outcomes with explicit state prices.
method Analyzes the finite mutually exclusive outcome version of risk-constrained Kelly optimization with explicit state prices, proving support invariance and developing a structured algorithm.
result Support is invariant across CRRA parameter and drawdown-surrogate parameter in the overround regime.

We provide a proof of backpropagation algorithm in matrix notation.

problem The lack of a full induction proof of backpropagation algorithm in matrix notation.
method We provide a full induction proof of the BP algorithm in matrix notation, situating it in the framework of matrix differential calculus.
result We prove the validity of the backpropagation algorithm in inductive form.

New algorithm reduces regret in private online learning with optimal gap-dependent rate.

problem Optimal gap-dependent regret rate for private stochastic decision-theoretic online learning.
method Horizon-free pure-DP algorithm with exponential block partitioning and softmax selection.
result Explicit regret bound of 1000(logKΔmin+logKε)1000 \cdot (\frac{\log K}{Δ_{\min}}+\frac{\log K}{\varepsilon}).

We calculate the Chern-Simons invariants of the hyperbolic orbifolds of the knot with Conway's notation C(2n,3)C(2n, 3) using the Schläfli formula for the generalized Chern-Simons function on the family of C(2n,3)C(2n,3) cone-manifold structures. We present the concrete and explicit formula of them. We apply the general instruct…

2016-01-05abs ↗pdf ↗

Improved bounds on acylindricity for right-angled Artin groups.

problem Bounding the acylindrical action of right-angled Artin groups on their extension graphs.
method Exploring lattice properties, studying prefixes of powers, and extending quasi-root uniqueness.
result Cardinality of rr-quasi-stabilizer is bounded by a linear function of rr.

As deep learning techniques advance more than ever, hyper-parameter optimization is the new major workload in deep learning clusters. Although hyper-parameter optimization is crucial in training deep learning models for high model performance, effectively executing such a computation-heavy workload still remains a chal…

2019-11-24abs ↗pdf ↗

We consider the problem of path inference: given a path prefix, i.e., a partially observed sequence of nodes in a graph, we want to predict which nodes are in the missing suffix. In particular, we focus on natural paths occurring as a by-product of the interaction of an agent with a network---a driver on the transporta…

2019-03-18abs ↗pdf ↗

We improve private training accuracy with learning rate schedules and matrix factorizations.

problem Private training with learning rate schedules and correlated noise.
method General upper and lower bounds for learning rate schedules, memory-efficient constructions, and schedule-aware factorizations.
result Schedule-aware factorizations improve accuracy in private training.

We introduce a matrix representation of a chord on a tangle which leads us to representing tangle chord diagrams as stacks of matrices that we call books. We show that band sum moves, Reidemeister moves as well as orientation changes are implemented on \widetilde{Z}_f - a framed link invariant constructed from the Kont…

2010-10-14abs ↗pdf ↗

A new fairness metric for decision-making algorithms, conditioning on known fair variables.

problem Fairness issues in decision-making systems.
method Conditional fairness metric, Derivable Conditional Fairness Regularizer (DCFR), adversarial representation.
result Traditional fairness notations are special cases of the new conditional fairness notation.

The paper identifies patterns in language model weights used for memorizing paragraphs.

problem Locating the specific mechanisms and weights used by language models to memorize paragraphs.
method Examined gradients and attention patterns in language models to identify memorized paragraphs.
result Gradients of memorized paragraphs have a distinguishable spatial pattern, and localized attention heads are involved in paragraph memorization.

This article contains general formulas for Tutte and Jones polynomials for families of knots and links given in Conway notation and "portraits of families"-- plots of zeroes of their corresponding Jones polynomials.

2010-04-24abs ↗pdf ↗

After defining reduced minimum braid word and criteria for a braid family representative, different braid family representatives are derived, and a correspondence between them and families of knots and links given in Conway notation is established.

2005-04-23abs ↗pdf ↗

The following discourse is inspired by the works on hyperbolic groups of Epstein, and Neumann/Reeves. Epstein showed that geometrically finite hyperbolic groups are biautomatic. Neumann/Reeves showed that virtually central extensions of word hyperbolic groups are biautomatic. We prove the following generalisation: Theo…

2003-02-20abs ↗pdf ↗

In this paper, we present a new comparative study on automatic essay scoring (AES). The current state-of-the-art natural language processing (NLP) neural network architectures are used in this work to achieve above human-level accuracy on the publicly available Kaggle AES dataset. We compare two powerful language model…

2019-09-18abs ↗pdf ↗

Probabilistic models can handle causal inference without special tools.

problem Confusion over necessary tools for causal inference.
method Demonstrated through concrete examples that causal questions can be answered using standard probabilistic models.
result Causal questions can be addressed using standard probabilistic modelling and inference.

Braids can be represented geometrically as laminations of punctured disks. The geometric complexity of a braid is the minimal complexity of a lamination that represents it, and tight laminations are representatives of minimal complexity. These laminations give rise to a normal form of braids, via a relaxation algorithm…

2015-07-12abs ↗pdf ↗

Penrose's two-spinor notation for 44-dimensional Lorentzian manifolds can be extended to two-component notation for quaternionic manifolds, which is a very useful tool for calculation. We construct a family of quaternionic complexes over unimodular quaternionic manifolds by elementary calculation. On complex quaternio…

2016-10-20abs ↗pdf ↗

PPT optimizes transformer behavior by steering its latent posterior using prior samples.

problem Eliciting desired behavior from transformers without backpropagation.
method Posterior Prefix Tuning (PPT) uses predictive Monte Carlo (PMC) samples and importance sampling to optimize the latent posterior.
result PPT optimizes transformer behavior without backpropagation, achieving high utility across different utility functions.

In this paper, we define a new special curve in Euclidean 3-space which we call {\it kk-slant helix} and introduce some characterizations for this curve. This notation is generalization of a general helix and slant helix. Furthermore, we have given some necessary and sufficient conditions for the kk-slant helix.

2009-09-13abs ↗pdf ↗

We give explicit formulae for the volumes of hyperbolic cone-manifolds of double twist knots, a class of two-bridge knots which includes twist knots and two-bridge knots with Conway notation C(2n,3)C(2n,3). We also study the Riley polynomial of a class of one-relator groups which includes two-bridge knot groups.

2015-12-27abs ↗pdf ↗

Power-SMC reduces inference latency for training-free LLM reasoning.

problem Training-free LLM reasoning with low latency.
method Power-SMC, a training-free Sequential Monte Carlo scheme targeting sequence-level power distribution.
result Power-SMC reduces inference latency from 16-28× to 1.4-3.3× over baseline decoding.