Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

2356 · May 202619922001200920172026
26 results for e-process

This paper shows how to combine optimal tests into log-optimal processes.

problem How to combine optimal sequential tests into log-optimal processes.
method Using a new class of WAIT e-processes, the paper aggregates asymptotically optimal sequential tests into asymptotically log-optimal processes.
result It is possible to aggregate asymptotically optimal sequential tests into asymptotically log-optimal e-processes.

In sequential anytime-valid inference, any admissible procedure must be based on e-processes: generalizations of test martingales that quantify the accumulated evidence against a composite null hypothesis at any stopping time. This paper proposes a method for combining e-processes constructed in different filtrations b…

2024-02-15abs ↗pdf ↗

A new method for backtesting ES forecasts in banking.

problem Designing a model-free backtesting procedure for Expected Shortfall forecasts.
method Use e-values and e-processes to introduce backtest e-statistics for VaR and ES.
result The proposed method can be applied to various risk measures and statistical quantities.

Develops new e-processes and confidence sequences for Gaussian means with unknown variance.

problem Constructing valid t-tests and confidence sequences for Gaussian means with unknown variance.
method Explores generalized nonintegrable martingales and extended Ville's inequality, developing two new e-processes and confidence sequences.
result Analyzes the width of resulting confidence sequences with a polynomial dependence on error probability, proving it to be unavoidable and even better than classical fixed-sample t-tests.

A new method for releasing AI workflows to avoid premature incorrect results.

problem Statistical challenges in releasing AI workflows with adaptive scoring.
method Wrapper that calibrates and accumulates evidence from high-scoring failures.
result Reduces premature incorrect release while still releasing on moderate evidence.

Extends FC-RAG to anytime-valid sequential coverage for language model swarms.

problem Maintain distribution-free coverage for a swarm of weak language models over time.
method Introduces Anytime-FC-RAG, a sequential extension with a summable calibration-deviation budget.
result Achieves time-uniform alarm validity and safety under predictable adaptive control.

This paper shows how to construct sequential tests with power one against weakly compact sets in Polish spaces.

problem Testing composite null hypotheses involving weakly compact sets in Polish spaces.
method Develops sequential tests for i.i.d. laws in Polish spaces, providing a sufficient condition for power one.
result Power-one sequential tests exist for weakly compact sets against their complements in i.i.d. laws in Polish spaces.

Unified framework controls false discovery rate in bandit multiple testing.

problem Designing adaptive algorithms to identify true discoveries in multiple hypothesis testing.
method Unified modular framework using e-processes for FDR control in arbitrary settings.
result Unified framework ensures FDR control for dependent and simultaneous arm queries.

CSA fills a gap in RLVR-trained LLM deployment by providing anytime-valid selective risk control.

problem Deployment of RLVR-trained LLMs in regulated organizations requires a safety certificate for every round without waiting for long-run averages.
method CSA uses a (test statistic, validity guarantee, deployment rule) framework to fill the gap, maintaining a Ville-type e-process per threshold on a Bonferroni grid.
result CSA provides the first anytime-valid selective risk control for RLVR-trained LLMs, matching the long-run average certification rate and satisfying pathwise validity and non-refusing deployment on every cell.

aLTT selects hyperparameters efficiently with statistical guarantees.

problem Statistical validity and efficiency in hyperparameter selection.
method Sequential data-dependent multiple hypothesis testing with early termination.
result Reduces testing rounds while maintaining statistical validity.

PITMonitor monitors model calibration over time with formal error guarantees.

problem Fixed-sample tests applied to models over time can lead to false alarms.
method PITMonitor uses mixture e-processes to detect distributional shifts in probability integral transforms.
result PITMonitor achieves competitive detection rates on river's FriedmanDrift benchmark.

CITE algorithm provides anytime-valid certification of model outputs.

problem Challenges in controlling error levels in LLM self-consistency.
method Certification by Intersection-union Testing with E-processes (CITE) algorithm.
result Provable control of false certification at any prescribed level under arbitrary stopping rules.

Study near-maturity convergence rates of American put prices in Lévy models.

problem Analyzing convergence rates of optimal exercise prices in Lévy models.
method Examined two settings: jumps of unbounded and bounded variation, deriving near-maturity expansions.
result Near-maturity convergence rate of optimal exercise price is of order √(T-t).

E-valuator converts verifier scores into reliable decision rules.

problem Ensuring the correctness of agent trajectories based on heuristic scores.
method Sequential hypothesis testing framework for online monitoring of agent trajectories.
result E-valuator provides better false alarm rate control and statistical power than other strategies.