New method detects inconsistencies in AHP matrices using triadic preference reversals.
problem Challenges in assessing consistency in AHP pairwise comparison matrices.
method Triadic preference reversals to detect inconsistencies between pairs of elements.
result 97% accuracy in detecting inconsistencies, significantly surpassing traditional methods.
Neural networks combining multiple data sources can reverse preferences, affecting decision reliability.
problem Preference reversals in neural networks under pooled data.
method Formalized through Case-Based Decision Theory, analyzed Gram geometry, introduced regularization, and developed auditing methods.
result Pooled refitting can reverse shared preferences, and conditions for preserving preferences are derived.
Unified framework for aligning LLMs from human feedback.
problem Lack of strong theoretical justification for RLHF and inability to compare methods.
method Reframed alignment as distribution learning from pairwise preferences, proposing three principled objectives.
result Proposed objectives achieve strong non-asymptotic convergence to target LM.
New RLHF framework handles general preference oracles without reward functions.
problem Handling general preference oracles without assuming a reward function.
method Developed a minimax game between two LLMs for RLHF under a general preference oracle, focusing on KL-regularized preference.
result Proposed algorithms for efficient offline and online RLHF learning.
This paper introduces f-DPO, a generalized approach to Direct Preference Optimization using diverse divergence constraints.
problem Aligning large language models with human preferences while mitigating safety risks.
method Incorporates diverse divergence constraints to simplify the relationship between reward and optimal policy, eliminating the need for estimating the normalizing constant.
result Optimizes LLMs to align with human preferences more efficiently and under a broader set of divergence constraints.
Paper introduces a novel framework for recognizing dynamic ranking structures in preference-based data.
problem Complex and noisy preference-based data often hide underlying homogeneous structures.
method Developed an approach to identify dynamic ranking groups using temporal penalties and spectral estimation. Introduced an objective function for detecting structural changes.
result Consistent recognition of ranking groups and structural changes in preference-based data.
Study on reinsurance decisions using mean-variance criterion with irreversible contracts.
problem Optimizing reinsurance premiums and contracts in a Stackelberg game with irreversible contracts.
method Unified singular control framework applied to both discrete and continuous time reinsurance contracts.
result A single once-for-all reinsurance contract is preferred over multiple contracts, and the signing time is crucial.
The paper analyzes how behavioral investors make portfolio decisions using Markowitz Stochastic Dominance criteria.
problem Understanding how behavioral investors make portfolio decisions.
method Developed stochastic optimization problems and MILP models to capture subjective decision weights and probability weighting functions.
result The developed models can be used to formulate computationally tractable portfolio analysis problems.
A new framework for adaptive behavior using reusable value profiles.
problem Adaptive behavior in changing environments requires switching among value-control regimes, but maintaining separate parameters for each situation is impractical.
method Introduces value profiles: reusable bundles of parameters assigned to hidden states, allowing for state-conditional strategy recruitment without independent parameters for each context.
result Profile-based models outperform simpler alternatives in probabilistic reversal learning, suggesting belief-dependent control of adaptive behavior.
Paper addresses RLHF alignment challenges with novel algorithms.
problem Challenges in RLHF alignment, especially in strategic exploration.
method Develops a reverse-KL regularized contextual bandit formulation and proposes efficient algorithms with theoretical guarantees.
result Proposed methods significantly outperform existing RLHF algorithms in real-world experiments.
Study shows marketable order routing to wholesalers benefits all traders, leading to lower market depth and price volatility.
problem Determining the preference of retail traders for marketable order routing.
method Two models: one for market makers competing for retail order flow (Bertrand model) and another for price-taking competitive liquidity providers (open exchange model).
result Routing marketable orders to wholesalers is preferred by all traders, leading to mean reverting inventories and lower market depth.
Paper proposes f-DPG for aligning language models with preferences.
problem Aligning language models with user preferences.
method Uses f-divergence to approximate target distributions and minimizes a forward KL from it using DPG.
result Jensen-Shannon divergence often outperforms forward KL divergence, leading to significant improvements.
ADPO optimizes relative advantage in reinforcement learning from human feedback.
problem Optimizing policy alignment in reinforcement learning from human preferences.
method ADPO explicitly parameterizes the optimal structure through anchored logits, decoupling response quality from prior popularity.
result Empirically, ADPO achieves state-of-the-art performance on reasoning tasks, outperforming GRPO by 30.9 percent.
New GP kernels avoid mean reversion without losing smoothness.
problem Pathological behavior in stationary GP regression.
method Improper Gaussian processes with non-positive kernels.
result Stationary, non-reverting covariance functions.
We introduce and establish the main properties of QHawkes ("Quadratic" Hawkes) models. QHawkes models generalize the Hawkes price models introduced in E. Bacry et al. (2014), by allowing all feedback effects in the jump intensity that are linear and quadratic in past returns. A non-parametric fit on NYSE stock data sho…
Methodology that recently lead us to predict to an amazing accuracy the date (July 11, 2008) of reverse of the oil price up trend is briefly summarized and some further aspects of the related oil price dynamics elaborated. This methodology is based on the concept of discrete scale invariance whose finance-prediction-or…
Classifies reversible and strongly reversible elements in quaternionic groups.
problem Classifying reversible and strongly reversible elements in quaternionic groups.
method Proves elements are reversible if and only if they are products of skew-involutions (resp. involutions).
result Proves elements are reversible if and only if they are products of skew-involutions (resp. involutions).
The paper classifies reversible and strongly reversible elements in Hermitian isometry groups.
problem Classifying reversible and strongly reversible elements in Hermitian isometry groups.
method Classification through group theory and algebraic manipulation.
result New classification of strongly reversible elements in Sp(n).
The paper classifies reversible and strongly reversible elements in quaternionic hyperbolic spaces.
problem Classifying reversible and strongly reversible elements in quaternionic hyperbolic spaces.
method Analyzing conjugacy classes and using properties of quaternionic hyperbolic spaces and their isometry groups.
result All elements of the isometry group of quaternionic hyperbolic spaces are strongly reversible.
SUMO provides unbiased log marginal likelihood estimation for latent variable models.
problem Biased estimates of log marginal likelihood in latent variable models.
method Randomized truncation of infinite series for unbiased estimation.
result Models trained with SUMO give better test-set likelihoods than standard methods.
Based on our "finance-prediction-oriented" methodology which involves such elements as log-periodic self-similarity, the universal preferred scaling factor lambda=2, and allows a phenomenon of the "super-bubble" we analyze the 2009 world stock market (here represented by the SP500, Hang Seng and WIG) development. We id…
This paper classifies reversible and strongly reversible elements in affine groups.
problem Classifying reversible and strongly reversible elements in affine groups.
method Identifying affine transformations and using conjugacy by involutions.
result Classification of reversible and strongly reversible elements in affine groups.
Study finds mean reversion strategies perform well on historical data but fail in recent market conditions.
problem Performance of mean reversion strategies in recent market data.
method Empirical investigation of three mean reversion strategies (PAMR, OLMAR, TCO) on historical S&P 500 data and benchmark datasets.
result Mean reversion strategies may fail in recent market conditions, especially with transaction costs.
Boundary-induced apparent risk aversion in non-ergodic growth models.
problem Risk aversion in multiplicative growth systems with absorbing boundaries.
method Exact lattice propagation and analysis of binary multiplicative processes.
result Optimal exposure is compressed near absorbing boundaries, mimicking risk aversion.
New neural net learns time-reversible symplectic dynamics.
problem Lack of time-reversibility in neural networks for symplectic systems.
method Proposes a new neural network architecture for time-reversible symplectic systems.
result Demonstrates learning of time-reversible symplectic dynamics from data.
A new trading strategy using reinforcement learning for statistical arbitrage.
problem Traditional statistical arbitrage models rely on model assumptions and price deviations from a long-term mean.
method Empirical reversion time metric, reinforcement learning framework, and state space optimization.
result Optimal mean reversion strategy identified through reinforcement learning.
Algebraic method reveals criterion for quaternionic Möbius group reversibility.
problem Characterizing reversibility in quaternionic Möbius group elements.
method Purely algebraic approach using matrix entries and conjugacy invariants.
result Explicit criterion for reversibility in terms of matrix entries.
A Finsler space is said to be geodesically reversible if each oriented geodesic can be reparametrized as a geodesic with the reverse orientation. A reversible Finsler space is geodesically reversible, but the converse need not be true. In this note, building on recent work of LeBrun and Mason, it is shown that a geodes…
Sharp stability results for reverse isoperimetric inequalities in 2D.
problem Reverse isoperimetric inequalities in the plane.
method Stability analysis of λ-convex bodies and convex bodies with smooth boundaries. result Sharp stability results for reverse isoperimetric inequalities, including inradius and Cheeger inequalities.
On-line portfolio selection has attracted increasing interests in machine learning and AI communities recently. Empirical evidences show that stock's high and low prices are temporary and stock price relatives are likely to follow the mean reversion phenomenon. While the existing mean reversion strategies are shown to …
Boundary rigidity proven for non-reversible Finsler metrics.
problem Recovering non-reversible Finsler metrics from boundary distance data.
method Sum of reversible Finsler norm and closed 1-form, boundary rigidity results.
result 1-form can be uniquely recovered from boundary distance data.
Conditional diffusion models can approximate target distributions well with Gaussian-mixture reverse kernels.
problem Approximating target distributions in conditional diffusion models.
method Using finite Gaussian mixtures with ReLU-network logits as reverse kernels, reducing the problem to static conditional density approximation.
result The resulting neural reverse-kernel class is dense in conditional KL divergence under exact terminal matching.
Reverse Experience Replay improves Deep Q-learning for sparse rewards.
problem Sparse rewards and reward-maximizing tasks in Deep Q-learning.
method Sampling transitions in reverse order for training.
result Significantly increased performance in tasks with limited experience and memory capacity.
New knots not rationally concordant to their reverses found.
problem Identifying knots not rationally concordant to their reverses.
method Infinite family of knots constructed, rational knot concordance group analyzed.
result Infinite rank subgroup in rational knot concordance group.
The paper classifies reversible elements in Seifert-fibered spaces and braid groups.
problem Classifying reversible elements in Seifert-fibered spaces and braid groups.
method Classification of reversible elements in Fuchsian groups, application to Seifert-fibered groups, and analysis of 3-torsion elements.
result Classification and analysis of reversible and 3-torsion elements in Seifert-fibered spaces and braid groups.
Reverse annealing boosts quantum matrix factorization performance.
problem Improving quantum matrix factorization performance.
method Combining forward and reverse annealing for nonnegative/binary matrix factorization.
result Combination of forward and reverse annealing significantly improves performance.
TRS-ODENs learn dynamics with time-reversal symmetry for more efficient learning.
problem Learning dynamics with time-reversal symmetry for more efficient learning.
method Proposed a loss function and a new framework (TRS-ODENs) to learn dynamics efficiently.
result TRS-ODENs can learn dynamics from noisy and complex trajectories efficiently.
Paper proves rigidity theorems for geodesically reversible Finsler metrics.
problem Understanding geodesically reversible Finsler metrics in closed manifolds.
method Applied theory of volumes and areas on Finsler spaces to establish rigidity theorems.
result Partial explanation of the scarcity of geodesically reversible Finsler metrics in closed manifolds.
Reduces identity testing of reversible Markov chains to simpler symmetric chain tests.
problem Testing identity of reversible Markov chains from a single trajectory.
method Using lumping-congruent Markov embeddings, the problem is simplified to testing symmetric chains over a larger state space.
result Achieves state-of-the-art sample complexity for identity testing.
Proof of reverse isoperimetric inequality for black holes.
problem Reverse isoperimetric inequality for black holes in Einstein gravity.
method Geometric-analytic approach.
result Reversal of the usual isoperimetric inequality is explained by curved backgrounds governed by Einstein's equations.
Optimizes molecular generation for chemist preferences.
problem Models lack inherent preferences for chemist-desired structures.
method Fine-tuning with Direct Preference Optimization.
result Approach is simple, efficient, and highly effective.
New method adapts to user preferences dynamically, improving recommendation models.
problem Current recommendation models lack dynamic adaptation to changing user preferences.
method Preference Discerning with LLM-Enhanced Generative Retrieval
result Mender achieves state-of-the-art performance in adapting to evolving user preferences.
New algorithm learns bridged diffusion processes without time-reversals.
problem Learning bridged diffusion processes efficiently and accurately.
method Score matching with Doob's h-transform, avoiding time-reversals.
result Outperforms existing methods in learning bridged diffusion processes.
Consider the problem of pricing options on forwards in energy markets, when spot prices follow a geometric multi-factor model in which several rates of mean reversion appear. In this paper we investigate the role played by slow mean reversion when pricing and hedging options. In particular, we determine both upper and …
DPS uses posterior sampling for preference-based RL, achieving a first regret guarantee.
problem Formal frameworks for preference-based RL with theoretical analysis.
method Preference-based posterior sampling, Bayesian credit assignment.
result First asymptotic Bayesian no-regret rate for preference-based RL.
Many real-world engineering problems rely on human preferences to guide their design and optimization. We present PrefOpt, an open source package to simplify sequential optimization tasks that incorporate human preference feedback. Our approach extends an existing latent variable model for binary preferences to allow f…
Enhances preference learning by incorporating response times into binary choices.
problem Limited information from binary choices about preference strength.
method Combines choices and response times using the EZ diffusion model.
result Response times improve utility estimation for strong preferences.
Bayesian optimization learns DM preferences for multi-outcome experiments.
problem Optimizing expensive experiments with unknown utility functions and multiple outcomes.
method Alternates preference learning and Bayesian optimization, using pairwise comparisons.
result Preference exploration strategies improve Bayesian optimization performance.