Reward collapse occurs when ranking-based reward models yield uniform rewards for different prompts.
problem Reward collapse in aligning large language models with human preferences.
method Introduced a prompt-aware optimization scheme to derive closed-form expressions for reward distributions.
result Our prompt-aware utility functions significantly alleviate reward collapse during training.
Fine-tunes diffusion models to generate diverse samples with high genuine rewards.
problem Reward collapse in finetuning diffusion models.
method Entropy-regularized control against pretrained diffusion models.
result Efficient generation of diverse samples with high genuine rewards.
New method prevents RLHF alignment collapse by accounting for policy's influence on reward model updates.
problem Iterative RLHF leads to alignment collapse where policies exploit RM's blind spots.
method Foresighted policy optimization (FPO) restores missing steering term via regularization.
result FPO prevents alignment collapse on LLM alignment pipelines using Llama-3.2-1B.
New metrics improve scRNA-seq perturbation modeling by reducing mode collapse.
problem Outperformed by simple mean prediction in scRNA-seq perturbation modeling.
method Introduce DEG-aware metrics (WMSE, Rw2(Δ)) and negative/positive baselines. result WMSE loss function reduces mode collapse and improves model performance.
Improved RL training for DMs reduces mode collapse and preserves diversity.
problem Mode collapse and training instability in RL fine-tuned diffusion models.
method Dynamic hierarchical RL training with sliding-window parameter regularisation.
result Models trained with HRF achieve better preservation of diversity in downstream tasks.
Proposes a new RL method to fine-tune flow-based models with arbitrary rewards.
problem Challenges in fine-tuning continuous flow-based generative models with arbitrary reward functions.
method Online Reward-Weighted Conditional Flow Matching with Wasserstein-2 Regularization (ORW-CFM-W2)
result Achieves optimal policy convergence with controllable trade-offs between reward maximization and diversity preservation.
SAGE enhances reinforcement learning by injecting hints to prevent model stagnation.
problem Sparse rewards cause large language models to stall under relative policy optimization.
method SAGE injects privileged hints during training to increase within-group outcome diversity.
result SAGE consistently outperforms GRPO on 6 benchmarks with LLMs, achieving significant improvements.
Text generation is a crucial task in NLP. Recently, several adversarial generative models have been proposed to improve the exposure bias problem in text generation. Though these models gain great success, they still suffer from the problems of reward sparsity and mode collapse. In order to address these two problems, …
New RLHF approach mitigates bias in aligning LLMs with human preferences.
problem Algorithmic bias in RLHF leading to preference collapse.
method Preference Matching (PM) RLHF, using PM regularizer and conditional variant.
result 29% to 41% improvement in alignment with human preferences.
Stable GFlowNets prevent loss spikes and mode collapse in training.
problem Unstable training of GFlowNets leading to loss spikes and mode collapse.
method Assessed sensitivity of GFlowNet objectives, derived loss-to-TV bounds, and proposed Stable GFlowNets.
result Stable GFlowNets improve training behavior and distributional fidelity.
The paper shows how curated synthetic data can optimize human preferences in generative models.
problem Contamination of web-scale datasets by synthetic data affects future model training.
method Theoretical study of iterated retraining of generative models with curated synthetic data.
result Data curation can be seen as an implicit preference optimization mechanism, maximizing expected reward.
RL enhances LLM planning but introduces spurious solutions and diversity collapse.
problem Theoretical understanding of RL's benefits and limitations in LLM planning.
method Graph-based abstraction, policy gradient, Q-learning, supervised fine-tuning.
result RL's exploration is crucial for generalization, but PG suffers from diversity collapse.
Curiosity-Critic improves world model training by focusing on cumulative prediction error.
problem Training world models with intrinsic rewards that consider cumulative prediction error.
method Curiosity-Critic uses a surrogate reward based on the difference between current and asymptotic prediction errors, estimated online by a co-trained critic.
result Curiosity-Critic outperforms other methods in training speed and final world model accuracy.
This paper shows RL with KL penalties is equivalent to Bayesian inference for fine-tuning LMs.
problem Fine-tuning large language models to avoid undesirable features.
method Analyzed KL-regularized RL and showed it's equivalent to variational inference.
result KL-regularized RL avoids distribution collapse and is more insightful as Bayesian inference.
UCPO improves diversity in reinforcement learning models, maintaining high accuracy.
problem RLVR objectives often lead to diversity collapse, reducing coverage of correct solutions.
method UCPO adds a conditional uniformity penalty to GRPO, redistributing probability mass.
result UCPO improves Pass@K and diversity while maintaining competitive Pass@1 accuracy.
AdaDEM decouples EM into two parts to improve class overlap and uncertainty.
problem Improper EM limits its effectiveness in various machine learning tasks.
method Decouple EM into CADF and GMC, and AdaDEM normalizes CADF reward and uses MEC.
result AdaDEM outperforms classical EM and improves performance in noisy and dynamic environments.
This paper proposes a cascading failure mitigation strategy based on Reinforcement Learning (RL) method. Firstly, the principles of RL are introduced. Then, the Multi-Stage Cascading Failure (MSCF) problem is presented and its challenges are investigated. The problem is then tackled by the RL based on DC-OPF (Optimal P…
The paper tackles statistical and computational challenges in learning correlated reward models.
problem The Independence of Irrelevant Alternatives (IIA) assumption collapses human preferences into a universal utility function, leading to coarse approximations.
method The paper investigates the statistical and computational challenges of learning a correlated probit model using best-of-three preference data.
result Best-of-three preference data overcomes the limitations of pairwise preference data, allowing for more fine-grained modeling of human preferences.
The paper analyzes RLVR's training dynamics, proving convergence depends on aligning update direction with Gradient Gap.
problem Understanding why RLVR works and its limitations.
method Analysis of RLVR's training process at trajectory and token levels, introducing Gradient Gap.
result Convergence depends on aligning update direction with Gradient Gap, with a sharp step-size threshold.
Paper explores limits and possibilities of aligning LLMs with human preferences.
problem Aligning LLMs with diverse human preferences to ensure fairness and informed outcomes.
method Analysis of probabilistic representation of human preferences and preservation of diverse preferences.
result LLMs can't fully align with human preferences using reward-based approaches due to Condorcet cycles, but mixed strategies are statistically possible.
Bayesian deep learning faces posterior collapse due to likelihood vs. prior competition.
problem Posterior collapse in Bayesian deep learning models.
method Identified competition between likelihood and prior regularization in a linear latent variable model.
result Posterior collapse is related to neural and dimensional collapse, suggesting a broader learning issue.
Proves weakly non-collapsed RCD spaces are strongly non-collapsed.
problem Proving the equivalence of weakly non-collapsed and strongly non-collapsed RCD spaces.
method Analyzes properties of RCD spaces and uses auxiliary results.
result Confirms conjecture about RCD spaces being strongly non-collapsed.
Study on Neural Collapse limits in deep learning.
problem Understanding the limits of Neural Collapse in deep learning.
method Investigated Neural Collapse in the context of generalization and feature learning, refining conjectures and conducting experiments.
result Neural Collapse primarily occurs on the train set and not on the test set, suggesting it is an optimization phenomenon with unclear connections to generalization.
Generalized dual discriminator GANs improve upon traditional GANs by using two discriminators and a flexible loss function.
problem Mode collapse in GANs.
method Introducing dual discriminator α-GANs and extending the approach to arbitrary functions. result The approach reduces the optimization problem to a linear combination of an f-divergence and a reverse f-divergence. Ricci flow smooths locally collapsing manifolds with controlled curvature.
problem Locally collapsing manifolds with controlled Ricci curvature.
method Ricci flow for a definite period of time, detecting collapsing infranil fiber bundles.
result Topological conditions detect collapsing infranil fiber bundles.
Generative Adversarial Networks (GANs) enjoy great success at image generation, but have proven difficult to train in the domain of natural language. Challenges with gradient estimation, optimization instability, and mode collapse have lead practitioners to resort to maximum likelihood pre-training, followed by small a…
Mathematical analysis shows annealing prevents mode collapse in Gaussian mixtures.
problem Mode collapse in variational inference for multimodal distributions.
method Analyzed annealing strategies for Gaussian mixtures, derived formulas, and tested on neural networks.
result Appropriately chosen annealing schemes can robustly prevent mode collapse.
The study characterizes and rules out collapsing in convex ancient mean curvature flow.
problem Characterizing and ruling out collapsing in convex ancient mean curvature flow.
method Characterization and counterexamples.
result Collapsing occurs if and only if the flow is asymptotic to at least one Grim hyperplane.
Despite excellent progress in recent years, mode collapse remains a major unsolved problem in generative adversarial networks (GANs).In this paper, we present spectral regularization for GANs (SR-GANs), a new and robust method for combating the mode collapse problem in GANs. Theoretical analysis shows that the optimal …
New method controls posterior collapse in VAEs without network architecture constraints.
problem Posterior collapse in VAEs reduces diversity of generated samples.
method Introduces Latent Reconstruction (LR) loss to control posterior collapse.
result Controls posterior collapse on various datasets without architectural constraints.
Study flat manifolds' collapsed limits as flat orbifolds.
problem Understanding collapsed limits of flat manifolds.
method Analyzing totally geodesic foliations and Gromov-Hausdorff limits.
result Identify collapsed limits as flat orbifolds and provide criteria for singularity.
Special Lagrangian submanifolds emerge from K3 surface collapse.
problem Understanding special Lagrangian submanifolds in K3 surface collapse.
method Lifting affine lines to degenerating sequences of special Lagrangian submanifolds.
result Constructing special Lagrangian two-spheres connecting Taub-NUT bubbles.
The torus cannot collapse to a segment under certain curvature conditions.
problem Impossibility of codimension-one collapse for surfaces of negative Euler characteristic.
method Analysis of Gaussian curvature and homology loops.
result The torus cannot collapse to a segment under similar conditions to surfaces of negative Euler characteristic.
Collapsibility is a combinatorial strengthening of contractibility. We relate this property to metric geometry by proving the collapsibility of any complex that is CAT(0) with a metric for which all vertex stars are convex. This strengthens and generalizes a result by Crowley. Further consequences of our work are: (1) …
Two-dimensional collapsed spaces with lower Ricci bounds are topological surfaces.
problem Topology of collapsed spaces with lower Ricci bounds
method Prove that collapsed spaces are topological surfaces
result Collapsed spaces are topological surfaces
Prove that collapsing CSC metrics can be perturbed to invariant collapsing CSC metrics.
problem Prove that collapsing constant scalar curvature metrics can be perturbed to invariant collapsing constant scalar curvature metrics.
method Prove that a sequence of constant scalar curvature metrics which is collapsing with bounded curvature to a manifold can be perturbed to a sequence of invariant collapsing constant scalar curvature metrics.
result Prove that a sequence of constant scalar curvature metrics which is collapsing with bounded curvature to a manifold can be perturbed to a sequence of invariant collapsing constant scalar curvature metrics.
We will simplify the earlier proofs of Perelman's collapsing theorem of 3-manifolds given by Shioya-Yamaguchi and Morgan-Tian. Among other things, we use Perelman's semi-convex analysis of distance functions to construct the desired local Seifert fibration structure on collapsed 3-manifolds. The verification of Perelma…
Estimate collapsibility of causal effects in CPDAGs via strong d-convex hulls.
problem Estimate causal effects in CPDAGs.
method Use strong d-convex hulls to characterize minimal collapsible sets.
result Efficient algorithm for obtaining collapsible sets in DAGs and CPDAGs.
We introduce the theory of strong homotopy types of simplicial complexes. Similarly to classical simple homotopy theory, the strong homotopy types can be described by elementary moves. An elementary move in this setting is called a strong collapse and it is a particular kind of simplicial collapse. The advantage of usi…
Lower Ricci curvature bound prevents first Betti number from dropping more than dimension in collapsing manifolds.
problem Understanding how the first Betti number behaves under manifold collapse with Ricci curvature bounds.
method Analyzing sequences of Riemannian manifolds with lower Ricci curvature bounds.
result The first Betti number cannot drop more than the dimension in collapsing manifolds.
Deep nets exhibit 'Neural Collapse' during training's final phase, simplifying decision-making.
problem Understanding and optimizing deep learning training phases.
method Direct measurements on three deepnet architectures across seven datasets.
result Deep nets exhibit 'Neural Collapse' during training's final phase, simplifying decision-making.
Study tackles criterion collapse in learning criteria, showing conditions for loss minimization.
problem Criterion collapse in optimization, focusing on error probability minimizers.
method Analyzes various learning criteria, including DRO, OCE risks, and non-monotonic criteria.
result Non-monotonic criteria can avoid collapse, while monotonic ones cannot.
Study on collapsing Calabi-Yau manifolds and their metrics.
problem Understanding degenerations of Calabi-Yau manifolds with Ricci-flat Kahler metrics.
method Survey of recent developments, focusing on volume collapsing metrics.
result New insights into the behavior of Calabi-Yau manifolds under volume collapse.
This is an expositiry article on collapsing theory in Riemannian geometry written for the Modern Encyclopedia of Mathematical Physics (MEMPhys). We focus on describing the geometric and topological structure of collapsed/non-collapsed regions in Riemannian manifold under various curvature assumptions. Numerous applicat…
In this paper we extend the works of Tancer and of Malgouyres and Francés, showing that (d,k)-collapsibility is NP-complete for d≥k+2 except (2,0). By (d,k)-collapsibility we mean the following problem: determine whether a given d-dimensional simplicial complex can be collapsed to some k-dimensional sub…
In this paper, we study collapsed manifolds with boundary, where we assume a lower sectional curvature bound, two sides bounds on the second fundamental forms of boundaries and upper diameter bound. Our main concern is the case when inradii of manifolds converge to zero. This is a typical case of collapsing manifolds w…
This paper examines how skip connections prevent rank collapse in sequence models.
problem Rank collapse in sequence models, leading to reduced expressivity and training instabilities.
method Analytical and ablation studies of lambda-skip connections in SSMs.
result A sufficient condition to prevent rank collapse across various architectures.
Survey on collapsing manifolds using group actions and foliations.
problem Collapsing manifolds with controlled curvature.
method Using group actions and singular Riemannian foliations.
result Recent extensions to singular Riemannian foliations.