MGDA converges under generalized smoothness for neural network optimization.
problem Optimizing neural networks with standard smoothness assumptions not holding.
method Revisited and analyzed MGDA and its stochastic version for generalized -smooth MOO problems.
result MGDA and its variants converge to Pareto stationary points with guaranteed CA distance.