Cautious Weight Decay modifies weight decay for better optimization.
problem Improving optimization in deep learning models.
method Applies weight decay selectively based on parameter sign alignment.
result Consistently improves model performance across various tasks and scales.
New algorithm improves fairness and robustness in federated learning.
problem Ensuring fairness and robustness in federated learning.
method Formulated federated learning as multi-objective optimization and proposed FedMGDA+.
result FedMGDA+ converges to Pareto stationary solutions, improving performance.
In this paper, we consider a simple kinetic model of economy involving both exchanges between agents and speculative trading. We show that the kinetic model admits non trivial quasi-stationary states with power law tails of Pareto type. In order to do this we consider a suitable asymptotic limit of the model yielding a…
Novel RL-based NPG improves multi-objective NAS efficiency and performance.
problem Discovering optimal neural architectures with multiple conflicting objectives.
method Non-stationary policy gradient with adaptive reward functions and shared model.
result Framework efficiently approximates full Pareto front and achieves superior performance.
MGDA converges under generalized smoothness for neural network optimization.
problem Optimizing neural networks with standard smoothness assumptions not holding.
method Revisited and analyzed MGDA and its stochastic version for generalized ℓ-smooth MOO problems. result MGDA and its variants converge to Pareto stationary points with guaranteed CA distance.
The so-called "Yard-Sale Model" of wealth distribution posits that wealth is transferred between economic agents as a result of transactions whose size is proportional to the wealth of the less wealthy agent. In recent work [B.M. Boghosian, "Kinetics of Wealth and the Pareto Law," {\it Phys. Rev. E} {\bf 89} (2014) 042…
Model predicts stationary equilibrium in investment decisions of firms in fluctuating markets.
problem Investment decisions in fluctuating markets with varying volatility and commodity prices.
method Mean-field model with Gaussian productivity shocks and two-state Markov chain for macroeconomic events.
result Existence, uniqueness, and characterization of stationary mean-field equilibrium with barrier-type investment strategy.
Proposes a new way to represent and analyze Pareto front surfaces.
problem Identifying and analyzing Pareto front surfaces in multi-objective optimization.
method Parameterizes Pareto front surfaces using polar coordinates and scalar-valued length functions.
result Derives statistics of Pareto front surfaces and develops visualisation techniques.
We consider an ideal closed stock market, in which 100 traders have economic activities. The assets of the traders change through buying and selling stocks. We simulate the assets under conservation of both total currency and total number of stocks. If the traders are identical, then the assets are distributed as a sta…
A computational model for the distribution of wealth among the members of an ideal society is presented. It is determined that a realistic distribution of wealth depends upon two mechanisms: an asymmetric flux of wealth in trading transactions that advantages the poorer of the two traders and a non-stationary creation …
Generalized Lotka-Volterra (GLV) models extending the (70 year old) logistic equation to stochastic systems consisting of a multitude of competing auto-catalytic components lead to power distribution laws of the (100 year old) Pareto-Zipf type. In particular, when applied to economic systems, GLV leads to power laws in…
Paper finds a method to compute fair risk-sharing rules.
problem Finding a fair and understandable risk-sharing rule.
method Established a one-to-one correspondence with a fixed point approach.
result Fast numerical method for computing AFPO risk-sharing rules.
New ABC method improves Bézier simplex fitting for noisy data.
problem Overfitting in Bézier simplex fitting when sample points are not on the Pareto set.
method Extended Bézier simplex model to a probabilistic one and proposed a new learning algorithm based on approximate Bayesian computation (ABC) with Wasserstein distance.
result The new algorithm converges on a finite sample and outperforms deterministic methods on noisy instances.
Proposes new stochastic algorithms for multi-objective optimization.
problem Multi-objective optimization in machine learning problems.
method Direction-oriented multi-objective formulation and Stochastic Direction-oriented Multi-objective Gradient descent (SDMGrad).
result Stochastic algorithms converge to Pareto stationary points with improved complexities.
This research extends the Pareto/NBD model using neural networks for better out-of-sample predictions.
problem The limitations of the Pareto/NBD model in predicting out-of-sample data.
method A neural network-based extension of the Pareto/NBD model.
result The proposed method shows extraordinary predictability on repeat purchases at individual and aggregate levels.
Optimization of conflicting functions is of paramount importance in decision making, and real world applications frequently involve data that is uncertain or unknown, resulting in multi-objective optimization (MOO) problems of stochastic type. We study the stochastic multi-gradient (SMG) method, seen as an extension of…
A new method for diverse Pareto solutions in multi-objective learning.
problem Maximizing diversity while maximizing hypervolume in Pareto solutions.
method Annealed Stein Variational Gradient Descent (SVGD) with diverse gradient directions.
result SVH-MOL achieves superior performance in multi-objective and multi-task learning.
We briefly review results on nonlinear kinetic equation of Boltzmann type which describe the evolution of wealth in a simple agents market. The mathematical structure of the underlying kinetic equations allows to use well-known techniques of wide use in kinetic theory of rarefied gases to obtain information on the proc…
PEF identifies the best subgroup performance balance for fairness.
problem Fairness constraints can degrade performance in skewed datasets.
method PEF identifies the closest operating point on the Pareto curve of subgroup performances.
result PEF achieves Pareto levels in accuracy for all subgroups.
The paper proposes a new method to learn choice functions using Pareto-embeddings.
problem Learning subset choices from feature vectors.
method Embedding choice alternatives into a higher-dimensional utility space and identifying choice sets with Pareto-optimal points. Minimizing a differentiable loss function.
result The feasibility of learning a Pareto-embedding demonstrated on benchmark datasets.
We present a multi-objective Bayesian optimisation algorithm that allows the user to express preference-order constraints on the objectives of the type "objective A is more important than objective B". These preferences are defined based on the stability of the obtained solutions with respect to preferred objective fun…
Paper develops sparse learning for heavy-tailed time series with locally stationary dynamics.
problem Sparse learning for high-dimensional heavy-tailed locally stationary time series.
method Additive modeling with kernel smoothing, sparsity-inducing penalized estimation.
result Prediction-error bounds and convergence rates for different sparsity structures.
Proposes Pareto efficient fairness for supervised learning models.
problem Ensuring fairness in machine learning models without sacrificing accuracy.
method Formulates a bilevel optimization problem to find Pareto efficient classifiers.
result Guaranteed solution on Pareto frontier for convex and non-convex objectives.
Proposes a new distribution for robust time series modeling with heavy tails.
problem Robust modeling of time series with heavy-tailed noise.
method Spliced Binned-Pareto distribution for non-stationary time series.
result Accurately models extreme events and captures time dependencies in higher moments.
Optimizing nonlinear systems involving expensive computer experiments with regard to conflicting objectives is a common challenge. When the number of experiments is severely restricted and/or when the number of objectives increases, uncovering the whole set of Pareto optimal solutions is out of reach, even for surrogat…
We studied the Bouchaud-Mézard(BM) model, which was introduced to explain Pareto's law in a real economy, on a random network. Using "adiabatic and independent" assumptions, we analytically obtained the stationary probability distribution function of wealth. The results shows that wealth-condensation, indicated by the …
New findings show Bregman proximal algorithms can get stuck near non-stationary points.
problem Bregman proximal algorithms can get stuck near non-stationary points, misleadingly suggesting convergence.
method Analysis of Bregman proximal algorithms and their behavior near non-stationary points.
result Bregman proximal algorithms can get stuck near spurious stationary points, even in convex problems.
We analyze stochastic gradient algorithms for optimizing nonconvex problems. In particular, our goal is to find local minima (second-order stationary points) instead of just finding first-order stationary points which may be some bad unstable saddle points. We show that a simple perturbed version of stochastic recursiv…
New method finds points for approximating distributions faster.
problem Approximating target probability distributions using finite points.
method Stationary MMD points computed via MMD gradient flows.
result Stationary MMD points converge faster than global minimizers.
We discuss a Pareto macro-economy (a) in a closed system with fixed total wealth and (b) in an open system with average mean wealth and compare our results to a similar analysis in a super-open system (c) with unbounded wealth. Wealth condensation takes place in the social phase for closed and open economies, while it …
Polyconvex energies with conformal invariance have smooth stationary points outside a discrete set.
problem Stationary points of conformally invariant polyconvex energies
method Proving smoothness of stationary points
result Smooth stationary points outside a discrete set
Paper proposes a new method for efficient Pareto Front modeling.
problem Efficient modeling of Pareto Front (PF) in decision making problems.
method Projection based active Gaussian process regression (P-aGPR) method.
result The proposed method can provide a generative PF model and examine new points efficiently.
Many real world applications can be framed as multi-objective optimization problems, where we wish to simultaneously optimize for multiple criteria. Bayesian optimization techniques for the multi-objective setting are pertinent when the evaluation of the functions in question are expensive. Traditional methods for mult…
Develops a deep non-stationary kernel for non-stationary spatio-temporal point processes.
problem Capturing non-stationary dependencies in point process data.
method Approximates the influence kernel with a novel low-rank decomposition and introduces a log-barrier penalty to maintain non-negativity.
result Demonstrates superior performance and computational efficiency compared to state-of-the-art methods.
Paper tackles entity matching over multi-source data, optimizing alignment and mitigating negative transfer.
problem Learning effective entity matching models over multi-source large-scale data with relaxed assumptions.
method Proposes a Relaxed Multi-source Large-scale Entity-matching (RMLE) problem and Incentive Compatible Pareto Alignment (ICPA) method.
result Optimized cross-source alignments and mitigated negative transfer, improving entity matching accuracy.
New method finds stationary points in bilevel optimization problems.
problem Solving nonconvex-strongly-convex bilevel optimization problems.
method Restarted Accelerated HyperGradient Descent (RAHGD) method.
result Achieves best-known theoretical guarantees for finding stationary points in bilevel optimization.
New algorithm finds approximate stationary points faster under differential privacy constraints.
problem Finding approximate stationary points of smooth and Lipschitz functions under differential privacy constraints.
method Developed an efficient algorithm that improves convergence rates to stationary points.
result Achieved faster rates of convergence to stationary points in both finite-sum and stochastic settings.
New method finds exact Pareto front for MO-MDPs efficiently.
problem Finding the exact Pareto front for MO-MDPs is challenging.
method Investigates geometric structure, develops efficient algorithm.
result Pareto front is on boundary of convex polytope of deterministic policies.
This paper tackles the computational complexity of finding approximate stationary points in non-convex optimization.
problem Finding approximate stationary points in non-convex optimization problems.
method PLS-completeness, zero-order algorithms, and gradient queries.
result The query complexity of finding approximate stationary points is Θ(1/ε) for d=2.
Paper proposes Adaptive Pareto Exploration for identifying Pareto optimal arms in multi-objective scenarios.
problem Identifying Pareto optimal arms in multi-objective scenarios with relaxed constraints.
method Adaptive Pareto Exploration strategy for different relaxations of Pareto Set Identification.
result Reduction in sample complexity when identifying at most k Pareto optimal arms.
Diversification improves profits for heavy-tailed investments.
problem Investment portfolios of Pareto-distributed returns.
method Stochastic dominance and majorization order.
result Diversification increases first-order stochastic dominance for heavy-tailed returns.
The paper tackles finding stationary points in stochastic convex optimization problems.
problem Finding stationary points for stochastic convex optimization problems.
method The approach relies on dimension theory to decompose the graph of the subdifferential of a convex function, showing how stochastic sampling preserves 'pieces' of these graphs, and allowing effective application of proximal-point-like methods.
result The paper provides convergence guarantees for finding stationary points in stochastic convex optimization problems.
We propose a strategy for approximating Pareto optimal sets based on the global analysis framework proposed by Smale (Dynamical systems, New York, 1973, pp. 531-544). The method highlights and exploits the underlying manifold structure of the Pareto sets, approximating Pareto optima by means of simplicial complexes. Th…
Neural networks' weights don't converge to stationary points but training loss stabilizes.
problem The disconnect between theoretical analyses and neural network training practice.
method An invariant measure perspective inspired by ergodic theory of dynamical systems.
result The distribution of weights converges to an approximate invariant measure, explaining loss stabilization.
Random variables of the generalized Pareto distribution, can be transformed to that of the Pareto distribution. Explicit expressions exist for the maximum likelihood estimators of the parameters of the Pareto distribution. The performance of the estimation of the shape parameter of generalized Pareto distributed using …
Gradient-based optimization methods are the most popular choice for finding local optima for classical minimization and saddle point problems. Here, we highlight a systemic issue of gradient dynamics that arise for saddle point problems, namely the presence of undesired stable stationary points that are no local optima…
In this article we study the regularity of stationary points of the knot energies Eα introduced by O'Hara in the range α∈(2,3). In a first step we prove that Eα is C1 on the set of all regular embedded closed curves belonging to H(α+1)/2,2 and calculate its derivative. After that we use the structure…
Paper introduces a neural network-based non-stationary influence kernel for complex event data.
problem Modeling complex, non-stationary, and dependent discrete event data.
method Neural Spectral Marked Point Processes (NSMPP) with a versatile non-stationary influence kernel.
result NSMPP outperforms state-of-the-art models on synthetic and real data.