Cautious Weight Decay modifies weight decay for better optimization.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New algorithm improves fairness and robustness in federated learning.
In this paper, we consider a simple kinetic model of economy involving both exchanges between agents and speculative trading. We show that the kinetic model admits non trivial quasi-stationary states with power law tails of Pareto type. In order to do this we consider a suitable asymptotic limit of the model yielding a…
Novel RL-based NPG improves multi-objective NAS efficiency and performance.
MGDA converges under generalized smoothness for neural network optimization.
The so-called "Yard-Sale Model" of wealth distribution posits that wealth is transferred between economic agents as a result of transactions whose size is proportional to the wealth of the less wealthy agent. In recent work [B.M. Boghosian, "Kinetics of Wealth and the Pareto Law," {\it Phys. Rev. E} {\bf 89} (2014) 042…
Model predicts stationary equilibrium in investment decisions of firms in fluctuating markets.
Proposes a new way to represent and analyze Pareto front surfaces.
We consider an ideal closed stock market, in which 100 traders have economic activities. The assets of the traders change through buying and selling stocks. We simulate the assets under conservation of both total currency and total number of stocks. If the traders are identical, then the assets are distributed as a sta…
A computational model for the distribution of wealth among the members of an ideal society is presented. It is determined that a realistic distribution of wealth depends upon two mechanisms: an asymmetric flux of wealth in trading transactions that advantages the poorer of the two traders and a non-stationary creation …
Generalized Lotka-Volterra (GLV) models extending the (70 year old) logistic equation to stochastic systems consisting of a multitude of competing auto-catalytic components lead to power distribution laws of the (100 year old) Pareto-Zipf type. In particular, when applied to economic systems, GLV leads to power laws in…
Paper finds a method to compute fair risk-sharing rules.
New ABC method improves Bézier simplex fitting for noisy data.
Proposes new stochastic algorithms for multi-objective optimization.
This research extends the Pareto/NBD model using neural networks for better out-of-sample predictions.
Optimization of conflicting functions is of paramount importance in decision making, and real world applications frequently involve data that is uncertain or unknown, resulting in multi-objective optimization (MOO) problems of stochastic type. We study the stochastic multi-gradient (SMG) method, seen as an extension of…
A new method for diverse Pareto solutions in multi-objective learning.
We briefly review results on nonlinear kinetic equation of Boltzmann type which describe the evolution of wealth in a simple agents market. The mathematical structure of the underlying kinetic equations allows to use well-known techniques of wide use in kinetic theory of rarefied gases to obtain information on the proc…
The paper proposes a new method to learn choice functions using Pareto-embeddings.
We present a multi-objective Bayesian optimisation algorithm that allows the user to express preference-order constraints on the objectives of the type "objective A is more important than objective B". These preferences are defined based on the stability of the obtained solutions with respect to preferred objective fun…
The potential for learned models to amplify existing societal biases has been broadly recognized. Fairness-aware classifier constraints, which apply equality metrics of performance across subgroups defined on sensitive attributes such as race and gender, seek to rectify inequity but can yield non-uniform degradation in…
Paper develops sparse learning for heavy-tailed time series with locally stationary dynamics.
Proposes Pareto efficient fairness for supervised learning models.
Proposes a new distribution for robust time series modeling with heavy tails.
Optimizing nonlinear systems involving expensive computer experiments with regard to conflicting objectives is a common challenge. When the number of experiments is severely restricted and/or when the number of objectives increases, uncovering the whole set of Pareto optimal solutions is out of reach, even for surrogat…
We studied the Bouchaud-Mézard(BM) model, which was introduced to explain Pareto's law in a real economy, on a random network. Using "adiabatic and independent" assumptions, we analytically obtained the stationary probability distribution function of wealth. The results shows that wealth-condensation, indicated by the …
New findings show Bregman proximal algorithms can get stuck near non-stationary points.
We analyze stochastic gradient algorithms for optimizing nonconvex problems. In particular, our goal is to find local minima (second-order stationary points) instead of just finding first-order stationary points which may be some bad unstable saddle points. We show that a simple perturbed version of stochastic recursiv…
New method finds points for approximating distributions faster.
We discuss a Pareto macro-economy (a) in a closed system with fixed total wealth and (b) in an open system with average mean wealth and compare our results to a similar analysis in a super-open system (c) with unbounded wealth. Wealth condensation takes place in the social phase for closed and open economies, while it …
Polyconvex energies with conformal invariance have smooth stationary points outside a discrete set.
Paper proposes a new method for efficient Pareto Front modeling.
Many real world applications can be framed as multi-objective optimization problems, where we wish to simultaneously optimize for multiple criteria. Bayesian optimization techniques for the multi-objective setting are pertinent when the evaluation of the functions in question are expensive. Traditional methods for mult…
Develops a deep non-stationary kernel for non-stationary spatio-temporal point processes.
Paper tackles entity matching over multi-source data, optimizing alignment and mitigating negative transfer.
New method finds stationary points in bilevel optimization problems.
New algorithm finds approximate stationary points faster under differential privacy constraints.
New method finds exact Pareto front for MO-MDPs efficiently.
This paper tackles the computational complexity of finding approximate stationary points in non-convex optimization.
Paper proposes Adaptive Pareto Exploration for identifying Pareto optimal arms in multi-objective scenarios.
Diversification improves profits for heavy-tailed investments.
The paper tackles finding stationary points in stochastic convex optimization problems.
We propose a strategy for approximating Pareto optimal sets based on the global analysis framework proposed by Smale (Dynamical systems, New York, 1973, pp. 531-544). The method highlights and exploits the underlying manifold structure of the Pareto sets, approximating Pareto optima by means of simplicial complexes. Th…
Neural networks' weights don't converge to stationary points but training loss stabilizes.
Random variables of the generalized Pareto distribution, can be transformed to that of the Pareto distribution. Explicit expressions exist for the maximum likelihood estimators of the parameters of the Pareto distribution. The performance of the estimation of the shape parameter of generalized Pareto distributed using …
Gradient-based optimization methods are the most popular choice for finding local optima for classical minimization and saddle point problems. Here, we highlight a systemic issue of gradient dynamics that arise for saddle point problems, namely the presence of undesired stable stationary points that are no local optima…
In this article we study the regularity of stationary points of the knot energies introduced by O'Hara in the range . In a first step we prove that is on the set of all regular embedded closed curves belonging to and calculate its derivative. After that we use the structure…
Paper introduces a neural network-based non-stationary influence kernel for complex event data.