Intersectional constraints improve selection outcomes by reducing inequality.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This work investigates implicit bias in multiclass separable data using a novel geometry-aware optimizer.
Paper introduces CageBO for optimizing complex public policy problems.
We consider whether algorithmic choices in over-parameterized linear matrix factorization introduce implicit regularization. We focus on noiseless matrix sensing over rank- positive semi-definite (PSD) matrices in , with a sensing mechanism that satisfies restricted isometry properties (RIP)…
In this paper, we make a generalization of Routh's reduction method for Lagrangian systems with symmetry to the case where not any regularity condition is imposed on the Lagrangian. First, we show how implicit Lagrange-Routh equations can be obtained from the Hamilton-Pontryagin principle, by making use of an anholonom…
We shed new insights on the two commonly used updates for the online -PCA problem, namely, Krasulina's and Oja's updates. We show that Krasulina's update corresponds to a projected gradient descent step on the Stiefel manifold of the orthonormal -frames, while Oja's update amounts to a gradient descent step using…
In this paper, we prove that the set of solutions of constraint equations for coupled Einstein and scalar fields in classical general relativity possesses Hilbert manifold structure. We follow the work of R. Bartnik [2] and use weighted Sobolev spaces and Implicit Function Theorem to prove our results.
We study the generalization properties of stochastic gradient methods for learning with convex loss functions and linearly parameterized functions. We show that, in the absence of penalizations or constraints, the stability and approximation properties of the algorithm can be controlled by tuning either the step-size o…
Study discretizes Dirac and port-Hamiltonian systems using manifolds.
Proposes a new model for image restoration combining deep learning and total variation.
Method estimates posterior model for boundary value problems with uncertain constraints.
Paper tackles efficient SGD methods for constrained bilevel optimization.
Two effective methods for writing the dynamical equations for non-holonomic systems are illustrated. They are based on the two types of representation of the constraints: by parametric equations or by implicit equations. They can be applied to linear as well as to non-linear constraints. Only the basic notions of vecto…
We propose graph-dependent implicit regularisation strategies for distributed stochastic subgradient descent (Distributed SGD) for convex problems in multi-agent learning. Under the standard assumptions of convexity, Lipschitz continuity, and smoothness, we establish statistical learning rates that retain, up to logari…
We present a unified approach to constrained implicit Lagrangian and Hamiltonian systems based on the introduced concept of Dirac algebroid. The latter is a certain almost Dirac structure associated with the Courant algebroid on the dual to a vector bundle . If this almost Dirac structure is integrable (Dir…
New framework improves robustness of implicit neural networks.
The mathematical problem concerning intrinsic storage optimisation is formulated and solved by means of variational analysis. The solution, though obtained in implicit form, still sheds light on many important features of the optimal exercise strategy. It is shown how the solution depends on different constraint types …
Representation learning systems typically rely on massive amounts of labeled data in order to be trained to high accuracy. Recently, high-dimensional parametric models like neural networks have succeeded in building rich representations using either compressive, reconstructive or supervised criteria. However, the seman…
A new method for optimizing non-decomposable metrics with constraints.
Although conservative Hamiltonian systems with constraints can be formulated in terms of Dirac structures, a more general framework is necessary to cover also dissipative systems such as gradient and metriplectic systems with constraints. We define Leibniz-Dirac structures which lead to a natural generalization of Dira…
New dual formulation reduces generalization error for ERM-fDR.
A core capability of intelligent systems is the ability to quickly learn new tasks by drawing on prior experience. Gradient (or optimization) based meta-learning has recently emerged as an effective approach for few-shot learning. In this formulation, meta-parameters are learned in the outer loop, while task-specific m…
In this note we describe how some objects from generalized geometry appear in the qualitative analysis and numerical simulation of mechanical systems. In particular we discuss double vector bundles and Dirac structures. It turns out that those objects can be naturally associated to systems with constraints -- we recall…
We consider a semilinear parabolic degenerated Hamilton-Jacobi-Bellman (HJB) equation with singularity which is related to a stochastic control problem with fuel constraint. The fuel constraint translates into a singular initial condition for the HJB equation. We first propose a transformation based on a change of vari…
The purpose of this paper is to define the concept of multi-Dirac structures and to describe their role in the description of classical field theories. We begin by outlining a variational principle for field theories, referred to as the Hamilton-Pontryagin principle, and we show that the resulting field equations are t…
The paper solves stochastic control problems with implicit objectives, finding equilibrium strategies.
This work addresses the instability in asynchronous data parallel optimization. It does so by introducing a novel distributed optimizer which is able to efficiently optimize a centralized model under communication constraints. The optimizer achieves this by pushing a normalized sequence of first-order gradients to a pa…
Most existing distance metric learning methods assume perfect side information that is usually given in pairwise or triplet constraints. Instead, in many real-world applications, the constraints are derived from side information, such as users' implicit feedbacks and citations among articles. As a result, these constra…
AdamW optimizes a constrained loss with norm constraint.
Semi-supervised learning is an important and active topic of research in pattern recognition. For classification using linear discriminant analysis specifically, several semi-supervised variants have been proposed. Using any one of these methods is not guaranteed to outperform the supervised classifier which does not t…
New framework improves reliability of learned representations by modeling uncertainty and structural constraints.
New method avoids failures in physics-constrained systems using active learning.
Images seen during test time are often not from the same distribution as images used for learning. This problem, known as domain shift, occurs when training classifiers from object-centric internet image databases and trying to apply them directly to scene understanding tasks. The consequence is often severe performanc…
Canonical Correlation Analysis (CCA) is a classic technique for multi-view data analysis. To overcome the deficiency of linear correlation in practical multi-view learning tasks, various CCA variants were proposed to capture nonlinear dependency. However, it is non-trivial to have an in-principle understanding of these…
While implicit feedback (e.g., clicks, dwell times, etc.) is an abundant and attractive source of data for learning to rank, it can produce unfair ranking policies for both exogenous and endogenous reasons. Exogenous reasons typically manifest themselves as biases in the training data, which then get reflected in the l…
In this paper, we propose a scalable algorithm for spectral embedding. The latter is a standard tool for graph clustering. However, its computational bottleneck is the eigendecomposition of the graph Laplacian matrix, which prevents its application to large-scale graphs. Our contribution consists of reformulating spect…
Let be a solution to the maximal constraint equations of general relativity on the unit ball of . We prove that if is sufficiently close to the initial data for Minkowski space, then there exists an asymptotically flat solution on that ext…
We introduce the implicitly constrained least squares (ICLS) classifier, a novel semi-supervised version of the least squares classifier. This classifier minimizes the squared loss on the labeled data among the set of parameters implied by all possible labelings of the unlabeled data. Unlike other discriminative semi-s…
Study reveals biases in gradient descent for GLNs, improving neural network performance.
Study develops a new method for creating fair models.
Muon optimizer improves deep learning with spectral norm constraints.
Random Matrix Theory (RMT) is applied to analyze the weight matrices of Deep Neural Networks (DNNs), including both production quality, pre-trained models such as AlexNet and Inception, and smaller models trained from scratch, such as LeNet5 and a miniature-AlexNet. Empirical and theoretical results clearly indicate th…
Human matting, high quality extraction of humans from natural images, is crucial for a wide variety of applications. Since the matting problem is severely under-constrained, most previous methods require user interactions to take user designated trimaps or scribbles as constraints. This user-in-the-loop nature makes th…
In a polynomial regression model, the divisibility conditions implicit in polynomial hierarchy give way to a natural construction of constraints for the model parameters. We use this principle to derive versions of strong and weak hierarchy and to extend existing work in the literature, which at the moment is only conc…
This paper shows how to train only the implicit layer of overparameterized implicit neural networks.
Autonomous cyber-physical agents and systems play an increasingly large role in our lives. To ensure that agents behave in ways aligned with the values of the societies in which they operate, we must develop techniques that allow these agents to not only maximize their reward in an environment, but also to learn and fo…
A four-dimensional Walker geometry is a four-dimensional manifold M with a neutral metric g and a parallel distribution of totally null two-planes. This distribution has a natural characterization as a projective spinor field subject to a certain constraint. Spinors therefore provide a natural tool for studying Walker …
Paper studies the theoretical equivalence between implicit and explicit neural networks in high dimensions.