The paper studies neural networks' convergence near origin and saddle points.
problem Directional convergence of neural networks near small initializations and saddle points.
method Gradient flow dynamics analysis of two-homogeneous neural networks.
result Neural networks' weights approximately converge in direction to KKT points for small initializations.
A new method solves SymNMF problems faster and more efficiently.
problem Symmetric nonnegative matrix factorization (SymNMF) for data analytics.
method Nonconvex variable splitting method.
result The method converges to KKT points and has a global sublinear convergence rate.
Early training of deep neural networks leads to small, directionally converging weights.
problem Training dynamics of deep homogeneous neural networks with small initializations.
method Gradient flow analysis and study of KKT points for neural correlation function.
result Weights converge in direction to KKT points during early training stages.
New estimator for tensor weights with improved bias.
problem Estimating tensor weights from noisy data.
method Random matrix theory and KKT conditions.
result Asymptotically unbiased estimator for tensor rank.
LCBO tackles constrained optimization in high dimensions, offering a polynomial convergence rate.
problem Bayesian optimization for high-dimensional constrained problems.
method LCBO uses local descent and uncertainty-driven exploration, proving polynomial convergence rate.
result LCBO achieves a polynomial convergence rate for KKT residuals in high dimensions.
Paper proves k k k -means clustering works on persistence diagrams.
problem Complex geometry of persistence diagram space.
method Proves convergence of k k k -means on persistence diagram space. result Performance of k k k -means on persistence diagrams and measures is superior. New insights into how linear classifiers and leaky ReLU networks can overfit without harming generalization.
problem Understanding conditions for benign overfitting in linear classifiers and leaky ReLU networks.
method Utilizing Karush--Kuhn--Tucker (KKT) conditions for margin maximization.
result Satisfaction of KKT conditions leads to benign overfitting in linear classifiers and leaky ReLU networks.
We embed KKT points in neural networks of different sizes.
problem Classifying data using homogeneous neural networks.
method Introducing KKT point embedding principle and proving it for different network types.
result KKT points of a smaller network can be mapped to those of a larger network via linear transformations.
New algorithm reduces regret in both adversarial and stochastic contexts.
problem Contextual combinatorial semi-bandits with adversarial and corrupted stochastic regimes.
method Follow-the-Regularized-Leader (FTRL) framework with Shannon entropy regularizer, accelerated by Karush-Kuhn-Tucker conditions.
result Achieves O ~ ( T ) \widetilde{\mathcal{O}}(\sqrt{T}) O ( T ) regret in adversarial and O ~ ( ln T ) \widetilde{\mathcal{O}}(\ln T) O ( ln T ) regret in corrupted stochastic regimes. Exact inference possible without observing latent variables in unknown domains.
problem Can exact inference be done without knowing latent variables or their domain?
method Semidefinite programming (SDP) approach based on Karush-Kuhn-Tucker (KKT) conditions and matrix spectrum.
result Exact inference can be achieved without knowing latent variables or their domain.
GD iterates for non-homogeneous deep nets increase margin and converge in direction.
problem Understanding implicit bias in non-homogeneous deep networks.
method Characterization of GD iterates' properties starting from small empirical risk.
result GD iterates converge in direction despite diverging norms, satisfying KKT conditions.
The paper refines NOTEARS for learning Bayesian networks, improving accuracy and efficiency.
problem Learning Bayesian networks from continuous optimization.
method Generalized algebraic characterizations and Karush-Kuhn-Tucker (KKT) conditions for optimization.
result Local search post-processing improves structural Hamming distance by a factor of 2 or more.
DCCNNs reduce computational overhead and ambiguity in convolutional neural networks.
problem Reducing computational overhead and ambiguity in convolutional neural networks.
method Introducing a primal learning problem and constructing a dual convex training program, using Fenchel conjugates and Karush-Kuhn-Tucker conditions.
result Eliminates ambiguity and reduces computational overhead in constructing a large kernel matrix.
Optimization with inequality constraints using embedded gradient vector field method
problem Optimization with inequality constraints
method Geometric framework using quadratic slack variables
result Derives Lagrange multiplier functions and second-order optimality conditions
Gradient ascent method successfully removes specific data points from neural networks without retraining.
problem Addressing privacy and ethical concerns by removing specific data points from trained models.
method Gradient ascent approach to unlearning, leveraging the implicit bias of gradient descent towards margin maximization conditions.
result Gradient ascent method can successfully unlearn specific data points from two-layer ReLU neural networks without retraining.
New K-SVD framework speeds up image denoising with active set algorithm.
problem Efficiently denoise images with high noise levels.
method Proposes K-SVD P _P P using Primal-dual active set (PDAS) algorithm. result Demonstrates comparable performance to state-of-the-art methods.
New attacks bypass data sanitization defenses, increasing model errors.
problem Data poisoning attacks corrupt machine learning models trained on external data.
method Developed three attacks that coordinate poisoned points and formulate as optimization problems.
result 3% poisoned data increases test error from 3% to 24% on Enron spam detection.
Paper proposes ExsdHawkes to model LOBs, capturing volatility dynamics.
problem Modeling volatility signature plots in LOBs with high-frequency trading dynamics.
method Extended State-Dependent Hawkes Process (ExsdHawkes) with relaxed constraints.
result ExsdHawkes uniquely reproduces volatility signature plots, identifying MLOs as catalysts.
We consider rules for discarding predictors in lasso regression and related problems, for computational efficiency. El Ghaoui et al (2010) propose "SAFE" rules that guarantee that a coefficient will be zero in the solution, based on the inner products of each predictor with the outcome. In this paper we propose strong …
Neural network discovers exact solutions to QP with linear constraints.
problem Discovering exact solutions to Quadratic Programs (QP) with linear constraints using neural networks.
method Proposes a neural network modeling approach that analytically derives model parameters from problem coefficients, ensuring closed-form solutions without training.
result The closed-form NN model produces exact solutions for every critical region of the QP solution function, outperforming DNNs and commercial solvers in terms of optimality and feasibility.
Paper proposes SNAP algorithm for finding approximate SOSPs efficiently.
problem Finding approximate second-order stationary points of non-convex problems with linear constraints.
method SNAP algorithm uses strict complementarity condition and negative curvature projections.
result SNAP and SNAP + ^+ + achieve polynomial per-iteration complexity and global sublinear rate for finding SOSPs. The paper tackles partial inference in structured prediction using a convex optimization approach.
problem Maximizing a score function with unary and pairwise potentials in graph label spaces.
method Generative model approach with two-stage convex optimization for label recovery.
result Conditions for recovering a majority of labels with provable guarantees.
New framework for large-scale SUMCOR GCCA with structural regularization.
problem Handling large-scale multiview data with structural regularization.
method Structured SUMCOR Multiview GCCA framework with scalable algorithms.
result Proposed algorithm converges to KKT point of regularized SUMCOR problem.
New method tackles non-smooth tensor data for better recovery.
problem Non-smooth changes in tensor data degrade traditional t-SVD methods.
method Learnable tensor nuclear norm, Alternating Proximal Multiplier Method (APMM), multi-objective tensor recovery framework.
result The proposed method effectively recovers tensor data with non-smooth changes.
Enhanced PC 2 ^2 2 improves surrogate modeling for high-dimensional problems.
problem Degrading performance and efficiency of PC 2 ^2 2 in high-dimensional parameter spaces. method Integrates SULM solver and D-optimal sampling strategy into PC 2 ^2 2 framework. result Enhanced PC 2 ^2 2 demonstrates better comprehensive capability and efficiency. Study KKT conditions for multi-objective optimization on Hadamard manifolds.
problem Optimizing multi-objective interval-valued functions on Hadamard manifolds.
method Developed KKT conditions for Pareto optimal solutions under different ordering and convexity notions.
result Results are more general than on Euclidean spaces.
Develops ADMM for deep neural networks with sigmoid activations to avoid saturation and improve approximation.
problem Gradient saturation in deep neural networks with sigmoid activations.
method Introduces sigmoid-ADMM pair for training deep sigmoid nets and proves its convergence.
result ADMM avoids saturation and improves approximation of deep sigmoid nets compared to ReLU nets.
Unified framework for clustering and learning causal graphs across subjects.
problem Bias and obscured subpopulation-specific dependencies in multivariate systems.
method Directed Acyclic Graph-based Dependency Clustering via Alternating Direction Method of Multipliers (DAG-DC-ADMM) integrated with Structural Equation Modeling (SEM).
result Unified framework recovers cluster-specific causal dependency structures with high true positive rate and low false discovery rate.
Boosted Difference of Convex Functions Algorithm solves VaR constrained portfolio optimization.
problem Designing VaR optimal portfolios under financial regulations.
method Boosted Difference of Convex Functions Algorithm (BDCA) with a novel line search framework.
result BDCA linearly converges to a Karush-Kuhn-Tucker point for VaR constrained portfolio problems.
AM-PPI uses multiple predictors to reduce label cost in healthcare AI.
problem Reduces label cost in post-deployment monitoring of healthcare AI.
method Combines model predictions with a small labeled sample, routing each instance to a cost-appropriate subset of predictors.
result Produces narrower confidence intervals than single-predictor methods.
The paper uses KKT conditions to reveal new insights into SVM behavior.
problem Understanding SVM behavior and tuning.
method Using Karush-Kuhn-Tucker conditions to explore SVM connections with other classifiers.
result SVM can be seen as a cropped version of mean difference and maximal data piling direction classifiers.
New single-loop algorithm tackles weakly convex constraints in stochastic optimization.
problem Optimization with weakly convex constraints in machine learning.
method Single-loop penalty-based stochastic algorithm using hinge-based penalty.
result Achieves state-of-the-art complexity for finding approximate KKT solutions.
The paper proves inequalities for isometries in loxodromic Kleinian groups.
problem Discreteness criteria for subgroups of PSL 2 ( C ) _2(\mathbb{C}) 2 ( C ) . method Generalization of discreteness criteria, using trace inequalities and optimization problems.
result Inequalities involving traces and hyperbolic displacements for loxodromic Kleinian groups.
Optimizes routing in decentralized exchanges with gas fees.
problem Routing in decentralized exchanges with fixed gas fees.
method General optimization framework with mixed-integer model, incorporating gas fees.
result Explicit Karush-Kuhn-Tucker system linking prices, fees, and activation.
New method solves constrained stochastic optimization problems efficiently.
problem Online statistical inference of constrained stochastic nonlinear optimization problems.
method Stochastic Sequential Quadratic Programming (StoSQP) with iterative sketching solver.
result The rescaled primal-dual sequence converges to a mean-zero Gaussian distribution.
We present new computations of approximately length-minimizing polygons with fixed thickness. These curves model the centerlines of "tight" knotted tubes with minimal length and fixed circular cross-section. Our curves approximately minimize the ropelength (or quotient of length and thickness) for polygons in their kno…
Algorithm improves SVM classification in non-Euclidean spaces.
problem Limitations of traditional SVM in non-Euclidean spaces.
method Covariance-adjusted SVM using Cholesky Decomposition.
result Cholesky-SVM outperforms traditional SVM in non-Euclidean spaces.
This paper introduces f f f -DPO, a generalized approach to Direct Preference Optimization using diverse divergence constraints.
problem Aligning large language models with human preferences while mitigating safety risks.
method Incorporates diverse divergence constraints to simplify the relationship between reward and optimal policy, eliminating the need for estimating the normalizing constant.
result Optimizes LLMs to align with human preferences more efficiently and under a broader set of divergence constraints.
Renet improves Elastic Net by dynamically selecting between convex blending and refitting, enhancing prediction accuracy.
problem Elastic Net's shrinkage bias limits its prediction accuracy in high-dimensional settings.
method Adaptive relaxation procedure that dynamically dispatches between convex blending and efficient sub-path refitting.
result Renet consistently outperforms standard Elastic Net and Adaptive Elastic Net in high-dimensional, low signal-to-noise ratio, and high-multicollinearity scenarios.
New manifolds found without interior conjugate points.
problem Existence of interior conjugate points in hyperbolic manifolds.
method Construction of non-trapping asymptotically hyperbolic manifolds.
result Found manifolds without interior conjugate points.
Example shows not all conjugate points are bifurcation points in semi-Riemannian geodesics.
problem Determining which conjugate points in semi-Riemannian geodesics are bifurcation points.
method Revisiting and correcting an example by Musso, Pejsachowicz, and Portaluri.
result Every conjugate point on the improved example is a bifurcation point.
Proposes model-based approach for MI learning using point process theory.
problem Lack of statistical point pattern models in MI learning.
method Develops framework using point process theory for principled extensions of MI learning tasks.
result Tractable point pattern models and solutions for MI learning and decision making.
Estimator calculates surface curvature from point cloud samples.
problem Accurately estimating curvature from limited point cloud data.
method Algorithm using probability distribution and nearby points control.
result Controlled number of points ensures accurate curvature estimation.
This paper analyzes saddle points and minimax points in non-convex smooth games.
problem Understanding local optimal points in non-convex smooth games.
method Comprehensive analysis of local minimax points, including their optimality conditions and stability.
result Local saddle points are uniformly local minimax points under mild continuity assumptions.
New tools for constructing fixed point sets in digital topology.
problem Constructing fixed point sets in digital topology.
method Defining excludable points and articulation points, and showing their exclusion from freezing sets.
result Excludable points and articulation points can be excluded from all freezing sets.
Study on connection points on double regular polygons, providing coordinates and proving non-connection points.
problem Identifying connection points on double regular polygons.
method Examined coordinates in trace field, provided constructive proof for prime n n n . result For n = 7 n=7 n = 7 , conjectured all remaining points are connection points; for n ≥ 7 n \geq 7 n ≥ 7 prime, provided explicit separatrix. New families of translation surfaces with multiple oblivious points discovered.
problem Identifying points on translation surfaces without nearby closed geodesics.
method Constructing new families of translation surfaces and proving existence in higher genera.
result Translation surfaces in every genus ≥3 have at least one oblivious point.
PoPPy simplifies point process modeling and analysis.
problem Efficient modeling and analysis of sequential data.
method Flexible design and efficient learning of point process models.
result PoPPy enables large-scale point process analysis, simulation, and prediction.