This work shows MLPs can approximate monotonic functions without bounded activations.
problem Optimizing MLPs with monotonic constraints and bounded activations.
method Generalized theoretical results showing MLPs with non-negative weights and saturating activations are universal approximators.
result MLPs with non-negative weights and saturating activations are universal approximators for monotonic functions.
Study learns a neuron with non-monotonic activation functions.
problem Learning a single neuron with non-monotonic activation functions.
method Gradient descent (GD) with conditions on activation function and input distribution.
result Learnability of non-monotonic activation functions is established without monotonicity assumption.
Stochastic neural networks use Gaussian noise with monotonic activation functions.
problem Training RBM with non-linearities.
method Laplace approximation with Gaussian noise for monotonic activation functions.
result Exp-RBM learns useful representations using stochastic units.
BCD algorithm finds global minima in neural networks.
problem Training deep neural networks to find global minima.
method Block coordinate descent with skip connections and non-negative projection.
result Proves convergence to global minima for strictly monotonic and ReLU activations.
Develops a new geometric framework for quantum metrics.
problem Quantum metric generalization for pure two-qubit states.
method Support-projected Petz monotone geometry for pure two-qubit families.
result Strictly generalizes SLD/Bures case and includes other metrics.
Improves active learning efficiency by warping input space based on observed outputs.
problem Insensitivity of Gaussian process uncertainty to actual observations.
method Input warping with learned monotone reparameterization to adjust acquisition function behavior.
result Significantly improved sample efficiency across various benchmarks, especially in non-stationary conditions.
New method learns SIMs with arbitrary monotone activations without strong distributional assumptions.
problem Learning Single-Index Models with arbitrary monotone activations.
method Based on omniprediction with calibrated multiaccuracy and Bregman divergences.
result First agnostic learning result for SIMs with arbitrary monotone activations.
Mish is a new activation function that improves neural network performance.
problem Improving the performance and training dynamics of neural networks.
method Mish is a self-regularized non-monotonic activation function defined as f(x)=xanh(softplus(x)). It outperforms other functions on benchmarks like ImageNet-1k and MS-COCO. result Mish outperforms Leaky ReLU and ReLU on benchmarks like MS-COCO and ImageNet-1k, respectively, with comparable network parameters.
UMNNs improve density estimation and variational inference without constraints.
problem Creating expressive invertible transformations without constraints.
method Proposed UMNN architecture enforcing monotonicity with a free-form neural network.
result UMNNs enhance autoregressive flows for density estimation and variational inference.
Additive Gaussian process framework handles monotonicity constraints in high dimensions.
problem Handling monotonicity constraints in high-dimensional data.
method Additive Gaussian process framework with MaxMod algorithm for dimension reduction.
result Framework enables to satisfy monotonicity constraints everywhere in the input space.
Monotonic neural additive models simplify machine learning for credit scoring.
problem Complex machine learning methods make models less transparent and fair.
method Introducing monotonic neural additive models that simplify neural networks while maintaining regulatory compliance.
result Monotonic neural additive models achieve similar accuracy to complex neural networks but are more transparent and fair.
Narrow neural networks have unbounded decision regions.
problem Understanding decision regions of narrow neural networks.
method Analyzing decision regions of neural networks with width ≤ input dimension.
result All connected components of decision regions are unbounded.
Robots gather information resiliently despite failures and attacks.
problem Resilient information gathering in adversarial or failure-prone environments.
method First scalable algorithm for minimal communication, system-wide resiliency, and provable approximation performance.
result Algorithm ensures optimal or near-optimal solutions for any number of failures and attacks.
Deep polynomial neural networks measure their expressiveness by the dimension of their functional space.
problem Measuring the expressiveness of deep polynomial neural networks.
method Analyzing the algebraic variety defined by the polynomial neural network's weights and activations.
result The dimension of the algebraic variety is a precise measure of the network's expressiveness.
New activation function SERLU improves neural network performance.
problem Improving neural network performance and avoiding overfitting.
method Introducing a new activation function (SERLU) that breaks monotonicity while preserving self-normalizing property and developing shift-dropout for regularization.
result SERLU-based neural networks provide consistently promising results compared to other activation functions.
MonoNet enforces monotonicity to improve model interpretability.
problem Lack of interpretability in complex machine learning models.
method Enforces monotonicity between features and outputs in deep learning models.
result Enforcing monotonicity allows for better understanding of model predictions.
Study fractal and regular geometry in deep neural networks.
problem Investigate geometric properties of neural networks.
method Analyze boundary volumes of excursion sets for different activations.
result Hausdorff dimension increases with depth for non-regular activations.
Gradient descent dynamics in quadratic regression models are analyzed, revealing five phases: monotonic, catapult, periodic, chaotic, and divergent.
problem Analyzing the dynamics of gradient descent in quadratic regression models.
method Fine-grained bifurcation analysis of gradient descent dynamics using a cubic map parameterized by the step-size.
result Gradient descent dynamics in quadratic regression models exhibit five distinct phases: monotonic, catapult, periodic, chaotic, and divergent.
New activations improve deep network reproducibility without sacrificing accuracy.
problem Deep networks' reproducibility issues, especially on distributed systems.
method Developed SmeLU activations, smoother than ReLU, to enhance reproducibility.
result SmeLU activations provide better accuracy-reproducibility tradeoffs.
Algorithm learns neural networks with two layers in polynomial time.
problem Learning neural networks with two nonlinear layers without assumptions.
method Isotonic regression combined with kernel methods.
result First provably efficient algorithm for two-layer neural networks.
Neural networks can approximate any L^p functions on R^n.
problem Approximating functions on unbounded domains with neural networks.
method Monotone sigmoid, ReLU, ELU, Softplus, LeakyReLU activation functions.
result Shallow neural networks can arbitrarily well approximate L^p functions on R^n.
Optimizes AI learning with limited human feedback budgets.
problem Optimizing allocation of a fixed annotation budget for AI learning.
method Preference-Calibrated Active Learning (PCAL) using semi-parametric inference.
result Proves asymptotic optimality and robustness of the PCAL estimator.
Introduces NQ network for non-crossing quantile learning.
problem Quantile crossing issue in distributional learning.
method Non-negative activation functions ensure monotonic distributions.
result Effective for distributional reinforcement learning and causal effect estimation.
We simplify neural networks to 3D to study their topological changes.
problem Understanding how neural network layers affect low-dimensional topological invariants.
method Limiting each layer to a width of 3D space, tracking changes in linking numbers.
result ResNets and transformers are equally powerful in changing linking numbers.
New algorithm reduces sample complexity for learning CNNs.
problem Learning one-hidden-layer CNNs with various activation functions.
method Approximate gradient descent algorithm for training CNNs.
result Sample complexity matches information-theoretic lower bound for linear activation functions.
Extends stability approach to BSDEs with jumps, providing criteria for existence and uniqueness.
problem Existence and uniqueness of solutions to BSDEs with jumps.
method Monotone stability approach, non-convex generator, non-global Lipschitz conditions.
result Concrete criteria for existence and uniqueness of solutions, comparison, and bounds.
A new KAN variant uses sinusoidal activations to approximate functions.
problem Approximating multivariable functions using neural networks.
method Replacing inner and outer functions in Kolmogorov-Arnold representation with weighted sinusoidal functions.
result The new KAN variant outperforms fixed-frequency Fourier transform and achieves comparable performance to MLPs.
New insights on eluder dimension for function approximation in machine learning.
problem Complexity measure for online bandits and reinforcement learning with function approximation.
method Study the relationship between eluder dimension and generalized rank for different activation functions.
result Eluder dimension can be exponentially smaller or larger than generalized rank depending on the activation function.
We solve ReLU regression with efficient approximations for various distributions.
problem Finding the best fitting ReLU function with square loss from unknown distributions.
method Introduced efficient constant-factor approximation algorithm and polynomial-time approximation scheme.
result First constant-factor approximation algorithm for ReLU regression with weak concentration conditions.
Large GD stepsizes improve margins and speed up training for non-homogeneous networks.
problem Training efficiency and margin improvement in non-homogeneous two-layer networks.
method Investigation of two distinct phases in GD training, showing margin growth and empirical risk decrease.
result Large GD stepsizes lead to faster convergence and improved margins in non-homogeneous networks.
New neural network with RePU activation approximates smooth functions and their derivatives.
problem Approximating smooth functions and their derivatives with neural networks.
method Differentiable neural networks with RePU activation functions.
result Improved approximation error bounds for RePU-activated neural networks.
Active learning improves subspace clustering with less labeled data.
problem Efficiently incorporating labeled data to improve subspace clustering models.
method Proposes an active learning framework for subspace clustering that queries informative points and updates the subspace model.
result Demonstrates the advantage of the proposed active strategy over state-of-the-art methods.
Active learning selects high-quality examples for text-to-SQL systems.
problem Efficiently annotate large language models for text-to-SQL systems.
method Formalizes example selection as a constrained experimental design problem over semantic query embeddings, proposing a stratified greedy algorithm that maximizes heteroscedastic mutual information.
result Proposed method significantly reduces labeling effort while maintaining high text-to-SQL retrieval accuracy.
A new method for disentangled latent spaces in VAEs that can manipulate attributes.
problem Disentangled representation of attributes in latent spaces of VAEs.
method Attribute-based regularization loss to enforce monotonic relationships between attributes and latent codes.
result Manipulation of attributes in latent spaces post-training.
SSFN self-estimates network size with low complexity and consistent performance.
problem Designing a self-estimating feed-forward network with low complexity and consistent performance.
method Joint optimization for layer and node estimation, low computational complexity, and use of lossless flow property and convex optimization.
result Consistent performance across Monte-Carlo trials and monotonically non-increasing cost with network growth.
SAFE detects fraudsters in advance by predicting survival probabilities.
problem Detecting fraudsters in time given their activity sequences.
method Survival analysis with RNN to map user activities to hazard values.
result SAFE outperforms existing models in fraud early detection.
Active-set algorithm improves Cox regression for shape-restricted covariates.
problem Improving Cox regression for shape-restricted covariates.
method Shape-restricted inference using active-set optimization for spline basis expansion.
result Active-set algorithm produces accurate linear covariate effect estimates.
The paper addresses monotonicity in machine learning models for fairness and accountability.
problem Ensuring fairness and accountability in transparent machine learning models.
method Study of three types of monotonicity (individual, weak pairwise, strong pairwise) and propose monotonic groves of neural additive models.
result Monotonic groves of neural additive models maintain transparency, accountability, and fairness.
New algorithm tackles stochastic optimization with inequality constraints.
problem Stochastic optimization with inequality constraints in various applications.
method Active-set stochastic sequential quadratic programming (StoSQP) with a differentiable exact augmented Lagrangian.
result Global convergence for any initialization, KKT residuals converge to zero almost surely.
Controller-Augmented Hidden Markov Models (CHMMs) are a framework for constrained sequential inference.
problem Hidden Markov models fail under pathwise constraints like precedence, visitation, or monotonic state progression.
method CHMMs compile constraints into finite-state controllers, then use standard forward-backward and Viterbi recursions to compute exact constrained posteriors and paths.
result CHMMs provide exact constrained inference, monotone ascent in constrained EM, and linear complexity in controller cardinality.
The paper tackles non-monotonic learning performance and proposes algorithms to make models more monotone.
problem Non-monotonic learning performance where more data does not always improve model quality.
method Proposes three algorithms to make supervised learning models more monotone, proving consistency and monotonicity with high probability.
result The algorithm MT-HT reduces less than 1% non-monotonic decisions on MNIST while maintaining competitive error rates.
Neural network memorizes external stimuli through synaptic strength changes.
problem Memory and classification in neural networks.
method One-to-one mapping between stimulus and synaptic strength under synaptic plasticity constraints.
result Neural network can memorize external stimuli through synaptic changes.
We investigate quotation and transaction activities in the foreign exchange market for every week during the period of June 2007 to December 2010. A scaling relationship between the mean values of number of quotations (or number of transactions) for various currency pairs and the corresponding standard deviations holds…
Probit Monotone BART estimates binary outcomes using monotonic functions.
problem Estimating conditional mean functions for binary outcomes with monotonicity constraints.
method Proposes a new BART variant that incorporates monotonicity constraints for binary outcomes.
result Allows for more precise estimation of monotonic functions in binary outcome models.
Develops a first-order interior-point method for solving constrained variational inequalities.
problem Solving constrained variational inequalities with nontrivial constraints.
method ADMM-based interior-point method for constrained VIs (ACVI).
result First-order interior-point method with global convergence guarantees for general cVI problems.
Monotone neural networks can approximate and interpolate functions efficiently.
problem Understanding the efficiency and expressiveness of monotone neural networks.
method Solving the monotone interpolation problem using depth-4 networks and comparing size bounds with arbitrary networks.
result Monotone neural networks can approximate and interpolate functions efficiently, but may require exponential size in high dimensions.
Improves k-NN for monotonic data with robustness against noise.
problem Class noise in real-life data violates monotonic constraints in k-NN.
method Monotonic Fuzzy k-NN (MonFkNN) with new fuzzy membership calculation.
result Significant accuracy improvements and robustness against monotonic noise.
Study examines explainable machine learning for monotonic models, finding Integrated gradients better for strong monotonicity.
problem Applying explainable machine learning to science-informed models.
method Proposed axioms for monotonicity, tested Shapley value and Integrated gradients methods.
result Integrated gradients provides better explanations for strong monotonicity.