Supervised learning is an active research area, with numerous applications in diverse fields such as data analytics, computer vision, speech and audio processing, and image understanding. In most cases, the loss functions used in machine learning assume symmetric noise models, and seek to estimate the unknown function …
Function-space MAP estimation leads to better generalization and robustness.
problem The mismatch between parameter posterior and function posterior in model training.
method Directly estimating the most likely function implied by the model and data.
result Function-space MAP estimation can lead to flatter minima, better generalization, and improved robustness.
Study families of Morse functions for manifolds with boundary.
problem Characterize degeneracies in 1-parameter families of Morse functions.
method List all possible degeneracies in generic 1-parameter families.
result Identified all degeneracies in generic 1-parameter families.
Deep neural networks with specific parameter sets can approximate smooth functions efficiently.
problem Approximating smooth functions with deep neural networks.
method Deep neural networks with ReLU activation and specific parameter sets {0,±21,±1,2} are used to approximate Cβ-smooth functions. result The constructed networks can approximate Cβ-smooth functions with parameters {0,±21,±1,2} efficiently, achieving the same convergence rate as sparse networks with parameters in [−1,1]. A new model approximates complex functions in parameter space.
problem Complex and nonlinear functional regression problems.
method Mapping-to-Parameter function model with B-spline free knot placement.
result Robust knot placement algorithms improve model performance.
Study introduces a new method for multiple parameter regularization in polynomial functional regression.
problem Handling varying regularization parameters in polynomial functional regression.
method Developed a theoretically grounded algorithm for multiple parameter regularization and model aggregation.
result Promising results from evaluations on synthetic and real-world data.
Optimized neural network approximates high-dimensional functions with minimal parameters.
problem Achieving optimal approximation of high-dimensional continuous functions with minimal parameters.
method Developed a neural network with a specific activation function and architecture to achieve super approximation property.
result A composed network with at most 10889d + 10887 nonzero parameters achieves super approximation property, suggesting optimality in parameter growth.
Efficient estimators for smooth Hilbert-valued parameters with theoretical guarantees.
problem Estimating smooth Hilbert-valued parameters with theoretical guarantees.
method Pathwise differentiable Hilbert-valued parameters, efficient influence functions, regularized one-step estimators.
result Theoretical guarantees for efficient estimators even when nuisance functions are arbitrary.
Improved Bayesian optimization for conditional parameter spaces.
problem Efficient global optimization of expensive-to-evaluate functions in conditional parameter spaces.
method Additive tree-structured covariance function for conditional parameter optimization.
result Significantly improved sample-efficiency and wider applicability compared to existing methods.
Deep networks can approximate functions with fewer learnable parameters than previously thought.
problem High computational costs due to large number of parameters in deep neural networks.
method Theoretical design of ReLU networks with a few intrinsic parameters and numerical experiments.
result ReLU networks with a small number of intrinsic parameters can achieve good approximations of functions.
Study minimal timelike surfaces in 3D Lorentz-Minkowski space using holomorphic functions.
problem Characterize minimal timelike surfaces in R13. method Use a Weierstrass-type formula with holomorphic functions in split-complex numbers to find canonical parameters and corresponding holomorphic functions.
result Enneper surfaces are the only minimal timelike surfaces with polynomial parametrization of degree 3 in isothermal parameters.
Functional dimension varies in ReLU networks, with implications for symmetry and connectivity.
problem Understanding the functional dimension of ReLU neural networks.
method Careful definition and analysis of functional dimension, study of quotient space and fibers.
result Functional dimension is inhomogeneous and can be non-constant, with implications for symmetry and connectivity.
We consider inference about a scalar parameter under a non-parametric model based on a one-step estimator computed as a plug in estimator plus the empirical mean of an estimator of the parameter's influence function. We focus on a class of parameters that have influence function which depends on two infinite dimensiona…
New method encodes function preferences into neural nets for better generalization.
problem Challenges in encoding explicit function preferences in neural network training.
method Function-space empirical Bayes (FSEB) regularization.
result FSEB leads to near-perfect semantic shift detection and improved generalization.
Paper introduces a new robust loss function for RL.
problem Heuristic selection of threshold parameters in quantile Huber loss.
method Derived from Wasserstein distance, captures noise in quantile values.
result Enhances robustness against outliers and enables parameter adjustment.
We provide adaptive inference methods, based on ℓ1 regularization, for regular (semi-parametric) and non-regular (nonparametric) linear functionals of the conditional expectation function. Examples of regular functionals include average treatment effects, policy effects, and derivatives. Examples of non-regular f…
Study biharmonic functions on vector bundles with spherical symmetry.
problem Investigate biharmonic functions on vector bundles with spherically symmetric metrics.
method Analyze vertical lifts and radial functions of functions on vector bundle manifolds.
result Construct an infinite two-parameter family of proper biharmonic functions.
The paper introduces canonical parameters for marginally trapped surfaces in Minkowski space.
problem Determining marginally trapped surfaces in Minkowski space.
method Introducing canonical parameters and proving existence and uniqueness theorems.
result Every marginally trapped surface is determined by three smooth functions.
Proposes a new method for estimating non-pathwise differentiable functional parameters.
problem Estimating dose-response curves for continuous exposure.
method Targeted Highly Adaptive Lasso (HAL) for non-pathwise differentiable functional parameters.
result The Targeted HAL-MLE achieves dimension-free rates up to log(n) factors and outperforms other methods in simulations.
A neural network with a single hidden layer can't represent certain multivariable functions.
problem Representing certain multivariable functions with a neural network having only one hidden layer.
method Developed a continuum version of a one-hidden-layer neural network with ReLU activation, and proved constraints on its parameters and second derivative.
result Existence of a smooth binary function that cannot be precisely represented by any such neural network.
Solves challenges in estimating parameters of softmax gating Gaussian mixture models.
problem Identifiability issues and complex interactions in Gaussian mixture of experts.
method Proposes novel Voronoi loss functions and establishes convergence rates of MLE.
result Connects convergence rate of MLE to a solvability problem of polynomial equations.
Novel covariance function improves Bayesian optimization efficiency.
problem Efficient global optimization of expensive black-box functions.
method Additive tree-structured covariance function and parallel optimization algorithm.
result Significantly outperforms state-of-the-art methods in conditional parameter optimization.
Proposes a generalized XGBoost method for nonconvex loss functions.
problem Limited to convex loss functions in XGBoost.
method Extends XGBoost to use nonconvex loss functions and multivariate loss functions.
result Generalized XGBoost method can model multiple parameters in various distributions.
MPF method improves parameter estimation in probabilistic models.
problem Difficulty in fitting probabilistic models due to intractable partition function.
method Minimum Probability Flow (MPF) method for parameter estimation.
result MPF outperforms existing techniques in convergence time and accuracy.
Estimates parameters in a deviated Gaussian mixture model.
problem Testing goodness-of-fit between a known function and a mixture of experts.
method Constructs novel Voronoi-based loss functions to estimate parameters.
result Characterizes local convergence rates of parameter estimation more accurately.
In this work, a method of random parameters generation for randomized learning of a single-hidden-layer feedforward neural network is proposed. The method firstly, randomly selects the slope angles of the hidden neurons activation functions from an interval adjusted to the target function, then randomly rotates the act…
New insights into neural network complexity reveal better generalization performance.
problem Mysterious generalization in deep models despite high parameter counts.
method Effective dimensionality as a measure of parameter space complexity.
result Double descent behavior in generalization as a function of parameters explained.
This paper introduces a family of local feature aggregation functions and a novel method to estimate their parameters, such that they generate optimal representations for classification (or any task that can be expressed as a cost function minimization problem). To achieve that, we compose the local feature aggregation…
Neural networks cannot approximate certain functions in Sobolev spaces, leading to unbounded parameter growth.
problem Non-closedness of sets of neural networks in Sobolev spaces.
method Construction of sequences of neural networks whose realizations converge to functions not realizable by neural networks.
result Sets of realized neural networks are not closed in order-(m−1) Sobolev spaces Wm−1,p for p∈[1,∞]. Develops methods to measure and set function-space learning rates in neural networks.
problem Measuring and optimizing changes in neural network output functions.
method Efficient methods to measure and set function-space learning rates, requiring minimal computational overhead.
result Demonstrates FLeRM (Function-space Learning Rate Matching) for hyperparameter transfer across model scales.
New method improves neural network robustness by identifying functions rather than parameters.
problem Neural networks' lack of robustness to distribution shifts.
method Identify the function represented by quadratic networks, not their parameters.
result Obtain robust generalization bounds for neural networks.
Develops methods for estimating constrained function-valued parameters in infinite-dimensional models.
problem Estimating function-valued parameters with structural constraints in complex models.
method Characterizes constrained solutions as minimizers of penalized population risk, using a Lagrange-type formulation and path through unconstrained space.
result Proposes estimators that achieve optimal risk and constraint satisfaction, applicable across various statistical learning approaches.
Quantum neural networks approximate periodic functions more efficiently.
problem Approximating periodic functions with quantum neural networks.
method Using Jackson's inequality to construct a QNN that approximates a trigonometric polynomial of the function.
result Quantum neural networks can achieve better approximation results with fewer parameters for smoother functions.
This paper improves neural network approximation for analytic functions with adjustable depth and width.
problem Approximating analytic functions using neural networks with depth and width parameters.
method Characterizes approximation rates as a joint function of width (N) and depth (L) for ReLU networks.
result Establishes upper bounds for analytic function approximation rates of O(N^(-CL^τ)) with τ influenced by N and L.
Due to the intractable partition function, the exact likelihood function for a Markov random field (MRF), in many situations, can only be approximated. Major approximation approaches include pseudolikelihood and Laplace approximation. In this paper, we propose a novel way of approximating the likelihood function throug…
New insights into the top-K sparse softmax gating function for deep learning.
problem Understanding the theoretical effects of the top-K sparse softmax gating function on density and parameter estimations.
method Using a Gaussian mixture of experts, novel loss functions, and theoretical analysis.
result The convergence rates of density and parameter estimations are parametric under certain conditions, but slow under over-specified models.
The paper establishes bounds on the smoothness parameter in Gaussian process interpolation.
problem Estimating the smoothness parameter in Gaussian process models.
method Approximation theory in Sobolev spaces and general theorems on parameter estimation.
result Maximum likelihood estimation recovers the true smoothness for certain classes of functions.
Multi-class classification methods based on both labeled and unlabeled functional data sets are discussed. We present a semi-supervised logistic model for classification in the context of functional data analysis. Unknown parameters in our proposed model are estimated by regularization with the help of EM algorithm. A …
When applying Machine Learning techniques to problems, one must select model parameters to ensure that the system converges but also does not become stuck at the objective function's local minimum. Tuning these parameters becomes a non-trivial task for large models and it is not always apparent if the user has found th…
New approach uses negative controls to estimate causal parameters without completeness conditions.
problem Estimating causal parameters when not all confounders are observed.
method Identification strategy based on minimax learning formulations for general function classes.
result Avoids completeness conditions and uniqueness assumptions on bridge functions.
Proposes debiasing strategy for ill-posed regression problems.
problem Estimating functions with conditional moment restrictions, especially when estimators are sensitive to misspecification.
method Debiased estimation using influence function of modified mean squared error.
result Demonstrates finite-sample convergence rate and robustness to misspecification.
In Divide & Recombine (D&R), big data are divided into subsets, each analytic method is applied to subsets, and the outputs are recombined. This enables deep analysis and practical computational performance. An innovate D\&R procedure is proposed to compute likelihood functions of data-model (DM) parameters for big dat…
We consider the Willmore functional on graphs, with an additional penalization of the area where the curvature is non-zero. Interpreting the penalization parameter as a Lagrange multiplier, this corresponds to the Willmore functional with a constraint on the area where the graph is flat. Sending the penalization parame…
New findings on hidden symmetries in ReLU networks.
problem Understanding the redundancy and symmetries in ReLU network parameter space.
method Analyzing parameter settings and function classes for various network architectures.
result For certain network architectures, there are no hidden symmetries.
Maxout networks show similar complexity issues as ReLU networks.
problem Understanding the complexity of maxout networks and decision boundaries.
method Analyzing the parameter space and decision boundaries, obtaining lower bounds, and investigating initialization procedures.
result Maxout networks exhibit a wide range of complexity, similar to ReLU networks.
Symmetry of neural network densities can be determined from correlation functions.
problem Determining symmetries of neural network densities without knowing the density itself.
method Symmetry-via-duality approach using invariance properties of correlation functions.
result Symmetries of neural network densities can be determined via dual computations of correlation functions.
Policy evaluation is a key process in reinforcement learning. It assesses a given policy using estimation of the corresponding value function. When using a parameterized function to approximate the value, it is common to optimize the set of parameters by minimizing the sum of squared Bellman Temporal Differences errors…
Training-free model learns SDE dynamics without training, accelerating parameter studies.
problem High computational cost of simulating parameter-dependent SDEs.
method Training-free conditional diffusion model with joint kernel-weighted Monte Carlo estimator.
result Accurate approximation of conditional distributions across varying parameter values.