This paper proposes a method to approximate non-Gaussian likelihoods in Gaussian Processes.
problem Approximating non-Gaussian likelihoods in Gaussian Processes.
method Proposes a piece-wise constant approximation for the inverse-link function.
result Yields a closed form solution for the SVGP lower bound.
Discrete approach proves a new version of Wirtinger's inequality.
problem Classical Wirtinger's inequality for piece-wise functions.
method Theorem of classification of isometries and Fourier series development.
result New elementary proof of Wirtinger's inequality.
XOFM explains attribute effects in ordinal regression using piece-wise linear functions.
problem Lack of detailed attribute contributions in existing ordinal regression models.
method XOFM uses piece-wise linear functions to approximate attribute contributions and introduces ordinal transformation.
result XOFM provides superior explainability and state-of-the-art prediction accuracy.
L*ReLU improves deep learning for fine-grained image classification.
problem Fine-grained image classification requires specific AFs.
method Proposes L*ReLU, piece-wise linear AFs for deep learning.
result L*ReLU achieves superior results on FGVC datasets.
This paper extends depth separation results to piece-wise oscillatory functions.
problem Approximating functions with piece-wise oscillatory structure using neural networks.
method Extends existing results to piece-wise oscillatory functions using proof strategy from (Eldan and Shamir, 2016).
result Approximation by one-hidden-layer networks holds at a poly(d) rate for functions with constant domain radius and oscillation rate.
We introduce a new activation function using Chebyshev-Lagrange polynomials for improved neural network performance.
problem Improving data efficiency and accuracy of neural networks.
method Parameterized piece-wise polynomial activation functions based on Chebyshev nodes and Lagrangian interpolation.
result Significant improvements in model capacity and accuracy, especially in linear extrapolation.
Paper investigates a new type of elastica formed during beam rupture.
problem Understanding the shape of a beam during sudden rupture.
method Developed a mathematical theory using elliptic ζ-function.
result Explicit shape of the new elastica (Λ-elastica) is described.
This work simplifies adversarial attacks using neural networks, reducing computation and improving training convergence.
problem Efficiently generating and training against ideal adversarial attacks with minimal computational overhead.
method Representing ideal adversarial attacks as smooth piece-wise functions and approximating them with neural networks. Using a mathematical game between an attack network and a defense network for adversarial training.
result Obtained convergence rates of adversarial loss in terms of sample size n for adversarial training. PyChEst detects changes in non-stationary time series without distributional assumptions.
problem Detecting changes in non-stationary time series data.
method Nonparametric algorithms for consistent detection of multiple changepoints in piece-wise stationary processes.
result PyChEst consistently detects changes without distributional assumptions.
Estimates change-points and graph structures in a time-varying Ising model.
problem Detecting and understanding changes in a time-varying Ising model.
method Maximizing a penalized conditional log-likelihood to estimate neighborhood of each node, enforcing sparsity and piece-wise constant graph structures.
result First change-points consistency theorems for unknown number of change-points in time-varying Ising model.
Deep nets' optima are surprisingly connected by simple paths.
problem Understanding the loss landscape of deep nets.
method Mathematical explanations based on generic properties of well-trained nets.
result Optima in deep nets are connected by piece-wise linear paths.
RUMBoost combines RUMs and deep learning for better choice modelling.
problem Creating interpretable and robust discrete choice models.
method Gradient Boosted Regression Trees for utility functions, with constraints for interpretability and monotonicity.
result RUMBoost outperforms ML and RUM benchmarks in predictive performance and interpretability.
Bayesian method detects change points and clusters in piece-wise constant signals.
problem Detecting change points and clustering in piece-wise constant signals.
method Nonparametric penalized least square model selection on partitions of design points, with an efficient algorithm.
result Oracle inequality and adaptive upper bound on expected square risk of the estimator.
We investigate the functional determinant of the laplacian on piece-wise flat two-dimensional surfaces, with conical singularities in the interior and/or corners on the boundary. Our results extend earlier investigations of the determinants on smooth surfaces with smooth boundaries. The differences to the smooth case a…
It is shown that most of the well-known basic results for Sobolev-Slobodeckii and Bessel potential spaces, known to hold on bounded smooth domains in Rn, continue to be valid on a wide class of Riemannian manifolds with singularities and boundary, provided suitable weights, which reflect the nature of the s…
Neural models can realize decision trees with parameter sharing and improved performance.
problem Training and optimizing oblique decision trees.
method Locally constant networks based on ReLU gradients, parameter sharing, and neural tools.
result Locally constant networks can implicitly model oblique decision trees with fewer neurons.
Efficiently infers switching nonlinear systems with collapsed amortized variational inference.
problem Inference in switching nonlinear dynamical systems with discrete latent variables.
method Learn an inference network as a proposal for continuous latent variables, performing exact marginalization of discrete variables.
result Successfully segments time series data into meaningful regimes using piece-wise nonlinear dynamics.
A new algorithm finds optimal solutions for constrained decision processes.
problem Optimizing state-value functions with constraints in CMDPs.
method Gradient-Aware Search (GAS) exploiting PWLC structure.
result GAS converges faster and more reliably than existing methods.
In this survey article, we review the relation between heat kernels and path integrals. In particular, we review recent results on the approximation of the Wiener measure on compact manifold by measures on (finite-dimensional) spaces of piece-wise geodesics.
Most of machine learning approaches have stemmed from the application of minimizing the mean squared distance principle, based on the computationally efficient quadratic optimization methods. However, when faced with high-dimensional and noisy data, the quadratic error functionals demonstrated many weaknesses including…
The functional determinant of an elliptic operator with positive, discrete spectrum may be defined as e−Z′(0), where Z(s), the zeta function, is the sum ∑n∞λn−s analytically continued to s around the origin. In this paper Z′(0) is calculated for the Laplace operator with Dirichlet boundary…
Method finds differential equations for integrable billiard tables.
problem Finding differential equations for integrable billiard tables.
method Introducing a method to find differential equations for functions defining tables.
result Illustrated method in three billiard systems.
DAMI uses interpretable regions to select informative samples for deep learning models.
problem Efficiently identifying informative samples for deep learning models with minimal annotation cost.
method Inspired by piece-wise linear interpretability in DNN, DAMI selects samples on different linearly separable regions.
result DAMI outperforms state-of-the-art approaches in tabular data.
CTR prediction in real-world business is a difficult machine learning problem with large scale nonlinear sparse data. In this paper, we introduce an industrial strength solution with model named Large Scale Piece-wise Linear Model (LS-PLM). We formulate the learning problem with L1 and L2,1 regularizers, leadin…
New methods for calculating curvature in graph theory.
problem Calculating curvature in graphs and random walks.
method Analyzing continuous and discrete-time Ollivier-Ricci curvatures of weighted graphs.
result Generalized existence and properties of Ollivier-Ricci curvature for various random walks.
Novel method converts time series data into functional data for high dimensional classification.
problem Small sample size problem in high dimensional time series data.
method Classwise Functional Principal Component Analysis (PCA) followed by Bayesian linear classifier.
result Demonstrated efficacy on synthetic and real data sets.
This paper studies a class of continuous-time scalar-state stochastic Linear-Quadratic (LQ) optimal control problem with the linear control constraints. Applying the state separation theorem induced from its special structure, we develop the explicit solution for this class of problem. The revealed optimal control poli…
The paper proves Sard's theorem for polynomial maps in infinite dimensions.
problem The validity of Sard's theorem for polynomial maps in infinite-dimensional Banach manifolds.
method Sharp quantitative criteria for the validity of Sard's theorem.
result The paper provides criteria for the validity of Sard's theorem in infinite-dimensional Banach manifolds.
Method to create rational Seifert surfaces for knots in Lens space.
problem Creating rational Seifert surfaces for knots in Lens space.
method Assuming a regular projection, construct rational Seifert surface on twist toroidal diagram.
result A method to construct rational Seifert surfaces for knots in Lens space.
Let S be a triangulated 2-sphere with fixed triangulation T. We apply the methods of thin position from knot theory to obtain a simple version of the three geodesics theorem for the 2-sphere [5]. In general these three geodesics may be unstable, corresponding, for example, to the three equators of an ellipsoid. Using a…
Given two points on a soup can or conical cup with lid, we find and classify all paths of minimal length connecting them. When the number of minimal paths is finite, there are at most four on a can and three on a cup. At worst, minimal paths are piece-wise smooth with three components, each of which is a classical geod…
This paper proposes a novel Gaussian process approach to fault removal in time-series data. Fault removal does not delete the faulty signal data but, instead, massages the fault from the data. We assume that only one fault occurs at any one time and model the signal by two separate non-parametric Gaussian process model…
Gradient descent can use larger step sizes to avoid strict saddle points.
problem Avoiding strict saddle points in non-convex optimization.
method Proving that gradient descent with step-size up to 2/L avoids strict saddle points with high probability.
result Gradient descent with step-size up to 2/L almost surely avoids strict saddle points.
AdaPID optimizes diffusion-based samplers by dynamically adjusting schedules.
problem Optimizing the intermediate-time dynamics in diffusion-based samplers.
method Develops a time-varying stiffness schedule using Piece-Wise-Constant (PWC) parametrizations and a hierarchical refinement approach.
result QoS-driven PWC schedules consistently improve sampling fidelity and accuracy.
Finding minimum distortion of adversarial examples and thus certifying robustness in neural network classifiers for given data points is known to be a challenging problem. Nevertheless, recently it has been shown to be possible to give a non-trivial certified lower bound of minimum adversarial distortion, and some rece…
Optimal order execution strategies for brokers under reference benchmarks.
problem Maximizing broker's utility of excess profit-and-loss subject to reference strategies.
method Formulated as a utility maximization problem, optimal strategies derived in closed form.
result General reference strategies can be approximated by piece-wise linear combinations of IS and TC orders.
Study tackles inverse problems on low-dimensional manifolds, proving stability and proposing a reconstruction algorithm.
problem Inverse problems in infinite-dimensional spaces with nonlinear and ill-posed nature.
method Assumption of low-dimensional manifold, proving stability, proposing Landweber-type algorithm.
result Global convergence of the proposed algorithm, Lipschitz stability for specific inverse problems.
UNIPoint universally approximates point process intensities.
problem How to precisely describe the flexibility of point process models.
method Proof using Stone-Weierstrass Theorem, transfer functions, and recurrent neural networks.
result UNIPoint performs better than other models on synthetic and real-world datasets.
The paper proposes new cross-correlators using Price's Theorem and piecewise-linear decomposition.
problem Optimal method for estimating cross-correlations using finite samples.
method General mathematical framework using Price's Theorem and piecewise-linear decomposition.
result Some cross-correlators based on Huber's loss functions, MP functions, and LSE functions have higher SNR.
The study examines generalization bounds for regression and classification tasks on adaptive input domains.
problem Understanding the generalization error in adaptive input domains for regression and classification.
method The analysis considers regression and classification separately, using Lipschitz continuity and 2-norm/0/1 loss for measurement. It also highlights the polynomial relationship between generalization bounds and network parameters.
result Generalization bounds for regression and classification are inversely proportional to a polynomial of the number of parameters, emphasizing the advantages of over-parameterized networks.
Motivated by an important insight from neural science, we propose a new framework for understanding the success of the recently proposed "maxout" networks. The framework is based on encoding information on sparse pathways and recognizing the correct pathway at inference time. Elaborating further on this insight, we pro…
Tree ensemble kernels improve Bayesian optimization for mixed features and constraints.
problem Optimizing over mixed-feature spaces with known constraints.
method Kernel interpretation of tree ensembles as Gaussian Process prior, compatible optimization formulation for acquisition function, integration of known constraints.
result Framework outperforms state-of-the-art methods for mixed-feature spaces and constraints.
Improves discrete latent representations using differentiable approximation bridges.
problem Improving discrete latent representations in neural networks.
method Training with a differentiable approximation bridge (DAB) neural network.
result Improves state-of-the-art performance in various domains.
In many applications we seek to maximize an expectation with respect to a distribution over discrete variables. Estimating gradients of such objectives with respect to the distribution parameters is a challenging problem. We analyze existing solutions including finite-difference (FD) estimators and continuous relaxatio…
Neural model accelerates SDDP for stochastic optimization.
problem Exponential complexity of SDDP limits its applicability to low-dimensional problems.
method Trainable neural model maps problem instances to a low-dimensional piecewise linear value function.
result ν-SDDP significantly reduces problem solving cost without sacrificing solution quality.
We propose a strategy for approximating Pareto optimal sets based on the global analysis framework proposed by Smale (Dynamical systems, New York, 1973, pp. 531-544). The method highlights and exploits the underlying manifold structure of the Pareto sets, approximating Pareto optima by means of simplicial complexes. Th…
We study in this paper a class of constrained linear-quadratic (LQ) optimal control problem formulations for the scalar-state stochastic system with multiplicative noise, which has various applications, especially in the financial risk management. The linear constraint on both the control and state variables considered…
The L1 loss landscape of neural nets near local minima behaves differently, revealing exponential decay and increased vertex density.
problem Understanding the L1 loss landscape of neural nets near local minima.
method Iterative minimization of the loss function on adjacent vertices of the Deep ReLU Simplex algorithm.
result Exponential decay of loss levels and increased vertex density around local minima.