A new loss function penalizes different types of errors in deep learning models.
problem Different types of errors in deep learning models are not equally harmful.
method Introducing the log bilinear loss to differentiate error penalties.
result The log bilinear loss method better contains error within the correct super-class.
Paper reduces sample complexity for bilinear systems identification to nearly constant.
problem Identifying discrete-time bilinear systems under bounded disturbances.
method Uses trajectory-dependent regressors and polynomial mean-square state growth analysis.
result Proves sample complexity of O ~ ( 1 / ε ) \widetilde{\mathcal O}(1/ε) O ( 1/ ε ) for estimation error ε ε ε . New methods improve GANs' conditional generation.
problem Improper conditioning in GANs limits their ability to generate images with specific attributes.
method Introducing two models: an information retrieving model and a spatial bilinear pooling model.
result Significantly enhanced log-likelihood of test data under conditional distributions.
New inequalities generalize Li's theorem on mixed Hodge structures.
problem Generalizing Li's theorem on mixed Hodge structures.
method Develop new Hodge-Riemann bilinear relations in mixed settings.
result New Khovanskii-Teissier type inequalities and log-concavity results.
Paper proposes method to recover quantized data with missing info.
problem Recovering quantized data with missing information.
method Regularized convex cost function with Bi-factorization and Augmented Lagrangian Method.
result The method finds global minimizer of the cost function.
Proposes a new CNN approach for multimodal biometric identification.
problem Improving biometric identification accuracy across multiple modalities.
method Uses a bank of modality-specific CNNs, fuses their outputs, and optimizes the system.
result Significantly outperforms unimodal systems and demonstrates reduction in parameters.
Efficient algorithms for large-scale multiclass classification with linear classifiers.
problem Training ℓ 1 \ell_1 ℓ 1 -regularized linear classifiers with high dimensionality and many classes. method Combines quasi-bilinear objective, stochastic mirror descent, and non-uniform sampling.
result Proposes a sublinear algorithm for multiclass hinge loss.
ECC compresses DNNs for energy-constrained devices like UAVs and smartphones.
problem Energy-constrained deep neural networks in vision applications.
method ECC uses a bilinear regression model to estimate DNN energy consumption and optimizes compression to meet energy constraints.
result ECC achieves higher accuracy under the same or lower energy budget compared to state-of-the-art techniques.
Proves tropical Hodge theory for smooth projective varieties, conditional on Laplacian regularity.
problem Proving log-concavity of characteristic polynomials of matroids.
method Combinatorial approach, conditional proof of Kähler package.
result Conditional proof of Kähler package for tropical cohomology.
New metrics defined on SPD matrices link to divergences and curvature.
problem Defining and characterizing metrics on SPD matrices.
method Developed a principle of deformed metrics and introduced balanced bilinear forms.
result Introduce Mixed-Euclidean metrics with negative sectional curvature.
In a multi-class classification problem, it is standard to model the output of a neural network as a categorical distribution conditioned on the inputs. The output must therefore be positive and sum to one, which is traditionally enforced by a softmax. This probabilistic mapping allows to use the maximum likelihood pri…
Paper analyzes statistical properties of log-cosh loss function.
problem No statistical analysis of log-cosh loss function in literature.
method Presented statistical properties of log-cosh loss function, compared to Cauchy distribution, and examined various statistical procedures.
result Characterized statistical properties of log-cosh loss function, including distribution, likelihood function, and Fisher information.
A new framework improves LSTM performance without adding more parameters.
problem Improving LSTM performance without increasing model complexity.
method A unifying framework of bilinear LSTMs that balances hidden state vector size and weight matrix approximation quality.
result Bilinear LSTMs achieve superior performance compared to linear LSTMs without additional parameters.
New algorithm reduces regret in graphical bilinear bandits.
problem Optimizing decisions in a network of agents playing bilinear games.
method Optimism in the face of uncertainty principle applied to combinatorial NP-hard problem.
result Upper bound of i l d e O ( T ) ilde{O}(\sqrt{T}) i l d e O ( T ) on α α α -regret demonstrated. Prod algorithm improves robustness and efficiency in log-loss prediction.
problem Efficient and robust algorithms for log-loss prediction under expert advice.
method Analysis of Prod algorithm for mixtures of experts with log-loss.
result Prod algorithm provides linear-time bound independent of largest loss and gradient.
New method proves fast regret bounds for online RLHF with generalized preferences.
problem Minimizing max-regret in online RLHF with general preferences and bandit feedback.
method Adopted Generalized Bilinear Preference Model (GBPM) to investigate polylogarithmic regret guarantees.
result Proved polylogarithmic regret bounds for Greedy Sampling and Explore-Then-Commit policies under GBPM.
The study generalizes twistor spinors to Kähler manifolds and finds bilinear form equations.
problem Generalizing twistor spinors to Kähler manifolds.
method Finding differential equations and reducing them to conformal Killing-Yano equations.
result Bilinear forms of Kählerian twistor spinors reduce to Kählerian conformal Killing-Yano equations under certain conditions.
Enhances ASC using time- and frequency-liked CNNs and bilinear pooling.
problem Improving acoustic scene classification accuracy.
method Harmonic and percussive source separation, two-stream CNN architecture, bilinear pooling.
result Improved accuracy on DCASE 2019 sub task 1a dataset.
A new loss function α α α -loss bridges log-loss and 0 0 0 - 1 1 1 loss for binary classification.
problem Improving binary classification performance using a tunable loss function.
method Introducing α α α -loss, proving its margin-based form and classification-calibration, and providing an upper bound on empirical risk. result Empirical and expected risk difference upper bound for logistic regression-based classification.
Paper proposes new loss functions for training energy networks.
problem Challenges in computing gradients for training energy networks.
method Proposes generalized Fenchel-Young losses for efficient gradient computation.
result Demonstrates the calibration of excess risk for linear-concave energies.
Proposes a new loss function for deep neural networks.
problem Deep neural networks lack a direct method to discriminate between correct and competing classes.
method Introduces a discriminative loss function based on negative log likelihood ratio.
result Significantly outperforms cross-entropy loss on image classification tasks.
The paper studies harmonic symmetric bilinear forms on Riemannian manifolds and proves properties of the Bourguignon Laplacian.
problem Analyzing harmonic symmetric bilinear forms on Riemannian manifolds.
method Developed the theory of harmonic symmetric bilinear forms and proved properties of the Bourguignon Laplacian.
result The kernel of the Bourguignon Laplacian is a finite-dimensional vector space of harmonic symmetric bilinear forms on a compact Riemannian manifold.
This note improves on universal portfolios by using bilinear strategies.
problem Improving on the best constant-rebalanced portfolio.
method Generalizing the best bilinear trading strategy using performance-weighted averaging.
result A universal bilinear portfolio that asymptotically dominates the original universal portfolio.
SGD's escape rate depends on log loss barrier, not linear loss barrier.
problem Understanding the escape rate of SGD from local minima.
method Derived a stochastic differential equation (SDE) with additive noise from SGD's multiplicative noise property.
result The log loss barrier determines the escape rate of SGD, not the linear loss barrier.
Algorithm identifies bilinear dynamical systems from noisy data.
problem Learning a realization of a partially observed bilinear dynamical system.
method Regression of outputs to highly correlated covariates for Markov-like parameters.
result High probability error bounds on identification algorithm under uniform stability assumption.
Gaussian-SVGD dynamics converge to Gaussian distributions under certain conditions.
problem Understanding the theoretical properties of SVGD, especially for Gaussian targets.
method Detailed theoretical study of Gaussian-SVGD dynamics, considering both mean-field PDE and discrete particle systems.
result Gaussian-SVGD dynamics converge linearly to the Gaussian distribution closest to the target in KL divergence.
In this paper, we extend Su-Zhang's Cheeger-Mueller type theorem for symmetric bilinear torsions to manifolds with boundary in the case that the Riemannian metric and the non-degenerate symmetric bilinear form are of product structure near the boundary. Our result also extends Bruening-Ma's Cheeger-Mueller type theorem…
Improved diffusion bridge sampling with rKL-LD loss.
problem Improving sampling from unnormalized distributions using diffusion bridges.
method Employing the rKL-LD loss instead of the Log Variance (LV) loss for diffusion bridges.
result rKL-LD consistently outperforms LV loss in diffusion bridges.
Bilinear MLPs offer a new way to interpret deep learning models without complex nonlinearities.
problem Lack of mechanistic understanding in how MLPs compute.
method Introduced bilinear MLPs without element-wise nonlinearities, analyzed their weights using tensor and eigendecomposition.
result Bilinear MLPs provide interpretable weight structures and enable adversarial attacks and overfitting analysis.
Identifies bilinear systems from a single trajectory with optimal sample complexity.
problem Learning bilinear systems from a single trajectory of states and inputs.
method Uses a mild marginal mean-square stability assumption and martingale small-ball condition.
result Sample complexity and statistical error rates are optimal.
Generalizes Riemann's results on flat coordinates for non-symmetric bilinear forms.
problem Finding flat coordinates for non-symmetric bilinear forms.
method Provides explicit necessary and sufficient conditions for a tensor field of type (0,2) to be flat.
result Explicit conditions for a tensor field to have constant entries in local coordinates.
The study explores loss functions for learning distributions, finding the log loss and others are sufficient under certain conditions.
problem Understanding loss functions for distribution learning and density estimation.
method An axiomatic approach to design loss functions, proposing criteria and showing that no single loss function satisfies all criteria.
result No loss function satisfies all criteria, but the log loss and others do under the condition of candidate distributions being calibrated.
MuLFA predicts drug interactions more accurately than existing methods.
problem Improving drug safety by predicting drug interactions.
method Proposes MuLFA, a factorization autoencoder that models nonlinear interactions between drug pairs.
result MuLFA outperforms state-of-the-art methods in predicting drug interactions.
Burq-Gérard-Tzvetkov and Hu established L p L^p L p estimates ( 2 ≤ p ≤ ∞ 2\le p\le \infty 2 ≤ p ≤ ∞ ) for the restriction of eigenfunctions to submanifolds. The estimates are sharp, except for the log loss at the endpoint L 2 L^2 L 2 estimates for submanifolds of codimension 2. It has long been believed that the log loss at the endpoint can be remov…
Separable losses are inconsistent for structured prediction models.
problem Inconsistency of separable losses in structured prediction models.
method Analysis of separable negative log-likelihood losses for structured prediction.
result Separable losses are not Bayes consistent and may not predict the most probable structure.
Invariants from bilinear forms help classify Lefschetz fibrations.
problem Classifying Lefschetz fibrations over the 2-sphere.
method Construct invariants from right G-modules and bilinear functions.
result Found infinitely many Lefschetz fibrations homeomorphic but not mutually isomorphic.
PACMAN provides bounds for classification tasks considering accuracy vs. negative log-loss mismatch.
problem Mismatch between accuracy and negative log-loss in classification tasks.
method Point-wise PAC approach over generalization gap, using likelihood ratio and concentration inequalities.
result PACMAN provides point-wise PAC bounds for the generalization problem.
Proposes a new Huber loss combining absolute and quadratic properties.
problem Improving robustness in learning models.
method Introduces a generalized Huber loss with a log-exp transform and provides an efficient minimization algorithm.
result Shows that the new loss function can be minimized efficiently.
New algorithm reduces prediction errors across various loss functions.
problem Online forecasting algorithms' inability to adapt to different loss functions.
method Design of a novel Follow-the-Perturbed-Leader (FTPL) algorithm with self-concordant noise.
result Simultaneously achieves i l d e O ( T ) ilde O(\sqrt{T}) i l d e O ( T ) regret for bounded proper losses and O ( log T ) O(\log T) O ( log T ) regret for bounded smooth proper losses. Improved COCO algorithms with better constraint control.
problem Achieving small regret and constraint violation in online convex optimization.
method Simple projection-based algorithm leveraging self-contraction geometry.
result Exponential improvement in cumulative constraint violation for strongly convex losses.
Enhances knot invariants using bilinear forms on vector spaces.
problem Improving classical and virtual knot invariants.
method Uses bilinear forms on vector spaces indexed by pairs of elements of a finite quandle.
result New enhanced invariants of knots and links.
Constructs a bilinear form from a quasimorphism on symplectic manifold groups.
problem Understanding symplectic group properties through quasimorphisms and bilinear forms.
method Develops machinery to construct a real-valued bilinear form from a quasimorphism on the commutator subgroup of symplectic group.
result The constructed bilinear form b \mathfrak{b} b controls extendability of quasimorphisms and triviality of characteristic classes. ASVGD accelerates SVGD for efficient sampling from Gaussian targets.
problem Efficient sampling from Gaussian distributions using SVGD.
method Accelerated gradient flow in a metric space of probability densities, including momentum and Wasserstein regularization.
result ASVGD achieves optimal convergence rate for Gaussian targets, independent of covariance.
We demonstrate that, in the classical non-stochastic regret minimization problem with d d d decisions, gains and losses to be respectively maximized or minimized are fundamentally different. Indeed, by considering the additional sparsity assumption (at each stage, at most s s s decisions incur a nonzero outcome), we derive…
The paper improves sparse Gaussian processes by optimizing predictive loss.
problem Optimizing predictive loss in sparse Gaussian processes.
method Direct loss minimization (DLM) for log-loss and square loss, with product sampling (uPS) and biased Monte Carlo (bMC) for non-conjugate cases.
result DLM shows significant performance improvement in both log-loss and square loss cases.
This work tackles Bayesian neural networks by addressing loss landscape symmetries.
problem Understanding and optimizing the loss landscape of Bayesian neural networks.
method The approach involves extending marginalized loss barrier formalism to BNNs, proposing a matching algorithm to search for linearly connected solutions using permutation matrices and combinatorial optimization.
result Nearly zero marginalized loss barriers for linearly connected solutions were found.
Despite being the standard loss function to train multi-class neural networks, the log-softmax has two potential limitations. First, it involves computations that scale linearly with the number of output classes, which can restrict the size of problems we are able to tackle with current hardware. Second, it remains unc…
New method optimizes clustering with better log-likelihood landscape.
problem Nonconvex log-likelihood optimization in model-based clustering.
method Entropic optimal transport loss for Sinkhorn-EM algorithm.
result New loss function avoids spurious local optima.