Gradient ascent method successfully removes specific data points from neural networks without retraining.
problem Addressing privacy and ethical concerns by removing specific data points from trained models.
method Gradient ascent approach to unlearning, leveraging the implicit bias of gradient descent towards margin maximization conditions.
result Gradient ascent method can successfully unlearn specific data points from two-layer ReLU neural networks without retraining.
The Joint Optimization of Fidelity and Commensurability (JOFC) manifold matching methodology embeds an omnibus dissimilarity matrix consisting of multiple dissimilarities on the same set of objects. One approach to this embedding optimizes the preservation of fidelity to each individual dissimilarity matrix together wi…
In the information-based paradigm of inference, model selection is performed by selecting the candidate model with the best estimated predictive performance. The success of this approach depends on the accuracy of the estimate of the predictive complexity. In the large-sample-size limit of a regular model, the predicti…
The paper proposes a method to reliably select design algorithms for machine learning-guided design tasks.
problem Choosing the right design algorithm for machine learning-guided design tasks.
method Combining designs' predicted property values with held-out labeled data to reliably forecast characteristics of the label distributions produced by different design algorithms.
result The method is guaranteed to return design algorithms that yield successful label distributions.
Proposes a new criterion for selecting Nash equilibria considering both utility and inequality.
problem Finding a fair Nash equilibrium in group decision-making.
method Introduces entropy-norm space for geometric selection of strict Nash equilibria.
result The closest entropy-norm pair to the largest entropy-norm pair in rescaled space is the most suitable equilibrium.
Paper introduces a flow-based framework for representation learning.
problem Capturing fine-grained structural details in complex data distributions.
method Zero-flow criterion for conditional independence and a tractable loss function.
result The zero-flow criterion enables learning of sufficient information from data.
Maximum mean discrepancy (MMD), also called energy distance or N-distance in statistics and Hilbert-Schmidt independence criterion (HSIC), specifically distance covariance in statistics, are among the most popular and successful approaches to quantify the difference and independence of random variables, respectively. T…
The success of convolutional neural networks (CNNs) in various applications is accompanied by a significant increase in computation and parameter storage costs. Recent efforts to reduce these overheads involve pruning and compressing the weights of various layers while at the same time aiming to not sacrifice performan…
We test three common information criteria (IC) for selecting the order of a Hawkes process with an intensity kernel that can be expressed as a mixture of exponential terms. These processes find application in high-frequency financial data modelling. The information criteria are Akaike's information criterion (AIC), the…
Risk and uncertainty will always be a matter of experience, luck, skills, and modelling. Leverage is another concept, which is critical for the investor decisions and results. Adaptive skills and quantitative probabilistic methods need to be used in successful management of risk, uncertainty and leverage. The author ex…
Two new criteria help understand the advantage of deep neural networks.
problem Understanding the advantage of deepening neural networks.
method Proposed two new criteria to evaluate the expressivity of functions computable by deep neural networks.
result Increasing layers is more effective than increasing units in improving the expressivity of deep neural networks.
If X is a compact set, a {\it topological contraction} is a self-embedding f such that the intersection of the successive images fk(X), k>0, consists of one point. In dimension 3, we prove that there are smooth topological contractions of the handlebodies of genus ≥2 whose image is essential. Our proof i…
This paper investigates Shampoo's heuristics and decouples preconditioner updates.
problem Improving Shampoo's heuristics for training neural networks.
method Decomposing preconditioner updates, correcting eigenvalues, and adapting eigenbasis computation frequency.
result Principled techniques to remove Shampoo's heuristics and improve training algorithms.
Optimizes non-linear outcomes from summed contributions.
problem Maximizing a non-linear function of summed small contributions.
method Derives a scalable descent algorithm leveraging concentration properties.
result Directly optimizes for stated objective, e.g., A/B test success criterion.
Penalized regression models are popularly used in high-dimensional data analysis to conduct variable selection and model fitting simultaneously. Whereas success has been widely reported in literature, their performances largely depend on the tuning parameters that balance the trade-off between model fitting and model s…
Paper studies M-estimators with derivatives and residual distribution for robust adaptive tuning.
problem Tackles robustness and adaptive tuning of M-estimators with heavy-tailed noise.
method Provides formulae for derivatives, characterizes residual distribution, proposes adaptive criterion.
result Characterizes distribution of residuals and proposes adaptive criterion as out-of-sample error proxy.
DE-QT detects optimal Q-learning stopping points.
problem Information loss in Q-learning during prolonged training.
method Introducing DE-QT to detect entropy changes in Q-tables.
result DE-QT identifies the best stopping point for Q-learning.
Neural Autoregressive Distribution Estimators (NADEs) have recently been shown as successful alternatives for modeling high dimensional multimodal distributions. One issue associated with NADEs is that they rely on a particular order of factorization for P(x). This issue has been recently addressed by a vari…
New method speeds up HSIC for multiple variables.
problem Quadratic computational complexity of HSIC for multiple variables.
method Nyström approximation to HSIC for M≥2. result Consistent Nyström HSIC estimator for M≥2. New convergence rates for shuffling gradient methods without strong convexity.
problem Theoretical gap between shuffling gradient methods' empirical success and established convergence rates.
method Proved last-iterate convergence rates for shuffling gradient methods using function value gap.
result First last-iterate convergence rates for shuffling gradient methods without strong convexity.
Active feature selection uses mutual information to choose fewer labels for better feature selection.
problem Selecting features with limited labeled data.
method Uses active feature selection with mutual information criterion, optimizing label selection for better feature quality.
result Algorithm selects features with higher mutual information using fewer labels than the data set size.
Colorectal cancer (CRC) is the third most common cancer and the second leading cause of cancer-related deaths worldwide. Most CRC deaths are the result of progression of metastases. The assessment of metastases is done using the RECIST criterion, which is time consuming and subjective, as clinicians need to manually me…
Entropy Search (ES) and Predictive Entropy Search (PES) are popular and empirically successful Bayesian Optimization techniques. Both rely on a compelling information-theoretic motivation, and maximize the information gained about the argmax of the unknown function; yet, both are plagued by the expensive computatio…
Study disproves a generalized numerical criterion for certain pairs.
problem Generalized numerical criterion for pairs
method Provided counterexamples
result Negative answer to the generalized numerical criterion problem
New polynomial criterion for periodic knots identified.
problem Identifying periodic knots efficiently.
method Examined HOMFLY-PT and Kauffman polynomials of periodic links.
result Criterion is stronger than existing methods.
New algorithms optimize a soft-robust criterion in reinforcement learning, reducing conservatism.
problem Computing robust policies for high-stakes decisions with limited data.
method Soft-robust criterion using risk measures, two algorithms for optimization.
result Our algorithms produce less conservative solutions than existing methods.
We introduce a new criterion to determine the order of an autoregressive model fitted to time series data. It has the benefits of the two well-known model selection techniques, the Akaike information criterion and the Bayesian information criterion. When the data is generated from a finite order autoregression, the Bay…
Study how actions affect perception in embodied agents using group theory.
problem Understanding how actions influence perception in autonomous agents.
method Mathematical formalism of group theory applied to sensory commutativity of action sequences.
result Introduced Sensory Commutativity Probability (SCP) to measure action effects on perception.
Modified Bakry-Émery criterion inequality for Tsallis entropy monotonicity.
problem Establishing improved logarithmic Sobolev inequalities and monotonicity of Tsallis entropy.
method Proving a one-parameter family of weighted Bakry-Émery Γ2 criterion inequalities and a modified inequality. result Yields a family of sharp Sobolev inequalities and monotonicity of Tsallis entropy.
A widely applicable Bayesian information criterion (Watanabe, 2013) is applicable for both regular and singular models in the model selection problem. This criterion tends to overestimate the log marginal likelihood. We identify an overestimating term of a widely applicable Bayesian information criterion. Adjustment of…
Ensemble approaches for uncertainty estimation have recently been applied to the tasks of misclassification detection, out-of-distribution input detection and adversarial attack detection. Prior Networks have been proposed as an approach to efficiently \emph{emulate} an ensemble of models for classification by paramete…
Adversarial learning is one of the most successful approaches to modelling high-dimensional probability distributions from data. The quantum computing community has recently begun to generalize this idea and to look for potential applications. In this work, we derive an adversarial algorithm for the problem of approxim…
GraphLIME explains GNN models by selecting key features locally.
problem Explaining the effectiveness of GNN models is challenging due to complex nonlinear transformations.
method GraphLIME uses HSIC Lasso for nonlinear feature selection in GNN models.
result GraphLIME provides more descriptive explanations than existing methods.
New criterion improves predictive evaluation in weighted inference scenarios.
problem Improving predictive evaluation in scenarios with different likelihoods for estimation and evaluation.
method Developed the posterior covariance information criterion (PCIC) to handle weighted likelihood inference.
result PCIC is asymptotically unbiased for quasi-Bayesian generalization error in weighted inference.
Criterion for stopping conjugacy class enumeration in triangle groups.
problem Enumerating all conjugacy classes in cocompact triangle groups.
method Encoding by P. Dehornoy and T. Pinsky; stopping criterion based on geometric length.
result Stopping criterion for the generation of conjugacy classes in cocompact triangle groups.
In [D.A. Fedoseev, V.O. Manturov, A sliceness criterion for odd free knots,arXiv:1707.04923], the authors proved a sliceness criterion for odd free knots: free knots with odd chords. In the present paper we give a similar criterion for stably odd free knots. Some additional results on knot sliceness and cobordism are g…
New criterion for solving inverse Hessian equations, including J-equation.
problem Existence of solutions to inverse Hessian equations, including J-equation.
method Stability of pairs in the sense of Paul, formulated in terms of GIT criterion.
result New numerical criterion for existence of solutions to inverse Hessian equations.
Criterion for nilpotent Lie groups to have nilsolitons.
problem Existence of nilsolitons in nilpotent Lie groups.
method Algebraic criterion for nilpotent Lie algebras, proving necessary and sufficient condition for nilsolitons.
result Criterion provides a necessary and sufficient condition for nilpotent Lie groups to admit nilsolitons.
Study proposes a stopping criterion for active learning based on error stability.
problem Improving predictive performance in active learning by adaptively annotating samples.
method Proposes a stopping criterion based on error stability for Bayesian active learning.
result Demonstrates the proposed criterion stops active learning at the appropriate timing for various models and datasets.
Clarifies boundary criterion for non-one-ended subgroups in cubulation theory.
problem Boundary criterion for relative cubulation in non-one-ended subgroups.
method Showed that if boundary criterion is satisfied for a relatively hyperbolic group, the group admits a relatively geometric action on a CAT(0) cube complex.
result The refinement of the boundary criterion is useful for constructing new relative cubulations.
Extends Kelly Criterion to more complex betting scenarios.
problem Maximizing long-term growth in complex betting models.
method Generalizes Kelly Criterion to Lévy processes and high-frequency limits.
result Improved strategies for high-frequency betting.
Derives criteria for Kähler structures on holomorphic submersions.
problem Criteria for Kähler structures on holomorphic submersions.
method Derives a criterion for Kähler structures using holomorphic submersions.
result Proves Kähler structures for certain holomorphic submersions.
Partial answer to affineness of entire Grauert tubes, with Stein manifold criterion.
problem Affineness of entire Grauert tubes
method Generalized Demailly's criterion for Stein manifolds
result Complement of a codimension-one subset is affine
Optimizes recommendation models using skew normal distribution.
problem Improving personalized recommendation systems.
method Develops a new optimization criterion based on skew normal distribution.
result Significantly outperforms state-of-the-art models.
Criterion for solvability of complex 2-Hessian equation on compact Kähler manifolds.
problem Solvability of complex 2-Hessian equation on compact Kähler manifolds.
method Nakai--Moishezon-type criterion associated with the complex 2-Hessian equation.
result Criterion equivalent to existence of a smooth 2-admissible representative in complex dimension three.
A new algorithm removes stale observations in dynamic Bayesian optimization.
problem Optimizing functions that change over time, keeping track of the optimum.
method Wasserstein distance-based criterion to quantify relevancy, removing stale observations.
result W-DBO maintains good predictive performance and high sampling frequency.
Paper explains why small-loss criterion works for learning from noisy labels.
problem Learning from noisy labels in deep learning with limited labeled data.
method Theoretical analysis and reformulation of the small-loss criterion.
result Theoretical explanation and reformulation of the small-loss criterion.
New criterion assesses cluster separability for validation.
problem Validating cluster analysis results and determining the number of clusters.
method Distinguishability criterion, combined loss function-based framework.
result Validated cluster configurations and determined the number of clusters.