Two CLEVER extensions improve neural network robustness evaluation.
problem Improving robustness evaluation of neural networks.
method Two extensions of CLEVER: second-order robustness guarantee and BPDA for non-differentiable inputs.
result Demonstrated effectiveness on a 121-layer Densenet model.
This paper tackles the 'Clever Hans' effect in anomaly detection models.
problem The 'Clever Hans' effect undermines the generalization capability of anomaly detection models.
method An explainable AI procedure to highlight relevant features used by anomaly detection models.
result The Clever Hans effect is widespread in anomaly detection and occurs in many forms.
The robustness of neural networks to adversarial examples has received great attention due to security implications. Despite various attack approaches to crafting visually imperceptible adversarial examples, little has been developed towards a comprehensive measure of robustness. In this paper, we provide a theoretical…
A key problem in research on adversarial examples is that vulnerability to adversarial examples is usually measured by running attack algorithms. Because the attack algorithms are not optimal, the attack algorithms are prone to overestimating the size of perturbation needed to fool the target model. In other words, the…
Scientists interact with deep learning models to avoid misleading results.
problem Deep neural networks can misinterpret data and achieve high performance by exploiting confounding factors.
method Introduce explanatory interactive learning (XIL) where scientists revise models based on explanations.
result XIL helps prevent misleading results and encourages model trust.
Unsupervised learning models can produce accurate but misleading predictions.
problem Widespread misleading predictions in unsupervised learning models.
method Developed Explainable AI techniques to detect misleading predictions.
result Widespread Clever Hans effects in unsupervised learning models.
Estimates solutions to Hessian quotient equations on HKT manifolds.
problem Finding C0 estimates for solutions to Hessian quotient equations. method Using the cone condition directly to show C0 estimates. result Showed C0 estimates for solutions to Hessian quotient equations on HKT manifolds. We consider the optimization of active extension portfolios. For this purpose, the optimization problem is rewritten as a stochastic programming model and solved using a clever multi-start local search heuristic, which turns out to provide stable solutions. The heuristic solutions are compared to optimization results o…
Frank-Wolfe algorithms for convex minimization have recently gained considerable attention from the Optimization and Machine Learning communities, as their properties make them a suitable choice in a variety of applications. However, as each iteration requires to optimize a linear model, a clever implementation is cruc…
New proof shows exotic 4-manifolds exist without complex calculations.
problem Existence of exotic 4-manifolds in topology.
method Used Beliakova and Wehrli's s-invariant for links in S3 and Stošić's induction scheme to simplify computations. result Existence of exotic compact, orientable 4-manifolds proven without skein lasagna modules.
SARGAN uses GANs to fill in missing radar frequency bands.
problem Missing spectral information in low frequency radar bands.
method Generative adversarial network (GAN) trained on paired original and missing band signals.
result Initial results show promise in recovering missing spectral information.
Clever sampling methods can be used to improve the handling of big data and increase its usefulness. The subject of this study is remote sensing, specifically airborne laser scanning point clouds representing different classes of ground cover. The aim is to derive a supervised learning model for the classification usin…
Assume that the compact Riemannian spin manifold (Mn,g) admits a G-structure with characteristic connection ∇ and parallel characteristic torsion (∇T=0), and consider the Dirac operator D1/3 corresponding to the torsion T/3. This operator plays an eminent role in the investigation of such man…
Paper proposes detecting video manipulation using stream descriptors.
problem Misuse of manipulated video content.
method Binary classifiers on multimedia stream descriptors.
result Scalable approach can detect high-quality manipulations.
We describe and document three mechanisms by which corporations can influence or even control stock prices. (i) Parent and holding companies wield control over other publicly traded companies. (ii) Through clever management of treasury stock based on buyback programs and stock issuance, stock price fluctuations can be …
Neural networks approximate likelihood ratios for complex models.
problem Difficulty in computing likelihood ratios for modern models.
method Applying the likelihood ratio trick with neural network classifiers.
result Different neural network setups can approximate likelihood ratios with varying performance.
BN^2MF identifies unknown exposure patterns in environmental mixtures.
problem Identifying unknown exposure patterns in environmental mixtures.
method Bayesian non-parametric non-negative matrix factorization (BN^2MF) with non-negative continuous priors and a non-parametric sparse prior.
result Estimates patterns of chemical exposures without specifying the number of patterns.
Paper introduces CLVR to reduce price volatility in AMM exchanges.
problem Intra-block price volatility in AMM exchanges.
method CLVR constructs an ordering to minimize price volatility with low computation cost.
result CLVR minimizes price volatility with a small computation cost and can be externally verified.
New analysis reveals diverse problem-solving behaviors in machine learning models.
problem Understanding and evaluating the diverse problem-solving behaviors of machine learning models.
method Spectral Relevance Analysis to characterize and validate machine learning models.
result Standard performance metrics fail to distinguish diverse problem-solving behaviors.
Unified theory for causal inference using various methods.
problem Estimating causal effects in ATE estimation.
method Riesz regression, covariate balancing, DRE, TMLE, matching estimator.
result Unified theory integrating multiple methods for ATE estimation.
A new algorithm for optimizing huge-scale black-box problems with reduced memory usage.
problem Optimizing huge-scale black-box problems with limited vector operations.
method ZO-BCD algorithm for zeroth-order optimization with reduced memory footprint.
result ZO-BCD achieves state-of-the-art adversarial attack success rate of 97.9%.
CrossNet uses cross-consistency to improve unpaired image translation.
problem Image-to-image translation without paired data.
method Introduced a novel architecture with cross-translators and latent cross-consistency constraints.
result CrossNet outperforms state-of-the-art on various image translation tasks.
BR-SNIS reduces bias in self-normalized IS without increasing variance.
problem Bias in self-normalized IS.
method Iterated sampling-importance resampling (ISIR) to form a bias-reduced estimator.
result Significant reduction in bias without increasing variance.
Paper introduces MultiTargeted testing, improving adversarial testing efficiency.
problem Improving efficiency and effectiveness of adversarial testing methods.
method Introduces MultiTargeted testing, using alternative surrogate losses.
result MultiTargeted outperforms other PGD-based methods, requiring fewer iterations.
IBP simplifies robust training for large networks.
problem Training robust neural networks with scalable methods.
method Interval Bound Propagation (IBP) for robust training.
result IBP enables training large provably robust networks.
This work shows cosine similarity is equivalent to Pearson correlation for word vectors, but not all vectors are suitable for cosine.
problem The use of cosine similarity for semantic textual similarity is often taken for granted, despite its limitations.
method Characterized cases where Pearson correlation is unfit and introduced rank correlation as an alternative.
result Pearson correlation is equivalent to cosine similarity for many word vectors but not all, and rank correlation can improve performance.
The design of neural network architectures is an important component for achieving state-of-the-art performance with machine learning systems across a broad array of tasks. Much work has endeavored to design and build architectures automatically through clever construction of a search space paired with simple learning …
WALNUTS improves sampling efficiency and robustness for multi-scale distributions.
problem Adapting leapfrog step size for multi-scale posterior distributions.
method Adapts leapfrog step size at fixed intervals of simulated time, selecting the largest step size to keep energy error below a threshold.
result Substantial improvements in sampling efficiency and robustness compared to standard NUTS.
Generalizes classifying spaces for topological groups with torsion.
problem Classifying spaces for topological group actions with non-Hausdorff spaces.
method Generalizes Milnor's, Gelfand-Fuks', and Segal's theorems to non-Hausdorff spaces.
result Existence and uniqueness theorems for G-spaces over metric spaces. Study proposes a new method to estimate bias-correction term for ATE estimation.
problem Estimating the bias-correction term for ATE estimation.
method Directly estimating the bias-correction term by minimizing Bregman divergence.
result Automatic covariate balancing property achieved through specific model choices.
This paper considers the multi-task learning problem and in the setting where some relevant features could be shared across few related tasks. Most of the existing methods assume the extent to which the given tasks are related or share a common feature space to be known apriori. In real-world applications however, it i…
This report is concerned with the Mondrian process and its applications in machine learning. The Mondrian process is a guillotine-partition-valued stochastic process that possesses an elegant self-consistency property. The first part of the report uses simple concepts from applied probability to define the Mondrian pro…
GRUwE improves irregular time series prediction with simpler, efficient RNN-based approach.
problem Irregularly sampled multivariate time series prediction challenges.
method Gated Recurrent Unit with Exponential basis functions (GRUwE).
result GRUwE achieves competitive or superior performance compared to recent state-of-the-art methods.
New algorithms improve generalization of Adam and AdamW optimizers.
problem Adam and AdamW optimizers generalize worse than SGD.
method Algorithmic stability analysis and clever momentum-based SGD integration.
result HomeAdam(W) algorithms achieve better generalization and faster convergence.
DAISYnt evaluates synthetic data quality and privacy in regulated domains.
problem Balancing data quality and privacy in regulated domains.
method Developed a suite of advanced tests (DAISYnt) to evaluate synthetic data quality and privacy.
result DAISYnt sets a de facto standard for synthetic data evaluation in regulated domains.
AI-Interpret transforms opaque policies into simple, interpretable decision rules.
problem Designing effective decision aids for professionals to mitigate decision-making biases.
method Combining imitation learning, program induction, and clustering to transform learned policies into interpretable descriptions.
result Providing interpretable decision rules as flowcharts significantly improves people's planning strategies and decisions.
The rise of deep learning in recent years has brought with it increasingly clever optimization methods to deal with complex, non-linear loss functions. These methods are often designed with convex optimization in mind, but have been shown to work well in practice even for the highly non-convex optimization associated w…
The Linear Attention Recurrent Neural Network (LARNN) is a recurrent attention module derived from the Long Short-Term Memory (LSTM) cell and ideas from the consciousness Recurrent Neural Network (RNN). Yes, it LARNNs. The LARNN uses attention on its past cell state values for a limited window size k. The formulas ar…
This paper revisits Differential Galois Theory using Hopf algebras for Lie pseudogroups.
problem Understanding the structure of algebraic Lie pseudogroups using differential algebra and geometry.
method Mixing differential algebra, differential geometry, and algebraic geometry; using Hopf algebras.
result Reveals confusion between prime differential ideals and maximal ideals in Vessiot's work.