SPIGOT bypasses gradients of argmax functions in neural nets with discrete latent variables.
problem Training neural networks with discrete latent variables.
method Structured projected intermediate gradient optimization technique (SPIGOT).
result SPIGOT bypasses gradients of argmax functions effectively.
New proofs of h-principles in contact 3-manifolds.
problem Characterizing contractible bypasses in contact 3-manifolds.
method Topological characterization and disjoint bypass construction.
result New proofs of h-principles in overtwisted contact 3-manifolds.
Direct proof shows adaptive gradient descent converges near-linearly for convex functions.
problem Proving near-linear convergence of adaptive gradient descent for convex functions.
method Direct Lyapunov-based argument for convex functions with unique minimizer.
result Direct proof of near-linear convergence for convex functions.
We use the generalized Pontryagin-Thom construction to analyze the effect of attaching a bypass on the homotopy class of the contact structure. In particular, given a 3-dimensional contact manifold with convex boundary, we show that the bypass triangle attachment changes the homotopy class of the contact structure rela…
Defensive distillation fails against targeted adversarial attacks, forcing a tradeoff between learning and security.
problem Defensive distillation's limitations in blocking targeted adversarial attacks.
method Systematic exploration of defensive distillation's effectiveness and limitations.
result Defensive distillation is effective against non-targeted attacks but fails against targeted attacks, necessitating a tradeoff between learning and security.
We give a necessary and sufficient condition for the addition of a collection of disjoint bypasses to a convex surface to be universally tight -- namely the nonexistence of a polygonal region which we call a virtual pinwheel.
Gradient-based meta-learning techniques are both widely applicable and proficient at solving challenging few-shot learning and fast adaptation problems. However, they have practical difficulties when operating on high-dimensional parameter spaces in extreme low-data regimes. We show that it is possible to bypass these …
Low-variance gradient estimation is crucial for learning directed graphical models parameterized by neural networks, where the reparameterization trick is widely used for those with continuous variables. While this technique gives low-variance gradient estimates, it has not been directly applicable to discrete variable…
Cheap methods improve uncertainty in SGD solutions.
problem Uncertainty quantification in SGD solutions.
method Two resampling-based methods: parallel resampling with replacement and online resampling.
result Significantly reduced computation effort in constructing confidence intervals.
New method detects adversarial images by exploiting their density.
problem Detecting adversarial images that are hard to find with gradient methods.
method Develop a test based on the density of adversarial directions.
result Unprecedented accuracy in detecting adversarial attacks under white-box setting.
In structured prediction problems where we have indirect supervision of the output, maximum marginal likelihood faces two computational obstacles: non-convexity of the objective and intractability of even a single gradient computation. In this paper, we bypass both obstacles for a class of what we call linear indirectl…
Novel approach simplifies VI problems with faster performance.
problem Black-box VI optimization problems.
method Sample Average Approximation (SAA) combined with quasi-Newton methods and line search.
result Achieves faster performance than existing methods.
Projective DP-SGD reduces privacy error by identifying low-dimensional gradient subspaces.
problem Differentially private SGD's error rate scales with model's dimensionality, problematic for over-parameterized models.
method Projective DP-SGD, projecting noisy gradients to a low-dimensional subspace identified from a public dataset.
result The method reduces the dependence on model dimensionality, improving accuracy in high privacy regimes.
Neural Replicator Dynamics improves deep RL performance in nonstationary environments.
problem Nonstationarity and instability in multiagent reinforcement learning.
method Derive a new algorithm using replicator dynamics to bypass softmax in policy gradient methods.
result Neural Replicator Dynamics (NeuRD) outperforms policy gradient methods in nonstationary environments.
Batch normalization (BN) is a technique to normalize activations in intermediate layers of deep neural networks. Its tendency to improve accuracy and speed up training have established BN as a favorite technique in deep learning. Yet, despite its enormous success, there remains little consensus on the exact reason and …
Proposes KDA to protect deep nets from adversarial attacks.
problem Machine learning system vulnerability to adversarial attacks.
method Key based diversified aggregation with pre-filtering.
result Demonstrates high robustness and universality against various attacks.
We study b-arc foliation change and exchange move of open book foliations which generalize the corresponding operations in braid foliation theory. We also define a bypass move as an analogue of Honda's bypass attachment operation. As applications, we study how open book foliations change under a stabilization of the op…
New implicit Krasulina's k-PCA update avoids QR-decomposition and improves convergence.
problem Online k-PCA problem with orthonormality constraint.
method Derived an implicit form of Krasulina's update that bypasses orthonormality constraint.
result The new update avoids costly QR-decomposition and yields superior convergence.
XGBoost detects unlawful insider trading with high accuracy.
problem Detecting unlawful insider trading from large volumes of transactions.
method Applying eXtreme Gradient Boosting (XGBoost) for identifying and ranking key features.
result XGBoost achieves 97% accuracy in detecting unlawful transactions.
A method for efficient statistical inference from online algorithms.
problem Computational constraints in online algorithms make traditional variance estimation difficult.
method HulC method that wraps around online algorithms to produce valid confidence regions.
result The HulC method produces asymptotically valid confidence regions for online algorithms.
Paper bypasses backdoor detection algorithms in deep learning models.
problem Adversaries can embed backdoors in deep learning models, making them behave differently on specific inputs.
method Adversarial training algorithm that optimizes original loss function and maximizes hidden representation indistinguishability.
result The paper presents an adversarial backdoor embedding algorithm that can bypass existing detection algorithms.
Feature Squeezing is a recently proposed defense method which reduces the search space available to an adversary by coalescing samples that correspond to many different feature vectors in the original space into a single sample. It has been shown that feature squeezing defenses can be combined in a joint detection fram…
Visual spoofing bypasses spam filters and plagiarism detection.
problem Vulnerability in spam filters that can be exploited by visually similar but differently encoded characters.
method Replaces characters with visually similar but differently encoded characters from a different alphabet.
result Spammers can create messages that bypass existing spam filters.
Optimization-based pruning eliminates backpropagation for large language models.
problem Suboptimal pruning performance due to heuristic metrics.
method Optimization of Bernoulli distribution to learn pruning masks without backpropagation.
result Efficient pruning of large language models with improved performance.
Cross-domain collaborative filtering (CF) aims to alleviate data sparsity in single-domain CF by leveraging knowledge transferred from related domains. Many traditional methods focus on enriching compared neighborhood relations in CF directly to address the sparsity problem. In this paper, we propose superhighway const…
Develops a new solver for optimizing with stochastic dominance constraints.
problem Optimizing with stochastic dominance constraints is computationally expensive and impractical.
method Introduces Light Stochastic Dominance Solver (light-SD) that uses Lagrangian properties and surrogate approximation.
result The light-SD solver demonstrates superior performance on various problems.
Develops a new reinforcement learning framework for complex control problems.
problem Continuous-time extended mean field control with deterministic policies.
method Model-free sensitivity formula, deterministic policy gradient, local value and advantage-rate representations.
result Demonstrates efficiency, stability, and robustness in solving complex control problems.
ALIAS uses RL to learn DAGs without acyclicity constraints.
problem Efficiently learning DAGs from observational data without acyclicity constraints.
method ALIAS employs RL to generate DAGs in a single step with optimal complexity, bypassing acyclicity constraints.
result ALIAS outperforms state-of-the-art methods in causal discovery.
Efficient diffusion model for symmetric manifolds reduces training and computation costs.
problem Heat kernel computations for manifold diffusion models are computationally expensive and infeasible.
method Spatially-varying covariance diffusion model, efficient objective derived via Ito's Lemma.
result Our model reduces training time and arithmetic operations by orders of magnitude.
New method bypasses statistical and classifier-based detection of adversarial examples.
problem Vulnerability of deep learning classifiers to adversarial examples.
method Classifier-based adaptation of statistical test method and Logit Mimicry Attack.
result Reduces detection performance to less than 2.2% TPR and 1.6% TPR for statistical test and classifier-based methods, respectively, even at 5% FPR.
A new method reduces both input and output dimensions for better goal-oriented analysis.
problem Simultaneous reduction of input and output dimensions for more accurate analysis.
method Coupled input-output dimension reduction, optimizing gradient-based bounds.
result Determine most informative sensors and influential parameters efficiently.
Sequential learning of tasks using gradient descent leads to an unremitting decline in the accuracy of tasks for which training data is no longer available, termed catastrophic forgetting. Generative models have been explored as a means to approximate the distribution of old tasks and bypass storage of real data. Here …
Systematic and multifactor risk models are revisited via methods which were already successfully developed in signal processing and in automatic control. The results, which bypass the usual criticisms on those risk modeling, are illustrated by several successful computer experiments.
Improved private AdaGrad achieves faster convergence rates for convex functions.
problem Private empirical risk minimization with differential privacy.
method Noisy AdaGrad with knowledge of gradient subspace geometry.
result Faster convergence rates for convex functions, bypassing traditional bounds.
Simplified analysis of SGD for linear regression with weight averaging.
problem Understanding SGD optimization in linear regression models.
method Simplified analysis using linear algebra tools, bypassing complex operator manipulations.
result Recovery of bias and variance bounds for SGD in linear regression.
We consider the problem of minimizing a convex objective function F when one can only evaluate its noisy approximation F^. Unless one assumes some structure on the noise, F^ may be an arbitrary nonconvex function, making the task of minimizing F intractable. To overcome this, prior work has often focu…
New attacks bypass authentication models using mouse dynamics data.
problem Adversarial attacks on behavioural mouse authentication.
method Generative approaches to create adversarial mouse trajectories.
result Imitation-based attacks often outperform surrogate-based attacks.
Paper introduces a PDE-free method for decomposing forces in any dimension.
problem Analyzing non-conservative forces in arbitrary dimensions.
method Geometric decomposition using homotopy operator and Frobenius theorem.
result Decomposes forces into gradient and antiexact components, characterizing curl forces.
New DP training ensures models behave similarly at training and test time.
problem Standard SGD training leads to inconsistent model behavior at training and test time.
method Differentially-Private (DP) training ensures WYSIWYG property through distributional generalization.
result DP training guarantees high-level WYSIWYG property, improving model robustness and privacy.
Discrete Flow Maps bypass sequential prediction limits for parallel text generation.
problem Sequential autoregressive prediction limits large language model speed.
method Flow Maps compress generative trajectories into single-step mappings.
result Discrete Flow Maps surpass previous state-of-the-art results in discrete flow modeling.
ESS improves MCMC efficiency for correlated & multimodal distributions.
problem Slice Sampling's sensitivity to initial length scale and difficulty with correlated distributions.
method Adaptive tuning and parallel walkers for efficient sampling.
result ESS improves efficiency by more than an order of magnitude on correlated distributions.
DDICA separates nonlinear mixed signals robustly.
problem Blind source separation of nonlinear mixed signals.
method Deep deterministic neural network with matrix-based entropy function.
result DDICA effectively separates independent components with high accuracy.
Randomized diversification defends machine learning models against adversarial attacks.
problem Vulnerability of machine learning systems to adversarial attacks.
method Multi-channel architecture with shared secret key for randomized transforms in a gray-box scenario.
result Increased robustness against various adversarial attacks.
In 1989, Y. Eliashberg proved that two overtwisted contact structures on a closed oriented 3-manifold are isotopic if and only if they are homotopic as 2-plane fields. We provide an alternative proof of this theorem using the convex surface theory and bypasses.
Improved differentially private deep learning with group-wise clipping techniques.
problem Efficiency and privacy trade-offs in deep learning models.
method Group-wise clipping techniques (per-layer and per-device) to reduce compute time and memory overhead.
result Private learning with group-wise clipping achieves similar or better performance than non-private learning with less wall time.
MetaCaDI learns causal graphs and unknown interventions from few data instances.
problem Discovering causal mechanisms in systems with high data costs and unknown interventions.
method MetaCaDI is a Bayesian meta-learning framework that optimizes for rapid adaptation to new intervention targets.
result MetaCaDI significantly outperforms state-of-the-art methods in causal graph recovery and intervention target prediction.
EnsLoss combines multiple loss functions to prevent overfitting in classification.
problem Preventing overfitting in classification models.
method EnsLoss is an ensemble method that combines loss functions, ensuring calibration and consistency.
result EnsLoss improves classification accuracy compared to fixed loss methods.
New dynamic backdoor attacks bypass current defenses.
problem Vulnerability of machine learning models to backdoor attacks.
method Random Backdoor, Backdoor Generating Network (BaN), and conditional Backdoor Generating Network (c-BaN).
result Dynamic backdoors can bypass current detection and defense mechanisms.