A new framework decouples CNN features into intra-class and semantic differences.
problem Learning visual representations in CNNs is challenging.
method Proposes a decoupled learning framework that models intra-class variation and semantic difference independently.
result Decoupled reparameterization leads to significant performance gains and easier convergence.
Probabilistic decoupling separates labels from classes for improved classification.
problem Improving classification accuracy with noisy or partially labeled data.
method Probabilistic decoupling of labels from underlying classes.
result Method enhances performance on various classification tasks, including noisy and partially labeled data.
Decouples RL components for better task transfer and robustness.
problem Poor generalization in reinforcement learning to task domain changes.
method Separates learning into task representation, dynamics, and reward.
result Decoupling improves performance and robustness to domain changes.
Spectral decoupling improves neural network generalization in medical imaging.
problem Poor generalization of neural networks trained on medical imaging data.
method Spectral decoupling, a regularization technique that encourages learning more features.
result Spectral decoupling increases network robustness and performance on external datasets.
Decoupled GCN is shown to be equivalent to label propagation.
problem Improving semi-supervised node classification in graph learning.
method The paper proves the equivalence of decoupled GCN and label propagation, and proposes a new method named PTA.
result Decoupled GCN is equivalent to two-step label propagation and can automatically assign weights to pseudo-labels.
AdaDEM decouples EM into two parts to improve class overlap and uncertainty.
problem Improper EM limits its effectiveness in various machine learning tasks.
method Decouple EM into CADF and GMC, and AdaDEM normalizes CADF reward and uses MEC.
result AdaDEM outperforms classical EM and improves performance in noisy and dynamic environments.
Recent variants improve knowledge distillation performance.
problem Improving the performance of knowledge distillation.
method Introducing additional components or changing the learning process.
result These variants have shown promising results.
dpVAEs improve VAEs by decoupling representation and generation.
problem VAEs struggle with both representation learning and sample generation.
method Introduce decoupled priors (dpVAEs) that separate representation and generation spaces.
result dpVAEs enable regularization without compromising sample generation.
New algorithms for interpreting complex multivariate functions.
problem Hard interpretation of multivariate functions due to many parameters.
method Filtered tensor decompositions of derivative information.
result Nonparametric estimates of smooth decoupled functions.
A new model decouples global and local image representations without supervision.
problem Learning decoupled global and local image representations without supervision.
method Variational auto-encoding framework with invertible generative flow.
result The model effectively learns decoupled representations of images.
This paper tackles efficient learning for factorial marked temporal point processes.
problem Efficient learning for factorial marked temporal point processes.
method Decoupled learning method with two procedures: ADM-M and Fast ISTA, and a reformulated Logistic Regression model.
result Empirical results show the efficiency of the decoupled and reformulated method.
DEKF maintains stability in LSTM learning with bounded perturbations.
problem Stability of DEKF in LSTM-based online learning.
method Modeling DEKF as a perturbed extended Kalman filter and deriving stability conditions.
result DEKF learns LSTM parameters with similar stability properties to the global extended Kalman filter.
Decoupled PFNs improve sequential decision-making by separating epistemic and aleatoric uncertainties.
problem Sequential decision-making requires distinguishing between epistemic uncertainty about latent signals and irreducible aleatoric observation noise.
method Developed a decoupled PFN architecture that uses query-level labels to train separate heads for latent signal and aleatoric noise.
result Empirically, decoupled PFNs mitigate the failure mode of total-variance exploration in noisy and heteroscedastic settings.
We investigate finite-time decoupled convergence in nonlinear two-time-scale stochastic approximation.
problem Achieving decoupled convergence in nonlinear two-time-scale stochastic approximation.
method Nested local linearity assumption, suitable step size selection, convergence analysis of matrix cross term, fourth-order moment convergence rates.
result Finite-time decoupled convergence rates can be achieved in nonlinear two-time-scale stochastic approximation with proper step size selection.
Unified q-learning for mean-field jump-diffusion models with unobservable population distribution.
problem Continuous-time q-learning in mean-field jump-diffusion models with unobservable population distribution.
method Proposed decoupled Iq-function for unified policy evaluation in MFG and MFC problems; unified q-learning algorithm based on test policies and averaged martingale orthogonality condition.
result Unified policy evaluation rule for MFG and MFC problems based on decoupled Iq-function.
This paper investigates the effectiveness of decoupled weight decay at the start of training.
problem The traditional approach to weight decay is not effective throughout training.
method The authors investigate decoupled weight decay, applying it only at the start of training.
result Applying weight decay only at the start of training stabilizes network weights and improves performance.
A new decoupled approach for Gaussian processes reduces complexity and improves performance.
problem Superlinear complexity in sparse variational inference methods for Gaussian processes.
method Orthogonally decoupling the mean and covariance functions of Gaussian processes to achieve linear complexity and expressive posterior mean functions.
result Our method achieves significantly faster convergence compared to state-of-the-art approaches.
Proposes a new backpropagation algorithm for deep learning with guaranteed convergence.
problem Backward locking in backpropagation limits parallel updates in deep neural networks.
method Decouples gradients and splits the network into modules for parallel updates, proving convergence for non-convex problems.
result The proposed algorithm achieves significant speedup without accuracy loss in training deep convolutional neural networks.
Successor Features improve transfer in RL by decoupling feature and reward.
problem Improving feature representation for task transfer in reinforcement learning.
method Decouples feature representation from reward function, allowing domain transfer.
result Advantages and limitations of Successor Features for transfer identified.
Study shows how large neural networks avoid overfitting through decoupling of feature learning and complexity growth.
problem Understanding inductive bias and generalization in large neural networks.
method Dynamical mean field theory applied to large two-layer networks.
result Training dynamics of large networks exhibit a separation of timescales, decoupling feature learning and overfitting.
Hierarchical decoupling improves sample efficiency for complex robots.
problem Learning long-range behaviors on complex robots.
method Two-part policy: low-level imitation and high-level transfer, with KL regularization.
result Hierarchical transfer significantly improves zero-shot high-level transfer and stabilizes learning.
Deep learning has gained great popularity due to its widespread success on many inference problems. We consider the application of deep learning to the sparse linear inverse problem encountered in compressive sensing, where one seeks to recover a sparse signal from a small number of noisy linear measurements. In this p…
A new method for training CNNs that decouples layer updates.
problem Update locking inefficiency in neural network training.
method Decoupled Greedy Learning (DGL) that relaxes joint training objective.
result DGL leads to better generalization than sequential greedy optimization.
Study off-policy evaluation in partially observable environments, reducing bias and errors.
problem Bias and large errors in off-policy evaluation for partially observable environments.
method Defined and solved off-policy evaluation for POMDPs, introduced Decoupled POMDP model.
result Demonstrated and compared off-policy evaluation methods, showing benefits of new approach.
Novel Adam-family method with decoupled weight decay for training neural networks.
problem Training nonsmooth neural networks with weight decay.
method Proposes a novel Adam-family method with decoupled weight decay, establishing convergence properties and demonstrating superior performance.
result Asymptotically approximates SGD and enhances generalization performance.
A new loss function HUG decouples and generalizes neural collapse.
problem Neural collapse limits in deep learning models.
method Hyperspherical uniformity gap (HUG) as a unified framework.
result HUG decouples and generalizes neural collapse, improving model flexibility and robustness.
The paper proposes a decoupled approach to efficiently estimate CoVaR, a measure of systemic financial risk.
problem Estimating CoVaR, a measure of systemic financial risk, is challenging due to zero-probability events and portfolio repricing.
method The paper introduces a decoupled approach using smoothing techniques and a functional perspective to model CoVaR.
result The decoupled estimator achieves a rate of convergence of approximately OmP(Γ−1/2). Paper tackles self-supervised learning for non-homophilous graphs.
problem Existing self-supervised learning methods assume homophilous graphs, but real-world graphs often lack this assumption.
method Develops a decoupled self-supervised learning (DSSL) framework that decouples different semantics between neighborhoods.
result DSSL framework achieves better performance on various graph benchmarks compared to competitive baselines.
Clarifies method of phase synchronization for decoupling linear differential equations.
problem Velocity-dependent transformations in linear second-order differential equations.
method Linear transformation of coordinates and velocities.
result Velocity-dependent transformations do not preserve second-order character and define their own system.
Two graph auto-encoders decouple feature propagation from graph convolution layers.
problem Designing efficient graph auto-encoders with fixed receptive fields.
method L-GAE and L-VGAE using linear matrix computation before auto-encoder input.
result Comparable performance to VGAEs with smaller, simpler networks.
Kahler geometry explains decoupling of Kerr perturbations.
problem Decoupling of curvature scalars in Kerr spacetime.
method Hidden Kahler structure in Kerr spacetime, showing decoupling as a consequence of Kahler geometry.
result Decoupling of Teukolsky equations on Kahler background.
A new method for faster multi-objective optimization by evaluating objectives separately.
problem Finding the Pareto front of trade-offs between multiple objectives efficiently.
method Knowledge Gradient with decoupled evaluations, accounting for different costs.
result The method significantly outperforms existing approaches in terms of evaluation cost.
SG-DNI outperforms standard neural interfaces in speed and accuracy.
problem Inflexible structure of neural networks limits parallelization.
method Introduces synthetic gradients with decoupled neural interfaces (SG-DNI).
result SG-DNI is over 3-fold faster with comparable accuracy.
New method combines spectral and sparse methods for Gaussian processes.
problem Efficiently fitting Gaussian processes to large datasets.
method Orthogonally decoupled variational Fourier features.
result Competitive performance on synthetic and real-world data.
We establish decoupled functional CLTs for two-time-scale stochastic approximation.
problem Understanding the asymptotic behavior of two-time-scale stochastic approximation.
method Martingale problem approach and auxiliary sequence.
result The limiting dynamics of two-time-scale SA are independent of each other.
A new algorithm speeds up deep learning training by decoupling computation and communication.
problem High communication cost limits the speedup of distributed SGD.
method CoCoD-SGD: runs computation and communication in parallel.
result Linear time speedup with respect to hardware resources.
AvaGrad optimizes vision tasks by decoupling learning rate and adaptability.
problem Improving optimization methods for vision tasks.
method Derives AvaGrad, a new optimizer that decouples learning rate and adaptability.
result AvaGrad outperforms SGD on vision tasks when adaptability is properly tuned.
In the artificial intelligence field, learning often corresponds to changing the parameters of a parameterized function. A learning rule is an algorithm or mathematical expression that specifies precisely how the parameters should be changed. When creating an artificial intelligence system, we must make two decisions: …
Novel algorithm for optimal control of nonlinear systems.
problem Optimal control of nonlinear stochastic dynamical systems with unknown dynamics.
method Decoupled data-based approach combining open-loop and closed-loop control.
result Performance of D2C algorithm is approximately optimal and significantly reduces training time.
This work studies how contrastive learning extracts features from unlabeled data.
problem How neural networks trained by contrastive learning can extract features from unlabeled data.
method Formal analysis of contrastive learning's feature learning process, considering two types of features: sparse and dense.
result Contrastive learning using ReLU networks can learn sparse features if proper augmentations are adopted.
Generative models learn useful representations for complex sequential data.
problem Sequence prediction for high-dimensional input sequences.
method Three models based on Generative Stochastic Networks (GSN) for unsupervised sequence learning.
result GSNs provide evidence as a viable framework for complex sequential data.
Stylized facts can be regarded as constraints for any modeling attempt of price dynamics on a financial market, in that an empirically reasonable model has to reproduce these stylized facts at least qualitatively. The dynamics of market prices is modeled on a macro-level as the result of the dynamic coupling of two dyn…
This paper investigates Shampoo's heuristics and decouples preconditioner updates.
problem Improving Shampoo's heuristics for training neural networks.
method Decomposing preconditioner updates, correcting eigenvalues, and adapting eigenbasis computation frequency.
result Principled techniques to remove Shampoo's heuristics and improve training algorithms.
In this paper, we study the generalization properties of online learning based stochastic methods for supervised learning problems where the loss function is dependent on more than one training sample (e.g., metric learning, ranking). We present a generic decoupling technique that enables us to provide Rademacher compl…
New equations simplify gauge-theoretic Khovanov homology solutions.
problem Solving the Haydys-Witten equations for Khovanov homology.
method Introduced decoupled version of Haydys-Witten equations; investigated asymptotic behavior.
result Decoupled equations simplify analysis of full equations on manifolds with ends and boundaries.
This work decouples language from math problems to enable cross-language learning.
problem Current machine learning representations are language dependent.
method Inspired by linguistics, the work learns language agnostic representations.
result Models trained on one language achieve similar accuracies in other languages.
Regularized training of an autoencoder typically results in hidden unit biases that take on large negative values. We show that negative biases are a natural result of using a hidden layer whose responsibility is to both represent the input data and act as a selection mechanism that ensures sparsity of the representati…
DAPS++ improves diffusion-based image restoration by decoupling prior and likelihood.
problem Decoupling prior and likelihood in diffusion-based inverse problems for better performance.
method Introducing DAPS++, which separates diffusion initialization from likelihood refinement.
result DAPS++ achieves high computational efficiency and robust reconstruction performance.