Introduces Causal Energy Minimization to understand Transformer layers.
problem Empirical parameterization of Transformer blocks remains largely unexplored.
method Causal Energy Minimization framework that recasts Transformer layers as optimization steps on conditional energy functions.
result Identifies design space for Transformer layers including weight sharing and energy-based interpretations.
Deep learning depends on tuning layers near critical points.
problem Understanding how deep learning architectures depend on tuning parameters.
method Random energy approach to analyze statistical dependence in deep belief networks.
result Statistical dependence can propagate only if layers are tuned near critical points.
Improves energy efficiency of neuromorphic hardware by optimizing memory organization and encoding schemes.
problem Energy inefficiency in neuromorphic hardware, especially in digital accelerators.
method Synthesized controller and memory for different encoding schemes, introduced functional encoding for structured connectivity.
result Functional encoding offers a 58% reduction in energy for weight updates in convolutional layers.
CERN improves activity recognition in videos with novel energy layer and p-values.
problem Recognizing group activities in videos at semantic levels.
method Two-level LSTM network with Confidence-Energy Recurrent Network (CERN).
result Superior performance compared to state-of-the-art approaches.
Paper optimizes neural network layers to reduce energy usage without sacrificing accuracy.
problem Energy consumption in deep neural networks during inference.
method Layerwise noise maximization to optimize reliability of memory elements.
result Reduces memory energy consumption by 3.3 times at equal accuracy.
This work relaxes energy constraints in self-attention layers for a more general analysis.
problem Understanding inherent biases and dynamics in self-attention layers without energy functions.
method Dynamical systems analysis and Jacobian matrix examination.
result Normalized dynamics are close to a critical state, indicating high inference performance.
New algorithm for signal reconstruction from multi-layered measurements.
problem Reconstructing signals from multi-layered non-linear measurements.
method Multi-Layer Approximate Message Passing (ML-AMP) algorithm and state evolution equations.
result Asymptotic free energy and minimal achievable error derived.
FixyNN splits CNN models into fixed and trainable parts for efficient on-device inference.
problem Energy inefficiency in on-device CNN inference for real-time computer vision.
method Co-designed hardware accelerator platform with transfer learning for training.
result Achieved nearly 2x better energy efficiency than a conventional accelerator.
A hybrid neural network optimizes AI deployment on edge and cloud for energy efficiency.
problem Energy and resource constraints in edge devices for deep learning models.
method Conditionally deep hybrid neural network with quantized layers at edge and full-precision layers at cloud.
result Early classification at the edge reduces energy consumption by 5.5x on CIFAR-10 dataset.
MI-GAN solves OPF with renewable uncertainty using model-informed layers.
problem Optimal Power Flow (OPF) under renewable uncertainty.
method Model-Informed Generative Adversarial Network (MI-GAN) framework with three layers.
result MI-GAN improves solution feasibility and optimality.
NeuralPower predicts and optimizes energy consumption of CNNs.
problem Energy efficiency of CNNs during inference.
method Sparse polynomial regression for layer-wise energy prediction.
result NeuralPower achieves up to 68.5% improvement in prediction accuracy.
RBM models reveal how hidden unit tail behavior affects pattern reconstruction.
problem Understanding how the tail behavior of hidden units in RBMs influences pattern reconstruction.
method Identified an effective energy function for RBMs and studied its local minima.
result The ability to reconstruct patterns depends on the tail behavior of the hidden unit prior distribution.
We study minimal energy problems for strongly singular Riesz kernels on a manifold. Based on the spatial energy of harmonic double layer potentials, we are motivated to formulate the natural regularization of such problems by switching to Hadamard's partie finie integral operator which defines a strongly elliptic pseud…
A new transformer model accelerates training with optimization techniques.
problem Training deep neural networks efficiently and effectively.
method Interprets transformer layers as optimization steps, applying Nesterov acceleration.
result The new model outperforms existing models on benchmark datasets.
This paper examines how energy in feature maps decays in deep neural networks.
problem Understanding energy decay in deep convolutional neural networks.
method Analyzes energy conservation and decay rates in various deep neural network architectures.
result Energy in feature maps decays polynomially or exponentially across layers.
Ternary MobileNets improve efficiency and accuracy on constrained devices.
problem Efficiently compressing MobileNets for real-time applications on constrained devices.
method Per-layer hybrid filter banks for ternary quantization of MobileNets.
result 27.98% energy savings and 51.07% reduction in model size with comparable accuracy.
This paper introduces a hierarchical associative memory model with multiple layers.
problem Limitations of traditional associative memory models with only one hidden layer.
method Develops a fully recurrent model with arbitrary layers, including locally connected ones, and a corresponding energy function.
result The model can dynamically assemble memories using weights from lower layers and higher layers' rules.
New method trains energy-based models faster and more stably.
problem Training efficiency and stability of energy-based models.
method EBFlow with score-matching objectives.
result EBFlow achieves significant speedup and better performance.
Graph neural networks over-smooth when layers increase, reducing discriminative power.
problem Over-smoothing in graph neural networks reduces model performance as the number of layers increases.
method Analyzed over-smoothing in general graph neural network architecture using Dirichlet energy.
result The Dirichlet energy of embeddings converges to zero, leading to loss of discriminative power.
Study identifies stable configurations of intertwined threads with repulsive interactions.
problem Stable configurations of entangled systems with repulsive interactions.
method Analysis of steepest descent flow of an energy functional.
result Existence and uniqueness of stable configuration of two layers drifting apart at t1/3 rate. This paper optimizes deep learning systems for high performance and low energy consumption.
problem Achieving ultra-high energy efficiency and performance for deep neural networks.
method Developed an algorithm-hardware co-optimization framework that reduces computational and storage complexity.
result Achieved at least 152X speedup and 71X energy efficiency gain compared to IBM TrueNorth processor.
Deep neural networks undergo hierarchical free-energy landscape transitions with increasing data size.
problem Understanding the design space and dynamics of deep neural networks.
method Statistical mechanical approach based on replica method.
result Hierarchical free-energy landscape transitions with ultrametricity, leading to simpler configurations in deeper layers.
A new method learns hierarchical EBM models with diffusion schemes.
problem Challenges in learning EBM models with multi-modal distributions.
method Proposes a diffusion probabilistic scheme to learn EBM models in hierarchical latent spaces.
result Demonstrates superior performance on various tasks with diffusion-learned EBM.
New FPI layers enable efficient backpropagation in deep networks.
problem Designing deep neural networks to handle complex constraints.
method Fixed-point iteration layers for forward and backward propagation.
result Backward FPI layer simplifies gradient calculation without explicit Jacobian.
Paper proposes energy-efficient DNN training methods.
problem Energy-constrained deployment of deep neural networks.
method Weighted sparse projection and layer input masking integrated into DNN training.
result Framework provides higher accuracy with same or lower energy budgets.
Energy Transformer integrates attention, energy models, and associative memory.
problem Lack of clear theoretical foundations in attention mechanisms and straightforward design of energy functions in energy-based models.
method Proposes Energy Transformer, a sequence of attention layers with a specifically engineered energy function.
result Obtained strong results on graph anomaly detection and classification tasks.
Study free energy in spherical spin glasses, proving universality dichotomy.
problem Analyzing free energy in spherical spin glass models with different tail exponents.
method Introduced a tail-adapted normalization and used universality dichotomy.
result Sharp universality dichotomy for free energy across different tail exponents.
SchNet models quantum interactions using continuous filters, outperforming traditional methods.
problem Capturing continuous atomic positions in molecules without losing physical information.
method Continuous-filter convolutional neural network architecture in SchNet.
result SchNet models both total energy and interatomic forces with rotationally invariant predictions and a smooth potential energy surface.
Study on feature learning in Leaky ResNets, explaining bottleneck structure.
problem Understanding feature learning in deep neural networks.
method Lagrangian and Hamiltonian reformulation of representation geodesics.
result Emergence of a bottleneck structure in large effective depth networks.
D-LinOSS models learn to dissipate energy, improving performance on long-range tasks.
problem Representational limitations of LinOSS models in long-range reasoning.
method Introducing Damped Linear Oscillatory State-Space models (D-LinOSS) that learn to dissipate latent state energy on arbitrary time scales.
result D-LinOSS consistently outperforms previous LinOSS methods on long-range learning tasks, achieving faster convergence and reducing hyperparameter search space.
AutoQ automatically optimizes quantization for CNNs, reducing latency and energy.
problem Efficiently quantizing CNN weights for low-power mobile devices.
method Hierarchical-DRL for kernel-wise quantization bitwidth selection.
result Reduces inference latency and energy consumption by 54.06% and 50.69% respectively.
An unsupervised learning algorithm trains capsule networks for generating realistic images.
problem Training capsule networks for generating realistic images without labeled data.
method Developed an unsupervised learning algorithm using dynamic routing and an energy function for capsule networks.
result The algorithm successfully generates realistic looking images from a learned distribution.
Graph signal processing detects hallucinations in large language models.
problem Detecting factual reasoning from hallucinations in large language models.
method Modeling transformer layers as dynamic graphs, using spectral analysis to define diagnostics.
result Spectral signatures can distinguish different types of hallucinations and achieve high accuracy.
EagerNet detects network attacks quickly with less resources.
problem Efficiently detecting network attacks with minimal resources.
method Proposes a new architecture that trades prediction speed for accuracy, evaluating only a subset of layers.
result Comparable accuracies to simple FCNNs achieved with early predictions, saving energy and computational efforts.
Energy-efficient detection of natural errors in deep networks.
problem Deep networks lack error detection capability without additional energy costs.
method Append RACs at hidden layers to detect natural errors with early classification termination.
result Early classification termination reduces energy consumption.
New metrics reveal oversmoothing in GNNs more accurately than traditional methods.
problem Oversmoothing in graph neural networks reduces model performance.
method Rank-based metrics to measure oversmoothing in GNNs.
result Rank-based metrics consistently capture oversmoothing, while energy-based metrics often fail.
Paper shows training can improve GCN performance without changing architecture.
problem Training difficulty of GCNs limits their performance.
method Identified and mitigated energy loss during training.
result Significant decrease in training difficulties and notable performance boost.
Meta learns low-rank covariance factors for better uncertainty estimation.
problem Sub-optimal covariance matrices in multi-task settings.
method Meta learns diagonal or diagonal plus low-rank factors using an attentive set encoder.
result Efficiently constructed task-specific covariance matrices improve uncertainty estimation.
E2-Train reduces training energy by 80%+ for state-of-the-art CNNs.
problem Efficient training of energy-hungry CNNs on edge devices.
method Selective layer update, stochastic mini-batch dropping, and sign prediction for low-precision backpropagation.
result Achieves >90% energy savings for training ResNet-74 on CIFAR-10.
Grid-scale batteries' bid patterns in price uncertainty markets
problem Interpreting bids from grid-scale batteries in wholesale electricity markets under price uncertainty
method Developing an asset-level model of a price-taking battery
result Empirical results deliver insights into withholding behavior, uncertainty effects, and risk management reshaping bid curves
Low-complexity spiking networks learn complex tasks with minimal trainable parameters.
problem Training complex reinforcement learning tasks with minimal resources.
method Reinforcement learning on simple networks of spiking neurons with random connections.
result Small random spiking networks achieve learning efficiency similar to humans on complex tasks.
Unified perspective on Hopfield networks with attention module.
problem Understanding and optimizing Hopfield networks with attention mechanisms.
method Study of BM counterparts of modern Hopfield networks and their salient properties.
result Introduction of AttnBM with tractable likelihood and gradient.
DIET-SNN optimizes SNNs for faster, lower-energy image classification.
problem High inference latency and inefficient input encoding in SNNs.
method End-to-end backpropagation to optimize membrane leak and firing threshold.
result Achieves top-1 accuracy of 69% on ImageNet with 5 timesteps and 12x less compute energy.
Loihi neuromorphic chip outperforms conventional hardware in keyword spotting efficiency.
problem Benchmarking keyword spotting efficiency on neuromorphic hardware.
method Comparative analysis of a two-layer neural network trained to recognize a single phrase on Intel's Loihi neuromorphic chip and conventional hardware devices.
result Loihi outperforms conventional hardware on energy cost per inference for this keyword spotting application.
A method estimates and prunes neural network filters to reduce computation and improve accuracy.
problem Reduction of neural network parameters to save computation and energy.
method Estimates each neuron's contribution to loss using first and second-order Taylor expansions; iteratively removes less important neurons.
result High (>93%) correlation between estimated and true importance; 40% FLOPS reduction with 0.02% top-1 accuracy loss.
Adversarial domain adaptation reduces sample bias in high energy physics classifier.
problem Sample bias in high energy physics classifier training.
method Adversarial domain adaptation using neural networks with gradient reversal layer.
result Successful bias removal on simulated events at the LHC.
Develops hyperparameter transfer methods for Dense Associative Memories.
problem Challenges in transferring hyperparameters for DenseAMs due to unique architecture and activation functions.
method Derives explicit prescriptions for hyperparameter transfer from small to large models.
result Excellent agreement between theoretical and empirical results.
Study on buckling of cylindrical shells using elastic energy scaling.
problem Buckling behavior of cylindrical shells under compression.
method Scaling analysis and solution of an obstacle problem for minimal elastic energy.
result Explicit bifurcation point between compression and buckling determined.