Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

70140210280 · Jun 202019922001200920182026
48 results for energy layer

Introduces Causal Energy Minimization to understand Transformer layers.

problem Empirical parameterization of Transformer blocks remains largely unexplored.
method Causal Energy Minimization framework that recasts Transformer layers as optimization steps on conditional energy functions.
result Identifies design space for Transformer layers including weight sharing and energy-based interpretations.

Improves energy efficiency of neuromorphic hardware by optimizing memory organization and encoding schemes.

problem Energy inefficiency in neuromorphic hardware, especially in digital accelerators.
method Synthesized controller and memory for different encoding schemes, introduced functional encoding for structured connectivity.
result Functional encoding offers a 58% reduction in energy for weight updates in convolutional layers.

This work relaxes energy constraints in self-attention layers for a more general analysis.

problem Understanding inherent biases and dynamics in self-attention layers without energy functions.
method Dynamical systems analysis and Jacobian matrix examination.
result Normalized dynamics are close to a critical state, indicating high inference performance.

FixyNN splits CNN models into fixed and trainable parts for efficient on-device inference.

problem Energy inefficiency in on-device CNN inference for real-time computer vision.
method Co-designed hardware accelerator platform with transfer learning for training.
result Achieved nearly 2x better energy efficiency than a conventional accelerator.

A hybrid neural network optimizes AI deployment on edge and cloud for energy efficiency.

problem Energy and resource constraints in edge devices for deep learning models.
method Conditionally deep hybrid neural network with quantized layers at edge and full-precision layers at cloud.
result Early classification at the edge reduces energy consumption by 5.5x on CIFAR-10 dataset.

RBM models reveal how hidden unit tail behavior affects pattern reconstruction.

problem Understanding how the tail behavior of hidden units in RBMs influences pattern reconstruction.
method Identified an effective energy function for RBMs and studied its local minima.
result The ability to reconstruct patterns depends on the tail behavior of the hidden unit prior distribution.

We study minimal energy problems for strongly singular Riesz kernels on a manifold. Based on the spatial energy of harmonic double layer potentials, we are motivated to formulate the natural regularization of such problems by switching to Hadamard's partie finie integral operator which defines a strongly elliptic pseud…

2016-02-27abs ↗pdf ↗

This paper introduces a hierarchical associative memory model with multiple layers.

problem Limitations of traditional associative memory models with only one hidden layer.
method Develops a fully recurrent model with arbitrary layers, including locally connected ones, and a corresponding energy function.
result The model can dynamically assemble memories using weights from lower layers and higher layers' rules.

Graph neural networks over-smooth when layers increase, reducing discriminative power.

problem Over-smoothing in graph neural networks reduces model performance as the number of layers increases.
method Analyzed over-smoothing in general graph neural network architecture using Dirichlet energy.
result The Dirichlet energy of embeddings converges to zero, leading to loss of discriminative power.

Study identifies stable configurations of intertwined threads with repulsive interactions.

problem Stable configurations of entangled systems with repulsive interactions.
method Analysis of steepest descent flow of an energy functional.
result Existence and uniqueness of stable configuration of two layers drifting apart at t1/3t^{1/3} rate.

This paper optimizes deep learning systems for high performance and low energy consumption.

problem Achieving ultra-high energy efficiency and performance for deep neural networks.
method Developed an algorithm-hardware co-optimization framework that reduces computational and storage complexity.
result Achieved at least 152X speedup and 71X energy efficiency gain compared to IBM TrueNorth processor.

Deep neural networks undergo hierarchical free-energy landscape transitions with increasing data size.

problem Understanding the design space and dynamics of deep neural networks.
method Statistical mechanical approach based on replica method.
result Hierarchical free-energy landscape transitions with ultrametricity, leading to simpler configurations in deeper layers.

Paper proposes energy-efficient DNN training methods.

problem Energy-constrained deployment of deep neural networks.
method Weighted sparse projection and layer input masking integrated into DNN training.
result Framework provides higher accuracy with same or lower energy budgets.

Energy Transformer integrates attention, energy models, and associative memory.

problem Lack of clear theoretical foundations in attention mechanisms and straightforward design of energy functions in energy-based models.
method Proposes Energy Transformer, a sequence of attention layers with a specifically engineered energy function.
result Obtained strong results on graph anomaly detection and classification tasks.

Study free energy in spherical spin glasses, proving universality dichotomy.

problem Analyzing free energy in spherical spin glass models with different tail exponents.
method Introduced a tail-adapted normalization and used universality dichotomy.
result Sharp universality dichotomy for free energy across different tail exponents.

SchNet models quantum interactions using continuous filters, outperforming traditional methods.

problem Capturing continuous atomic positions in molecules without losing physical information.
method Continuous-filter convolutional neural network architecture in SchNet.
result SchNet models both total energy and interatomic forces with rotationally invariant predictions and a smooth potential energy surface.

D-LinOSS models learn to dissipate energy, improving performance on long-range tasks.

problem Representational limitations of LinOSS models in long-range reasoning.
method Introducing Damped Linear Oscillatory State-Space models (D-LinOSS) that learn to dissipate latent state energy on arbitrary time scales.
result D-LinOSS consistently outperforms previous LinOSS methods on long-range learning tasks, achieving faster convergence and reducing hyperparameter search space.

An unsupervised learning algorithm trains capsule networks for generating realistic images.

problem Training capsule networks for generating realistic images without labeled data.
method Developed an unsupervised learning algorithm using dynamic routing and an energy function for capsule networks.
result The algorithm successfully generates realistic looking images from a learned distribution.

Graph signal processing detects hallucinations in large language models.

problem Detecting factual reasoning from hallucinations in large language models.
method Modeling transformer layers as dynamic graphs, using spectral analysis to define diagnostics.
result Spectral signatures can distinguish different types of hallucinations and achieve high accuracy.

EagerNet detects network attacks quickly with less resources.

problem Efficiently detecting network attacks with minimal resources.
method Proposes a new architecture that trades prediction speed for accuracy, evaluating only a subset of layers.
result Comparable accuracies to simple FCNNs achieved with early predictions, saving energy and computational efforts.

Energy-efficient detection of natural errors in deep networks.

problem Deep networks lack error detection capability without additional energy costs.
method Append RACs at hidden layers to detect natural errors with early classification termination.
result Early classification termination reduces energy consumption.

Meta learns low-rank covariance factors for better uncertainty estimation.

problem Sub-optimal covariance matrices in multi-task settings.
method Meta learns diagonal or diagonal plus low-rank factors using an attentive set encoder.
result Efficiently constructed task-specific covariance matrices improve uncertainty estimation.

Grid-scale batteries' bid patterns in price uncertainty markets

problem Interpreting bids from grid-scale batteries in wholesale electricity markets under price uncertainty
method Developing an asset-level model of a price-taking battery
result Empirical results deliver insights into withholding behavior, uncertainty effects, and risk management reshaping bid curves

Low-complexity spiking networks learn complex tasks with minimal trainable parameters.

problem Training complex reinforcement learning tasks with minimal resources.
method Reinforcement learning on simple networks of spiking neurons with random connections.
result Small random spiking networks achieve learning efficiency similar to humans on complex tasks.

DIET-SNN optimizes SNNs for faster, lower-energy image classification.

problem High inference latency and inefficient input encoding in SNNs.
method End-to-end backpropagation to optimize membrane leak and firing threshold.
result Achieves top-1 accuracy of 69% on ImageNet with 5 timesteps and 12x less compute energy.

Loihi neuromorphic chip outperforms conventional hardware in keyword spotting efficiency.

problem Benchmarking keyword spotting efficiency on neuromorphic hardware.
method Comparative analysis of a two-layer neural network trained to recognize a single phrase on Intel's Loihi neuromorphic chip and conventional hardware devices.
result Loihi outperforms conventional hardware on energy cost per inference for this keyword spotting application.

A method estimates and prunes neural network filters to reduce computation and improve accuracy.

problem Reduction of neural network parameters to save computation and energy.
method Estimates each neuron's contribution to loss using first and second-order Taylor expansions; iteratively removes less important neurons.
result High (>93%) correlation between estimated and true importance; 40% FLOPS reduction with 0.02% top-1 accuracy loss.

Adversarial domain adaptation reduces sample bias in high energy physics classifier.

problem Sample bias in high energy physics classifier training.
method Adversarial domain adaptation using neural networks with gradient reversal layer.
result Successful bias removal on simulated events at the LHC.

Develops hyperparameter transfer methods for Dense Associative Memories.

problem Challenges in transferring hyperparameters for DenseAMs due to unique architecture and activation functions.
method Derives explicit prescriptions for hyperparameter transfer from small to large models.
result Excellent agreement between theoretical and empirical results.