Carbontracker tracks and predicts training DL models' carbon footprint.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Sustainability became the most important component of world development, as countries worldwide fight the battle against the climate change. To understand the effects of climate change, the ecological footprint, along with the biocapacity should be observed. The big part of the ecological footprint, the carbon footprin…
This paper analyzes energy and carbon footprints in distributed and federated learning.
Persistent homology reveals geometric features of metric spaces, especially geodesic circles.
Compact DNNs increase memory footprint and reduce energy efficiency.
A new Input-Output model, called the Multi-Entity Input-Output (MEIO) model, is introduced to estimate the responsibility of entities of an ecosystem on the footprint of each other. It assumed that the ecosystem is comprised of end users, service providers, and utilities. The proposed MEIO modeling approach can be seen…
SDAMI enhances interpretable high-dimensional regression with sparse deep learning and footprint principle.
Study analyzes carbon footprint of 1,417 ML models on Hugging Face.
We propose a statistical model to understand people's perception of their carbon footprint. Driven by the observation that few people think of CO2 impact in absolute terms, we design a system to probe people's perception from simple pairwise comparisons of the relative carbon footprint of their actions. The formulation…
We construct Peano curves whose "footprints" , , have boundaries and are tangent to a common continuous line field on the punctured plane . Moreover, these boundaries can be taken -close to any prescribed smooth family…
Serenity optimizes neural network execution for edge devices by scheduling with optimal memory footprint.
This paper optimizes decarbonized indices for financial tracking, balancing risk and environmental impact.
The Long-Short-Term-Memory Recurrent Neural Networks (LSTM RNNs) are a popular class of machine learning models for analyzing sequential data. Their training on modern GPUs, however, is limited by the GPU memory capacity. Our profiling results of the LSTM RNN-based Neural Machine Translation (NMT) model reveal that fea…
Paper proposes a KGE framework that reduces training time and carbon footprint.
Automates building structural design with reduced mass and carbon footprint.
New method predicts political ideology from online activity.
This paper proposes an alternative to E2E training for deep networks, reducing memory footprint.
SGQuant reduces GNN memory usage without significant accuracy loss.
Despite showing state-of-the-art performance, deep learning for speech recognition remains challenging to deploy in on-device edge scenarios such as mobile and other consumer devices. Recently, there have been greater efforts in the design of small, low-footprint deep neural networks (DNNs) that are more appropriate fo…
Estimates Mozambique's population using remote sensing and microcensus data.
Quantum-inspired tensor network speeds up financial risk assessment.
We propose a max-pooling based loss function for training Long Short-Term Memory (LSTM) networks for small-footprint keyword spotting (KWS), with low CPU, memory, and latency requirements. The max-pooling loss training can be further guided by initializing with a cross-entropy loss trained network. A posterior smoothin…
Neural network models are resource hungry. It is difficult to deploy such deep networks on devices with limited resources, like smart wearables, cellphones, drones, and autonomous vehicles. Low bit quantization such as binary and ternary quantization is a common approach to alleviate this resource requirements. Ternary…
In many learning situations, resources at inference time are significantly more constrained than resources at training time. This paper studies a general paradigm, called Differentiable ARchitecture Compression (DARC), that combines model compression and architecture search to learn models that are resource-efficient a…
Study measures investment funds' climate transition risk, finds moderate losses.
NVIDIA cuDNN is a low-level library that provides GPU kernels frequently used in deep learning. Specifically, cuDNN implements several equivalent convolution algorithms, whose performance and memory footprint may vary considerably, depending on the layer dimensions. When an algorithm is automatically selected by cuDNN,…
GENRE retrieves entities autoregressively, improving efficiency and accuracy.
Localized Multidirectional Correction improves non-refusal target-response behavior in foundation models.
Optimal design portfolios improve energy efficiency and reduce risk in uncertain reservoirs.
Machine learning predicts greenhouse gas emissions for undisclosed companies.
We propose a novel self-attention mechanism that can learn its optimal attention span. This allows us to extend significantly the maximum context size used in Transformer, while maintaining control over their memory footprint and computational time. We show the effectiveness of our approach on the task of character lev…
ActNN reduces neural network training memory by 2-bit quantization.
Paper proposes a new method to select memory data for online class-incremental learning.
Modeling wind dynamics in Saudi Arabia using deep learning and stochastic PDEs.
Recognizing written domain numeric utterances (e.g. I need $1.25.) can be challenging for ASR systems, particularly when numeric sequences are not seen during training. This out-of-vocabulary (OOV) issue is addressed in conventional ASR systems by training part of the model on spoken domain utterances (e.g. I need one …
Bayesian histograms achieve optimal distribution estimation with minimal memory usage.
BiQGEMM efficiently multiplies quantized DNN weights using lookup tables.
We propose NovoGrad, an adaptive stochastic gradient descent method with layer-wise gradient normalization and decoupled weight decay. In our experiments on neural networks for image classification, speech recognition, machine translation, and language modeling, it performs on par or better than well tuned SGD with mom…
Deep neural networks (DNNs) are shown to be promising solutions in many challenging artificial intelligence tasks. However, it is very hard to figure out whether the low precision of a DNN model is an inevitable result, or caused by defects. This paper aims at addressing this challenging problem. We find that the inter…
The wavelet scattering transform is an invariant signal representation suitable for many signal processing and machine learning applications. We present the Kymatio software package, an easy-to-use, high-performance Python implementation of the scattering transform in 1D, 2D, and 3D that is compatible with modern deep …
The study examines label smoothing to improve confidence calibration in fine-tuned LLMs.
Decision-tree-based ensemble classification methods (DTEMs) are a prevalent tool for supervised anomaly detection. However, due to the continued growth of datasets, DTEMs result in increasing drawbacks such as growing memory footprints, longer training times, and slower classification latencies at lower throughput. In …
We introduce graph normalizing flows: a new, reversible graph neural network model for prediction and generation. On supervised tasks, graph normalizing flows perform similarly to message passing neural networks, but at a significantly reduced memory footprint, allowing them to scale to larger graphs. In the unsupervis…
Thanks to the access to labeled orders on the Cac40 index future provided by Euronext, we are able to quantify market participants contributions to the volatility in the diffusive limit. To achieve this result we leverage the branching properties of Hawkes point processes. We find that fast intermediaries (e.g., market…
Improved recommendation systems using multi-layer embeddings reduce model size while maintaining accuracy.
DNN pruning reduces memory footprint and computational work of DNN-based solutions to improve performance and energy-efficiency. An effective pruning scheme should be able to systematically remove connections and/or neurons that are unnecessary or redundant, reducing the DNN size without any loss in accuracy. In this p…
New GP-VAE model improves scalability and performance.
Inducing sparseness while training neural networks has been shown to yield models with a lower memory footprint but similar effectiveness to dense models. However, sparseness is typically induced starting from a dense model, and thus this advantage does not hold during training. We propose techniques to enforce sparsen…