Optimized neural networks for Edge TPU achieve high accuracy in real-time image classification.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
LAD-BNet improves real-time energy forecasting on edge devices.
Google's Cloud TPUs are a promising new hardware architecture for machine learning workloads. They have powered many of Google's milestone machine learning achievements in recent years. Google has now made TPUs available for general use on their cloud platform and as of very recently has opened them up further to allow…
Monte Carlo methods are critical to many routines in quantitative finance such as derivatives pricing, hedging and risk metrics. Unfortunately, Monte Carlo methods are very computationally expensive when it comes to running simulations in high-dimensional state spaces where they are still a method of choice in the fina…
Training deep learning models is compute-intensive and there is an industry-wide trend towards hardware specialization to improve performance. To systematically benchmark deep learning platforms, we introduce ParaDnn, a parameterized benchmark suite for deep learning that generates end-to-end models for fully connected…
The study shows how trade uncertainty affects stock-bond correlations over time.
The paper models and analyzes faults in TPU-based neural networks.
BlackJAX simplifies Bayesian inference with modular, fast implementations.
In a recent paper, we have demonstrated how the affinity between TPUs and multi-dimensional financial simulation resulted in fast Monte Carlo simulations that could be setup in a few lines of python Tensorflow code. We also presented a major benefit from writing high performance simulations in an automated differentiat…
Framework for precise recall control in spatial conflation tasks.
We propose an sorting algorithm by Machine Learning method, which shows a huge potential sorting big data. This sorting algorithm can be applied to parallel sorting and is suitable for GPU or TPU acceleration. Furthermore, we discuss the application of this algorithm to sparse hash table.
Efficiently optimizes orthogonal and Stiefel matrices on parallel units.
GShard enables scaling of large neural networks with automatic sharding and lightweight APIs.
New attacks exploit neural network energy and latency, increasing costs by 10-200x.
This paper optimizes deep learning training by efficiently sharding weight updates across replicas.
Training machine learning (ML) models on large datasets requires considerable computing power. To speed up training, it is typical to distribute training across several machines, often with specialized hardware like GPUs or TPUs. Managing a distributed training job is complex and requires dealing with resource contenti…
Deep neural networks (DNN) are increasingly being accelerated on application-specific hardware such as the Google TPU designed especially for deep learning. Timing speculation is a promising approach to further increase the energy efficiency of DNN accelerators. Architectural exploration for timing speculation requires…
NS-RGS improves orthogonal group synchronization with faster convergence.
The impact of the maximally possible batch size (for the better runtime) on performance of graphic processing units (GPU) and tensor processing units (TPU) during training and inference phases is investigated. The numerous runs of the selected deep neural network (DNN) were performed on the standard MNIST and Fashion-M…
RETINA Benchmark evaluates Bayesian deep learning on diabetic retinopathy detection.
A cost-effective method to generate high-resolution images using wavelet-based super-resolution.
Efficiently trains BERT on academic GPUs in 12 days.
AdaNet is a lightweight TensorFlow-based (Abadi et al., 2015) framework for automatically learning high-quality ensembles with minimal expert intervention. Our framework is inspired by the AdaNet algorithm (Cortes et al., 2017) which learns the structure of a neural network as an ensemble of subnetworks. We designed it…
Deep learning is extremely computationally intensive, and hardware vendors have responded by building faster accelerators in large clusters. Training deep learning models at petaFLOPS scale requires overcoming both algorithmic and systems software challenges. In this paper, we discuss three systems-related optimization…
Sparse deep neural networks(DNNs) are efficient in both memory and compute when compared to dense DNNs. But due to irregularity in computation of sparse DNNs, their efficiencies are much lower than that of dense DNNs on regular parallel hardware such as TPU. This inefficiency leads to poor/no performance benefits for s…
We describe TF-Replicator, a framework for distributed machine learning designed for DeepMind researchers and implemented as an abstraction over TensorFlow. TF-Replicator simplifies writing data-parallel and model-parallel research code. The same models can be effortlessly deployed to different cluster architectures (i…
Neural Tangents is a library designed to enable research into infinite-width neural networks. It provides a high-level API for specifying complex and hierarchical neural network architectures. These networks can then be trained and evaluated either at finite-width as usual or in their infinite-width limit. Infinite-wid…
By introducing several improvements to the AlphaZero process and architecture, we greatly accelerate self-play learning in Go, achieving a 50x reduction in computation over comparable methods. Like AlphaZero and replications such as ELF OpenGo and Leela Zero, our bot KataGo only learns from neural-net-guided Monte Carl…
VeLO learns versatile optimizers from deep learning tasks.
Efficient algorithm finds fast Transformer models.
Paper detects anomalous edges in social networks using edge exchangeability.
Can we automatically design a Convolutional Network (ConvNet) with the highest image classification accuracy under the latency constraint of a mobile device? Neural Architecture Search (NAS) for ConvNet design is a challenging problem due to the combinatorially large design space and search time (at least 200 GPU-hours…
We study parallel surfaces and dual surfaces of cuspidal edges. We give concrete forms of principal curvature and principal direction for cuspidal edges. Moreover, we define ridge points for cuspidal edges by using those. We clarify relations between singularities of parallel and dual surfaces and differential geometri…
OL4EL optimizes edge learning on resource-constrained servers.
New GPs model edge functions on complex networks, capturing divergence and curl.
In L^3, cuspidal edges can have bounded mean curvature under specific conditions.
Along cuspidal edge singularities on a given surface in Euclidean 3-space, which can be parametrized by a regular space curve, a unit normal vector field is well-defined as a smooth vector field of the surface. A cuspidal edge singular point is called generic if the osculating plane of the cuspidal edge (as a regul…
Bundling of graph edges (node-to-node connections) is a common technique to enhance visibility of overall trends in the edge structure of a large graph layout, and a large variety of bundling algorithms have been proposed. However, with strong bundling, it becomes hard to identify origins and destinations of individual…
Edge augmentation connects disconnected graphs by elevating eigenvalues.
We prove several results about chordal graphs and weighted chordal graphs by focusing on exposed edges. These are edges that are properly contained in a single maximal complete subgraph. This leads to a characterization of chordal graphs via deletions of a sequence of exposed edges from a complete graph. Most interesti…
Under what conditions is an edge present in a social network at time t likely to decay or persist by some future time t + Delta(t)? Previous research addressing this issue suggests that the network range of the people involved in the edge, the extent to which the edge is embedded in a surrounding structure, and the age…
In the emerging advancement in the branch of autonomous robotics, the ability of a robot to efficiently localize and construct maps of its surrounding is crucial. This paper deals with utilizing thermal-infrared cameras, as opposed to conventional cameras as the primary sensor to capture images of the robot's surroundi…
Study of cuspidal edges on focal surfaces of regular surfaces.
Method certifies edge predictions with cloud-level reliability.
Defense against user shilling attacks in collaborative filtering using edge reweighting.
A hybrid neural network optimizes AI deployment on edge and cloud for energy efficiency.
CoMGNN models heterogeneous graphs with evolving nodes and edges.
Study relates Gaussian curvature signs to cuspidal edge types and geometric invariants.