A new method directly encodes data into latent space using gradient flow.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper considers the problem of implementing large-scale gradient descent algorithms in a distributed computing setting in the presence of {\em straggling} processors. To mitigate the effect of the stragglers, it has been previously proposed to encode the data with an erasure-correcting code and decode at the maste…
Gradient flow autoencoder improves data efficiency over traditional autoencoders.
Enhances GBDT robustness with one-hot encoding and regularization.
Unbiased gradient estimation improves VAE performance.
Conformer encoder reverses sequence in time dimension, affecting decoder training.
New method for scalable set encoding with unbiased gradient approximation.
This paper improves deep learning models' accuracy with differential privacy using gradient encoding and denoising.
Paper analyzes dataset distillation for efficient encoding of task-relevant information.
Improves shared encoder representations for better multi-task learning performance.
Paper formalizes and analyzes a new bound for variational inference.
Transformers learn causal structure through gradient descent on self-attention mechanisms.
Novel PCA method for high-dimensional inverse problems.
NES optimizes discrete structured VAEs effectively without gradient propagation.
PE-SVI reduces SVI inference complexity by finding a suitable start point.
We prove that the evolution of weight vectors in online gradient descent can encode arbitrary polynomial-space computations, even in very simple learning settings. Our results imply that, under weak complexity-theoretic assumptions, it is impossible to reason efficiently about the fine-grained behavior of online gradie…
Federated learning is a distributed learning method to train a shared model by aggregating the locally-computed gradient updates. In federated learning, bandwidth and privacy are two main concerns of gradient updates transmission. This paper proposes an end-to-end encrypted neural network for gradient updates transmiss…
New hashing method improves document retrieval precision.
This paper will explore the use of autoencoders for semantic hashing in the context of Information Retrieval. This paper will summarize how to efficiently train an autoencoder in order to create meaningful and low-dimensional encodings of data. This paper will demonstrate how computing and storing the closest encodings…
Image compression is an essential approach for decreasing the size in bytes of the image without deteriorating the quality of it. Typically, classic algorithms are used but recently deep-learning has been successfully applied. In this work, is presented a deep super-resolution work-flow for image compression that maps …
This research compares two encoding methods for categorical attributes in machine learning, affecting model fairness.
Regularized target encoding beats traditional methods for high cardinality features in ML.
We develop a functional encoder-decoder approach to supervised meta-learning, where labeled data is encoded into an infinite-dimensional functional representation rather than a finite-dimensional one. Furthermore, rather than directly producing the representation, we learn a neural update rule resembling functional gra…
Performance of distributed optimization and learning systems is bottlenecked by "straggler" nodes and slow communication links, which significantly delay computation. We propose a distributed optimization framework where the dataset is "encoded" to have an over-complete representation with built-in redundancy, and the …
We extend contrastive learning theory for multiway classification and prove convergence guarantees.
Stochastic gradient descent (SGD) is a key ingredient in the training of deep neural networks and yet its geometrical significance appears elusive. We study a deterministic model in which the trajectories of our dynamical systems are described via geodesics of a family of metrics arising from the diffusion matrix. Thes…
A new method for target propagation using iterative approximations converges fast and is more biologically plausible.
We provide theoretical and empirical evidence that using tighter evidence lower bounds (ELBOs) can be detrimental to the process of learning an inference network by reducing the signal-to-noise ratio of the gradient estimator. Our results call into question common implicit assumptions that tighter ELBOs are better vari…
This paper proposes an unsupervised learning method to solve heat equations on chips.
We introduce the Mutual Information Machine (MIM), a probabilistic auto-encoder for learning joint distributions over observations and latent variables. MIM reflects three design principles: 1) low divergence, to encourage the encoder and decoder to learn consistent factorizations of the same underlying distribution; 2…
Variational Auto-Encoders (VAEs) have become very popular techniques to perform inference and learning in latent variable models as they allow us to leverage the rich representational power of neural networks to obtain flexible approximations of the posterior of latent variables as well as tight evidence lower bounds (…
Gradient boosting method enforced with individual fairness.
Affine spiking neural networks learn efficiently and generalize well.
Gradient boosted models are a fundamental machine learning technique. Robustness to small perturbations of the input is an important quality measure for machine learning models, but the literature lacks a method to prove the robustness of gradient boosted models. This work introduces VeriGB, a tool for quantifying the …
Quantum networks offer exponential communication savings for large machine learning models.
We present the perceptor gradients algorithm -- a novel approach to learning symbolic representations based on the idea of decomposing an agent's policy into i) a perceptor network extracting symbols from raw observation data and ii) a task encoding program which maps the input symbols to output actions. We show that t…
A graph VAE framework optimizes neural architectures in a continuous space.
Inference models are a key component in scaling variational inference to deep latent variable models, most notably as encoder networks in variational auto-encoders (VAEs). By replacing conventional optimization-based inference with a learned model, inference is amortized over data examples and therefore more computatio…
To address the challenge of backpropagating the gradient through categorical variables, we propose the augment-REINFORCE-swap-merge (ARSM) gradient estimator that is unbiased and has low variance. ARSM first uses variable augmentation, REINFORCE, and Rao-Blackwellization to re-express the gradient as an expectation und…
Gradient descent on LSE objectives implicitly performs EM, leading to collapse without volume control.
In this note we present a generative model of natural images consisting of a deep hierarchy of layers of latent random variables, each of which follows a new type of distribution that we call rectified Gaussian. These rectified Gaussian units allow spike-and-slab type sparsity, while retaining the differentiability nec…
We study the dynamics of the vector field on an open surface given by the gradient of a Green's function. This dynamical approach enables us to show that this field induces an invariant decomposition of the surface as the union of a disk and a 1-skeleton that encodes the topology of the surface. We analyze the structur…
Conventional embedding methods directly associate each symbol with a continuous embedding vector, which is equivalent to applying a linear transformation based on a "one-hot" encoding of the discrete symbols. Despite its simplicity, such approach yields the number of parameters that grows linearly with the vocabulary s…
StructureBoost improves gradient boosting for complex categorical variables efficiently.
New estimator reduces variance in discrete random variables.
New methods improve Fisher Matrix approximations for neural networks at low cost.
A new method learns graph compression from data.
It is a generally shared opinion that significant information about the topology of a bounded domain of a riemannian manifold is encoded into the properties of the distance, , %, , from the boundary of . To confirm such an idea we propose an approach based on the in…