A new neural network model predicts multi-symbol tokens over multiple scales.
problem Language modeling with improved flexibility and performance.
method A learned dictionary of multi-symbol tokens using BPE compression.
result The model outperforms LSTM on language modeling tasks, especially for smaller models.
BPE improves text-to-SQL generation by reducing training time and improving accuracy.
problem Improving text-to-SQL generation accuracy with neural models.
method Adapted Byte-Pair Encoding (BPE) for SQL generation, introduced a novel stopping criterion, and used AST BPE.
result Improved accuracy on 5 out of 6 English text-to-SQL tasks, reduced training time by 50%.
Improved online AED models with multi-stage training and multi-task learning.
problem Enhance performance of online attention-based encoder-decoder models.
method Three-stage training with character encoder, BPE encoder, and attention decoder; multi-task learning at character and BPE levels; transfer learning from bidirectional encoder.
result 35% and 10% relative improvement over baselines for smaller and bigger models, respectively.
Synthetic noise training improves machine translation robustness to spelling mistakes.
problem Making machine translation robust to spelling mistakes and natural noise.
method Training on synthetic noise to improve robustness to natural noise.
result Training on synthetic noise improves robustness to natural noise without diminishing performance on clean text.
Analyzes how BPE tokenisation affects corpus statistics and model entropy in transformer models.
problem Understanding how natural language properties relate to tokenisation schemes in transformer models.
method Analyzes Shannon entropy of corpora under Zipfian distribution, investigates BPE transformations, trains language models, and uses attention diagnostics.
result Transformer models trained on BPE-tokenised corpora increasingly agree with Zipfian predictions as BPE depth increases, indicating reduced local token dependencies.
DLGNet improves dialogue response generation by leveraging transformer architecture.
problem Lack of relevance, diversity, and coherence in dialogue responses.
method Transformer-based model for dialogue response generation, incorporating long-range structures and random paddings.
result Significant improvements over state-of-the-art models on multiple datasets.
Stochastic encoders outperform deterministic ones in 'perfect perceptual quality'.
problem Understanding when stochastic encoders outperform deterministic ones.
method Provided a toy example to illustrate performance.
result Stochastic encoders can significantly outperform deterministic ones in 'perfect perceptual quality'.
BiPE blends intra-segment and inter-segment encodings for better length extrapolation.
problem Improving length extrapolation in language models.
method Bilevel Positional Encoding (BiPE) that separates intra-segment and inter-segment encodings.
result BiPE enhances length extrapolation across various text modalities.
Enhanced Bayesian target encoding uses sampling techniques to improve model performance.
problem Improving target encoding for better model performance in machine learning.
method Using sampling techniques in Bayesian target encoding to extract intra-category distribution information.
result Improves generalization and reduces target leakage in machine learning models.
Wasserstein auto-encoders benefit from random latent spaces.
problem The role of latent space dimensionality in WAEs.
method Experimentation on synthetic and real datasets.
result Random encoders outperform deterministic ones.
Paper proposes LCP for structural encodings, outperforming existing methods.
problem Improving Graph Neural Networks performance through effective structural encodings.
method Geometric perspective, Local Curvature Profiles (LCP) for structural encodings, combining with global positional encodings, comparing with rewiring techniques.
result LCP significantly outperforms existing structural encodings and combining LCP with global positional encodings improves performance.
New definition reveals encoding explanations that retain predictive power.
problem Challenges in evaluating and identifying encoding explanations.
method Developed a definition of encoding based on conditional dependence.
result Existing evaluation scores do not rank non-encoding explanations correctly, but STRIPE-X does.
Study on encoding neural architectures for NAS, showing impact on performance.
problem Impact of different encodings on NAS performance.
method Formal definition and characterization of encodings, empirical study of encodings' effectiveness.
result Different encodings can significantly affect NAS performance.
Improved OOD detection across various shifts using multi-encoder fusion of RDMs.
problem Out-of-distribution detection across multiple types of distribution shifts.
method Statistical identification of encoder sensitivity, EncMin2L fusion, and Tippett minimum combination.
result Achieves AUROC ≥ 0.94 across four shift types, outperforming state-of-the-art detectors.
Single auto-encoder learns cross-domain image translation.
problem Cross-domain image-to-image translation using a single encoder-decoder architecture.
method Single auto-encoder with independent domain and content encodings.
result Cross-domain mapping achieved without separate encoders.
Binary encoding enables neural networks to extrapolate periodic functions.
problem Extrapolating periodic functions without prior knowledge of their form.
method Normalized Base-2 Encoding (NB2E) for continuous numerical values.
result MLPs using NB2E can successfully extrapolate diverse periodic signals.
Similarity encoding improves learning from messy categorical data.
problem Learning from categorical variables with high cardinality and redundancy.
method Similarity encoding, a generalization of one-hot encoding that uses similarities between categories.
result Similarity encoding significantly outperforms traditional encoding methods in prediction accuracy.
Wasserstein Auto-Encoders minimize Wasserstein distance for better data modeling.
problem Building a generative model of data distributions.
method Minimizes a penalized Wasserstein distance between model and target distributions.
result Generates better quality samples compared to other techniques.
Framework learns encodings to protect private attributes from inference.
problem Protect private attributes from inference in image encodings.
method Adversarial training of deep neural networks to inhibit classifier learning.
result Stable optimization approach yields encoders resistant to privacy inference.
This research compares two encoding methods for categorical attributes in machine learning, affecting model fairness.
problem The impact of encoding protected categorical attributes on fairness in machine learning models.
method Comparison of one-hot encoding and target encoding methods.
result Target encoding can lead to more unfair models compared to one-hot encoding due to induced bias.
Proposes a method to encode dynamical systems for efficient data compression.
problem Efficiently encoding sequences generated by dynamical systems.
method Introduces a local information criterion for model selection in dynamical systems.
result The proposed method can select the best model closer to the true dynamics.
Quasi-orthonormal encoding reduces high dimensionality for categorical data.
problem High dimensionality and low sample size issues in categorical data encoding.
method Quasi-orthonormal encoding (QOE) for categorical data.
result QOE reduces dimensionality and improves machine learning performance.
The paper identifies saddlepoints in unsupervised auto-encoding neural nets.
problem The risk landscape of unsupervised least squares in auto-encoding neural nets.
method Established an equivalence between unsupervised least squares and principal manifolds, discussed regularization strategies for auto-encoders.
result All non-trivial critical points in auto-encoding are saddlepoints, which are degenerate in overcomplete auto-encoding.
Regularized target encoding beats traditional methods for high cardinality features in ML.
problem Efficiently encoding high cardinality categorical variables for ML algorithms.
method Regularized target encoding compared to traditional encodings like integer and one-hot encoding.
result Regularized target encoding consistently provided the best results in a large-scale benchmark experiment.
Graph attention auto-encoder reconstructs graph structure and attributes.
problem Lack of methods to reconstruct graph structure and node attributes in graph auto-encoders.
method Stacked encoder/decoder layers with self-attention mechanisms, regularized node representations to reconstruct graph structure.
result Competitive performance on node classification benchmarks, including inductive learning.
CRsAE auto-encoder recovers convolutional dictionary from noisy signals.
problem Recovering a convolutional dictionary from noisy signals.
method Constrained recurrent sparse auto-encoder (CRsAE) architecture.
result CRsAE successfully recovers the underlying dictionary in the presence of noise.
A new encoding framework predicts brain activity from visual stimuli and intrinsic brain connections.
problem Traditional encoding models ignore brain inner states, limiting their performance in natural image identification.
method Proposes a novel encoding framework combining external stimuli and brain inner states, using a forward encoding model and an inner state model.
result The framework achieves better performance on natural image identification from fMRI responses than traditional models.
Study efficient neural operator learning using variation spaces.
problem Operator learning using encoder-decoder neural networks.
method Introduce variation space for nonlinear operators, establish approximation bounds.
result Algebraic approximation and learning rates for polynomially decaying input and output encoding errors.
PointPillars improves object detection speed and accuracy in point clouds.
problem Encoding point clouds for efficient object detection.
method PointPillars uses PointNets to learn pillar representations of point clouds, combined with a lean downstream network.
result PointPillars outperforms previous encoders in both speed and accuracy.
New research shows encoder-decoder GANs can still fail even on real data.
problem Theoretical limitations of Encoder-Decoder GAN architectures.
method Rigorous analysis of Encoder-Decoder GAN training objectives.
result The training objectives cannot prevent mode collapse or learning meaningless codes.
New methods encode high-cardinality string variables efficiently.
problem Efficient encoding of high-cardinality string categorical variables.
method Two approaches: Gamma-Poisson matrix factorization and min-hash encoder.
result Improves supervised learning with high-cardinality categorical variables.
MUTE improves neural network performance with efficient target encoding.
problem Improving neural network performance with limited resources.
method MUTE optimizes Hamming distances among target encoding by understanding class confusion.
result MUTE offers better generalization and robustness with minimal overhead.
DRN compactly encodes functions for distribution regression.
problem Challenges in encoding functions compactly in neural networks.
method Designs a compact network representation to encode and propagate functions in single nodes.
result Achieves higher prediction accuracies with fewer parameters.
A new model SEQ clusters and classifies encoded features for better interpretability.
problem Lack of interpretability in classical supervised classification tasks.
method Proposes a novel supervised learning model named Supervised-Encoding Quantizer (SEQ) that applies a quantizer to cluster and classify encoded features.
result The quantizer provides an interpretable graph where each cluster represents a class with a particular style.
New method enforces encoder sparsity in HPF for more interpretable feature selection.
problem Lack of encoder sparsity in HPF leads to lack of column-clustering property.
method Enforces encoder sparsity using a generalized additive model (GAM).
result Gains ability to perform feature selection and relates each representation to original features.
The paper shows homeomorphic encoders are impossible for non-trivial manifolds.
problem Designing a homeomorphic encoder for non-trivial manifolds.
method Topological arguments and constraints derived for manifold-specific encoders.
result Homeomorphic encoders are impractical for non-trivial manifolds.
DGA and DVGA learn disentangled graph representations to improve graph analysis.
problem Holistic graph auto-encoders fail to capture latent factors effectively.
method Design disentangled graph convolutional network and component-wise flow, impose independence constraints.
result Improved disentangled graph representations enhance graph analysis tasks.
A simple encoder and complex decoder for secure image encryption and decryption.
problem Secure and efficient image encryption and decryption.
method Uses a shallow encoder neural network for encryption and a deep decoder for decryption, trained independently.
result Decrypted images are nearly identical to the original, demonstrating the effectiveness of the framework.
STRING improves 2D and 3D position encodings for better performance.
problem Efficient and accurate position encoding for 2D and 3D applications.
method STRING extends Rotary Position Encodings with a unifying theoretical framework, maintaining translation invariance and low computational cost.
result STRING shows substantial gains in open-vocabulary object detection and robotics.
A novel feedbackward approach for semantic segmentation reduces parameter count.
problem Semantic segmentation challenges in terms of model complexity and performance.
method Reverse-direction encoder-decoder architecture using the same network for both encoding and decoding.
result Significantly outperforms other models using only 13 VGG-16 layers, achieving higher IoU scores.
Neural networks encode inputs deterministically and categorically, behaving like hash encoders.
problem Understanding the encoding properties of neural networks.
method Analyzed the input space partitioned by ReLU-like activations in neural networks.
result Neural networks can be represented by unique activation patterns, similar to hash encoders.
Conformer encoder reverses sequence in time dimension, affecting decoder training.
problem Reversal of sequence in Conformer encoder impacts decoder training.
method Analyzed initial behavior of decoder cross-attention and proposed methods to avoid flipping.
result Self-attention module of Conformer starts dominating, allowing only reversed information to pass.
Two graph auto-encoders decouple feature propagation from graph convolution layers.
problem Designing efficient graph auto-encoders with fixed receptive fields.
method L-GAE and L-VGAE using linear matrix computation before auto-encoder input.
result Comparable performance to VGAEs with smaller, simpler networks.
New insights into how encoder-decoder networks generate attention matrices.
problem Understanding how encoder-decoder networks use attention matrices.
method Decomposing hidden states into temporal and input-driven components.
result Attention matrices are formed based on task requirements, not architecture type.
Randomized positional encodings boost transformer performance on longer sequences.
problem Transformers struggle with generalizing to sequences of arbitrary length.
method Introduced randomized positional encodings that simulate longer sequences and randomly select positions.
result Randomized positional encodings increase test accuracy by 12.0% on average for sequences of unseen length.
Causal terminology is often introduced in the interpretation of encoding and decoding models trained on neuroimaging data. In this article, we investigate which causal statements are warranted and which ones are not supported by empirical evidence. We argue that the distinction between encoding and decoding models is n…
Efficient method for vertex embedding and community detection.
problem Vertex embedding and community detection.
method Normalized one-hot graph encoder and rank-based cluster size measure.
result Excellent numerical performance of graph encoder ensemble algorithm.
A new method uses persistent homology to assess auto-encoders' latent manifold quality.
problem Chaos in auto-encoders' latent manifold and failure of current distance measures.
method Persistent Homology for Wasserstein Auto-Encoders (PHom-WAE).
result PHom-WAE improves auto-encoders' performance in credit card transaction data.