The paper introduces models to learn generalized transformation equivariant representations.
problem Capturing intrinsic visual structures equivariant to various transformations.
method Deterministic and probabilistic AutoEncoding Transformations (AET and AVT) models trained to learn visual representations from generic groups of transformations.
result Generalized TERs (GTERs) that are equivariant to transformations in a more general fashion.
New Feedback Transformer architecture improves model performance by exposing past representations to future.
problem Limitations of Transformers in fully exploiting sequential input.
method Proposes Feedback Transformer exposing all past representations to future.
result Demonstrates improved performance with smaller, shallower models.
Unified framework for learning function representations using INRs and Transformers.
problem Scalability and efficiency limitations in existing generative models.
method Integrates INRs and Transformer-based hypernetworks into latent variable models.
result Improved scalability, expressiveness, and generalization over existing models.
Transformer learns graph structure better with subgraph info.
problem Transformer struggles with structural similarity in graph learning.
method Structure-Aware Transformer with subgraph attention.
result Improves graph prediction benchmarks significantly.
The Weyl transform is introduced as a rich framework for data representation. Transform coefficients are connected to the Walsh-Hadamard transform of multiscale autocorrelations, and different forms of dyadic periodicity in a signal are shown to appear as different features in its Weyl coefficients. The Weyl transform …
New definition of disentangled representations using symmetry transformations.
problem Lack of a generally agreed-upon definition of disentangled representations.
method Focus on transformation properties of the world using symmetry transformations and group representation theory.
result First formal definition of disentangled representations.
Relational data representations have become an increasingly important topic due to the recent proliferation of network datasets (e.g., social, biological, information networks) and a corresponding increase in the application of statistical relational learning (SRL) algorithms to these domains. In this article, we exami…
Unified description of Weierstrass-type representations.
problem Creating surfaces with special curvature properties.
method Unified description using classical transformation theory of Ω-surfaces.
result Unified understanding of Weierstrass-type representations.
PatchGT uses non-trainable graph patches to improve graph representation learning.
problem Learning high-level information in graph tasks with direct Transformer models.
method PatchGT segments graphs into non-trainable patches, uses GNN for patch-level learning, and Transformer for graph-level learning.
result PatchGT achieves higher expressiveness and competitive performance on benchmark datasets.
Transformative machine learning improves model accuracy and explainability with limited data.
problem Improving model accuracy and interpretability with limited data in scientific tasks.
method Transforming intrinsic data representations to extrinsic ones based on model predictions.
result Transformative machine learning significantly outperforms intrinsic representations in drug-design, gene expression prediction, and meta-learning.
We present a signal representation framework called the sparse manifold transform that combines key ideas from sparse coding, manifold learning, and slow feature analysis. It turns non-linear transformations in the primary sensory signal space into linear interpolations in a representational embedding space while maint…
Improves contrastive learning invariance with novel training objectives and feature averaging.
problem Contrastive learning's implicit invariance is insufficient for robust performance.
method Introduces a novel training objective and feature averaging approach to enforce invariance.
result Improved performance and robustness to transformations on downstream tasks.
UGformer uses transformers to learn graph representations.
problem Graph representation learning for various tasks.
method UGformer is a transformer-based GNN model that samples or considers all neighbors for each node.
result UGformer achieves state-of-the-art accuracy on graph classification and text classification tasks.
Transformer learns representations from time series data for money laundering detection.
problem Detecting money laundering using structured time series data.
method Contrastive learning for representation learning, followed by scoring and thresholding.
result Transformer outperforms rule-based and LSTM methods in detecting money laundering with controlled false positives.
Researchers analyze the geometric and statistical properties of transformer model representations.
problem Understanding the semantic structure of large transformer models across various data types.
method Characterization of geometric and statistical properties through analysis of intrinsic dimension and neighbor composition.
result The semantic information of the dataset is better expressed at the end of the first peak in transformer models.
A neural scene representation framework enforcing 3D transformations.
problem Learning 3D scene representations from images without 3D supervision.
method Introducing a loss enforcing equivariance of the scene representation with 3D transformations.
result Real-time neural rendering with comparable results to models requiring minutes for inference.
TCT learns multimodal sequence representations by translating from related sequences.
problem Challenges in learning semantic representations from multimodalities.
method Transformer based Cross-modal Translator (TCT) combined with Multimodal Transformer Network (MTN).
result Proposed method achieves new state-of-the-art performance on video-grounded dialogue.
A framework uses free probability to analyze Transformer models.
problem Understanding the dynamics and complexity of Transformer-based language models.
method Formal operator-theoretic analysis using free probability theory.
result Entropy-based generalization bounds derived under freeness assumptions.
Enhanced Transformer solves math problems better with explicit relation encoding.
problem Improving Transformer models for solving math word problems.
method Integrates Tensor-Product Representations and TP-Attention mechanism.
result Sets new state of the art on the Mathematics Dataset.
Researchers present and compare different representations of dissipative Hamiltonian DAE systems.
problem Understanding and transforming dissipative Hamiltonian DAE systems.
method Global geometric and algebraic points of view, translations between representations, characterizations, and numerical methods for computing structural information.
result A general DAE system can be transformed into a dissipative Hamiltonian or port-Hamiltonian DAE system.
Transforms solutions of Davey-Stewartson II equation geometrically.
problem Solving the Davey-Stewartson II equation.
method Moutard transform and spinor representation of surfaces.
result Constructs examples of solutions with smooth initial data losing regularity.
Representations in the auditory cortex might be based on mechanisms similar to the visual ventral stream; modules for building invariance to transformations and multiple layers for compositionality and selectivity. In this paper we propose the use of such computational modules for extracting invariant and discriminativ…
This study investigates abrupt learning dynamics in Transformers, revealing plateau formation and internal representation collapse.
problem Abrupt learning in Transformers, particularly during the loss plateau.
method Investigates mechanisms of abrupt learning in shallow Transformers, focusing on attention maps and hidden states.
result Reveals plateau formation, internal representation collapse, and strong repetition bias in outputs.
Learning invariant representations is an important problem in machine learning and pattern recognition. In this paper, we present a novel framework of transformation-invariant feature learning by incorporating linear transformations into the feature learning algorithms. For example, we present the transformation-invari…
Neural network models transform physical systems into latent Gaussian distributions.
problem Simplifying and solving classical Hamiltonian systems.
method Symplectic neural networks for canonical transformations.
result Captures nonlinear collective modes in latent space.
New model learns content and transformation separately from data.
problem Learning disentangled representations from data without explicit labels.
method Group-based variational autoencoders, assuming content and transformation groups.
result Model learns generalizable content representations from unseen data.
The paper reinterprets knot group invariants using affine transformations.
problem Alexander invariants of knots and their geometric interpretation.
method Representation varieties of knot groups into extrmAGL1(C). result Alexander polynomial as the singular locus of a coherent sheaf.
CRATE-MAE learns structured representations from unlabeled data.
problem Learning structured representations from unlabeled data.
method Structured Diffusion with White-Box Transformers.
result CRATE-MAE achieves highly promising performance on large-scale imagery datasets.
Transforms data into separable subspaces for clustering.
problem Data is not always separable into subspaces.
method Embeds subspace clustering techniques into transform learning.
result Improves upon state-of-the-art clustering techniques.
IGT learns graph representations without supervision.
problem Building deep unsupervised graph representations.
method Generic complex-valued spectral graph architecture from Fourier transform generalization, greedy concave objective for discriminative and invariant features.
result IGT learns both discriminative and invariant features from graph topology.
We present a 2x2 Lax representation for discrete circular nets of constant negative Gauß curvature. It is tightly linked to the 4D consistency of the Lax representation of discrete K-nets (in asymptotic line parametrization). The description gives rise to Bäcklund transformations and an associated family. All the membe…
New representations for discrete surfaces derived from dual transforms.
problem Constructing discrete surfaces in differential geometry.
method Using Ω-dual transform and lightlike Gauss maps in Laguerre geometry. result All discrete linear Weingarten surfaces arise via Weierstrass-type representations.
Normality equations describe Newtonian dynamical systems admitting normal shift of hypersurfaces. These equations were first derived in Euclidean geometry. Then very soon they were rederived in Riemannian and in Finslerian geometry. Recently I have found that normality equations can be derived in geometry given by clas…
The paper proposes efficient dictionary learning algorithms that avoid multiplications for sparse representations.
problem Sparse representation with reduced computational complexity.
method Factorizations of the dictionary into binary orthonormal, scaling, and shear transformations with closed-form solutions.
result The proposed methods are effective and can be compared to well-known transforms like FFT and DCT.
Different neural networks learn similar mappings with different weights.
problem Understanding shared representations across neural networks with varying weights.
method Shared response model and orthogonal transformations.
result Different neural networks encode the same input examples as different orthogonal transformations of an underlying shared representation.
GTNs learn new graph structures and improve node representation learning.
problem Learning node representations on misspecified or heterogeneous graphs.
method Graph Transformer Networks (GTNs) that generate new graph structures and learn effective node representations.
result GTNs achieve state-of-the-art performance in node classification tasks without predefined meta-paths.
Improved BERT model with latent persona and topic variables.
problem Improving BERT's domain-specific utility while maintaining generalization.
method Combining BERT with Universal Transformer, adding latent persona and topic variables.
result Pre-trained model for social texts outperforms baseline.
Enhances autoencoders to represent transformations explicitly.
problem Lack of explicit representation of transformations in autoencoders.
method Extended variational autoencoders to include latent transformations, using hierarchical graphical models.
result Inferred latent transformations reflect interpretable properties in the observation space.
Vision Transformers show different internal representations compared to CNNs.
problem Understanding how Vision Transformers solve image classification tasks.
method Comparative analysis of ViT and CNN architectures on image classification benchmarks.
result ViT has more uniform representations across all layers, while CNNs have more varied representations.
New IT representation improves symbolic regression approximations.
problem Finding better approximations to real-world data sets.
method Evolutionary Algorithm with IT representation using only mutation.
result IT representation finds better approximations than traditional methods.
Paper recovers latent causal structure and linear transformation from indirect observations.
problem Recovering latent causal structure and linear transformation from indirect observations.
method Established sufficient conditions for DAG recovery, leveraged score function properties, and used soft/hard interventions.
result Perfect recovery of latent DAG structure and linear transformation up to scaling using soft interventions, hard interventions with additional hypothesis testing.
We construct a certain cross product of two copies of the braided dual H~ of a quasitriangular Hopf algebra H, which we call the elliptic double EH, and which we use to construct representations of the punctured elliptic braid group extending the well-known representations of the planar braid group attache…
CascadeXML improves multi-resolution learning for XMC with transformer features.
problem Learning subset labels from millions of choices with trade-offs between performance and computation.
method End-to-end multi-resolution learning pipeline using transformer multi-layer architecture.
result Significantly outperforms existing approaches on benchmark datasets.
Transformers reduce redundancy by focusing on invariant relational quantities.
problem Substantial internal redundancy in Transformer models due to coordinate-dependent representations and continuous symmetries.
method Reformulate representations, attention mechanisms, and optimization dynamics in terms of invariant relational quantities, eliminating redundant degrees of freedom by construction.
result Architectures that operate directly on relational structures, providing a principled geometric framework for reducing parameter redundancy and analyzing optimization.
Transformer autoencoder learns musical style from performances.
problem Learning high-level controls over symbolic music generation.
method Aggregates encodings of input data across time to obtain global style representation.
result Improves control over performance style and melody in music generation tasks.
Tool visualizes Transformer model attention for better understanding.
problem Understanding complex attention mechanisms in deep learning models.
method Developed an open-source tool to visualize attention at three levels.
result Visualization helps interpret and analyze Transformer models.
Transformer-M learns molecular data in 2D or 3D formats.
problem Learning models for molecules are limited to specific data formats.
method Developed a Transformer-based model that can handle 2D and 3D molecular data.
result Transformer-M achieves strong performance on both 2D and 3D molecular tasks.
Study geometric and representation theory of statistical transformation models.
problem Understand relationships between induced structures and actions on measure spaces.
method Investigate geometric properties and symplectic actions on induced structures.
result Show equivariance of action and relationships between tangent bundles and projectivizations.