Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

100200299399 · Jun 202019922001200920182026
48 results for layer similarity

Task loss matching misrepresents similarity between neural network layers.

problem Measuring similarity between neural network layers using task loss matching.
method Task loss matching vs. direct matching; comparison with CCA and CKA.
result Direct matching provides a better similarity index than task loss matching.

Modified RV-coefficient reveals how training affects neural network representations.

problem Understanding how training affects intermediate representations in convolutional neural networks.
method Experimented with modified RV-coefficient (RV2) to compare activation patterns in deep networks trained on varying amounts of data and layers.
result RV2 successfully recovered expected similarity patterns and provided interpretable similarity matrices.

Single-layer GCN model improves recommendation performance with less complexity.

problem Severe computational burden and excessive model parameters in existing GCN models.
method Proposes a single-layer GCN architecture with a simplified aggregation step using DA similarity.
result Significantly outperforms existing GCN models and achieves up to a few orders of magnitude speedup.

ALMA improves clustering of multilayer networks.

problem Clustering multilayer networks with distinct layers and communities.
method Alternating minimization algorithm (ALMA) for simultaneous layer partition and community estimation.
result ALMA achieves higher accuracy than TWIST in clustering multilayer networks.

Cosine normalization uses cosine similarity to reduce neuron variance in neural networks.

problem Large variance in neuron outputs leads to poor generalization and internal covariate shift.
method Replace dot product with cosine similarity or centered cosine similarity in neural networks.
result Cosine normalization improves model performance on various datasets.

A new method for verifying deep learning architectures on FPGAs is proposed.

problem Design-time verification of deep learning architectures on FPGAs.
method 2-Level 3-Way (2L-3W) hardware-software co-verification methodology.
result Layer-by-layer similarity scores of 99% accuracy for successful mappings.

NSL layer improves convolutional networks' ability to recognize novel appearances.

problem Convolutional networks struggle with recognizing novel appearances not seen in training data.
method NSL layer uses neighborhood similarity to induce appearance invariance.
result NSL layer enhances network's ability to generalize to novel appearances.

Self-supervised and supervised methods learn similar intermediate visual representations but diverge in final layers.

problem Comparing self-supervised and supervised methods for visual learning.
method Comparison of contrastive self-supervised and supervised methods on simple image data.
result Contrastive and supervised methods learn similar intermediate representations but diverge in final layers.

Proposes SimPool for graph pooling using structural similarity features.

problem Challenges in graph pooling due to lack of spatial locality.
method Integrates structural similarity features with a revised pooling layer to propose SimPool.
result SimPool produces node cluster assignments resembling CNN's locality preserving pooling.

Binary autoencoder with sparse hidden layer preserves information and zero reconstruction error.

problem Preserving information and zero reconstruction error in binary neural networks.
method Binary autoencoder with random binary weights, sparse hidden layer, and varying neuron thresholds.
result Zero reconstruction error for any input with a large hidden layer and varying neuron thresholds.

The paper explores how AI trading agents' similar information representation can cause financial market instability.

problem Systemic instability in AI-dominated financial markets due to similar information representation.
method Structural multi-agent market model with two-layer decision architecture for AI agents.
result Representation homogeneity can lead to systemic instability in financial markets.

Proposes a new model to cluster layers in multilayer networks.

problem Identifying meaningful similarities in community structure across layers.
method Strata Multilayer Stochastic Block Model (sMLSBM) with algorithms for layer separation and parameter estimation.
result Joint clustering of nodes and layers provides insights into underlying relational patterns.

Coherent Multiplex analyzes real-time wavelet coherence among multiple signals.

problem Identifying and visualizing coherence among multiple time series.
method Fast spectral similarity based on cosine similarity metrics of Fourier-transformed signals and sparse time-frequency wavelet coherence.
result Scalable real-time system for low-latency inference and monitoring of inter-signal relationships.

The paper studies how neural networks evolve representations, finding a unique fixed point for nonlinear activations.

problem Understanding how neural networks transform input data across layers.
method Theoretical framework for the evolution of the kernel sequence, using mean-field regime and Hermite polynomials.
result For nonlinear activations, the kernel sequence converges globally to a unique fixed point.

LART detects communities across layers in multiplex networks.

problem Detecting shared communities across layers in multiplex networks.
method Locally Adaptive Random Transitions (LART) using random walks with layer similarity-dependent transition probabilities.
result LART outperforms existing algorithms in detecting communities across layers.

New multi-layer algorithm improves CNN performance.

problem Efficiently modeling and processing information with parsimonious representations.
method Generalized Basis Pursuit to multi-layer setting, proposing ML-ISTA and ML-FISTA algorithms.
result Nested first order algorithms converge to solve the multi-layer problem.

Study reveals a log-periodic structure in ETF sizes and finds large ETFs outperform small ones.

problem Understanding the size distribution and performance of ETFs.
method Detailed statistical analyses of ETF size distribution and performance metrics.
result Large ETFs outperform small ones, with a log-periodic structure in size distribution.

Proposes a simple neural network model similar to gradient boosted decision trees.

problem Building a neural network equivalent to gradient boosted decision trees.
method Converts an ensemble of decision trees to a neural network, relaxes properties, and trains a simple neural network model.
result The proposed Hammock model achieves similar performance to gradient boosted decision trees.

Single wide layer followed by a pyramidal structure ensures global convergence in deep networks.

problem Ensuring global convergence in deep neural networks with limited width constraints.
method Proves that a single wide layer followed by a pyramidal structure guarantees global convergence for over-parameterized networks.
result Single wide layer of width NN suffices for global convergence in deep networks with constant-width remaining layers.

IEA improves CNN models by averaging multiple convolutional layers.

problem Improving CNN model accuracy through ensemble learning.
method Replacing single convolutional layers with Inner Average Ensembles (IEA) of multiple convolutional layers.
result CNN models using IEA outperform those with regular convolutional layers.

The Soft Nearest Neighbor Loss improves representation quality and generalization.

problem Improving the quality of representations in machine learning models.
method Explored and expanded the Soft Nearest Neighbor Loss to measure class entanglement.
result Maximizing the entanglement of different class representations leads to better generalization and uncertainty estimation.

The paper explains deep learning models for recommendations using layer-wise relevance propagation.

problem Explainable recommendations in deep learning models.
method Layer-wise relevance propagation applied to a Deep Convolutional Neural Network.
result Demonstrates the effectiveness of the method on an Amazon products dataset.

Deconfounds neural network representation similarity metrics to improve consistency and accuracy.

problem Confounding by population structure in similarity metrics like RSA and CKA.
method Covariate adjustment regression to adjust for confounders.
result Improves detection of semantically similar neural networks and consistency in transfer learning.

A new method aligns convolution filters for temporal sequences using Dynamic Time Warp.

problem Improving deep learning models' ability to handle temporal sequence data.
method Integrates Dynamic Time Warp algorithm into 1-D convolution layers for better alignment of input and filter.
result Exceeds or matches standard 1-D convolution layers in time series classification tasks.

Adaptive RNN using mixture layer for multi-pattern sequences.

problem Inadequate RNN performance on sequences with multiple patterns.
method Introducing a mixture layer to partition and store prototype vectors, enabling adaptive state updates.
result M-RNN outperforms traditional RNN in assimilating sequences with multiple patterns.

The study compares feed-forward and attention layers in language models.

problem Understanding the role of feed-forward and attention layers in language models.
method Empirical and theoretical analysis in a synthetic setting.
result Feed-forward layers learn simple distributional associations, while attention layers focus on in-context reasoning.

This work analyzes when contrastive models are close to PCA or kernel methods.

problem Understanding when contrastive models are equivalent to kernel methods or PCA.
method Analyzing the training dynamics of two-layer contrastive models with non-linear activation.
result Wide contrastive models with cosine similarity based losses are close to PCA.

Layered graphical models improve discriminative learning efficiency.

problem Improving discriminative learning efficiency in graphical models.
method Designing layered graphical models (LGMs) in analogy to neural networks, using tensorized truncated variational inference and backpropagation.
result LGMs achieve competitive results in image classification, comparable to neural networks.

New similarity index avoids limitations of CCA in neural networks.

problem Limitations of existing methods in measuring neural network representation similarity.
method Introducing a similarity index based on centered kernel alignment (CKA) to measure representational similarity matrices.
result CKA reliably identifies correspondences between representations in networks trained from different initializations.

Deep networks trained with Hebbian updates perform similarly to back-propagation on image datasets.

problem Training deep networks with realistic asymmetric connections and updates.
method Use Hebbian updates with separate feedforward and feedback weights, and local rule for updates.
result Similar performance to back-propagation achieved with Hebbian updates on challenging image datasets.

This paper investigates how forgetting affects neural network representations and stabilizes deeper layers.

problem Catastrophic forgetting in machine learning models trained on sequential tasks.
method Representational analysis techniques and empirical studies on CIFAR-10 and CIFAR-100 datasets.
result Deeper layers are disproportionately the source of forgetting, and methods to mitigate forgetting stabilize these layers.

A network architecture improves classification performance on various datasets.

problem Improving classification performance on diverse datasets.
method Divide fully connected layer into three levels, learn existing layers, cluster similar classes, reclassify using clustering masks.
result Achieved state-of-the-art performance with an error rate of 11.56% on Cifar-100.

New metric shows how different regularization methods affect deep linear networks.

problem Understanding the training dynamics of deep linear networks.
method Introduced a new metric called layer imbalance to analyze training dynamics. Demonstrated behavior of different regularization methods and stochastic gradient descent.
result Different regularization methods behave similarly, leading to a flat minima.

Lean 2-layer RBMs achieve similar representational power as single-layer RBMs with fewer parameters.

problem Understanding and quantifying the representational power of multi-layer RBMs.
method Inherent Structure Capacity (ISC) and Lean RBMs.
result 2-layer RBMs can achieve the same representational power as single-layer RBMs with fewer parameters.

SBSS uses similarity to split data for better classifier training.

problem Training better classifiers with realistic performance estimation.
method SBSS uses both input and output space information to split data using similarity functions.
result SBSS outperformed ordinary stratified 10-fold cross-validation in 75% of scenarios.

Transformers can learn spectral methods and perform unsupervised learning.

problem Learning spectral methods using unsupervised learning.
method Using multi-layered Transformers, pre-trained on a large set of instances, to learn and perform statistical estimation tasks.
result Proven that pre-trained Transformers can learn spectral methods and perform tasks like PCA and clustering.

A new tensor network method for image classification reduces computation cost.

problem Efficiently classifying images in high-dimensional spaces.
method Proposes a multi-layered tensor network (MLTN) that performs one MPS operation per layer, reducing computation cost.
result Reduces computation cost without degrading performance.

SDP approach recovers communities in multilayer hypergraphs from aggregated similarity matrices.

problem Community recovery in multilayer hypergraphs using aggregated similarity matrices.
method Semidefinite programming (SDP) approach.
result Information-theoretic conditions for exact recovery in both assortative and disassortative cases.