Knowledge distillation (KD) is a popular method for reducing the computational overhead of deep network inference, in which the output of a teacher model is used to train a smaller, faster student model. Hint training (i.e., FitNets) extends KD by regressing a student model's intermediate representation to a teacher mo…
Self-supervised and supervised methods learn similar intermediate visual representations but diverge in final layers.
problem Comparing self-supervised and supervised methods for visual learning.
method Comparison of contrastive self-supervised and supervised methods on simple image data.
result Contrastive and supervised methods learn similar intermediate representations but diverge in final layers.
Neural networks benefit from intermediate representations, reducing sample complexity.
problem Understanding how neural networks leverage intermediate representations for hierarchical learning.
method Fixed, randomly initialized neural network as a representation function, compared with raw inputs and other trainable networks.
result Neural representations can achieve improved sample complexities compared to raw inputs, especially for low-rank polynomials.
This paper introduces a novel approach to measuring privacy risks in deep computer vision models based on intermediate outputs.
problem The exposure of intermediate results in hidden layers of deep computer vision models poses significant privacy concerns.
method The approach leverages Degrees of Freedom (DoF) to evaluate the amount of information retained in each layer and combines this with the rank of the Jacobian matrix to assess sensitivity to input variations.
result The proposed framework provides deeper insights into privacy risks associated with intermediate representations without requiring adversarial attack simulations.
To reduce the large computation and storage cost of a deep convolutional neural network, the knowledge distillation based methods have pioneered to transfer the generalization ability of a large (teacher) deep network to a light-weight (student) network. However, these methods mostly focus on transferring the probabili…
New model-independent compact representations of imaginary-time data are presented in terms of the intermediate representation (IR) of analytical continuation. This is motivated by a recent numerical finding by the authors [J. Otsuki et al., arXiv:1702.03056]. We demonstrate the efficiency of the IR through continuous-…
End-to-end learnable network for safer self-driving with interpretable intermediate representations.
problem Safe motion planning for self-driving vehicles.
method Differentiable semantic occupancy representation for cost calculation in motion planning.
result Significantly outperforms state-of-the-art planners in imitating human behaviors and producing safer trajectories.
New analysis of annealing paths in sampling and estimation.
problem Sampling from complex distributions and estimating normalization constants.
method Extending known results on Bregman divergence to quasi-arithmetic means under monotonic embedding.
result Analogous result for quasi-arithmetic means, highlighting the interplay between means, parametric families, and divergence functionals.
TIPRDC anonymizes data features to protect privacy while retaining useful information.
problem Privacy concerns from crowdsourced data hinder deep learning applications.
method Hybrid training method combining adversarial and mutual information estimation.
result Feature extractor hides private information while preserving original data features.
We study Anosov representations whose limit set has intermediate regularity, namely is a Lipschitz submanifold of a flag manifold. We introduce an explicit linear functional, the unstable Jacobian, whose orbit growth rate is integral on this class of representations. We prove that many interesting higher rank represent…
We present a method for feature interpretation that makes use of recent advances in autoregressive density estimation models to invert model representations. We train generative inversion models to express a distribution over input features conditioned on intermediate model representations. Insights into the invariance…
In this paper, we diagnose deep neural networks for 3D point cloud processing to explore utilities of different intermediate-layer network architectures. We propose a number of hypotheses on the effects of specific intermediate-layer network architectures on the representation capacity of DNNs. In order to prove the hy…
The deep layers of modern neural networks extract a rather rich set of features as an input propagates through the network. This paper sets out to harvest these rich intermediate representations for quantization with minimal accuracy loss while significantly reducing the memory footprint and compute intensity of the DN…
Neural networks are vulnerable to adversarial examples, malicious inputs crafted to fool trained models. Adversarial examples often exhibit black-box transfer, meaning that adversarial examples for one model can fool another model. However, adversarial examples are typically overfit to exploit the particular architectu…
New approach decouples skill learning and language grounding for autonomous agents.
problem Autonomous acquisition of skills without external instructions and feedback.
method Language-Goal-Behavior (LGB) architecture with semantic representation.
result Decouples skill learning and language grounding, enabling diversity and strategy switching.
We review the current state of automatic differentiation (AD) for array programming in machine learning (ML), including the different approaches such as operator overloading (OO) and source transformation (ST) used for AD, graph-based intermediate representations for programs, and source languages. Based on these insig…
Researchers analyze the geometric and statistical properties of transformer model representations.
problem Understanding the semantic structure of large transformer models across various data types.
method Characterization of geometric and statistical properties through analysis of intrinsic dimension and neighbor composition.
result The semantic information of the dataset is better expressed at the end of the first peak in transformer models.
Study finds more flood risk strategies can improve outcomes in NYC.
problem Managing future flood risks with complex models.
method Used an intermediate complexity model to analyze flood risk strategies.
result More combinations of risk mitigation strategies expand the solution set and improve outcomes.
AP-Calculus offers a new framework for causal inference in Bayesian networks.
problem Causal inference in Bayesian networks with complex architectures.
method Introduces Attribution Projection Calculus (AP-Calculus) to determine causal relationships.
result Proves that for each label, exactly one intermediate node acts as a deconfounder.
Several intrinsic topological ways to encode connections on vector bundles on smooth complex algebraic curves will be described. In particular the notion of {\em Stokes decompositions} will be formalised, as a convenient intermediate category between the Stokes filtrations and the Stokes local systems/wild monodromy re…
Paper analyzes why deeper layers of ViTs perform worse on out-of-distribution tasks.
problem Performance degradation of intermediate layers in ViTs under distribution shift.
method Extensive linear probing experiments across various benchmarks and fine-grained analysis of transformer modules.
result Probing feedforward network activations yields best performance under significant distribution shift.
Neural networks are vulnerable to adversarial examples, malicious inputs crafted to fool trained models. Adversarial examples often exhibit black-box transfer, meaning that adversarial examples for one model can fool another model. However, adversarial examples may be overfit to exploit the particular architecture and …
Quantized-TinyLLaVA reduces communication costs in split learning for multimodal models.
problem High communication costs in split learning for multimodal models.
method Integrates a compression module that quantizes intermediate features into discrete representations before transmission.
result Achieves an approximate 87.5% reduction in communication overhead with 2-bit quantization.
RAE improves image representation learning with simplified design choices.
problem Improving image representation learning using pretrained vision encoders.
method Generalized RAE formulation, complementary working mechanisms of RAE and REPA, and free CFG guidance.
result RAEv2 achieves state-of-the-art results with 10x faster convergence and less training time.
A new method aggregates generative classifiers to resist adversarial attacks.
problem Adversarial attacks on deep neural networks.
method Rank-aggregating ensemble of generative classifiers trained on intermediate layer responses.
result The ensemble of generative classifiers shows robustness to adversarial attacks.
Framework learns stochastic dynamics from endpoint and intermediate distributions using soft energy constraints.
problem Learning stochastic dynamics from endpoint and intermediate distributional observations.
method Formulates generation as a McKean-Vlasov control problem with soft energy constraints, solving it through FBSDE.
result Model learns coherent stochastic trajectories matching prescribed marginal laws.
Study on rigidity with non-negative intermediate curvature on low-dimensional manifolds.
problem Extending non-existence theorem of positive scalar curvature to product manifolds.
method Introduced intermediate curvature and studied rigidity conditions.
result Rigidity when intermediate curvature is non-negative in low dimensions.
Preserves positive intermediate curvature on manifolds.
problem Obstructs positive intermediate curvature on partial tori.
method Shows smooth interpolation of metrics with positive intermediate curvature.
result Proves non-existence of certain manifolds with positive intermediate curvature.
NKI integrates obfuscated datasets using nonlinear kernels for improved data collaboration.
problem Privacy-preserving data collaboration with reduced reconstruction risk.
method Formulates linear kernel integration, kernelizes it, and introduces graph regularization and centering constraints.
result NKI improves classification accuracy over existing linear integration methods under nonlinear dimensionality reduction.
Statistical methods protecting sensitive information or the identity of the data owner have become critical to ensure privacy of individuals as well as of organizations. This paper investigates anonymization methods based on representation learning and deep neural networks, and motivated by novel information theoretica…
The study connects manifold topology to metrics with positive intermediate curvature.
problem Understanding the relationship between manifold topology and metrics with positive intermediate curvature.
method Formulated a conjecture and proved it for specific dimensions and conditions.
result Closed, aspherical 6-manifolds cannot admit metrics with positive 4-intermediate curvature.
For a pair of points in a smooth closed convex planar curve γ, its mid-line is the line containing its mid-point and the intersection point of the corresponding pair of tangent lines. It is well known that the envelope of the mid-lines (EML) is formed by the union of three affine invariants sets: Affine Envelope Sy…
New rigidity results for manifolds with maximal symmetry rank and positive intermediate Ricci curvature.
problem Understanding the structure of manifolds with maximal symmetry rank and positive intermediate Ricci curvature.
method Recovering stronger topological rigidity results using higher intermediate Ricci curvatures and nontrivial fundamental groups.
result Stronger topological rigidity results for manifolds with maximal symmetry rank and positive intermediate Ricci curvature.
The EMNLP 2018 workshop BlackboxNLP was dedicated to resources and techniques specifically developed for analyzing and understanding the inner-workings and representations acquired by neural models of language. Approaches included: systematic manipulation of input to neural networks and investigating the impact on thei…
ie-HGCN addresses HIN challenges by efficiently learning node representations.
problem Lack of flexibility in exploring meta-paths and high computational complexity in HIN GCN methods.
method Hierarchical aggregation architecture that automatically extracts useful meta-paths and reduces computational cost.
result ie-HGCN outperforms state-of-the-art methods on real network datasets.
Study on spaces of metrics with intermediate curvature bounds.
problem Understanding spaces of metrics with lower bounds on intermediate curvatures.
method Analyzing spaces of Riemannian metrics with specific curvature bounds on high-dimensional Spin-manifolds.
result Spaces of metrics with positive p-curvature and k-positive Ricci curvature have non-trivial homotopy groups.
Proves metrics with positive intermediate Ricci curvature on complex manifolds.
problem Establishing metrics with positive intermediate Ricci curvature on complex manifolds.
method Canonical variation and surgery techniques.
result Existence of metrics with positive intermediate Ricci curvature on various examples.
Extends Perelman's theorem to positive intermediate curvature conditions.
problem Positive intermediate curvature conditions and their implications.
method Generalization of Perelman's gluing theorem to positive intermediate curvature conditions.
result Observer moduli space can have non-trivial higher homotopy groups.
The study proves that certain manifolds with boundary cannot have metrics with positive intermediate curvatures.
problem Proving the nonexistence of metrics with positive intermediate curvatures on manifolds with boundary.
method Curvature obstruction theorems for manifolds with boundary.
result Topologically nontrivial compact manifolds with boundary cannot have metrics of positive m-intermediate curvature if the boundary is m-convex. Neural network feature optimization for causal inference.
problem Estimating heterogeneous treatment effects from data.
method Genetic algorithm optimization of intermediate neural network layers for feature representations.
result Retains useful features for outcome prediction even if related to treatment assignment.
Reduces data leakage in distributed deep learning models.
problem Prevents reconstruction of sensitive raw data patterns during client communications.
method Reduces distance correlation between raw data and learned representations.
result Resilient to reconstruction attacks while maintaining model accuracy.
New model learns graph neural networks equivariant to various transformations.
problem Learning equivariant graph neural networks for complex transformations.
method E(n)-Equivariant Graph Neural Networks (EGNNs) that are computationally efficient and scalable.
result Achieves competitive or better performance without higher-order representations.
Study finds metrics with positive intermediate Ricci curvature on specific low-dimensional manifolds.
problem Existence of invariant metrics with positive intermediate Ricci curvature on low-dimensional cohomogeneity one manifolds.
method Construction of invariant metrics with positive intermediate Ricci curvature on specific manifolds.
result Invariant metrics with positive 4th-intermediate Ricci curvature exist but not for 3rd-intermediate Ricci curvature on certain manifolds.
In this paper, we propose a new method called ProfWeight for transferring information from a pre-trained deep neural network that has a high test accuracy to a simpler interpretable model or a very shallow network of low complexity and a priori low test accuracy. We are motivated by applications in interpretability and…
Sharp dimension constraints for positive intermediate curvature metrics are established.
problem Proving sharp dimension constraints for metrics with positive intermediate curvature.
method Constructing counterexamples and extending rigidity results.
result Sharp dimension constraints for positive intermediate curvature metrics are established.
The paper proves manifold splitting theorems with nonnegative intermediate curvature.
problem Proving rigidity results for manifolds with nonnegative intermediate curvatures.
method New recursion theorem for spectral intermediate curvatures and cylindrical splitting theorems.
result Smooth metrics with uniformly positive intermediate curvature constructed.
A framework to explain decoder-only sequence classification models using intermediate predictions.
problem Explaining predictions of decoder-only sequence classification models.
method Progressive Inference framework with Single Pass-Progressive Inference and Multi Pass-Progressive Inference methods.
result Significantly better attributions compared to prior work on text classification tasks.
The paper shows how to answer future and past questions from high-dimensional time series data.
problem Challenges in answering probabilistic inference questions from high-dimensional time series data.
method Temporal contrastive learning to learn Gaussian representations that enable compact closed-form solutions.
result Representations learned via contrastive learning follow a Gauss-Markov chain, enabling efficient inference and planning.