This research quantifies neural networks using magnitude, a topological invariant.
problem Understanding the generalization capabilities of neural networks.
method Using a novel topological invariant called magnitude to study neural network representations.
result Magnitude dimension is theoretically connected to generalisation error and can predict it.
This paper introduces new invariants for time series analysis.
problem Analyzing the diversity and invariants of time series data.
method Introduces new invariants derived from the continuity of magnitude and maximum diversity.
result Demonstrates improved performance in machine learning experiments with real-world data.
Magnitude is a real-valued invariant of metric spaces, analogous to the Euler characteristic of topological spaces and the cardinality of sets. The definition of magnitude is a special case of a general categorical definition that clarifies the analogies between various cardinality-like invariants in mathematics. Altho…
Magnitude study on manifolds using fractional Laplacian.
problem Magnitude invariant of compact metric spaces via fractional Laplacian.
method Semiclassical analysis of nonlocal boundary value problem related to fractional Laplacian.
result Asymptotic expansion of magnitude in terms of curvature invariants.
Magnitude of manifolds linked to Riesz energies and beta functions.
problem Magnitude invariant and its geometric significance.
method Relating magnitude invariant to Brylinski's beta function and pseudodifferential analysis.
result Precise relation between magnitude invariant and beta function for closed manifolds.
Magnitude is not continuous but may be stable for most finite metric spaces.
problem Stability of magnitude invariant in finite metric spaces.
method Investigates the continuity properties of magnitude with respect to Gromov-Hausdorff topology.
result Magnitude is nowhere continuous but may be generically continuous.
Magnitude homology reveals that graphs can have torsion subgroups.
problem Understanding torsion in magnitude homology of graphs.
method Analysis of magnitude homology defined by Hepworth and Willerton.
result Torsion of any prime order can appear in graphs' magnitude homology.
New measures quantify diversity of latent representations using metric space magnitude.
problem Evaluating the diversity of latent representations in machine learning models.
method Developed magnitude-based measures for latent representations, stable under data perturbations.
result Demonstrated superior performance across various domains and tasks.
A new pruning criterion reduces model size and improves performance.
problem Overparameterized neural networks are computationally and memory intensive, leading to overfitting.
method Introduces a magnitude and uncertainty (M&U) pruning criterion inspired by statistical Wald test.
result Our M&U pruning criterion leads to more compressed models with less loss in predictive power.
Magnitude-based features capture interactions between different entities in multispecies spatial data.
problem Capturing interactions between different entities in multispecies spatial data.
method Developing magnitude-based features for multispecies spatial data.
result Identifies distinct neighbourhood types and spatial heterogeneity.
Order-flow entropy predicts price magnitude without directionality.
problem Predicting price magnitude in financial markets.
method Real-time order-flow entropy computed from a 15-state Markov transition matrix.
result Order-flow entropy predicts the magnitude of intraday returns with high accuracy.
Most learning algorithms are not invariant to the scale of the function that is being approximated. We propose to adaptively normalize the targets used in learning. This is useful in value-based reinforcement learning, where the magnitude of appropriate value approximations can change over time when we update the polic…
Performance of neural networks can be significantly improved by encoding known invariance for particular tasks. Many image classification tasks, such as those related to cellular imaging, exhibit invariance to rotation. We present a novel scheme using the magnitude response of the 2D-discrete-Fourier transform (2D-DFT)…
Pruned neural networks' error scales predictably with architecture and task.
problem Understanding the predictability of pruning across different scales and architectures.
method Functionally approximated the error of pruned networks, showing it is predictable in terms of invariant tying width, depth, and pruning level.
result The error of pruned networks follows a scaling law with interpretable coefficients that depend on architecture and task.
Defines magnitude for length spaces with measures, agreeing with finite spaces' magnitude.
problem Defining magnitude for non-finite metric spaces with measures.
method Integrals over geodesics, using counting and weight measures.
result Magnitude agrees with finite spaces' magnitude and volume under specific conditions.
Lookahead pruning extends single-layer optimization to multi-layer, outperforming magnitude-based pruning.
problem Pruning neural networks to reduce computational cost and memory usage.
method Developed a multi-layer optimization approach extending the single-layer optimization of magnitude-based pruning.
result Consistently outperforms magnitude-based pruning on various networks, especially in high sparsity.
Magnitude of Euclidean domains predicts Willmore energy in odd dimensions.
problem Magnitude function of compact domains in odd dimensions.
method Asymptotic expansion of magnitude function at infinity.
result Magnitude function determines Willmore energy of boundary in odd dimensions.
Are expansions and recessions more likely to end as their magnitude increases? In this paper we apply parametric hazard models to investigate this issue in a sample of 16 countries from 1881 to 2000. For the total sample we find evidence of positive magnitude dependence for recessions, while for expansions we are not a…
Hepworth, Willerton, Leinster and Shulman introduced the magnitude homology groups for enriched categories, in particular, for metric spaces. The purpose of this paper is to describe the magnitude homology group of a metric space in terms of order complexes of posets. In a metric space, an interval (the set of points b…
Magnitude of geometric shapes studied for smooth manifolds, revealing spectral geometry insights.
problem Understanding the geometric significance of Leinster's magnitude for smooth manifolds.
method Investigation of magnitude function for various distance functions, including submanifolds and Riemannian manifolds, with asymptotic analysis in the limit.
result Magnitude function is well-defined and meromorphically continued for large distances, revealing volume, surface area, and curvature integrals.
A new method models financial returns by separating sign and magnitude, improving forecasting accuracy.
problem Capturing nonlinear predictability in financial return dynamics.
method Decomposes returns into sign and magnitude components, using a joint distribution model.
result Significantly outperforms traditional linear models in forecasting U.S. stock market returns.
Novel metric space magnitude and weighting vectors improve machine learning tasks.
problem Improving machine learning algorithms using novel metric space concepts.
method Metric space magnitude and weighting vectors for better machine learning.
result The weighting vector effectively detects boundaries and improves classic machine learning tasks.
We prove that ``almost generically'' for a one-relator group Delzant's T-invariant (which measures the smallest size of a finite presentation for a group) is comparable in magnitude with the length of the defining relator. The proof relies on our previous results regarding isomorphism rigidity of generic one-relator …
Previous studies indicate that nonlinear properties of Gaussian time series with long-range correlations, ui, can be detected and quantified by studying the correlations in the magnitude series ∣ui∣, i.e., the ``volatility''. However, the origin for this empirical observation still remains unclear, and the exact …
In this paper we establish the existence of periodic orbits belonging to any σ-atoroidal free homotopy class for Hamiltonian systems in the twisted disc bundle, provided that the compactly supported time-dependent Hamiltonian function is sufficiently large over the zero section and the magnitude of the weakly exact $…
SPADE-S improves time series forecasting accuracy for low-magnitude and sparse data.
problem Challenges in forecasting time series with strong heterogeneity in magnitude and sparsity.
method SPADE-S is a robust forecasting architecture that reduces biases and improves overall prediction accuracy.
result SPADE-S outperforms existing state-of-the-art approaches across diverse use cases, improving forecast accuracy by up to 15%.
In this paper we define the magnitude of metric spaces using measures rather than finite subsets as had been done previously and show that this agrees with earlier work with Leinster in arXiv:0908.1582. An explicit formula for the magnitude of an n-sphere with its intrinsic metric is given. For an arbitrary homogeneous…
Accelerated magnetic resonance (MR) scan acquisition with compressed sensing (CS) and parallel imaging is a powerful method to reduce MR imaging scan time. However, many reconstruction algorithms have high computational costs. To address this, we investigate deep residual learning networks to remove aliasing artifacts …
This paper continues a series of studies devoted to analysis of the bivariate probability distribution P(x,y) of two consecutive price increments x (push) and y (response) at intraday timescales for a group of stocks. Besides the asymmetry properties of P(x,y) such as Market Mill dependence patterns described in preced…
New pruning method retains model expressiveness for NLP tasks.
problem Pruning large pretrained transformer models for real-world deployment.
method Mixture Gaussian Prior Pruning (MGPP) algorithm.
result MGPP outperforms existing pruning methods in high sparsity settings.
Our work connects parameter magnitudes and Hessian eigenspaces in deep neural nets.
problem Understanding the relationship between parameter magnitudes and Hessian curvature in deep learning models.
method Developed a matrix-free algorithm based on sketched SVDs to measure similarity between parameter masks and Hessian eigenspaces.
result Top Hessian eigenvectors tend to be concentrated around larger parameters, indicating a connection between parameter magnitudes and loss curvature.
Study non-asymptotic bounds on correlation in high-dimensional linear systems, revealing invariant subspaces and bottlenecks.
problem Understanding correlation and mixing in high-dimensional linear systems with Gaussian noise.
method Sampling from sub-trajectories, using Talagrand's inequality, and analyzing invariant subspaces.
result Large discrepancy between algebraic and geometric multiplicity leads to bottlenecks between invariant subspaces.
A new method recovers latent potentials from graph flows, preserving ordering and stability.
problem Recovering latent potentials from graph flows is ill-posed and standard methods collapse the ordering.
method Gauge-invariant, parameter-insensitive regularization using Dirichlet energy.
result The method preserves ordering and stability across different regularization strengths.
Thermalizer stabilizes autoregressive models for long-term predictions in chaotic systems.
problem Long-term predictions in chaotic spatiotemporal systems are unreliable due to trajectory divergence.
method Diffusion models are used to implicitly estimate the score of an invariant measure, which stabilizes autoregressive emulators by applying denoising during inference.
result Thermalization extends the time horizon of stable predictions by an order of magnitude in chaotic systems.
New Lipschitz bound for ReLU networks resists weight rescaling.
problem Lack of robustness guarantees for ReLU networks under weight perturbations.
method Rescaling-invariant Lipschitz bound based on path-metrics.
result The new bound applies to various ReLU-DAG architectures and resists neuron-wise rescalings.
This paper proposes an approach to the joint modeling of the short-time Fourier transform magnitude and phase spectrograms with a deep generative model. We assume that the magnitude follows a Gaussian distribution and the phase follows a von Mises distribution. To improve the consistency of the phase values in the time…
Triangulation filters spurious circuits in multilingual models.
problem Unreliable explanations of multilingual models across languages.
method Formalizes reference families and introduces triangulation as a causal acceptance rule.
result Triangulation provides a falsifiable standard for mechanistic claims.
Upper bound derived for informed traders' gains in a model, akin to thermodynamics.
problem Informed traders' gains in a financial model with finite horizon.
method Bayesian inference and entropic inequality.
result Upper bound for expected gain, analogous to thermodynamics.
This paper proposes an end-to-end approach for single-channel speaker-independent multi-speaker speech separation, where time-frequency (T-F) masking, the short-time Fourier transform (STFT), and its inverse are represented as layers within a deep network. Previous approaches, rather than computing a loss on the recons…
Light neural network detects modulation in noisy signals.
problem Efficiently detecting modulation in noisy signals.
method Light neural network architecture invariant to impairments.
result Network achieves accuracy under realistic impairments.
Causal invariance can improve finite-sample domain adaptation, but only when the target risk margins are large.
problem Finite-sample domain adaptation
method Linear regression with causal knowledge
result Adaptive aggregation can match best candidate predictor while avoiding negative transfer
Many loss functions in representation learning are invariant under a continuous symmetry transformation. For example, the loss function of word embeddings (Mikolov et al., 2013) remains unchanged if we simultaneously rotate all word and context embedding vectors. We show that representation learning models for time ser…
A new energy-efficient pruning method for federated learning.
problem Energy inefficiency in gradient sparsification for federated learning.
method Formalized energy-constrained projection problem and proposed Cost-Weighted Magnitude Pruning (CWMP).
result CWMP optimally balances performance and energy efficiency in federated learning.
SAM improves generalization in overparameterized models, but its behavior in tensorized models is less understood.
problem Understanding the implicit regularization of SAM in tensorized models.
method Scale-invariance analysis and gradient flow analysis to derive Norm Deviation as a measure of core norm imbalance, and propose Deviation-Aware Scaling (DAS).
result DAS achieves competitive or improved performance over SAM, while offering reduced computational overhead.
This paper considers magnitude, asymptotics and duration of drawdowns for some Lévy processes. First, we revisit some existing results on the magnitude of drawdowns for spectrally negative Lévy processes using an approximation approach. For any spectrally negative Lévy process whose scale functions are well-behaved at …
New method identifies whether equity return predictability is due to magnitude shrinkage or directional reversal.
problem Determining the nature of equity return predictability (directional reversal vs magnitude shrinkage).
method Developed the Fourier-Residue Identity (FRI) to decompose return autocorrelation into sign and magnitude channels.
result The lag-1 autocorrelation in SPY is driven entirely by magnitude shrinkage, not directional reversal.
Distributed training of massive machine learning models, in particular deep neural networks, via Stochastic Gradient Descent (SGD) is becoming commonplace. Several families of communication-reduction methods, such as quantization, large-batch methods, and gradient sparsification, have been proposed. To date, gradient s…
ISALT uses inference to simulate SDEs with large time-steps, improving efficiency.
problem Efficiently simulating ergodic SDEs with large time-steps.
method Inference-based schemes adaptive to large time-steps (ISALT) from data.
result ISALT achieves significant time reduction and optimal accuracy.