Sparse Transformers degrade semantic information first, with early layers encoding more.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The worldwide trade network has been widely studied through different data sets and network representations with a view to better understanding interactions among countries and products. Here we investigate international trade through the lenses of the single-layer, multiplex, and multi-layer networks. We discuss diffe…
This paper analyzes and improves the effectiveness of BN techniques in controlling Internal Covariate Shift.
Paper presents an algorithm to analyze DGNNs by identifying generative boundaries and sampling.
Random feature maps improve forecasting with cheaper computation.
Many real-world complex systems across natural, social, and economical domains consist of manifold layers to form multiplex networks. The multiple network layers give rise to nonlinear effect for the emergent dynamics of systems. Especially, weak layers that can potentially play significant role in amplifying the vulne…
Layered neural networks have greatly improved the performance of various applications including image processing, speech recognition, natural language processing, and bioinformatics. However, it is still difficult to discover or interpret knowledge from the inference provided by a layered neural network, since its inte…
We develop a fast, tractable technique called Net-Trim for simplifying a trained neural network. The method is a convex post-processing module, which prunes (sparsifies) a trained network layer by layer, while preserving the internal responses. We present a comprehensive analysis of Net-Trim from both the algorithmic a…
Based on the approach of flow distances, the international trade flow system is studied from the perspective of multi-layer flow network. A model of multi-layer flow network is proposed for modelling and analyzing multiple types of flows in flow systems. Then, flow distances are introduced, and symmetric minimum flow d…
While the authors of Batch Normalization (BN) identify and address an important problem involved in training deep networks-- Internal Covariate Shift-- the current solution has certain drawbacks. Specifically, BN depends on batch statistics for layerwise input normalization during training which makes the estimates of …
Transformer attention layers solve single-location regression tasks.
Spain uses DEA to select international markets for exports.
Deep learning models can overfit noisy data without losing generalization.
Study neural networks by mapping correlations, revealing essential statistics.
A new framework detects anomalous inputs to DNNs.
Binary encoding enables neural networks to extrapolate periodic functions.
Nestedness has traditionally been used to detect assembly patterns in meta-communities and networks of interacting species. Attempts have also been made to uncover nested structures in international trade, typically represented as bipartite networks in which connections can be established between countries (exporters o…
A common assumption about neural networks is that they can learn an appropriate internal representations on their own, see e.g. end-to-end learning. In this work we challenge this assumption. We consider two simple tasks and show that the state-of-the-art training algorithm fails, although the model itself is able to r…
To meet the Basel II regulatory requirements for the Advanced Measurement Approaches, the bank's internal model must include the use of internal data, relevant external data, scenario analysis and factors reflecting the business environment and internal control systems. Quantification of operational risk cannot be base…
Proposes a new model for clustering multiplex networks with compositional data.
Transformer models align words through attention weights, closely approximating Optimal Transport.
Deep learning is a form of machine learning for nonlinear high dimensional pattern matching and prediction. By taking a Bayesian probabilistic perspective, we provide a number of insights into more efficient algorithms for optimisation and hyper-parameter tuning. Traditional high-dimensional data reduction techniques, …
Study uses Google matrix analysis to show how COVID-19 changed international trade flows.
A fundamental question in deep learning concerns the role played by individual layers in a deep neural network (DNN) and the transferable properties of the data representations which they learn. To the extent that layers have clear roles, one should be able to optimize them separately using layer-wise loss functions. S…
Develops a new framework for temporal anchoring in deep embedding spaces.
A new method generates natural-looking adversarial examples by bounding internal activation values.
We propose a novel learning method for multilayered neural networks which uses feedforward supervisory signal and associates classification of a new input with that of pre-trained input. The proposed method effectively uses rich input information in the earlier layer for robust leaning and revising internal representat…
Study uses AI to predict changes in international public finances based on US markets.
A new method, InfoGuide, improves automatic clustering analysis.
To meet the Basel II regulatory requirements for the Advanced Measurement Approaches in operational risk, the bank's internal model should make use of the internal data, relevant external data, scenario analysis and factors reflecting the business environment and internal control systems. One of the unresolved challeng…
We characterize a prevalent weakness of deep neural networks (DNNs)---overthinking---which occurs when a DNN can reach correct predictions before its final layer. Overthinking is computationally wasteful, and it can also be destructive when, by the final layer, a correct prediction changes into a misclassification. Und…
Physarum-inspired solver speeds up LP solving in neural networks.
The challenge of assigning importance to individual neurons in a network is of interest when interpreting deep learning models. In recent work, Dhamdhere et al. proposed Total Conductance, a "natural refinement of Integrated Gradients" for attributing importance to internal neurons. Unfortunately, the authors found tha…
Modeling influenza spread using feature engineering and international flow deconvolution.
We study the topological properties of the multinetwork of commodity-specific trade relations among world countries over the 1992-2003 period, comparing them with those of the aggregate-trade network, known in the literature as the international-trade network (ITN). We show that link-weight distributions of commodity-s…
For Portugal there are few or none works about the international trade of fruits between Portugal and the other countries. In this work it aims to analyze the more recent data for the Portuguese international trade of fruits. They were used data for the years from 2006 to 2010, available by the INE (Statistics Portugal…
CodNN uses error-correcting codes to make neural networks more resilient to noise.
Vision Transformers show different internal representations compared to CNNs.
Softmax attention approximates complex functions and subsumes many known universal approximators.
Recently, there is increasing interest and research on the interpretability of machine learning models, for example how they transform and internally represent EEG signals in Brain-Computer Interface (BCI) applications. This can help to understand the limits of the model and how it may be improved, in addition to possi…
Tools of the theory of critical phenomena, namely the scaling analysis and universality, are argued to be applicable to large complex web-like network structures. Using a detailed analysis of the real data of the International Trade Network we argue that the scaled link weight distribution has an approximate log-normal…
Study quantifies how LLMs capture higher-order statistical structure using cumulant expansion.
Batch Normalization (BatchNorm) is a widely adopted technique that enables faster and more stable training of deep neural networks (DNNs). Despite its pervasiveness, the exact reasons for BatchNorm's effectiveness are still poorly understood. The popular belief is that this effectiveness stems from controlling the chan…
We develop a new theoretical framework to analyze the generalization error of deep learning, and derive a new fast learning rate for two representative algorithms: empirical risk minimization and Bayesian deep learning. The series of theoretical analyses of deep learning has revealed its high expressive power and unive…
Achieving international food security requires improved understanding of how international trade networks connect countries around the world through the import-export flows of food commodities. The properties of food trade networks are still poorly documented, especially from a multi-network perspective. In particular,…
Algorithm improves model performance on shifted concepts without retraining.
The increasing integration of world economies, which organize in complex multilayer networks of interactions, is one of the critical factors for the global propagation of economic crises. We adopt the network science approach to quantify shock propagation on the global trade-investment multiplex network. To this aim, w…
We analyse a multiplex of networks between OECD countries during the decade 2002-2010, which consists of five financial layers, given by foreign direct investment, equity securities, short-term, long-term and total debt securities, and five environmental layers, given by emissions of N O x, P M 10 SO 2, CO 2 equivalent…