Paper proposes a new method to measure model sensitivity using final model only.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In complex transfer learning scenarios new tasks might not be tightly linked to previous tasks. Approaches that transfer information contained only in the final parameters of a source model will therefore struggle. Instead, transfer learning at a higher level of abstraction is needed. We propose Leap, a framework that …
New attacks reveal membership in label-only ML models.
A new method for distilling predictions from a teacher model to a student model without original training data.
New findings on plane waves in 3D spacetimes, showing non-unimodular elliptic plane waves are unique.
Softmax is a standard final layer used in Neural Nets (NNs) to summarize information encoded in the trained NN and return a prediction. However, Softmax leverages only a subset of the class-specific structure encoded in the trained model and ignores potentially valuable information: During training, models encode an ar…
We introduce a new class of graphical models that generalizes Lauritzen-Wermuth-Frydenberg chain graphs by relaxing the semi-directed acyclity constraint so that only directed cycles are forbidden. Moreover, up to two edges are allowed between any pair of nodes. Specifically, we present local, pairwise and global Marko…
Analyzes word2vec-like models revealing linear subspaces learned during training.
Attribute-aware CF models aims at rating prediction given not only the historical rating from users to items, but also the information associated with users (e.g. age), items (e.g. price), or even ratings (e.g. rating time). This paper surveys works in the past decade developing attribute-aware CF systems, and discover…
New attacks can infer model training membership using only label predictions, not confidence.
Convolutional neural networks are commonly used to control the steering angle for autonomous cars. Most of the time, multiple long range cameras are used to generate lateral failure cases. In this paper we present a novel model to generate this data and label augmentation using only one short range fisheye camera. We p…
Out of nearly 70,000 bills introduced in the U.S. Congress from 2001 to 2015, only 2,513 were enacted. We developed a machine learning approach to forecasting the probability that any bill will become law. Starting in 2001 with the 107th Congress, we trained models on data from previous Congresses, predicted all bills …
We show that many spin 6-manifolds have the homotopy type but not the homeomorphism type of a Kaehler manifold. Moreover, for given Betti numbers, there are only finitely many deformation types and hence topological types of smooth complex projectve spin threefolds of general type. Finally, on a fixed spin 6-manifold, …
Sound source separation has attracted attention from Music Information Retrieval(MIR) researchers, since it is related to many MIR tasks such as automatic lyric transcription, singer identification, and voice conversion. In this paper, we propose an intuitive spectrogram-based model for source separation by adapting U-…
We address the problem of graph classification based only on structural information. Inspired by natural language processing techniques (NLP), our model sequentially embeds information to estimate class membership probabilities. Besides, we experiment with NLP-like variational regularization techniques, making the mode…
We present a simple proof for the universality of invariant and equivariant tensorized graph neural networks. Our approach considers a restricted intermediate hypothetical model named Graph Homomorphism Model to reach the universality conclusions including an open case for higher-order output. We find that our proposed…
The paper distinguishes between conditional and marginal processes in language models and discusses conditions for usefulness.
We study in detail and explicitly solve the version of Kyle's model introduced in a specific case in \cite{BB}, where the trading horizon is given by an exponentially distributed random time. The first part of the paper is devoted to the analysis of time-homogeneous equilibria using tools from the theory of one-dimensi…
Framework certifies fairness of machine learning models interactively and privately.
There are many surprising and perhaps counter-intuitive properties of optimization of deep neural networks. We propose and experimentally verify a unified phenomenological model of the loss landscape that incorporates many of them. High dimensionality plays a key role in our model. Our core idea is to model the loss la…
This paper revisits the Bayesian CMA-ES and provides updates for normal Wishart. It emphasizes the difference between a normal and normal inverse Wishart prior. After some computation, we prove that the only difference relies surprisingly in the expected covariance. We prove that the expected covariance should be lower…
We consider a document classification problem where document labels are absent but only relevant keywords of a target class and unlabeled documents are given. Although heuristic methods based on pseudo-labeling have been considered, theoretical understanding of this problem has still been limited. Moreover, previous me…
In treatment allocation problems the individuals to be treated often arrive sequentially. We study a problem in which the policy maker is not only interested in the expected cumulative welfare but is also concerned about the uncertainty/risk of the treatment outcomes. At the outset, the total number of treatment assign…
Localized Multidirectional Correction improves non-refusal target-response behavior in foundation models.
For any strictly positive martingale for which has a characteristic function, we provide an expansion for the implied volatility. This expansion is explicit in the sense that it involves no integrals, but only polynomials in the log strike. We illustrate the versatility of our expansion by computing t…
We propose a new analytical method to study stochastic, binary-state models on complex networks. Moving beyond the usual mean-field theories, this alternative approach is based on the introduction of an annealed approximation for uncorrelated networks, allowing to deal with the network structure as parametric heterogen…
Non-equilibrium phenomena occur not only in physical world, but also in finance. In this work, stochastic relaxational dynamics (together with path integrals) is applied to option pricing theory. A recently proposed model (by Ilinski et al.) considers fluctuations around this equilibrium state by introducing a relaxati…
Improves model robustness to shifts in subpopulations.
We found an easy and quick post-learning method named "Icing on the Cake" to enhance a classification performance in deep learning. The method is that we train only the final classifier again after an ordinary training is done.
Fraud detection is a difficult problem that can benefit from predictive modeling. However, the verification of a prediction is challenging; for a single insurance policy, the model only provides a prediction score. We present a case study where we reflect on different instance-level model explanation techniques to aid …
Federated learning is a method of training models on private data distributed over multiple devices. To keep device data private, the global model is trained by only communicating parameters and updates which poses scalability challenges for large models. To this end, we propose a new federated learning algorithm that …
Survey of methods for learning from observation without requiring expert actions.
Federated learning enables the creation of a powerful centralized model without compromising data privacy of multiple participants. While successful, it does not incorporate the case where each participant independently designs its own model. Due to intellectual property concerns and heterogeneous nature of tasks and d…
Simple model-based reinforcement learning outperforms model-free methods in complex tasks.
Twitter has been proven to be a notable source for predictive modelling on various domains such as the stock market, the dissemination of diseases or sports outcomes. However, such a study has not been conducted in football (soccer) so far. The purpose of this research was to study whether data mined from Twitter can b…
Paper compares RL models for finance, finding Reward Clipping best.
Regularization is typically understood as improving generalization by altering the landscape of local extrema to which the model eventually converges. Deep neural networks (DNNs), however, challenge this view: We show that removing regularization after an initial transient period has little effect on generalization, ev…
Exposure bias refers to the train-test discrepancy that seemingly arises when an autoregressive generative model uses only ground-truth contexts at training time but generated ones at test time. We separate the contributions of the model and the learning framework to clarify the debate on consequences and review propos…
In this paper we study the distributional properties of a vector of lifetimes in which each lifetime is modeled as the first arrival time between an idiosyncratic shock and a common systemic shock. Despite unlike the classical multidimensional Marshall-Olkin model here only a unique common shock affecting all the lifet…
As demonstrated during the recent financial crisis, regulators require additional analytical tools to assess systemic risk in the financial sector. This paper describes one such tool; namely a novel market modeling and analysis capability. Our model builds upon two leading market models: one which emphasizes market mic…
State-space models win a forecasting competition for unstable data.
Rigidity of Wasserstein spaces over Riemannian manifolds
We present the theory of tensors with Young tableau symmetry as an efficient computational tool in dealing with the polynomial first integrals of a natural system in classical mechanics. We relate a special kind of such first integrals, already studied by Lundmark, to Beltrami's theorem about projectively flat Riemanni…
Paper proposes a robust test for high-dimensional models with large covariates and instruments.
Single neural network predicts ImageNet model parameters for faster training.
Understanding how users navigate in a network is of high interest in many applications. We consider a setting where only aggregate node-level traffic is observed and tackle the task of learning edge transition probabilities. We cast it as a preference learning problem, and we study a model where choices follow Luce's a…
Economic systems, traditionally analyzed as almost independent national systems, are increasingly connected on a global scale. Only recently becoming available, the World Input-Output Database (WIOD) is one of the first efforts to construct the multi-regional input-output (MRIO) tables at the global level. By viewing t…
Improves dialogue response model interpretability using attention and regularization.