NPO method improves LLM unlearning without catastrophic collapse.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper introduces a method to make deep neural networks more robust to adversarial attacks.
New method uses machine learning to estimate sensitivity without binning.
Karate Club simplifies graph mining for unsupervised learning.
A new method calculates optimal decisions from classifier outputs, improving predictions in drug discovery.
Algorithms for hyperparameter optimization abound, all of which work well under different and often unverifiable assumptions. Motivated by the general challenge of sequentially choosing which algorithm to use, we study the more specific task of choosing among distributions to use for random hyperparameter optimization.…
Modeling curvature-sensitive cells in visual cortex using manifold geometry.
Tree Index evaluates cluster quality by creating decision trees from data.
Random investment strategies outperform sensible ones, even with forecasts.
New analysis identifies key factors in wildfire-generated thunderstorms.
This paper presents analytical solutions to the problem of how to calculate sensible VaR (Value-at-Risk) and ES (Expected Shortfall) contributions in the CreditRisk+ methodology. Via the ES contributions, ES itself can be exactly computed in finitely many steps. The methods are illustrated by numerical examples.
INNs produce interval-valued uncertainty scores for DNNs.
Unified model combines scores and rankings for grant panel review.
Proposes SDE framework for uncertainty quantification in graph neural networks.
The development of algorithms for hierarchical clustering has been hampered by a shortage of precise objective functions. To help address this situation, we introduce a simple cost function on hierarchies over a set of points, given pairwise similarities between those points. We show that this criterion behaves sensibl…
New findings show neural networks can be fooled by adversarial data, but a simple bias fix works.
Although a key driver of Earth's climate system, global land-atmosphere energy fluxes are poorly constrained. Here we use machine learning to merge energy flux measurements from FLUXNET eddy covariance towers with remote sensing and meteorological data to estimate net radiation, latent and sensible heat and their uncer…
Graham's formula simplifies stock valuation for growth stocks.
Word embeddings have demonstrated strong performance on NLP tasks. However, lack of interpretability and the unsupervised nature of word embeddings have limited their use within computational social science and digital humanities. We propose the use of informative priors to create interpretable and domain-informed dime…
Proposes a black-box attack to test clustering algorithms' robustness.
We propose definitions of fairness in machine learning and artificial intelligence systems that are informed by the framework of intersectionality, a critical lens arising from the Humanities literature which analyzes how interlocking systems of power and oppression affect individuals along overlapping dimensions inclu…
New method combines score lists using joint CDFs, improving computation.
We propose a principled method for gradient-based regularization of the critic of GAN-like models trained by adversarially optimizing the kernel of a Maximum Mean Discrepancy (MMD). We show that controlling the gradient of the critic is vital to having a sensible loss function, and devise a method to enforce exact, ana…
Conditional independence testing is a key problem required by many machine learning and statistics tools. In particular, it is one way of evaluating the usefulness of some features on a supervised prediction problem. We propose a novel conditional independence test in a predictive setting, and show that it achieves bet…
We reconsider the multivariate Kyle model in a risk-neutral setting with a single, perfectly informed rational insider and a rational competitive market maker, setting the price of n correlated securities. We prove the unicity of a symmetric, positive definite solution for the impact matrix and provide insights on its …
AdaNet is a lightweight TensorFlow-based (Abadi et al., 2015) framework for automatically learning high-quality ensembles with minimal expert intervention. Our framework is inspired by the AdaNet algorithm (Cortes et al., 2017) which learns the structure of a neural network as an ensemble of subnetworks. We designed it…
Structural Causal Models (SCMs) provide a popular causal modeling framework. In this work, we show that SCMs are not flexible enough to give a complete causal representation of dynamical systems at equilibrium. Instead, we propose a generalization of the notion of an SCM, that we call Causal Constraints Model (CCM), an…
The issue of model risk in default modeling has been known since inception of the Academic literature in the field. However, a rigorous treatment requires a description of all the possible models, and a measure of the distance between a single model and the alternatives, consistent with the applications. This is the pu…
In recent years, data have become increasingly higher dimensional and, therefore, an increased need has arisen for dimension reduction techniques for clustering. Although such techniques are firmly established in the literature for multivariate data, there is a relative paucity in the area of matrix variate, or three-w…
We consider an agent's uncertainty about its environment and the problem of generalizing this uncertainty across observations. Specifically, we focus on the problem of exploration in non-tabular reinforcement learning. Drawing inspiration from the intrinsic motivation literature, we use density models to measure uncert…
In this paper, we reproduce the experiments of Artetxe et al. (2018b) regarding the robust self-learning method for fully unsupervised cross-lingual mappings of word embeddings. We show that the reproduction of their method is indeed feasible with some minor assumptions. We further investigate the robustness of their m…
We propose a novel approach for the generation of polyphonic music based on LSTMs. We generate music in two steps. First, a chord LSTM predicts a chord progression based on a chord embedding. A second LSTM then generates polyphonic music from the predicted chord progression. The generated music sounds pleasing and harm…
One of the most ambitious use cases of computer-assisted learning is to build a recommendation system for lifelong learning. Most recommender algorithms exploit similarities between content and users, overseeing the necessity to leverage sensible learning trajectories for the learner. Lifelong learning thus presents un…
We present a new model for prediction markets, in which we use risk measures to model agents and introduce a market maker to describe the trading process. This specific choice on modelling tools brings us mathematical convenience. The analysis shows that the whole market effectively approaches a global objective, despi…
Purely data driven approaches for machine learning present difficulties when data is scarce relative to the complexity of the model or when the model is forced to extrapolate. On the other hand, purely mechanistic approaches need to identify and specify all the interactions in the problem at hand (which may not be feas…
As the loop space of a Riemannian manifold is infinite-dimensional, it is a non-trivial problem to make sense of the "top degree component" of a differential form on it. In this paper, we show that a formula from finite dimensions generalizes to assign a sensible "top degree component" to certain composite forms, obtai…
In this paper, we consider the problem of fair statistical inference involving outcome variables. Examples include classification and regression problems, and estimating treatment effects in randomized trials or observational data. The issue of fairness arises in such problems where some covariates or treatments are "s…
Many recent papers address reading comprehension, where examples consist of (question, passage, answer) tuples. Presumably, a model must combine information from both questions and passages to predict corresponding answers. However, despite intense interest in the topic, with hundreds of published papers vying for lead…
A common assumption in causal modeling posits that the data is generated by a set of independent mechanisms, and algorithms should aim to recover this structure. Standard unsupervised learning, however, is often concerned with training a single model to capture the overall distribution or aspects thereof. Inspired by c…
We present Meena, a multi-turn open-domain chatbot trained end-to-end on data mined and filtered from public domain social media conversations. This 2.6B parameter neural network is simply trained to minimize perplexity of the next token. We also propose a human evaluation metric called Sensibleness and Specificity Ave…
We discuss a simple extension of the Ho and Lee model with generic time-dependent drift in which: 1) we compute bond prices analytically; 2) the yield curve is sensible and the asymptotic yield is positive; and 3) our analytical solution provides a clean and simple way of separating volatility from the drift in the sho…
We propose a novel Shapley value approach to help address neural networks' interpretability and "vanishing gradient" problems. Our method is based on an accurate analytical approximation to the Shapley value of a neuron with ReLU activation. This analytical approximation admits a linear propagation of relevance across …
A motivating question in this paper is whether a sensible investment strategy may systematically contain long positions in out-of-the-money European calls with short expiry. Here we consider a very simple trading strategy for calls. The main points of this note are the following. First, the presented trading strategy a…
Testing for conditional independence is a core aspect of constraint-based causal discovery. Although commonly used tests are perfect in theory, they often fail to reject independence in practice, especially when conditioning on multiple variables. We focus on discrete data and propose a new test based on the notion of …
We introduce a notion of measuring scales for quantum abelian gauge systems. At each measuring scale a finite dimensional affine space stores information about the evaluation of the curvature on a discrete family of surfaces. Affine maps from the spaces assigned to finer scales to those assigned to coarser scales play …
Advanced inference techniques allow one to reconstruct the pattern of interaction from high dimensional data sets. We focus here on the statistical properties of inferred models and argue that inference procedures are likely to yield models which are close to a phase transition. On one side, we show that the reparamete…
Researchers extend the concept of metric spaces to Lorentzian spaces and prove the feasibility of their c-completion.
Proposes GDTW for aligning time series on different, incomparable spaces.