Logit-GFN accelerates GFlowNets training by scaling logits based on temperature.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New method distills cloud models into edge-friendly ones.
Logit distance bounds representational similarity of models.
Logit models are usually applied when studying individual travel behavior, i.e., to predict travel mode choice and to gain behavioral insights on traveler preferences. Recently, some studies have applied machine learning to model travel mode choice and reported higher out-of-sample predictive accuracy than traditional …
Logit regularization induces logit clustering, affecting classifier performance.
The paper calibrates uncertainty in dropout variational inference models.
We consider neural network training, in applications in which there are many possible classes, but at test-time, the task is a binary classification task of determining whether the given example belongs to a specific class, where the class of interest can be different each time the classifier is applied. For instance, …
Improves adversarial robustness by constraining logits with a bounded function.
A new logit model derived from the Weibull manifold.
Proposes a new model for context-dependent decision-making.
Proposes SOVR loss to improve adversarial robustness by increasing logit margins.
Auxiliary Tuning adapts pre-trained models for novel tasks efficiently.
Generative classifiers have been shown promising to detect illegal inputs including adversarial examples and out-of-distribution samples. Supervised Deep Infomax~(SDIM) is a scalable end-to-end framework to learn generative classifiers. In this paper, we propose a modification of SDIM termed SDIM-\emph{logit}. Instead …
In this paper, we develop improved techniques for defending against adversarial examples at scale. First, we implement the state of the art version of adversarial training at unprecedented scale on ImageNet and investigate whether it remains effective in this setting - an important open scientific question (Athalye et …
Logit dynamics formula reveals self-regulation in softmax policy gradient methods.
Proposes a convex model for mixed logit to handle individual heterogeneity.
Recently, Kannan et al. [2018] proposed several logit regularization methods to improve the adversarial robustness of classifiers. We show that the computationally fast methods they propose - Clean Logit Pairing (CLP) and Logit Squeezing (LSQ) - just make the gradient-based optimization problem of crafting adversarial …
MANO normalizes logits to estimate test accuracy without labels.
New method uses low logit rank to simplify complex language models.
LAWN normalizes logits to improve deep network adaptability and generalization.
New algorithm reduces switching costs in multinomial logit bandit problems.
We evaluate the robustness of Adversarial Logit Pairing, a recently proposed defense against adversarial examples. We find that a network trained with Adversarial Logit Pairing achieves 0.6% accuracy in the threat model in which the defense is considered. We provide a brief overview of the defense and the threat models…
Unified framework suppresses model bias in semi-supervised learning with decoupled sampling control.
ELM improves neural model embeddings for long-tail learning.
Detecting adversarial examples currently stands as one of the biggest challenges in the field of deep learning. Adversarial attacks, which produce adversarial examples, increase the prediction likelihood of a target class for a particular data point. During this process, the adversarial example can be further optimized…
Paper tackles long-tailed labels in classification problems.
Deep learning classifiers are known to be vulnerable to adversarial examples. A recent paper presented at ICML 2019 proposed a statistical test detection method based on the observation that logits of noisy adversarial examples are biased toward the true class. The method is evaluated on CIFAR-10 dataset and is shown t…
The standard Gibbs sampler of Mixed Multinomial Logit (MMNL) models involves sampling from conditional densities of utility parameters using Metropolis-Hastings (MH) algorithm due to unavailability of conjugate prior for logit kernel. To address this non-conjugacy concern, we propose the application of Pólygamma data a…
The paper proposes an efficient method to scale Bayesian inference for mixed multinomial logit models to very large datasets.
This work improves interpretability and calibration of complex-valued neural networks using Newton-Puiseux analysis.
Logit correction improves model performance by correcting spurious correlations.
This article presents a proof of the existence of Bertrand-Nash equilibrium prices with multi-product firms and under the Logit model of demand that does not rely on restrictive assumptions on product characteristics, firm homogeneity or symmetry, product costs, or linearity of the utility function. The proof is based …
In this paper, we study the assortment optimization problem faced by many online retailers such as Amazon. We develop a \emph{cascade multinomial logit model}, based on the classic multinomial logit model, to capture the consumers' purchasing behavior across multiple stages. Different from existing studies, our model a…
The paper models network formation using mixed logit models.
In discrete choice modeling (DCM), model misspecifications may lead to limited predictability and biased parameter estimates. In this paper, we propose a new approach for estimating choice models in which we divide the systematic part of the utility specification into (i) a knowledge-driven part, and (ii) a data-driven…
This paper provides a theoretical and computational justification of the long held claim that of the similarity of the probit and logit link functions often used in binary classification. Despite this widespread recognition of the strong similarities between these two link functions, very few (if any) researchers have …
Analyzes how inclusion/exclusion from STOXX Europe 600 Index affects company prices.
Unified framework for critical scaling of inverse temperature in self-attention.
A new network model combines features of DCBM, LSM, and β-model, using a cancellation trick for parameter estimation.
Paper predicts M&A deal success using ML and DL techniques.
As datasets capturing human choices grow in richness and scale -- particularly in online domains -- there is an increasing need for choice models that escape traditional choice-theoretic axioms such as regularity, stochastic transitivity, and Luce's choice axiom. In this work we introduce the Pairwise Choice Markov Cha…
SLED improves factuality in LLMs without external knowledge.
Generating and eliminating adversarial examples has been an intriguing topic in the field of deep learning. While previous research verified that adversarial attacks are often fragile and can be defended via image-level processing, it remains unclear how high-level features are perturbed by such attacks. We investigate…
The recent success of generative adversarial networks and variational learning suggests training a classifier network may work well in addressing the classical two-sample problem. Network-based tests have the computational advantage that the algorithm scales to large samples. This paper proposes a two-sample statistic …
This paper explains how low-precision arithmetic causes loss spikes in deep learning models.
LLMs can be influenced by unseen dataset subtexts, revealing new ways to select data subsets.
Optimal design for multinomial logit models improves assortment selection efficiency.
Graph neural networks improve residential location choice predictions.