The paper tackles learning mixtures of two multinomial logits, showing identifiability and presenting an algorithm.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New algorithms reduce regret in reinforcement learning with MNL approximations.
Paper tackles combinatorial reinforcement learning with preference feedback.
New model optimizes assortment and pricing with dynamic customer arrivals.
We present a mixed multinomial logit (MNL) model, which leverages the truncated stick-breaking process representation of the Dirichlet process as a flexible nonparametric mixing distribution. The proposed model is a Dirichlet process mixture model and accommodates discrete representations of heterogeneity, like a laten…
Multinomial logit bandit is a sequential subset selection problem which arises in many applications. In each round, the player selects a -cardinality subset from candidate items, and receives a reward which is governed by a {\it multinomial logit} (MNL) choice model considering both item utility and substitution…
Improved online confidence bounds for multinomial logistic models in bandits.
Two algorithms optimize assortment selection for user choices in unknown MNL models.
New algorithms minimize risk in MNL bandits, achieving near-optimal performance.
Motivated by generating personalized recommendations using ordinal (or preference) data, we study the question of learning a mixture of MultiNomial Logit (MNL) model, a parameterized class of distributions over permutations, from partial ordinal or preference data (e.g. pair-wise comparisons). Despite its long standing…
Study MNL-Bandit in non-stationary settings with optimal regret bound.
New algorithms learn MNL weights efficiently for any slate size.
The paper achieves nearly optimal regret bounds for contextual multinomial logit bandits.
In this short note we consider a dynamic assortment planning problem under the capacitated multinomial logit (MNL) bandit model. We prove a tight lower bound on the accumulated regret that matches existing regret upper bounds for all parameters (time horizon , number of items and maximum assortment capacity )…
Improved regret bound for MNL MDPs with variance-aware approach.
Optimal design for multinomial logit models improves assortment selection efficiency.
New algorithm tackles non-linear utility in MNL bandits with regret.
Two algorithms achieve optimal regret with limited adaptivity in multinomial logistic bandits.
The paper models network formation using mixed logit models.
Study optimal product assortment using historical data, proving item coverage suffices.
DMNL bandits optimize assortment choices balancing relevance and diversity.
Travel providers such as airlines and on-line travel agents are becoming more and more interested in understanding how passengers choose among alternative itineraries when searching for flights. This knowledge helps them better display and adapt their offer, taking into account market conditions and customer needs. Som…
The geometric non-linear Schrodinger equation (GNLS) on the complex Grassmannian manifold M is the Hamiltonian equation for the energy functional on C(R,M) with respect to the symplectic form induced from the Kahler form on M. It has a Lax pair that is gauge equivalent to the Lax pair of the matrix non-linear Schroding…
We study the dynamic assortment planning problem, where for each arriving customer, the seller offers an assortment of substitutable products and customer makes the purchase among offered products according to an uncapacitated multinomial logit (MNL) model. Since all the utility parameters of MNL are unknown, the selle…
In discrete choice modeling (DCM), model misspecifications may lead to limited predictability and biased parameter estimates. In this paper, we propose a new approach for estimating choice models in which we divide the systematic part of the utility specification into (i) a knowledge-driven part, and (ii) a data-driven…
New algorithm for maximizing revenue in multinomial logistic bandits.
When tracking user-specific online activities, each user's preference is revealed in the form of choices and comparisons. For example, a user's purchase history is a record of her choices, i.e. which item was chosen among a subset of offerings. A user's preferences can be observed either explicitly as in movie ratings …
CRS model improves ranking data modeling with theoretical guarantees.
New algorithm reduces reinforcement learning regret by adapting to interaction variability.
Paper presents a privacy-preserving method for dynamic assortment selection.
This paper explores the adaptive (active) PAC (probably approximately correct) top- ranking (i.e., top- item selection) and total ranking problems from -wise () comparisons under the multinomial logit (MNL) model. By adaptively choosing sets to query and observing the noisy output of the most favored …
New algorithm tackles dynamic assortment optimization with knapsack constraints.
New algorithm reduces regret in dynamic assortment selection.
New model improves website ranking by considering user choices as a whole.
In this paper, we study the dynamic assortment optimization problem under a finite selling season of length . At each time period, the seller offers an arriving customer an assortment of substitutable products under a cardinality constraint, and the customer makes the purchase among offered products according to a d…
The question of aggregating pair-wise comparisons to obtain a global ranking over a collection of objects has been of interest for a very long time: be it ranking of online gamers (e.g. MSR's TrueSkill system) and chess players, aggregating social opinions, or deciding which product to sell based on transactions. In mo…
The Multinomial Logit (MNL) model and the axiom it satisfies, the Independence of Irrelevant Alternatives (IIA), are together the most widely used tools of discrete choice. The MNL model serves as the workhorse model for a variety of fields, but is also widely criticized, with a large body of experimental literature cl…
As datasets capturing human choices grow in richness and scale -- particularly in online domains -- there is an increasing need for choice models that escape traditional choice-theoretic axioms such as regularity, stochastic transitivity, and Luce's choice axiom. In this work we introduce the Pairwise Choice Markov Cha…
Many applications in preference learning assume that decisions come from the maximization of a stable utility function. Yet a large experimental literature shows that individual choices and judgements can be affected by "irrelevant" aspects of the context in which they are made. An important class of such contexts is t…
Assortment optimization is an important problem that arises in many industries such as retailing and online advertising where the goal is to find a subset of products from a universe of substitutable products which maximize seller's expected revenue. One of the key challenges in this problem is to model the customer su…
New algorithms reduce matching regret by limiting frequent updates.
A new model uses neural networks for consistent discrete choice analysis.
Study optimizes dynamic product selection and pricing using censored preference feedback.
We consider the problem of multi-product dynamic pricing, in a contextual setting, for a seller of differentiated products. In this environment, the customers arrive over time and products are described by high-dimensional feature vectors. Each customer chooses a product according to the widely used Multinomial Logit (…
Graph neural networks improve residential location choice predictions.
We consider the dynamic assortment optimization problem under the multinomial logit model (MNL) with unknown utility parameters. The main question investigated in this paper is model mis-specification under the -contamination model, which is a fundamental model in robust statistics and machine learning. In…
The paper tackles interpreting DCM with image data by addressing data isomorphism.
Algorithm stabilizes queues in asymmetric systems with unknown service rates.