Many applications in preference learning assume that decisions come from the maximization of a stable utility function. Yet a large experimental literature shows that individual choices and judgements can be affected by "irrelevant" aspects of the context in which they are made. An important class of such contexts is t…
New method discovers context effects in choice data.
problem Identifying context effects from choice data is challenging.
method Automatic discovery of context effects from observed choices.
result Automatic discovery of context effects from observed choices.
Develops deep learning models for choice modeling.
problem Computational intractability and sample inefficiency in existing choice model learning methods.
method Deep learning-based choice models in two settings: feature-free and feature-based.
result Demonstrates improved recovery of existing choice models and reduced sample complexity.
This paper uses Bayesian ARD to automatically determine utility functions for discrete choice models.
problem Challenging and time-consuming task in identifying optimal utility function specifications.
method Bayesian framework and automatic relevance determination (ARD) for data-driven utility function specification.
result The proposed DCM-ARD model accurately recovers true utility function specifications and outperforms previous methods.
Paper characterizes MDM for consumer choice modeling and prediction.
problem Modeling consumer choice behavior with parsimonious models.
method Establishes necessary and sufficient conditions for MDM consistency.
result Characterization leads to exact set of representable choice probabilities.
Optimizes choice sets to influence group decisions.
problem Maximizing agreement or disagreement in group decisions.
method Discrete choice modeling to develop optimization framework.
result Promoting a choice can be easier than encouraging consensus or discord.
Proposes new methods for Markov chain choice models with panel data.
problem Dependence among transactions for the same customer in historical data.
method Expectation-maximization (EM) algorithms incorporating partial-ordering preference information.
result EM algorithms outperform traditional methods on synthetic and real datasets.
Bayesian optimisation framework for multi-objective decision-making from choice data.
problem Optimizing multi-objective functions via choice judgements.
method Gaussian process prior and novel likelihood model for choice data.
result Proposes a novel Bayesian framework for learning latent functions from choice data.
Choice models, which capture popular preferences over objects of interest, play a key role in making decisions whose eventual outcome is impacted by human choice behavior. In most scenarios, the choice model, which can effectively be viewed as a distribution over permutations, must be learned from observed data. The ob…
Neural networks approximate random utility models for choice prediction.
problem Approximating random utility models with neural networks.
method RUMnets, a neural network-based model inspired by RUM framework.
result RUMnets can approximate any RUM model arbitrarily closely and vice versa.
Revealed preference theory studies the possibility of modeling an agent's revealed preferences and the construction of a consistent utility function. However, modeling agent's choices over preference orderings is not always practical and demands strong assumptions on human rationality and data-acquisition abilities. Th…
Study binary choice with asymmetric loss, offering simple solutions.
problem Binary choice with asymmetric loss in data-rich environments.
method Loss-based reweighting of logistic regression or machine learning techniques.
result Valid decisions on binary outcomes with general loss functions.
This paper explores the impact of metric choice on Fréchet regression.
problem Choosing the right metric for Fréchet regression in complex data.
method Review and extensive numerical studies of existing dimension reduction methods.
result Different metrics significantly affect the estimation of central and central mean space.
Ranking data arises in a wide variety of application areas but remains difficult to model, learn from, and predict. Datasets often exhibit multimodality, intransitivity, or incomplete rankings---particularly when generated by humans---yet popular probabilistic models are often too rigid to capture such complexities. In…
Active learning recovers choice model from noisy data.
problem Identifying non-parametric choice models from noisy data.
method Directed acyclic graph (DAG) representation and inclusion-exclusion approach.
result Algorithm more accurately recovers frequent preferences.
Bayesian methods detect significant IIA violations in similarity choice data.
problem Detecting IIA violations in similarity choice data complicates classical models.
method Proposed two statistical methods: classical goodness-of-fit test and Bayesian PPC.
result Significant IIA violations confirmed in both datasets, driven by context effects.
Binary choice forests model customer choices in retailing.
problem Estimating DCMs using transaction data is challenging and prone to misspecification.
method Random forest of binary decision trees to represent DCMs, interpretable and consistent predictions.
result Random forest can predict choice probabilities and assortments unseen in training data.
We introduce sparse random projection, an important dimension-reduction tool from machine learning, for the estimation of discrete-choice models with high-dimensional choice sets. Initially, high-dimensional data are compressed into a lower-dimensional Euclidean space using random projections. Subsequently, estimation …
Simultaneously estimates travel times and route choice model parameters.
problem Interdependent estimation of arc travel times and route choice model parameters.
method Maximum likelihood estimation for any differentiable route choice model.
result Strong performance in real-world data, even compared to arc travel time estimation methods.
Bayesian DL model improves DCMs for better predictive and inferential performance.
problem Limited interpretability and predictive underperformance of traditional DCMs.
method Integrates deep learning with approximate Bayesian inference (SGLD).
result Improves predictive and inferential metrics in discrete choice models.
Random Machines improves SVM performance with free kernel choice.
problem Efficiency and accuracy in solving classification and regression problems.
method Bagged-weighted support vector model with free kernel choice.
result Improved accuracy and reduced computational time.
A distributed method for Bayesian model choice using marginal likelihood and Monte Carlo sampling.
problem Bayesian model choice in large datasets with limited communication.
method Split data into subsets, locally compute model evidence, combine results using summary statistics.
result The method enables model choice in large datasets with speed-ups and theoretical error bounds.
Graph neural nets improve discrete choice modeling with network effects.
problem Modeling network effects in discrete choice problems.
method Graph Convolutional Neural Network (GCNN) architecture.
result Higher predictive performance than standard models with interpretability.
New framework predicts choice with set-related invariances.
problem Accurately predicting choice behavior in large-scale data.
method A learning framework that captures set-related invariances, derived from economics.
result Demonstrated utility on three large choice datasets.
A new method reduces high-dimensional state space for dynamic choice models.
problem Estimation of dynamic discrete choice models is computationally intensive and infeasible in high-dimensional settings.
method Recursive partitioning algorithm to reduce dimensionality of high-dimensional state space.
result Our method reduces estimation bias and makes estimation feasible.
Diffusion models adapt to low-dimensional data regardless of coefficient choices.
problem Understanding how diffusion models adapt to low-dimensional data structures.
method Analysis of diffusion models with flexible coefficient choices.
result Proven that O ~ ( k / ε ) \widetilde{O}(k/\varepsilon) O ( k / ε ) iterations suffice for accurate sampling in total variation distance. Study presents MMC model for better fitting multiple choice data.
problem Improving accuracy of latent trait estimates in IRT models.
method Fit autoencoders to MMC model, demonstrating better fit than nominal response model.
result MMC model outperforms traditional IRT models in fit.
Study improves choice model accuracy and heterogeneity representation using mixture models.
problem Improving prediction accuracy and heterogeneity representation in choice models.
method Semi-nonparametric Latent Class Choice Model with mixture models and EM algorithm.
result Mixture models enhance prediction accuracy and heterogeneity representation without sacrificing interpretability.
Paper extends top-k Mallows model for better user preference analysis.
problem Capturing real-world user preferences focusing on a limited set of items.
method Generalized top-k Mallows model, novel sampling scheme, efficient algorithm, active learning.
result New tools for analysis and prediction in decision-making scenarios.
Proposes robust assortment optimization from observational data.
problem Real-world scenarios often violate assumptions of stable customer preferences and correct choice models.
method Develops a robust framework that accounts for potential distributional shifts in customer choice behavior.
result Uncovered the notion of ``robust item-wise coverage'' as the minimal data requirement for sample-efficient robust assortment learning.
Paper uses stats to predict treatment choice based on illness probability.
problem Improving treatment decision-making in personalized medicine.
method Statistical decision theory with maximum regret evaluation.
result Estimates illness probability for better treatment choice.
The study examines how experimental design choices affect machine learning model performance.
problem Lack of guidelines on choosing experimental designs and machine learning models.
method 12 experimental designs, 7 families of predictive models, 7 test functions, 8 noise settings.
result Guidelines for practical applications of DOE and ML are provided.
Paper introduces Functional Effects Models to account for individual heterogeneity in panel data.
problem Accounting for preference heterogeneity in panel data with machine learning.
method Functional Effects Models using gradient boosting decision trees and deep neural networks to learn individual-specific preference parameters.
result Functional Effects Models outperform traditional models in learning inter-individual heterogeneity and predictive performance.
Paper improves anomaly detection by using non-uniform random choices in isolation forests.
problem Detecting clustered diverse outliers more effectively.
method Comparing different split guiding criteria in isolation forests.
result Non-uniform random choices improve outlier discrimination for certain outlier classes.
New research on Shapley values for feature attribution in machine learning, considering model vs. data fidelity.
problem Controversy in connecting machine learning models to coalitional games, differing approaches.
method Investigates two approaches: interventional vs. observational conditional expectation Shapley values for linear models.
result The choice between model and data fidelity depends on the specific application.
We present a novel method for obtaining high-quality, domain-targeted multiple choice questions from crowd workers. Generating these questions can be difficult without trading away originality, relevance or diversity in the answer options. Our method addresses these problems by leveraging a large corpus of domain-speci…
While a user's preference is directly reflected in the interactive choice process between her and the recommender, this wealth of information was not fully exploited for learning recommender models. In particular, existing collaborative filtering (CF) approaches take into account only the binary events of user actions …
New statistical models for predicting ranked preferences from partial orders.
problem Statistical models overlook information in list length.
method Composite and augmented ranking models for joint modeling of partial orders and list lengths.
result Augmented ranking models best predict both length and preferences.
Inherent risk scoring is an important function in anti-money laundering, used for determining the riskiness of an individual during onboarding before \textit{before} before fraudulent transactions occur. It is, however, often fraught with two challenges: (1) inconsistent notions of what constitutes as high or low risk by experts a…
Study improves posterior inference in neural processes with limited data.
problem Improving posterior predictive inference in probabilistic models with scarce conditioning data.
method Examined effects of pooling operators and variational families on posterior quality in neural processes.
result Novel neural process architectures lead to superior posterior predictive samples in image completion/in-painting tasks.
RCPO uses ranked choice modeling for better LLM alignment.
problem Pairwise preference optimization limits LLM alignment.
method Unified framework combining preference optimization and ranked choice modeling.
result RCPO outperforms competitive baselines in LLM alignment.
The paper tackles interpreting DCM with image data by addressing data isomorphism.
problem Interpreting DCM with image data due to isomorphic information.
method Proposes and benchmarks two methodologies: architectural adjustments and data source mitigation.
result Direct data source mitigation is more effective for maintaining DCM's interpretability.
Proposes a new model to predict irrational customer behavior.
problem Irrational customer behavior in decision-making.
method Nonparametric choice model using decision trees and probability distributions.
result Decision forest model accurately predicts non-rational customer behavior.
Discrete choice models are commonly used by applied statisticians in numerous fields, such as marketing, economics, finance, and operations research. When agents in discrete choice models are assumed to have differing preferences, exact inference is often intractable. Markov chain Monte Carlo techniques make approximat…
CRS model improves ranking data modeling with theoretical guarantees.
problem Lack of rich, multimodal models for ranking data.
method Contextual Repeated Selection (CRS) model for multimodal ranking data.
result CRS model significantly outperforms existing methods in various ranking contexts.
A new model uses neural networks for consistent discrete choice analysis.
problem Difficulties in specifying utility functions in RUM models.
method Alternative-Specific and Shared weights Neural Network (ASS-NN) model.
result ASS-NN provides consistent outcomes without specifying utility form.
Many applied settings in empirical economics involve simultaneous estimation of a large number of parameters. In particular, applied economists are often interested in estimating the effects of many-valued treatments (like teacher effects or location effects), treatment effects for many groups, and prediction models wi…
Gaussians as noise in NCE lead to exponentially bad conditioning, hindering its efficiency.
problem Exponential conditioning of Hessian in NCE with Gaussian noise.
method Using Gaussian as the noise distribution in NCE.
result Gaussian noise in NCE leads to exponentially bad conditioning of the loss Hessian.