Paper proposes a combined model for better recommendation by integrating explicit and implicit feedbacks.
problem Improve recommendation accuracy by considering both explicit and implicit feedbacks.
method Developed three models (RHC-PMF, RV-PMF, RHCV-PMF) that incorporate users' explicit and implicit feedbacks for better rating prediction.
result RHCV-PMF model outperforms other models in cold start scenarios for both users and items.
User preferences for items can be inferred from either explicit feedback, such as item ratings, or implicit feedback, such as rental histories. Research in collaborative filtering has concentrated on explicit feedback, resulting in the development of accurate and scalable models. However, since explicit feedback is oft…
Agents learn user preferences with less explicit feedback via spatial interface valuing.
problem Learning user preferences with high cognitive load feedback.
method Spatial Interface Valuing for reduced explicit feedback.
result Agents learn faster with spatial interface valuing compared to explicit feedback.
NCAE tackles collaborative filtering for both explicit and implicit feedback.
problem Existing models struggle with both explicit and implicit feedback, overfitting, and lack of deep learning potential.
method NCAE uses a neural collaborative autoencoder with a three-stage pre-training mechanism and error reweighting.
result NCAE significantly outperforms state-of-the-art models on real-world datasets.
Improves NMT with user feedback from eBay ratings and search tasks.
problem Improving neural machine translation quality with user feedback.
method Offline bandit learning of NMT parameters using real user feedback from eBay.
result Implicit task-based feedback from cross-lingual search tasks improves NMT quality.
A new learning method for prosthetic arms without explicit rewards.
problem Learning a prosthetic arm to interact with users without explicit reward signals.
method Interaction-Grounded Learning, observing multidimensional context and feedback vectors, discovering latent reward signal.
result The algorithm can discover a latent reward signal and ground its policies for successful interaction.
New insights into cascade feedback linearization of control systems.
problem Obtaining a cascade feedback linearization for invariant control systems.
method Introducing truncated versions of operators from the calculus of variations to prove new theorems.
result Established new geometry and foundational theorems for future work.
Paper introduces consumed item packs for better recommendation.
problem Personalizing web content using implicit feedback.
method Introduces consumed item packs (CIP) to link users/items based on consumption behavior.
result CIP-U, CIP-I, DEEPCIP, and FISM provide competitive recommendation quality.
New algorithms reduce regret in online learning with partial feedback.
problem Reducing regret in online learning with partial feedback.
method Combining chaining and auction theory, designed algorithms for different feedback models.
result Improved regret bounds for semi-Lipschitz losses and second-price auctions.
The abstract explores connections between reinforcement learning, scaling, and diffusion.
problem Aligning reinforcement learning with human feedback and scaling techniques.
method Clarifying connections between reinforcement learning, scaling, and diffusion.
result Introducing a resampling approach for alignment and reward-directed diffusion models.
The paper analyzes matrix completion with unlabeled implicit feedback and provides error bounds.
problem Matrix completion with shared low-rank ground truth and sampling distribution.
method Combining subspace recovery theory and matrix completion bounds.
result Error bounds showing contributions from estimating the sampling distribution and ground truth.
New method for online learning IC models with node-level feedback.
problem Learning IC models with node-level feedback in social networks.
method Detailed analysis and online algorithm with O ( T ) \mathcal{O}( \sqrt{T}) O ( T ) cumulative regret. result First confidence-region result and online algorithm for IC models with node-level feedback.
New algorithm allows IGL to work with action-inclusive feedback.
problem IGL's failure in scenarios with action-inclusive feedback.
method Developed an algorithm and provided theoretical guarantees.
result Demonstrated effectiveness on large-scale experiments.
Reduces user feedback needed for accurate recommender systems.
problem Limited user feedback in recommender systems.
method Partial Bandit and Semi-Bandit approach for efficient user feedback retrieval.
result Similar global accuracy and learning efficiency with reduced feedback.
Proposes CNN-based analog CSI feedback for FDD MIMO-OFDM systems.
problem High CSI feedback overhead in FDD MIMO systems.
method AnalogDeepCMC: maps downlink CSI to uplink channel input, reconstructs channel estimate.
result Significantly improves downlink spectral efficiency and simplifies operation.
Robot learns user preferences from brain signals.
problem Decoding user preferences for robot motions from brain signals.
method Proposes a novel approach using electroencephalography to decode user preferences from brain signals.
result Brain signals can reliably infer user preferences for robot trajectories.
Develops a new model for RLHF accounting for partially observed states and intermediate feedback.
problem Lack of models for partially observed states and intermediate feedback in RLHF.
method PORRL model with cardinal and dueling feedback methods.
result Demonstrates improved learning and alignment with new model-based and model-free methods.
Families of explicit solutions are found to a nonlinear Black-Scholes equation which incorporates the feedback-effect of a large trader in case of market illiquidity. The typical solution of these families will have a payoff which approximates a strangle. These solutions were used to test numerical schemes for solving …
New oracle uses uncertainty for active classification with noisy feedback.
problem Improving query complexity in interactive binary classifier learning.
method Proposes a new pairwise comparison oracle that considers uncertainty and an adaptive labeling algorithm.
result Demonstrates improved performance and efficiency compared to existing methods.
Proposes a new theoretical framework for PbRL that requires less human feedback.
problem Lack of theoretical work capturing practical PbRL frameworks.
method Introduces a reward-agnostic PbRL framework that acquires exploratory trajectories before human feedback.
result Demonstrates improved sample complexity for learning optimal policies in linear and low-rank MDPs.
Hybrid approach combines user feedback and machine learning for predicting user satisfaction.
problem Measuring user satisfaction in large-scale conversational agent systems.
method Fusion of explicit user feedback and predictions from two machine-learned models trained on different data types.
result Hybrid approach significantly improves user satisfaction predictions.
Model analyzes OTC market making with reputation feedback.
problem Optimizing electronic OTC liquidity provision considering reputation.
method Developed a stochastic-control model with feedback loops.
result Policy alternates between reputation-building and franchise monetization phases.
Model analyzes how reputation feedback affects OTC market making.
problem Understanding and optimizing OTC market making strategies.
method Developed a stochastic-control model with feedback loops.
result Policy alternates between reputation building and franchise monetization.
Improved item recommendation using VAEs with user-dependent priors and text feedback.
problem Improving recommendation quality by integrating user ratings and text feedback.
method Extended VAEs to incorporate user-dependent priors in a multimodal latent space.
result Model outperforms existing VAE models for collaborative filtering (up to 29.41% relative improvement).
ITAL uses mutual information for active learning in image retrieval.
problem Acquiring meaningful user feedback for content-based image retrieval.
method Information-Theoretic Active Learning (ITAL) maximizing mutual information between predicted relevance and user feedback.
result ITAL achieves state-of-the-art performance across various datasets.
Efficient algorithm for learning from indirect feedback in complex decision-making scenarios.
problem Learning from indirect feedback in realistic scenarios with personalized mechanisms.
method IGW algorithm for policy optimization, extending reward-estimator construction from single-step to multi-step.
result Achieves sublinear regret guarantee for contextual episodic MDPs with personalized feedback.
SetRank tackles collaborative ranking from implicit feedback using setwise Bayesian approach.
problem Challenges in pairwise and listwise approaches for implicit feedback.
method SetRank is a novel setwise Bayesian approach that accommodates implicit feedback characteristics.
result SetRank outperforms state-of-the-art baselines on real-world datasets.
SCR diversifies recommendations by applying style transfer to user profiles.
problem Diversify personalized recommendations without losing relevance.
method Style injection using Conditional Variational Autoencoder (CVAE) architecture.
result 12% improvement in NDCG@20 and 22% improvement in AUC across all classes.
New RLHF algorithm identifies optimal policies from human feedback without explicit reward inference.
problem Training large language models with human feedback without reward inference.
method Model-free RLHF algorithm B S A D \mathsf{BSAD} BSAD that identifies optimal policies directly from human preference. result Provable, instance-dependent sample complexity i l d e O ( c M S A 3 H 3 M log 1 δ ) ilde{\mathcal{O}}(c_{\mathcal{M}}SA^3H^3M\log\frac{1}δ) i l d e O ( c M S A 3 H 3 M log δ 1 ) . This paper extends a Kyle model to include price-responsive traders, revealing new dynamics and equilibria.
problem Real-world market dynamics involve price-responsive traders, affecting market equilibrium and insider profits.
method Developed a continuous-time Kyle model with two types of price-responsive traders (momentum and contrarian), leading to a forward-backward Riccati system for equilibrium.
result The model shows that feedback effects can lead to multiple equilibria and amplify price informativeness.
QS-BO optimizes functions using only rank-based feedback.
problem Optimizing expensive functions with unreliable or unavailable metric values.
method Quantile-scaling pipeline to convert ranks into Gaussian targets.
result QS-BO consistently achieves lower objective values and is statistically significant.
A new algorithm for conversational recommendation systems using dueling bandits in GLMs.
problem Limited user feedback in existing conversational bandit methods.
method Integrates dueling bandits with relative feedback in generalized linear models.
result Theoretical and empirical validation of ConDuel's efficacy.
Algorithm identifies best arm in linked bandits with reduced feedback.
problem Best arm identification in linked bandits with reduced feedback.
method Combines uniform sampling with regular bandit algorithm.
result Almost matching upper and lower bounds on sample complexity.
Proposes a meta-learning method to improve recommender systems with biased feedback.
problem Learning from biased feedback in recommender systems.
method Asymmetric tri-training framework for meta-learning, using three predictors.
result Minimizes the upper bound of true performance metric, improving robustness to selection bias.
AutoStan improves Bayesian models via predictive feedback.
problem Improving Bayesian models written in Stan.
method Iterative improvement of Stan models using NLPD and sampler diagnostics feedback.
result AutoStan can autonomously improve diverse Bayesian models across various structures.
The paper constructs optimal hedging strategies for options with price impact.
problem Optimal hedging strategies for options with temporary price impact.
method Combining analytic and probabilistic tools to establish feedback representation of the optimal strategy and derive utility indifference price.
result Explicit asymptotic expansion of utility indifference price quantifying price impact.
We consider a model of optimal investment and consumption with both habit formation and partial observations in incomplete Itô processes market. The investor chooses his consumption under the addictive habits constraint while only observing the market stock prices but not the instantaneous rate of return. Applying the …
Paper tackles offline preference-based RL with human feedback.
problem Offline Preference-based Reinforcement Learning with preference feedback.
method Two-step approach: MLE for reward estimation and distributionally robust planning.
result First guarantee for learning any target policy with polynomial samples.
The paper tackles PAC ranking with subset-wise feedback, achieving optimal sample complexity.
problem Probably Approximately Correct (PAC) ranking of items with subset-wise preference feedback.
method Adaptive subset-wise preference feedback, Plackett-Luce model, pivot trick for score estimates.
result Achieves optimal sample complexity for PAC ranking with subset-wise feedback.
Study controlled contagion with state-dependent killing, proving a comparison principle.
problem Analyzing controlled McKean--Vlasov contagion with state-dependent killing.
method Proof of a comparison principle using Wasserstein smooth-gauge comparison and killing-jump absorption estimates.
result Established a comparison principle for the two-population killed-particle HJB.
Plug-in method improves performative prediction accuracy.
problem Learning under performative feedback with slow convergence rates.
method Plug-in performative optimization using models.
result Plug-in method can be superior to model-agnostic strategies.
Algorithm maximizes influence in networks with limited feedback.
problem Maximizing influence in networks with limited feedback.
method Bandit algorithm using local node degree observations.
result Local observations are sufficient for maximizing global influence.
Local PBO methods improve preferential BO in high-dimensional problems.
problem Efficiently optimizing preferential BO in high-dimensional settings.
method Adapting high-dimensional BO techniques to preferential feedback.
result Local PBO methods reduce cumulative regret compared to global baselines.
Solves infinite horizon portfolio problem with path-dependent labor income.
problem Infinite horizon portfolio choice with path-dependent labor income.
method Solves an infinite dimensional stochastic optimal control problem using explicit solutions to the HJB equation.
result Explicit solutions to the optimal controls in feedback form are found.
Study learns optimal bidding strategy in auctions with dynamic values and aggregated feedback.
problem Optimizing bidding in auctions with time-dependent values and limited feedback.
method Combines plug-in estimators with differential-equation characterization of optimal policy.
result Achieves near optimal regret bounds for learning optimal policy.
The paper enhances a virtual assistant's humor to improve user satisfaction.
problem Improving a virtual assistant's ability to deliver humorous responses.
method Combines traditional NLP techniques with self-attentional networks and multi-task learning, using implicit feedback for labeling.
result Deep-learning models outperform heuristic methods in real-world user satisfaction.
Advanced and effective collaborative filtering methods based on explicit feedback assume that unknown ratings do not follow the same model as the observed ones (\emph{not missing at random}). In this work, we build on this assumption, and introduce a novel dynamic matrix factorization framework that allows to set an ex…
Quantum-enhanced metrology aims to estimate an unknown parameter such that the precision scales better than the shot-noise bound. Single-shot adaptive quantum-enhanced metrology (AQEM) is a promising approach that uses feedback to tweak the quantum process according to previous measurement outcomes. Techniques and form…