Hotel2vec learns hotel embeddings from multiple data sources.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New methods explain NE embeddings by identifying key variables.
One of the first things to do while planning a trip is to book a good place to stay. Booking a hotel online can be an overwhelming task with thousands of hotels to choose from, for every destination. Motivated by the importance of these situations, we decided to work on the task of recommending hotels to users. We used…
Recognizing a hotel from an image of a hotel room is important for human trafficking investigations. Images directly link victims to places and can help verify where victims have been trafficked, and where their traffickers might move them or others in the future. Recognizing the hotel from images is challenging becaus…
PriceAggregator optimizes hotel price fetching to increase Agoda's bookings.
In this paper, we present a real-world conversational AI system to search for and book hotels through text messaging. Our architecture consists of a frame-based dialogue management system, which calls machine learning models for intent classification, named entity recognition, and information retrieval subtasks. Our ch…
Many-to-one RNN predicts user hotel clicks from browsing history.
Study analyzes Hotelling-type tensor deflation for spiked tensors, providing insights into signal and noise.
Study suggests variable selection may not significantly reduce power in multivariate tests.
This paper analyzes how errors accumulate in PCA's deflation method.
Proposes RTL model for sentiment classification and key word detection in online reviews.
Researchers solved a model of an exhaustible resource with stochastic discoveries.
H. Hotelling proved that in the n-dimensional Euclidean or spherical space, the volume of a tube of small radius about a curve depends only on the length of the curve and the radius. A. Gray and L. Vanhecke extended Hotelling's theorem to rank one symmetric spaces computing the volumes of the tubes explicitly in these …
The growth of the modern knowledge-based economy is becoming less and less dependent on tangible assets and more on intangible ones. In this context, the role of human capital in the value creation process has become central. Despite the large amount of scientific work on human capital phenomena, little research has re…
Canonical correlation analysis was proposed by Hotelling [6] and it measures linear relationship between two multidimensional variables. In high dimensional setting, the classical canonical correlation analysis breaks down. We propose a sparse canonical correlation analysis by adding l1 constraints on the canonical vec…
Study analyzes accuracy of tensor deflation in noisy conditions.
New algorithm speeds up fair clustering by 12x.
When data analysts train a classifier and check if its accuracy is significantly different from chance, they are implicitly performing a two-sample test. We investigate the statistical properties of this flexible approach in the high-dimensional setting. We prove two results that hold for all classifiers in any dimensi…
Our society has been computerised and globalised due to emergence and spread of information and communication technology (ICT). This enables us to investigate our own socio-economic systems based on large amounts of data on human activities. In this article, methods of treating complexity arising from a vast amount of …
It is widely accepted that optimization of medical imaging system performance should be guided by task-based measures of image quality (IQ). Task-based measures of IQ quantify the ability of an observer to perform a specific task such as detection or estimation of a signal (e.g., a tumor). For binary signal detection t…
SIMPLE method quantifies uncertainty in network membership profiles.
The paper analyzes deflation for estimating a low-rank spike in large tensors with noise.
We introduce a semi-supervised discrete choice model to calibrate discrete choice models when relatively few requests have both choice sets and stated preferences but the majority only have the choice sets. Two classic semi-supervised learning algorithms, the expectation maximization algorithm and the cluster-and-label…
A method uses CG to create efficient channels for ideal observers.
Recently, an extension of independent component analysis (ICA) from one to multiple datasets, termed independent vector analysis (IVA), has been the subject of significant research interest. IVA has also been shown to be a generalization of Hotelling's canonical correlation analysis. In this paper, we provide the ident…
The asymptotic distribution of the Markowitz portfolio is derived, for the general case (assuming fourth moments of returns exist), and for the case of multivariate normal returns. The derivation allows for inference which is robust to heteroskedasticity and autocorrelation of moments up to order four. As a side effect…
We consider the hypothesis testing problem of detecting a shift between the means of two multivariate normal distributions in the high-dimensional setting, allowing for the data dimension p to exceed the sample size n. Specifically, we propose a new test statistic for the two-sample test of means that integrates a rand…
Statistical tests that compare classification algorithms are univariate and use a single performance measure, e.g., misclassification error, measure, AUC, and so on. In multivariate tests, comparison is done using multiple measures simultaneously. For example, error is the sum of false positives and false negatives…
The paper identifies universal features for high-dimensional data inference.
Paper presents a new framework for optimal asset and signal combination.
Paper presents a privacy-preserving method for dynamic assortment selection.
A new algorithm improves efficiency and robustness of heuristic optimization in simulation-based problems.
Explicit Taylor series for the volume of tubes in Lie groups
It is common in modern prediction problems for many predictor variables to be counts of rarely occurring events. This leads to design matrices in which many columns are highly sparse. The challenge posed by such "rare features" has received little attention despite its prevalence in diverse areas, ranging from natural …
Paper proposes a new KPI for early fault detection in hydropower plants.
Study optimizes pricing under uncertainty and capacity constraints.
New method detects bearing faults using multivariate statistical process control.
A new RL method improves revenue management with delayed feedback.
AI classifies tourist events for better service.
Paper tackles dynamic assortment with dual contexts, improving revenue in e-commerce.
Generative model learns object variability from MRI measurements.
FRONT optimizes decisions with interference, reducing regret over time.
This paper considers the problem of canonical-correlation analysis (CCA) (Hotelling, 1936) and, more broadly, the generalized eigenvector problem for a pair of symmetric matrices. These are two fundamental problems in data analysis and scientific computing with numerous applications in machine learning and statistics (…
We study a general problem of allocating limited resources to heterogeneous customers over time under model uncertainty. Each type of customer can be serviced using different actions, each of which stochastically consumes some combination of resources, and returns different rewards for the resources consumed. We consid…
Nonparametric two sample testing is a decision theoretic problem that involves identifying differences between two random variables without making parametric assumptions about their underlying distributions. We refer to the most common settings as mean difference alternatives (MDA), for testing differences only in firs…
Generative model for hypergraph clustering improves detection of higher-order structure.
Support Vector Data Description (SVDD) is a machine learning technique used for single class classification and outlier detection. SVDD based K-chart was first introduced by Sun and Tsung for monitoring multivariate processes when underlying distribution of process parameters or quality characteristics depart from Norm…
Embeddings are ubiquitous in machine learning, appearing in recommender systems, NLP, and many other applications. Researchers and developers often need to explore the properties of a specific embedding, and one way to analyze embeddings is to visualize them. We present the Embedding Projector, a tool for interactive v…