Smart phone apps that enable users to easily track their diets have become widespread in the last decade. This has created an opportunity to discover new insights into obesity and weight loss by analyzing the eating habits of the users of such apps. In this paper, we present diet2vec: an approach to modeling latent str…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
DIET tests conditional independence using marginal dependence measures of residual information.
Social media provide a platform for users to express their opinions and share information. Understanding public health opinions on social media, such as Twitter, offers a unique approach to characterizing common health issues such as diabetes, diet, exercise, and obesity (DDEO), however, collecting and analyzing a larg…
DIET-SNN optimizes SNNs for faster, lower-energy image classification.
We propose a network independent, hand-held system to translate and disambiguate foreign restaurant menu items in real-time. The system is based on the use of a portable multimedia device, such as a smartphones or a PDA. An accurate and fast translation is obtained using a Machine Translation engine and a context-speci…
Social media based digital epidemiology has the potential to support faster response and deeper understanding of public health related threats. This study proposes a new framework to analyze unstructured health related textual data via Twitter users' post (tweets) to characterize the negative health sentiments and non-…
New method estimates treatment-response curves with covariate and timing measurement errors.
Differential Calculus is a staple of the college mathematics major's diet. Eventually one becomes tired of the same routine, and wishes for a more diverse meal. The college math major may seek to generalize applications of the derivative that involve functions of more than one variable, and thus enjoy a course on Multi…
Social media analytics allows us to extract, analyze, and establish semantic from user-generated contents in social media platforms. This study utilized a mixed method including a three-step process of data collection, topic modeling, and data annotation for recognizing exercise related patterns. Based on the findings,…
We consider the problem of making machine translation more robust to character-level variation at the source side, such as typos. Existing methods achieve greater coverage by applying subword models such as byte-pair encoding (BPE) and character-level encoders, but these methods are highly sensitive to spelling mistake…
A recipe recommendation system suggests missing ingredients using collaborative filtering.
New algorithm improves causal discovery in biomedical data.
We present Vector-Space Markov Random Fields (VS-MRFs), a novel class of undirected graphical models where each variable can belong to an arbitrary vector space. VS-MRFs generalize a recent line of work on scalar-valued, uni-parameter exponential family and mixed graphical models, thereby greatly broadening the class o…
Machine learning predicts obesity causes using genetic and imaging data.
Tree-Query uses LLMs to discover causal relationships in a transparent, interpretable manner.
Enhances content moderation with culturally-aware models.
Proposes a convex method to estimate GGMs with covariates.
Enhanced latent spaces improve collider simulation precision.
Gaussian Process model improves blood glucose prediction using contextual data.
New framework maximizes perturbed samples for inverse classification with budget constraints.
The paper compares two methods for handling missing data in causal discovery.
Lottery tickets find good initializations for IMP with sparse training.
Raman spectroscopy's capability to provide meaningful composition predictions is heavily reliant on a pre-processing step to remove insignificant spectral variation. This is crucial in biofluid analysis. Widespread adoption of diagnostics using Raman requires a robust model which can withstand routine spectra discrepan…
Proposes a new approach to MSDA by introducing latent covariate shift to handle varying label distributions.
Learning tasks such as those involving genomic data often poses a serious challenge: the number of input features can be orders of magnitude larger than the number of training examples, making it difficult to avoid overfitting, even when using the known regularization techniques. We focus here on tasks in which the inp…
CAST predicts distribution-valued time series by stabilizing and transporting simplex-supported successors.
Paper proposes a method to create smaller, more efficient deep generative audio models.