Deep learning predicts availability of mobile crowdsourced services spatially and temporally.
problem Predicting the availability of mobile crowdsourced services in space and time.
method Two-stage prediction model: clustering services into regions, then forecasting availability duration using time series.
result Effectiveness validated through multiple experiments.
SHARE predicts city-wide parking availability using a hierarchical graph neural network.
problem Predicting city-wide parking availability is challenging due to spatial and temporal autocorrelation.
method SHARE uses a hierarchical graph convolution structure with contextual and soft clustering blocks, a recurrent neural network, and a parking availability approximation module.
result SHARE outperforms state-of-the-art baselines in predicting city-wide parking availability.
Models for recommender systems show similar results in item availability.
problem Model uncertainty in recommender systems.
method Examined different variations of a model for item availability using predictive multiplicity.
result Most models produce similar results in terms of item availability discrepancy.
A stability-based method selects the most desirable conformal prediction set.
problem Selecting the most desirable conformal prediction set from multiple valid sets invalidates coverage guarantees.
method A stability-based approach that ensures coverage for the selected prediction set.
result The stability-based approach maintains coverage guarantees for the selected prediction set.
Foundation models improve time series prediction reliability, especially with limited data.
problem Improving time series prediction reliability with limited data.
method Comparison of Time Series Foundation Models (TSFMs) with traditional methods in conformal prediction.
result TSFMs provide more reliable conformalized prediction intervals and more stable calibration with limited data.
Motivation: Prediction of ligands for proteins of known 3D structure is important to understand structure-function relationship, predict molecular function, or design new drugs. Results: We explore a new approach for ligand prediction in which binding pockets are represented by atom clouds. Each target pocket is compar…
Survey on ML advances for personalized prediction considering entity characteristics.
problem Sub-optimal performance in personalized prediction due to heterogeneity in data across entities.
method Organized current literature on entity-aware modeling based on characteristics availability and training data amount.
result Recent innovations in uncertainty quantification, fairness, and knowledge-guided machine learning can improve entity-aware modeling.
We have developed a strategy for the analysis of newly available binary data to improve outcome predictions based on existing data (binary or non-binary). Our strategy involves two modeling approaches for the newly available data, one combining binary covariate selection via LASSO with logistic regression and one based…
The paper introduces a knowledge score for GPR predictions to assess their reliability.
problem Uncertainty in probabilistic predictions from Gaussian process regression models.
method A knowledge score quantifying the reduction of uncertainty in GPR predictions.
result The knowledge score improves prediction accuracy in various tasks.
This study predicts parking availability using multi-source data and a self-supervised learning enhanced transformer.
problem Accurate parking availability prediction to support urban planning and management.
method Proposes SST-iTransformer, a self-supervised learning enhanced spatio-temporal inverted transformer, integrating multi-source data.
result SST-iTransformer achieves state-of-the-art performance in parking availability prediction.
We examine two fundamental tasks associated with graph representation learning: link prediction and semi-supervised node classification. We present a novel autoencoder architecture capable of learning a joint representation of both local graph structure and available node features for the multi-task learning of link pr…
Study improves engagement prediction in educational videos.
problem Addressing cold-start problem in educational recommenders.
method Introduced VLE dataset with content and engagement features, conducted experiments.
result VLE dataset leads to better engagement prediction models.
Two prediction models improve supply-demand forecasting for autonomous vehicles.
problem Improving accuracy and stability of supply-demand predictions for autonomous vehicles.
method Two prediction models based on residual network, LSTM, attention mechanism, and multi-attention mechanism.
result Our frameworks provide more accurate and stable prediction results than existing methods.
Researchers propose improved multivariate prediction models for HIV drug resistance.
problem Predicting drug resistances of HIV from mutation information.
method Revised stacking algorithms to borrow information among multiple prediction tasks.
result Proposed methods outperform other multivariate prediction methods.
Proposes a recursive MPC scheme with probabilistic safety guarantees for uncertain dynamic systems.
problem Probabilistic safety guarantees for MPC in dynamic environments with unknown stochastic agents.
method Uses conformal prediction to derive high-confidence prediction regions and gradually relax safety constraints online.
result Ensures recursive feasibility of MPC schemes by relaxing safety constraints over time.
Study shows cross-domain X-ray prediction performance discrepancies and label shifts.
problem Quantifying generalization limits across different X-ray datasets.
method Large-scale study on multiple X-ray datasets, focusing on performance and label shifts.
result Interesting discrepancies found between model performance and agreement, and concept similarity across tasks.
A new autoencoder learns graph representations for link prediction and node classification.
problem Link prediction and node classification on graph data.
method Multi-task graph autoencoder architecture.
result Significant improvement over three baselines on five graph datasets.
Study evaluates how limited training data affects streamflow predictions.
problem Limited historical meteorological and streamflow data affects streamflow prediction accuracy.
method Evaluated tree- and LSTM-based models on CAMELS dataset with varying training data sizes and time spans.
result Tree- and LSTM-based models provide similarly accurate predictions on small datasets, but LSTMs are superior with more training data.
A new method uses RF's out-of-bag errors for multiple imputation.
problem Missing data in biomedical studies and lack of prediction uncertainty.
method Constructs conditional distributions from the empirical distribution of out-of-bag prediction errors.
result Valid multiple imputation results achieved without parametric assumptions.
Model predicts multiple material properties with reduced error.
problem Limited materials data and lack of universal material descriptors.
method Integrates CGCNN with multi-task learning.
result Reduces test error by up to 8% for correlated properties.
Risk prediction is central to both clinical medicine and public health. While many machine learning models have been developed to predict mortality, they are rarely applied in the clinical literature, where classification tasks typically rely on logistic regression. One reason for this is that existing machine learning…
Use network embeddings to correct for unobserved confounding.
problem Causal inference in the presence of unobserved confounding.
method Use network embeddings to semi-supervised predict treatments and outcomes.
result Valid causal inferences under suitable conditions on predictive model quality.
MetaPred uses meta-learning to improve clinical risk prediction from limited EHR data.
problem Clinical risk prediction from sparse patient EHR data.
method Meta-learning approach to train a meta-learner from related tasks, then fine-tune for target risk prediction.
result MetaPred achieves better performance for target risk prediction with limited data.
XGBoost outperforms other models in predicting housing prices.
problem Accurate housing price prediction for socio-economic development.
method Employed XGBoost and other machine learning algorithms on housing price datasets.
result XGBoost outperformed other models in predicting housing prices.
Proposes a method to improve rare event prediction in healthcare.
problem Rare event classification in healthcare with low prevalence labels.
method Variational disentanglement approach to semi-parametric learning.
result Outperforms existing alternatives in mortality prediction on COVID-19 cohort.
This paper catalogs R packages for explaining AI models.
problem Ensuring safe and effective functioning of predictive models.
method Taxonomy of model explanations and comparison of R packages.
result 27 R packages for XAI identified and compared.
Radiomics aims to extract and analyze large numbers of quantitative features from medical images and is highly promising in staging, diagnosing, and predicting outcomes of cancer treatments. Nevertheless, several challenges need to be addressed to construct an optimal radiomics predictive model. First, the predictive p…
Enhanced travel time prediction using deep neural networks and road network information.
problem Improving travel time estimation using deep learning models.
method Proposes incorporating road network information into deep learning models for travel time prediction.
result Improved travel time prediction, especially with limited training data.
Deep learning predicts crop prices with improved accuracy.
problem Accurate prediction of agricultural crop prices for better decision-making.
method Innovative deep learning approach using GNNs and CNN models.
result At least 20% better performance than previous literature.
Anytime-valid confirmation of label-shift corrections
problem Small-batch scientific deployments with scarce labeled outcomes
method Conditional e-value and martingale-based rule
result Nonnegative martingale and anytime-valid confirmation rule
We propose that predictability is a prerequisite for profitability on financial markets. We look at ways to measure predictability of price changes using information theoretic approach and employ them on all historical data available for NYSE 100 stocks. This allows us to determine whether frequency of sampling price c…
Develops a test to assess if human experts add value to predictions.
problem Detecting if human expertise adds value to predictions.
method Statistical framework and hypothesis test to assess independence of expert predictions from outcomes.
result Physicians' decisions for AGIB patients incorporate information not available to a screening tool.
New models improve choice prediction accuracy.
problem Model misspecifications in discrete choice models lead to limited predictability and biased estimates.
method Proposes a new approach to estimate choice models by dividing the systematic part into knowledge-driven and data-driven components, which learns a new representation from available variables.
result The new models (L-MNL and L-NL) outperform traditional models in predictive performance and parameter estimation.
LSTM predicts COVID-19 growth patterns from global data.
problem Forecasting the spread of COVID-19 using big data and machine learning.
method Used multivariate long short-term memory (LSTM) to learn correlations over time.
result LSTM outperformed RNN in predicting COVID-19 growth with lower validation error.
Modeling shared mobility demand considering supply limitations.
problem Inaccurate demand predictions due to limited supply.
method Censored Gaussian Processes for demand modeling.
result Taking supply limitations into account improves demand predictions.
Quantified limits of nuclear stability beyond drip lines.
problem Predicting nuclear stability beyond known isotopes.
method Microscopic nuclear mass models, Bayesian methodology, Gaussian processes.
result Quantified predictions of one- and two-nucleon separation energies.
Predicts COVID-19 spread using dictionary learning and online NMF.
problem Limited daily case data for accurate prediction.
method Joint dictionary learning and online NMF for short evolution instances.
result Learned dictionary patterns improve predictions over time.
Deep learning models (aka Deep Neural Networks) have revolutionized many fields including computer vision, natural language processing, speech recognition, and is being increasingly used in clinical healthcare applications. However, few works exist which have benchmarked the performance of the deep learning models with…
AI virtual doctor predicts diabetes from non-invasive data.
problem Limited access to primary medical care in rural areas.
method Interactive AI system with speech recognition and synthesis, deep neural networks.
result System accurately predicts type 2 diabetes.
The paper proposes a method to integrate prior information into penalized regression.
problem Improving predictive performance in high-dimensional tasks with prior information.
method Integrating multiple sources of prior information into penalized regression.
result The method improves predictive performance, as shown by simulations and applications.
Agricultural research has been profited by technical advances such as automation, data mining. Today, data mining is used in a vast areas and many off-the-shelf data mining system products and domain specific data mining application soft wares are available, but data mining in agricultural soil datasets is a relatively…
Large dataset released for ITE and UM research.
problem Estimating causal impact of actions in various sectors.
method Release of a large dataset, formalization of UM, synthetic response surfaces, heterogeneous treatment assignment.
result Validation of ITE prediction and UM methods with high statistical significance.
PPI++ uses machine learning predictions to improve inference from small datasets.
problem Efficient inference from small labeled datasets with high-quality predictions.
method Adapts prediction-powered inference (PPI) to compute confidence sets for any parameter dimensionality.
result Improves classical intervals using only labeled data, always yielding better results.
Paper tackles multi-task learning for molecular property prediction with limited data.
problem Limited labeled data for each molecular property task in drug discovery.
method Proposes SGNN-EBM method to utilize relation graph between tasks and improve multi-task learning performance.
result Empirical results show the effectiveness of SGNN-EBM.
We introduce 'mixed LICORS', an algorithm for learning nonlinear, high-dimensional dynamics from spatio-temporal data, suitable for both prediction and simulation. Mixed LICORS extends the recent LICORS algorithm (Goerg and Shalizi, 2012) from hard clustering of predictive distributions to a non-parametric, EM-like sof…
A new method uses vector embeddings to improve analytics model performance.
problem Challenges in selecting high-quality datasets for enhanced analytics performance.
method Transform datasets into vector embeddings using NumTabData2Vec, then use similarity search for model inference.
result The proposed method accurately predicts analytics outcomes and increases speedup.
Theory explains how AI models can predict unseen tasks without labeled data.
problem Understanding how AI models can generalize to unseen tasks.
method Developed a theoretical framework to analyze zero-shot prediction.
result Identified key quantities and independence relationships for generalization.
Motivation: In a predictive modeling setting, if sufficient details of the system behavior are known, one can build and use a simulation for making predictions. When sufficient system details are not known, one typically turns to machine learning, which builds a black-box model of the system using a large dataset of in…