Evaluates transfer learning methods in dynamic data availability scenarios.
problem Real-world data availability varies over time, leading to unrealistic TL method evaluations.
method Proposes a data manipulation framework to simulate varying data availability and domain transformations.
result Demonstrates the usefulness of the framework on proprietary and publicly available datasets.
Interactive book integrates probability models with big data analytics.
problem Lack of integration between classical loss data models and modern analytic tools.
method Combines classical loss data models with modern analytic tools and big data.
result Promotes deeper learning through interactive elements and multiple language support.
Batch Reinforcement Learning (RL) algorithms attempt to choose a policy from a designer-provided class of policies given a fixed set of training data. Choosing the policy which maximizes an estimate of return often leads to over-fitting when only limited data is available, due to the size of the policy class in relatio…
Health care is one of the most exciting frontiers in data mining and machine learning. Successful adoption of electronic health records (EHRs) created an explosion in digital clinical data available for analysis, but progress in machine learning for healthcare research has been difficult to measure because of the absen…
Unified framework uses all data to improve multiple testing efficiency.
problem Improving predictive uncertainty control in decision-making.
method Uses all available data (null, alternative, unlabelled) for score construction and calibration.
result Significantly improves power and adaptability across diverse scenarios.
This study predicts parking availability using multi-source data and a self-supervised learning enhanced transformer.
problem Accurate parking availability prediction to support urban planning and management.
method Proposes SST-iTransformer, a self-supervised learning enhanced spatio-temporal inverted transformer, integrating multi-source data.
result SST-iTransformer achieves state-of-the-art performance in parking availability prediction.
DKT transfers biomarker information between neurodegenerative diseases.
problem Estimating biomarker trajectories in rare neurodegenerative diseases with limited data.
method DKT is a joint-disease generative model that transfers biomarker progressions from common neurodegenerative diseases to rare ones.
result DKT estimates plausible multimodal biomarker trajectories in rare diseases like PCA using only unimodal data.
New attacks prevent both supervised and contrastive learning from private data.
problem Preventing unauthorized use of private data and commercial datasets.
method Contrastive-like data augmentations in supervised error minimization or maximization frameworks.
result Achieve state-of-the-art worst-case unlearnability across SL and CL algorithms.
Accurate real-time tracking of influenza outbreaks helps public health officials make timely and meaningful decisions that could save lives. We propose an influenza tracking model, ARGO (AutoRegression with GOogle search data), that uses publicly available online search data. In addition to having a rigorous statistica…
Foundation models improve time series prediction reliability, especially with limited data.
problem Improving time series prediction reliability with limited data.
method Comparison of Time Series Foundation Models (TSFMs) with traditional methods in conformal prediction.
result TSFMs provide more reliable conformalized prediction intervals and more stable calibration with limited data.
Sourcerer uses deep learning to map land cover from limited labeled data.
problem Producing accurate land cover maps with scarce labeled data.
method Bayesian-inspired, deep learning approach with a novel regularizer.
result Sourcerer outperforms other methods, even with minimal labeled target data.
A new parallel clustering method improves speed and accuracy for single cell transcriptomic data.
problem Challenges in clustering single cell transcriptomic data, including poor quality, lack of prior knowledge, and slow computation.
method Parallel Split Merge Sampling on Dirichlet Process Mixture Model (Para-DPMM).
result The Para-DPMM model outperforms existing methods in clustering quality and computational speed.
Combining data from multiple speakers improves neural TTS quality, especially with imbalanced data.
problem Training high-quality TTS systems with imbalanced speaker data.
method Combine data from multiple speakers, train multi-speaker models, and use ensemble methods.
result Ensemble multi-speaker models improve synthetic speech quality for underrepresented speakers.
Protein-protein interaction (PPI) prediction is an important problem in machine learning and computational biology. However, there is no data set for training or evaluation purposes, where all the instances are accurately labeled. Instead, what is available are instances of positive class (with possibly noisy labels) a…
Agricultural research has been profited by technical advances such as automation, data mining. Today, data mining is used in a vast areas and many off-the-shelf data mining system products and domain specific data mining application soft wares are available, but data mining in agricultural soil datasets is a relatively…
Private release of sensitive data enables fair learning.
problem Learning fair predictors with restricted sensitive data.
method Private release of sensitive demographic data, adapting non-discriminatory learners.
result The approach provides theoretical guarantees on performance for fair predictors.
Extracts StarCraft II tournament data for AI and ML studies.
problem Lack of accessible esports data for scientific use.
method Gathered and processed StarCraft II tournament replays using an API parser library.
result The largest publicly available StarCraft II esports dataset.
Method estimates causal effects from incremental data, overcoming missing data challenges.
problem Estimating causal effects from non-stationary, incrementally available observational data.
method Continual Causal Effect Representation Learning
result Method achieves continual causal effect estimation without compromising original data.
Study compares new audio representation methods for limited data music retrieval.
problem Improving machine learning for audio data with limited training data.
method Investigated mel-spectrogram and Mel scattering representations, and augmented target loss function.
result All proposed methods outperform standard mel-spectrogram when using limited data.
VICatMix clusters categorical biomedical data efficiently and selects relevant variables.
problem Efficient clustering of high-dimensional categorical biomedical data.
method Variational Bayesian finite mixture model with variational inference.
result Improves clustering accuracy and variable selection on noisy, high-dimensional data.
A new method uses vector embeddings to improve analytics model performance.
problem Challenges in selecting high-quality datasets for enhanced analytics performance.
method Transform datasets into vector embeddings using NumTabData2Vec, then use similarity search for model inference.
result The proposed method accurately predicts analytics outcomes and increases speedup.
This paper reviews sentiment analysis on Indian languages.
problem Understanding sentiment in multilingual web data.
method Reviews and discusses approaches for sentiment analysis on Indian languages.
result Challenges in sentiment analysis on indigenous languages.
Torchmeta simplifies meta-learning evaluation across multiple datasets.
problem Inconsistent evaluation of meta-learning algorithms across different datasets.
method Introduces a library that provides data-loaders for standard benchmarks and simplifies model compatibility.
result Seamless and consistent evaluation of meta-learning algorithms on multiple datasets.
Assessing systemic risk in financial markets is of great importance but it often requires data that are unavailable or available at a very low frequency. For this reason, systemic risk assessment with partial information is potentially very useful for regulators and other stakeholders. In this paper we consider systemi…
MapLUR uses deep learning on map images to estimate NO2 pollution, outperforming traditional methods.
problem Limited availability of data for traditional LUR models makes them hard to adapt to new areas.
method Data-driven, open-source approach using convolutional neural networks trained on map data.
result MapLUR significantly outperforms traditional LUR models, including those with manually engineered features.
Federated learning algorithm improves with intermittent client availability.
problem Performance degradation in Federated Averaging due to client availability changes.
method Federated Latest Averaging (FedLaAvg) uses latest gradients from all clients, even when unavailable.
result FedLaAvg achieves sublinear speedup compared to classical Federated Averaging.
Develops workflow for synthetic insurance datasets.
problem Lack of realistic publicly available insurance datasets.
method Uses CTGAN neural network architecture to generate tabular data.
result Synthesized datasets evaluated positively in multiple aspects.
Human trafficking is among the most challenging law enforcement problems which demands persistent fight against from all over the globe. In this study, we leverage readily available data from the website "Backpage"-- used for classified advertisement-- to discern potential patterns of human trafficking activities which…
A new method controls risk for set predictors using cross-validation.
problem Inefficient set predictors when data limited.
method Cross-validation conformal risk control (CV-CRC).
result CV-CRC offers theoretical guarantees and reduces set size.
New method uses synthetic data to validate financial agent classification.
problem Validation of machine learning methods for financial agent classification.
method Agent-based model to generate synthetic data for validation.
result Unsupervised clustering may give incorrect results for financial agents.
Deep learning improves singing processing tasks.
problem Lack of data and computing resources for singing processing.
method State-of-the-art deep learning techniques.
result Advances in accuracy and sound quality.
New algorithms discover and utilize 'voids' in data to improve machine learning models.
problem Improving machine learning models by considering the unknown aspects of data.
method Developed algorithms to discover and utilize 'voids' in data, creating ignorance-aware prototypes.
result Improved performance of nearest neighbor classifiers through ignorance-aware prototype selection.
Proposes a method to forecast spatial-temporal data with limited training data.
problem Forecasting with nodes having no temporal training data.
method Temporal data augmentation and spatial graph topology learning.
result Improves forecasting performance on nodes without training data.
Bayesian model predicts patient survival from sparse EHR data.
problem Analyzing EHR data with few samples and diverse information.
method Nonparametric probabilistic model using Bayesian trees.
result Improved survival trajectory predictions on patient data.
Thanks to the growing availability of spoofing databases and rapid advances in using them, systems for detecting voice spoofing attacks are becoming more and more capable, and error rates close to zero are being reached for the ASVspoof2015 database. However, speech synthesis and voice conversion paradigms that are not…
Theoretical limits show experimental data can falsify but not validate causal estimates from observational studies.
problem Fundamental limits on validating causal estimates using experimental data in observational studies.
method Impossible inference framework, Gaussian Process based approach.
result Experimental data can falsify but not validate causal estimates from observational studies.
PyHHMM is a Python library for HHMMs with advanced features.
problem Handling heterogeneous observation models and missing data in HMMs.
method Object-oriented Python implementation with advanced features.
result PyHHMM supports a heterogeneous observation model and missing data inference.
Modern machine learning systems such as image classifiers rely heavily on large scale data sets for training. Such data sets are costly to create, thus in practice a small number of freely available, open source data sets are widely used. We suggest that examining the geo-diversity of open data sets is critical before …
Study shows publicly available news impacts financial markets.
problem Impact of publicly available news on financial markets.
method Extracted news from Common Crawl, identified relevant companies, used sentiment analysis and information theory.
result Publicly available news has significant impact on financial markets.
We have developed a strategy for the analysis of newly available binary data to improve outcome predictions based on existing data (binary or non-binary). Our strategy involves two modeling approaches for the newly available data, one combining binary covariate selection via LASSO with logistic regression and one based…
The task of clustering a set of objects based on multiple sources of data arises in several modern applications. We propose an integrative statistical model that permits a separate clustering of the objects for each data source. These separate clusterings adhere loosely to an overall consensus clustering, and hence the…
Unified method for multi-defect microscopy image restoration with limited training data.
problem Challenges in applying deep learning methods due to limited training data for multi-defect microscopy images.
method Two-stage approach: data augmentation with GAN and conditional GAN training.
result Proposed method gives comparable or superior results to existing methods in image quality restoration.
In this paper we propose a framework for automated forecasting of energy-related time series using open access data from European Network of Transmission System Operators for Electricity (ENTSO-E). The framework provides forecasts for various European countries using publicly available historical data only. Our solutio…
The paper uses CNNs on Sentinel-2 imagery to assess landslide risks.
problem Landslide risk assessment and prediction.
method Image augmentation, 3-D CNNs, satellite imagery.
result CNNs achieve significantly better accuracy than baseline.
Improved GAN performance with incomplete data using factorised discriminators.
problem Limited availability of labelled data for GAN training.
method Factorising data distribution into sub-distributions and training sub-discriminators.
result Improved performance in image generation, segmentation, and audio separation tasks.
Study develops ML emulators for generator models from terminal bus data.
problem Reconstruct generator models from terminal bus measurements.
method Used machine learning techniques, including VAR and LSTM models.
result Established trade-offs between linear AR and powerful LSTM models.
Federated Cox model handles non-proportional hazards in siloed data.
problem Handling non-proportional hazards in federated healthcare data.
method Developed a federated Cox model that relaxes proportional hazards assumption.
result Federated model performs similarly to standard models on clinical datasets.
Paper improves geographic location embeddings using Flickr tags and structured data.
problem Lack of integration between Flickr metadata and structured scientific data.
method Learning vector space embeddings of geographic locations.
result Improved predictions of ecological features using the new method.