This paper reviews transfer learning for financial data predictions, highlighting its potential.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Differences in data size per class, also known as imbalanced data distribution, have become a common problem affecting data quality. Big Data scenarios pose a new challenge to traditional imbalanced classification algorithms, since they are not prepared to work with such amount of data. Split data strategies and lack o…
Study examines challenges and applications of machine learning in finance.
A new boosting method reduces overfitting and negative transfer in transfer learning.
Activation functions influence behavior and performance of DNNs. Nonlinear activation functions, like Rectified Linear Units (ReLU), Exponential Linear Units (ELU) and Scaled Exponential Linear Units (SELU), outperform the linear counterparts. However, selecting an appropriate activation function is a challenging probl…
Rubin LSST DESC uses AI/ML for dark energy research.
Modern electronic health records (EHRs) provide data to answer clinically meaningful questions. The growing data in EHRs makes healthcare ripe for the use of machine learning. However, learning in a clinical setting presents unique challenges that complicate the use of common machine learning methodologies. For example…
This paper critiques flawed MVTS anomaly detection evaluation methods and proposes a simple baseline.
Ranking recommendation algorithms across datasets using Bradley-Terry model
Many complex systems can be represented as networks, and the problem of network comparison is becoming increasingly relevant. There are many techniques for network comparison, from simply comparing network summary statistics to sophisticated but computationally costly alignment-based approaches. Yet it remains challeng…
Introduces CCR for constructing confidence regions from conformal predictions.
Since the advent of the horseshoe priors for regularization, global-local shrinkage methods have proved to be a fertile ground for the development of Bayesian methodology in machine learning, specifically for high-dimensional regression and classification problems. They have achieved remarkable success in computation, …
Incremental learning from non-stationary data poses special challenges to the field of machine learning. Although new algorithms have been developed for this, assessment of results and comparison of behaviors are still open problems, mainly because evaluation metrics, adapted from more traditional tasks, can be ineffec…
Causal discovery improves fMRI analysis, but faces challenges.
Paper proposes a novel method to assess treatment effect estimators using cross-validation.
Study categorizes time series anomaly detection metrics based on evaluation challenges.
Fast accumulation of large amounts of complex data has created a need for more sophisticated statistical methodologies to discover interesting patterns and better extract information from these data. The large scale of the data often results in challenging high-dimensional estimation problems where only a minority of t…
This paper considers the mean variance portfolio management problem. We examine portfolios which contain both primary and derivative securities. The challenge in this context is due to portfolio's nonlinearities. The delta-gamma approximation is employed to overcome it. Thus, the optimization problem is reduced to a we…
Survey examines distillation methods for large language models.
Paper integrates ML with physics models for engineering and environmental challenges.
Develops a new Bayesian inference method for discrete data.
Proposes I-prior extension for additive interaction models.
This study optimizes currency arbitrage using quantum computing methods.
We use the theory of normal variance-mean mixtures to derive a data augmentation scheme for models that include gamma functions. Our methodology applies to many situations in statistics and machine learning, including Multinomial-Dirichlet distributions, Negative binomial regression, Poisson-Gamma hierarchical models, …
Unified Bayesian framework for PTA data analysis tackles hierarchical model issues.
We describe our methods that achieved the 3rd and 4th places in tasks 1 and 2, respectively, at ISIC challenge 2019. The goal of this challenge is to provide the diagnostic for skin cancer using images and meta-data. There are nine classes in the dataset, nonetheless, one of them is an outlier and is not present on it.…
Automated digital twin discovery from biological data improves drug discovery and personalized medicine.
Bayesian method improves sparse CCA for multi-view data.
Sherpa.ai framework combines federated learning and differential privacy for edge AI services.
Review of methods enabling causal predictions under hypothetical interventions.
Editorial discusses nine challenges in modern algorithmic trading.
This paper presents a methodology for creating streaming, distributed inference algorithms for Bayesian nonparametric (BNP) models. In the proposed framework, processing nodes receive a sequence of data minibatches, compute a variational posterior for each, and make asynchronous streaming updates to a central model. In…
Deep Learning (DL) methods have emerged as one of the most powerful tools for functional approximation and prediction. While the representation properties of DL have been well studied, uncertainty quantification remains challenging and largely unexplored. Data augmentation techniques are a natural approach to provide u…
The Affective Behavior Analysis in-the-wild (ABAW) 2020 Competition is the first Competition aiming at automatic analysis of the three main behavior tasks of valence-arousal estimation, basic expression recognition and action unit detection. It is split into three Challenges, each one addressing a respective behavior t…
The analysis of mixed data has been raising challenges in statistics and machine learning. One of two most prominent challenges is to develop new statistical techniques and methodologies to effectively handle mixed data by making the data less heterogeneous with minimum loss of information. The other challenge is that …
The notion of uncertainty is of major importance in machine learning and constitutes a key element of machine learning methodology. In line with the statistical tradition, uncertainty has long been perceived as almost synonymous with standard probability and probabilistic predictions. Yet, due to the steadily increasin…
The paper explores using machine learning for yield curve calibration in multiple markets.
NASirt automates CNN architecture design for spectral data.
Develops Bayesian inference methods for gamma models.
New method infers causal relationships from nonstationary time series data.
Bayesian method estimates coverage from sketching imperfect data.
New simulation technique speeds up Lévy-driven OU process pricing.
The development of molecular signatures for the prediction of time-to-event outcomes is a methodologically challenging task in bioinformatics and biostatistics. Although there are numerous approaches for the derivation of marker combinations and their evaluation, the underlying methodology often suffers from the proble…
City Logistics is characterized by multiple stakeholders that often have different views of such a complex system. From a public policy perspective, identifying stakeholders, issues and trends is a daunting challenge, only partially addressed by traditional observation systems. Nowadays, social media is one of the bigg…
The aim of this paper is to propose a new methodology that allows forecasting, through Vasicek and CIR models, of future expected interest rates (for each maturity) based on rolling windows from observed financial market data. The novelty, apart from the use of those models not for pricing but for forecasting the expec…
Systematic review of multimodal data challenges and solutions.
Paper improves basket option pricing for log-normal models.
Differential Privacy (DP) provides strong guarantees on the risk of compromising a user's data in statistical learning applications, though these strong protections make learning challenging and may be too stringent for some use cases. To address this, we propose element level differential privacy, which extends differ…