This study revisits Fama-French models using sample innovations to address misinterpretation of high R-squared values.
problem Misinterpretation of high R-squared values in Fama-French models due to serial dependence and volatility clustering.
method Use of sample innovations to derive standard econometrics time series models to overcome misinterpretation.
result Suggests the Fama-French model should consider heavy-tail distributions due to relevant tail behavior in financial data.
Gradient-free ensemble learns sector forecasts from diverse models.
problem Predicting sector returns in a volatile market.
method Dynamic model combination using out-of-sample R-squared.
result Ensemble outperforms individual models in sector rotation.
Deep learning model reduces food waste by stabilizing online food delivery supply chains.
problem Wastage and bullwhip effect in online food delivery services.
method Two-phase LSTM network for demand forecasting, newsvendor model for inventory management.
result Significant reduction in bullwhip effect and food waste, improved forecasting accuracy.
Improved stock price prediction model using generalized order flow imbalance.
problem Improving stock price prediction models using new order flow imbalance indicators.
method Proposed a generalized order flow imbalance construction method and applied it to CSI 500 stocks.
result Generalized Stationarized Order Flow Imbalance (log-GOFI) shows significant improvement in explaining stock price changes.
Satellite imagery improves house price prediction models.
problem Improving accuracy of housing price estimation models.
method Transfer learning from ImageNet-pretrained Inception-v3 model to satellite images.
result Achieved a 10% improvement in R-squared score.
The uncertainties in future Bitcoin price make it difficult to accurately predict the price of Bitcoin. Accurately predicting the price for Bitcoin is therefore important for decision-making process of investors and market players in the cryptocurrency market. Using historical data from 01/01/2012 to 16/08/2019, machin…
Study evaluates different mathematical models for three case studies using statistical fitting.
problem Estimating outcomes in population dynamics, temperature variations, and market equilibrium.
method Applied various statistical equations (e.g., fractional exponential, sinusoidal) to three case studies.
result Optimal models differ by case study (fractional exponential for population dynamics, sinusoidal for temperature and market equilibrium).
XGBoost predicts NEPSE Index log returns with low error and high directional accuracy.
problem Forecasting daily log-returns in the NEPSE Index with high accuracy.
method XGBoost machine learning, feature engineering, hyperparameter optimization, walk-forward validation.
result Optimal XGBoost configuration achieves lowest log-return RMSE and MAE.
Proposes a new method to estimate variable importance in black box models, mitigating correlation effects.
problem Correlation between covariates affects the interpretation of variable importance parameters.
method Develops a modified LOCO (Leave Out COvariates) method and uses semiparametric models for estimation.
result Shows how to estimate a modified LOCO method that mitigates correlation effects.
Study predicts soccer player market values using machine learning and SHAP for interpretability.
problem Predicting accurate market values for professional soccer players.
method Ensemble machine learning models, SHAP for interpretability, Boruta for feature selection.
result GBDT model achieved high predictive accuracy (R-squared 0.901, RMSE 3,221,632.175).
Hedonic models predict 84-92% of U.S. real estate prices, highlighting environmental factors' impact.
problem Predicting real estate prices using hedonic models with environmental factors.
method P-spline generalized additive models for real estate prices, contrasting with linear and polynomial models.
result GAM models explain 84-92% of U.S. real estate price variance, with environmental factors contributing minimally.
Study compares non-parametric models for predicting medical insurance reimbursement delays.
problem Estimating the time-lapse between medical insurance reimbursement.
method Comparative study of four non-parametric regression models (KNNs, SVMs, Decision Trees, Random Forests) using R-squared metric.
result Each model's performance varies with training data size, feature space, and hyperparameters.
GPR ensemble method predicts stock returns efficiently.
problem Predicting stock returns using machine learning.
method Ensemble Gaussian Process Regression (GPR) for online learning.
result Method outperforms existing models in R-squared and Sharpe ratio. The study compares different models for predicting factor premiums and finds neural networks perform better but have unstable weights.
problem Predicting and timing the CMA factor premium using machine learning models.
method Compared regression models (OLS, Ridge, Random Forest, Neural Network) and tested factor timing strategies.
result Neural networks outperform linear models in explaining factor premium variance, but weights are unstable.
Paper uses neural networks to predict NOx emissions from gas turbines.
problem Predicting NOx emissions from degrading gas turbines.
method Applied neural network algorithm to model NOx emissions from nine process variables.
result Neural network model optimizes process variables for minimal NOx emissions.
Filter or screening methods are often used as a preprocessing step for reducing the number of variables used by a learning algorithm in obtaining a classification or regression model. While there are many such filter methods, there is a need for an objective evaluation of these methods. Such an evaluation is needed to …
The paper explores high-dimensional learning in finance, proving key aspects and setting lower bounds.
problem Understanding when and how large, over-parameterized models achieve predictive success in finance.
method Theoretical foundations and empirical validation of two key aspects: standardization and information-theoretic lower bounds.
result Empirical validation shows that high-dimensional learning in finance often relies on lower-complexity artefacts rather than the intended mechanism.
The paper proposes a method to test features selected by SeqFS-DA with controlled FPR.
problem Ensuring reliability of feature selection after domain adaptation in high-dimensional regression.
method Proposes a novel method to test features selected by SeqFS-DA with controlled FPR.
result The proposed method controls FPR below a significance level α (e.g., 0.05) and enhances statistical power. We construct genomic predictors for heritable and extremely complex human quantitative traits (height, heel bone density, and educational attainment) using modern methods in high dimensional statistics (i.e., machine learning). Replication tests show that these predictors capture, respectively, ∼40, 20, and 9 perc…
Style Miner generates stable and significant style factors for time series analysis.
problem Finding significant and stable explanatory factors in high-dimensional time series data.
method Proposes a reinforcement learning method to balance explanatory power and stability constraints.
result Outperforms existing methods by a large margin and achieves a 10% gain in R-squared explanatory power.
In this paper, we present machine learning approaches for characterizing and forecasting the short-term demand for on-demand ride-hailing services. We propose the spatio-temporal estimation of the demand that is a function of variable effects related to traffic, pricing and weather conditions. With respect to the metho…
PNNs model aleatoric uncertainty in scientific machine learning with high accuracy.
problem Aleatoric uncertainty in scientific systems with unequal variance.
method Developed a probabilistic distance metric to optimize PNN architecture and used it in material science applications.
result PNNs yield remarkably accurate output mean estimates and high correlation in predicted intervals.
Proposes TgNN-LD to improve neural network effectiveness and efficiency.
problem Limits in maintaining tradeoff between data and domain knowledge.
method Converts loss function to constrained form with PDEs, ECs, and EK as constraints, incorporating Lagrangian variables for equitable tradeoff.
result Improves prediction accuracy and conserves resources.
Study examines downsizing impact on Indian construction firms' profitability.
problem Impact of downsizing layoffs on construction firms' profitability in India.
method Used Co-integration test, OLS, and VAR models on secondary data of 15 companies.
result Employee Expenses and Number of Employees have significant impact on profitability.
Deep learning maps tongue movements to speech sounds for voiceless individuals.
problem Developing silent speech interfaces for individuals without a larynx.
method Hybrid spatio-temporal 3D convolutions and feature shuffling for formant estimation and tracking from ultrasound tongue images.
result Best model achieves R-squared of 99.96% for vowel formant regression.
Study investigates how preprocessing, feature selection, and model selection affect performance on imbalanced genetic data.
problem Challenges in using machine learning on imbalanced genetic datasets.
method Comparative analysis of data preprocessing, feature selection techniques, and machine learning models on imbalanced genetic data.
result Class-imbalanced target variables and skewed predictors have little to no impact on classification performance.
Study forecasts vegetable prices in Nepal using a novel index and ensemble model.
problem High volatility and cultural influences on agricultural commodity prices.
method Developed KVPI, created features, evaluated multiple models, introduced Momentum-Corrected Online Stacking Ensemble.
result Achieved RMSE of 1.771, MAPE of 0.68%, and R-squared of 0.845 at 90-day horizon.
High-performing equity factor with Sharpe ratio above 13 out-of-sample.
problem Hidden cross-sectional predictability in stock returns.
method Regime-conditional signal activation combining value and short-term reversal signals.
result Annualized returns of 158.6% with 12.0% volatility, strong performance out-of-sample.
This paper measures the information quantity in paintings using entropy.
problem Traditional art pricing models lack variables capturing painting content.
method Extends Shannon entropy to measure painting information using pixel-level variances of line, color, value, shape/form, and space.
result Variance measurements significantly explain sales prices, improving traditional models.
This paper explores integration and contagion among US metropolitan housing markets. The analysis applies Federal Housing Finance Agency (FHFA) house price repeat sales indexes from 384 metropolitan areas to estimate a multi-factor model of U.S. housing market integration. It then identifies statistical jumps in metrop…
Empirical study of CAPM and Fama-French model in Chinese A-share market.
problem Testing and validating CAPM and Fama-French model in Chinese A-share market.
method Used Fama-MacBeth regression and Fama-French three-factor model to analyze Chinese A-share trading data from 2000 to 2019, adjusting for IPO shell value contamination.
result Fama-French model captures most of A-share market returns, with adjusted R-squared > 0.88.
This paper uses alternative data to forecast Japanese real estate performance.
problem Accurate rent and price forecasting in Japanese real estate markets.
method Created a comprehensive house price index using over 5 million transactions and economic factors.
result Alternative data variables can forecast real estate performance effectively.
Machine learning helps estimate risk premiums of stocks without knowing their factors.
problem Estimate risk premiums of stocks without knowing their underlying factors.
method Used elastic-net machine learning to project stock returns onto peers and construct replicate portfolios.
result Unique stocks have higher SARP and excess returns than ubiquitous stocks.
Current clinical practice guidelines for managing Coronary Artery Disease (CAD) account for general cardiovascular risk factors. However, they do not present a framework that considers personalized patient-specific characteristics. Using the electronic health records of 21,460 patients, we created data-driven models fo…
The study analyzes how cross-chain interoperability affects decentralized lending protocols' performance.
problem Understudied cross-chain elements in DeFi lending risk management.
method Panel regression fixed effects and OLS models applied to empirical analysis.
result Cross-chain activity impacts protocol performance, with bridge volume being a critical driver.
Study finds unsupervised imputation before cross-validation can reduce computational costs without significantly degrading model performance.
problem High computational costs in pipeline modeling algorithms with imputation steps.
method Empirical assessment of unsupervised imputation before vs during cross-validation.
result Reduced variance of imputation before cross-validation leads to lower overall root mean squared error.
Quantitative model predicts Sri Lankan stock market using NLP, clustering, and time-series forecasting.
problem Predicting economic regimes and market signals in Sri Lankan stock indices.
method Integrates NLP, clustering, and time-series forecasting; uses FinBERT for sentiment analysis, UMAP/HDBSCAN for clustering, and GRU/LSTM for forecasting.
result GRU model achieves 80.1% R-squared for daily closing price forecasts.
A visualization aids in comparing regression models by highlighting errors and correlations.
problem Comparing regression models is difficult due to varying hyper-parameters and metrics.
method Introduces a novel visualization approach using 2D residual space, Mahalanobis distance, and colormaps.
result Enhanced understanding of regression model performance differences and error distributions.
This paper improves prediction uncertainty estimation by inferring variation from neuron activation strength.
problem Estimating prediction uncertainty from ensemble methods is expensive and inaccurate.
method Introduced randomness into model training and inferred prediction variation from neuron activation strength.
result Average R squared on MovieLens is 0.56 and on Criteo is 0.81, with strong performance in variation detection.
VisitHGNN predicts visit probabilities between neighborhoods and POIs using graph neural networks.
problem Estimating visit probabilities between neighborhoods and POIs for urban planning.
method Heterogeneous, relation-specific graph neural network (VisitHGNN) trained on mobility data.
result Strong predictive performance with high fidelity to observed travel behavior.
RegPred Net forecasts foreign exchange rates with improved accuracy and interpretability.
problem Multi-step forecasting of Foreign Exchange (FX) rates.
method Bayesian optimization for hyperparameter tuning of a multi-layered regression network.
result RegPred Net significantly outperforms other models in terms of RMSE and correlation metrics.
Credit risk management in Italy is characterized, in the period June 2008 to June 2012, by frequent (frequency=0.5 cycles per year) and intense (peak amplitude: mean=39.2 billion Euros, s.e.=2.83 billion Euros) quarterly contractions and expansions around the mean (915.4 billion Euros, s.e.=3.59 billion Euros) of the n…
Optimizes QoS in FSO links over South Africa using ensemble learning.
problem Impact of weather on QoS in FSO links.
method Ensemble learning models (Random Forest, ADaBoost Regression, Stacking Regression, Gradient Boost Regression, Multilayer Neural Network) applied to meteorological data.
result Significant enhancement in QoS with RMSE and R-squared values of 0.0032 and 0.9906 respectively at George.
Deep learning extracts terrain texture covariates for geostatistical modeling.
problem Improving prediction accuracy in geostatistical modeling using terrain texture data.
method Deep learning approach to automatically derive optimal terrain texture covariates from SRTM 90m DEM.
result Deep learning-derived covariates have strong explanatory power (R-squared around 0.6) for geochemical data.
Study investigates micro-event detection on FLOSS version releases from Stack Overflow.
problem Detecting micro-events in FLOSS version release events from textual messages.
method Developed pipelines using LDA topic modeling, hSBM topics, and sentiment analysis; optimized feature spaces with RFECV; evaluated models with statistical analysis.
result Found characteristic changes in topics or sentiment features before or after FLOSS version releases.
Deep learning improves CVD risk prediction from health records.
problem Predicting cardiovascular disease risk from administrative health data.
method Combined survival analysis and deep learning models.
result Deep learning models outperform traditional Cox models in accuracy and explained time-to-event occurrence.