Bayesian network predicts corn yields at county level using historical data and expert knowledge.
problem Accurate corn yield prediction for policy and market impact.
method Data-driven Bayesian network approach using historical data and expert knowledge.
result Maximized likelihood model structure connecting predictors and yield.
Deep LSTM models predict corn yields at county level.
problem Predicting county-level corn yields to address information asymmetry and improve price efficiency.
method Employed Long Short-Term Memory (LSTM) neural networks for time series prediction.
result Deep LSTM models show promising predictive power for county-level corn yields.
Enhanced geographical features improve predictive models for colorectal cancer survival curves.
problem Predicting colorectal cancer survival curves in Iowa.
method Used neural networks to explore feature representations, comparing ABC performance.
result Spectral analysis-based representations improve predictive performance by approximately 40%.
Improved county-level COVID-19 forecasting model using LSTM and data augmentation.
problem Accurately forecasting county-level COVID-19 cases to optimize medical resources.
method Adapted TDEFSI-LONLY model, utilized LSTM, data augmentation, and inter-county mixing.
result CLEIR-Net model provides better forecasts than TDEFSI-LONLY.
Study predicts U.S. county COVID-19 growth using demographic and social distancing data.
problem Predicting county-level COVID-19 growth during the pandemic.
method Spectral clustering, correlation matrix, demographic features, social distancing scores, LSTM model.
result Effective prediction of future county growth using demographic and social distancing data.
Paper analyzes factors affecting COVID-19 risk in US counties.
problem Identifying factors influencing COVID-19 risk in US counties.
method Combines unsupervised (K-means clustering) and supervised learning models.
result Mean temperature, poverty, obesity, and other factors are most significant.
We analyze the cumulative distribution of total personal income of USA counties, and gross domestic product of Brazilian, German and United Kingdom counties, and also of world countries. We verify that generalized exponential distributions, related to nonextensive statistical mechanics, describe almost the whole spectr…
Satellite images predict U.S. county mortality rates.
problem Predicting mortality rates in U.S. counties using satellite imagery.
method Convolutional neural network trained on crude mortality rates, learned features interpreted using Shapley Additive Feature Explanations.
result Predicted mortality from satellite images correlated strongly with true mortality rates (Pearson r=0.72).
Paper presents a machine learning framework for corn yield forecasting.
problem Accurate and timely prediction of corn yields in the US Corn Belt.
method Machine learning ensembles considering complete and partial in-season weather data.
result Ensemble models outperform individual models, achieving best prediction accuracy.
Algorithm calculates distances between US counties based on people living there.
problem Defining a network graph for large, non-i.i.d. data sets.
method Co-clustering diffusion metric and data adaptive transportation cost.
result Approximate Earth Mover's Distance for US county comparisons.
Model predicts US COVID-19 deaths with quantile estimates.
problem Predicting US COVID-19 deaths at county level.
method Hybrid machine learning and epidemiological approach, minimizing pinball loss.
result Quantile estimates accurately forecast deaths for different forecast periods.
Satellite imagery improves house price prediction models.
problem Improving accuracy of housing price estimation models.
method Transfer learning from ImageNet-pretrained Inception-v3 model to satellite images.
result Achieved a 10% improvement in R-squared score.
A new machine learning model forecasts COVID-19 incidence at county level in the USA.
problem Inaccurate disease spread forecasting due to spatiotemporal homogeneity assumptions.
method Spatiotemporal machine learning using LSTM architecture with spatial and temporal features.
result COVID-LSTM outperforms COVID-19 Forecast Hub's Ensemble model in accuracy.
Study predicts cricket match outcomes using machine learning.
problem Predicting the outcome of English county cricket matches.
method Used machine learning to analyze team and player statistics.
result Optimal model significantly outperformed gambling industry benchmarks.
Anomaly detection for high-dimensional data using large deviations principle.
problem Challenges in anomaly detection for high-dimensional data.
method Large Deviations Anomaly Detection (LAD) algorithm.
result Outperforms state-of-the-art methods on high-dimensional data sets.
Count data, for example the number of observed cases of a disease in a city, often arise in the fields of healthcare analytics and epidemiology. In this paper, we consider performing regression on multivariate data in which our outcome is a count. Specifically, we derive log-likelihood functions for finite mixtures of …
Research explores how local communities and corporations interact in finance.
problem Impact of local government subsidies and corporate bankruptcy on bond yields.
method Difference-in-differences analysis, econometric models, deep-learning model.
result Corporate subsidies and bankruptcy filings affect bond yields significantly.
Machine learning predicts mask mandates reduce COVID-19 deaths.
problem Effectiveness of mask mandates on reducing COVID-19 deaths.
method Machine learning classification algorithms applied to survey data.
result Mask mandates correlated with reduced death ratios in most counties.
Religious adherence reduces corporate greenwashing behavior.
problem Greenwashing behavior by corporations.
method Analysis of a large US firm sample (2005-2019), focusing on selective disclosure.
result Religious adherence correlates with lower greenwashing behavior.
Study improves fraud detection in e-commerce with a stacked model combining CNNs, GNNs, and confidence gating.
problem Detecting credit card fraud in online transactions.
method Stacking approach with attention and confidence-driven layers, using DOWA and IOWA operators.
result The method achieves high accuracy and robust generalization in CCF detection.
Model predicts COVID-19 progression with interpretability.
problem Accurate and credible forecasting of COVID-19 progression.
method Integrates machine learning into disease modeling, uses interpretable encoders.
result More accurate forecasts than state-of-the-art alternatives.
Study examines equity in post-Snow Uri recovery, finds disparities.
problem Disproportionate impacts on vulnerable populations during recovery.
method County and census tract level data analysis, satellite imagery, statistical procedures.
result Negative associations between non-Hispanic whites and outages, positive associations with certain demographic variables.
Study evaluates how changes in mobility affect COVID-19 case rates.
problem Mixed evidence on mobility-COVID-19 case rate associations.
method Modified treatment policy (MTP) approach with TMLE and Super Learner ensemble.
result Shifts in mobility do not consistently affect subsequent case rates after adjusting for confounders.
TLRF improves timely COVID-19 outbreak detection with small sample size counties.
problem Balancing accuracy and speed in estimating COVID-19 case growth rates.
method Transfer Learning Random Forest (TLRF) framework for growth rate estimation.
result TLRF outperforms existing methods in predicting case growth rates and timely outbreak detection.
condLSTM-Q predicts COVID-19 deaths at county level with quantile forecasts.
problem Predicting COVID-19 mortality at fine geographical scales.
method Conditional Long Short-Term Memory networks with quantile output.
result Fine-scale quantile predictions inform about death toll distribution.
The aim of this paper is to get an overview of the online buyer profile, and also some key aspects in the way the online shopping is conducted. In this project we conducted a quantitative research, consisting of a questionnaire based survey. For data processing and interpretation we used SPSS statistical software and E…
New results for modeling voter probabilities in elections.
problem Modeling voter preferences with aggregate data and individual covariates.
method Maximum likelihood estimation for Poisson binomial distribution, approximated with heteroscedastic Gaussian.
result Existence and curvature results for the MLE of the Poisson binomial likelihood.
Labor productivity in Turkey, Spain, Belgium, Austria, Switzerland, and New Zealand has been analyzed and modeled. These counties extend the previously analyzed set of the US, UK, Japan, France, Italy, and Canada. Modelling is based on the link between the rate of labor participation and real GDP per capita. New result…
We develop a model of how information flows into a market, and derive algorithms for automatically detecting and explaining relevant events. We analyze data from twenty-two "political stock markets" (i.e., betting markets on political outcomes) on the Iowa Electronic Market (IEM). We prove that, under certain efficienc…
Machine learning detects drug overdose trends, aiding prevention.
problem Detecting subtle overdose patterns in spatio-temporal data.
method Gaussian Process Subset Scan and Multidimensional Tensor Scan.
result Identifies previously unknown overdose patterns and demographic clusters.
MLP outperforms other deep learning models for groundwater prediction.
problem Accurate groundwater level predictions under changing climatic conditions.
method Optimized hyperparameters of deep learning models using surrogate models.
result MLP performs best in terms of prediction accuracy and time-to-solution.
Study improves flood loss risk models using historical data and rainfall data.
problem Predicting financial losses from flooding events.
method Used neural networks, decision trees, and kernel-based regressors on NFIP dataset, incorporating rainfall data.
result Extreme Gradient Boosting provided the best results, and bias correction improved model performance.
Paper presents a spatio-temporal Bayesian model for early detection of COVID-19 hotspots.
problem Understanding spatio-temporal dynamics of COVID-19 hotspots to prevent outbreaks.
method Spatio-temporal Bayesian framework with a zero-mean Gaussian process and non-stationary kernel function enhanced by deep neural networks.
result Model demonstrates superior hotspot-detection performance compared to baseline methods.
New spectral clustering for directed graphs reveals socio-economic patterns.
problem Spectral clustering for directed graphs is unsatisfactory due to edge directionality.
method Proposes a complex-valued matrix representation and analysis for directed graphs.
result Our approach reveals socio-economic patterns in internal migration data.
Study finds non-adherence to schizophrenia meds leads to earlier adverse events.
problem Impact of medication non-adherence on adverse outcomes in schizophrenia patients.
method Survival analysis, causal inference methods (T-learner, S-learner, nearest neighbor matching), different amounts of longitudinal information.
result Non-adherence to schizophrenia meds advances adverse events by 1 to 4 months.
Study compares analytical and bootstrap DML confidence intervals across various machine learning algorithms.
problem Impact of machine learning algorithm choice on DML confidence intervals.
method Comprehensive simulation study comparing analytical and bootstrap DML confidence intervals across different machine learning algorithms.
result Substantial variability in coverage performance across analytical and bootstrap confidence intervals, highlighting the importance of learner choice.
Study forecasts cardiology admissions from cath lab using ARIMA models.
problem Complexity in managing cardiology admissions from cath lab.
method Retrospective data analysis with ARIMA, Holts method, and other models.
result ARIMA (2,0,2) (1,1,1) model selected as best fit.
Study analyzes climate impact on agricultural prices, offering insurance solutions.
problem Financial risk from climate-induced agricultural price volatility.
method Historical and future climate projections, EGARCH and SARIMAX models, Black-Scholes framework.
result Improved agricultural risk modeling and insurance mechanisms.
Study examines stock price reactions to Texas winter storm power outages.
problem Impact of natural disasters on stock market values.
method Used four benchmark models to measure abnormal returns.
result Firms experienced significant stock price drops after the Texas winter storm.
Study identifies and measures biases in legal case data.
problem Addressing representation biases and sentencing disparities in legal case data.
method Utilizes two regression models: a baseline and a fair judge model.
result Quantifies biases across demographic groups in criminal data from Cook County (Illinois).
Unified framework detects overfitting in crash classification models.
problem Evaluation metrics fail to detect overfitting in crash classification models.
method Random Matrix Theory and Heavy-Tailed Self-Regularization framework applied to various model types.
result Power-law exponent α reliably distinguishes well-regularized from overfit models.
Optimal rule lists for categorical data are created with guaranteed accuracy.
problem Creating interpretable yet accurate models for categorical data.
method Custom discrete optimization technique producing rule lists with optimal training performance.
result Optimal rule lists are constructed in seconds with guaranteed accuracy.
Study models weather index insurance pricing by insurers and farmers, finding flexible pricing kernels boost profits.
problem Monopoly pricing of weather index insurance with risk and flexibility considerations.
method Bowley-type sequential game with insurer and farmer, using neural networks for farmer's payoff.
result Flexible pricing kernels increase insurer profits closer to indemnity insurance levels.
DiD-BCF model improves causal inference in panel data with robust non-parametric methods.
problem Challenges in Difference-in-Differences (DiD) estimation, especially heterogeneous treatment effects and non-linearities.
method Difference-in-Differences Bayesian Causal Forest (DiD-BCF) with PTA-based reparameterization.
result DiD-BCF provides superior performance and uncovers significant heterogeneity in treatment effects.
WeatherFormer learns robust weather features from small datasets.
problem Modeling complex weather dynamics from limited data.
method Pretrained transformer encoder on large satellite dataset, with spatiotemporal encoding.
result State-of-the-art performance in county-level soybean yield prediction and influenza forecasting.
Unified model explains income inequality across countries.
problem Understanding and explaining income inequality across different countries.
method Analytical model based on a master equation with growth and reset terms, tested on real data.
result Income distributions collapse on a master-curve when normalized, suggesting a universal pattern.
Bayesian model tackles spatial count data issues with flexible non-parametric techniques.
problem Challenges in traditional parametric models for spatial count data with unbalanced distributions and complex dependencies.
method Bayesian semi-parametric spatial dispersed count model combining non-parametric techniques and adapted count models.
result Demonstrates superior performance in managing dispersion and capturing intricate spatial patterns.
Paper proposes a method to improve autonomous vehicle performance using synthetically generated images.
problem Limited access to real-world datasets for autonomous vehicle training in countries with scarce data.
method Synthetically generated images to augment and train neural networks on small datasets.
result About 10% improvement in model performance observed.