Model predicts drug overdose hotspots using EMS and toxicology data.
problem Predicting drug overdose hotspots to focus limited services.
method Spatial-temporal point process model integrating EMS and toxicology data.
result Model improves prediction accuracy by integrating heterogeneous data.
Bayesian framework improves minority class performance in class-imbalanced data.
problem Class imbalance in predictive toxicology models.
method Weighted likelihood approach modifying likelihood function weights inversely proportional to class proportions.
result Improves balanced accuracy and sensitivity for minority class (toxic compounds).
Everyday we are exposed to various chemicals via food additives, cleaning and cosmetic products and medicines -- and some of them might be toxic. However testing the toxicity of all existing compounds by biological experiments is neither financially nor logistically feasible. Therefore the government agencies NIH, EPA …
hyperSBINN improves drug cardiosafety assessment by efficiently modeling cardiac action potentials.
problem Complexity and limited data in modeling cardiac effects of drugs.
method Combining meta-learning with SBINNs to solve parameterized cardiac action potential models.
result hyperSBINN outperforms traditional solvers in speed and accuracy for predicting APD90 values.
Decision tool helps manage biofouling risks for ships in the Baltic Sea.
problem Biofouling of ships causes environmental and economic issues.
method Bayesian networks to identify biofouling management strategies.
result Optimal biofouling management includes biocidal-free coating and in-water cleaning.
CLARA generates clinical reports from raw inputs, improving accuracy and efficiency.
problem Generating accurate and detailed clinical reports from raw inputs is time-consuming and error-prone.
method Interactive method that generates reports sentence by sentence based on doctors' anchor words and partially completed sentences.
result CLARA achieves significant improvements in report generation accuracy and efficiency.
This study uses NLP to predict stock performance based on analyst reports.
problem Predicting stock performance using textual information from analyst reports.
method Natural language processing (NLP) and a customized BERT deep learning model for Chinese text.
result Strong positive sentiment in analyst reports increases excess return and intraday volatility, while strong negative sentiment increases volatility and trading volume but decreases excess return.
A Pathology report is arguably one of the most important documents in medicine containing interpretive information about the visual findings from the patient's biopsy sample. Each pathology report has a retention period of up to 20 years after the treatment of a patient. Cancer registries process and encode high volume…
Statistical hypothesis testing serves as statistical evidence for scientific innovation. However, if the reported results are intentionally biased, hypothesis testing no longer controls the rate of false discovery. In particular, we study such selection bias in machine learning models where the reporter is motivated to…
Corporate distress models typically only employ the numerical financial variables in the firms' annual reports. We develop a model that employs the unstructured textual data in the reports as well, namely the auditors' reports and managements' statements. Our model consists of a convolutional recurrent neural network w…
Predictive policing models can be biased by differential crime reporting rates.
problem Bias in predictive policing models due to differential crime reporting.
method Simulation based on Bogotá, Colombia's victimization and crime reporting data.
result Differential crime reporting rates can lead to misallocation of police patrols.
Study improves risk evaluation timing with right-censored reporting delays.
problem Improving risk evaluation under short observation windows due to administrative censoring.
method Jointly models parametric hazards for event and reporting processes, uses Monte Carlo expectation-maximization algorithm, and proposes transfer-learning procedure.
result Improves accuracy of timely risk evaluation under administrative censoring.
Private anchors affect how information is communicated and can improve or distort transmission.
problem How private anchors influence strategic communication and information transmission.
method Analyzed a sender-receiver game with costly reports and privately observed anchors.
result Small positive reporting costs can lead to full revelation, even with zero costs.
Paper develops a BERT-based classifier to reduce pathology report annotation workload.
problem Manual annotation of pathology reports is labor-intensive and time-consuming.
method Developed an automatic text classifier using BERT and introduced a human-centric metric to identify low-confidence cases.
result The model reduces manual annotation workload by 80% to 98%.
Typically, operational risk losses are reported above a threshold. Fitting data reported above a constant threshold is a well known and studied problem. However, in practice, the losses are scaled for business and other factors before the fitting and thus the threshold is varying across the scaled data sample. A report…
This report reviews the Edinburgh tram project's risk management. Projects frequently overrun their cost and timelines and fall short on intended benefits. Cost, schedule, and benefit risk of projects need to be carefully considered to avoid this. The report describes and evaluates risk assessment and management for th…
New method protects whistleblowers from retaliation by ensuring their reports remain private.
problem Whistleblowers face retaliation, and current protections are insufficient.
method Formalizes protection against strong-adversary threat model as per-report (0,δ)-differential privacy, and provides a generic mechanism to reduce private auditing to private continual counting. result Demonstrates a reduction in selection error and improved utility over randomized response.
Automated generation of medical reports from chest x-rays using expert annotations.
problem Generating long, unstructured text from medical images with context and consistency.
method First learn visually-informative medical concepts from raw reports, then use these concepts to auto-generate structured reports from images.
result Validation on OpenI dataset shows auto-generated reports are consistent with manual annotations.
Sell-side analysts' reports explain 10% of stock returns, with income statement analyses most impactful.
problem The value of sell-side analysts' information in predicting stock returns.
method Analysis of large language model embeddings and Shapley value decomposition.
result Income statement analyses contribute most to explaining stock returns.
In retrospective assessments, internet news reports have been shown to capture early reports of unknown infectious disease transmission prior to official laboratory confirmation. In general, media interest and reporting peaks and wanes during the course of an outbreak. In this study, we quantify the extent to which med…
Framework integrates financial and annual report data for better corporate credit ratings.
problem Lack of insights from non-financial data in credit rating models.
method Uses FinBERT to extract features from annual reports and combines them with financial data.
result Improves credit rating accuracy by 8-12%.
Paper uses LLMs to analyze annual reports for stock investment, improving efficiency.
problem Manual analysis of annual reports is time-consuming and requires expertise.
method Leverages Large Language Models to extract and analyze annual reports.
result Machine Learning model trained on LLM outputs outperforms S&P500 returns.
Analyst reports contain valuable information for investment decisions.
problem Investment value in analyst reports is not fully understood or utilized.
method Embedded analyst reports with LLMs and ML forecasts of future returns.
result Portfolios formed on analyst report narratives outperform numerical forecasts and established factors.
FinBERT-XRC model assesses financial report risk, offering transparent explanations.
problem Assessing post-event return volatility risk in financial reports.
method Deep-learning model FinBERT-XRC with explainability at word, sentence, and corpus levels.
result FinBERT-XRC outperforms state-of-the-art models in predictive accuracy.
Analysts use vague language in reports to convey useful information about future payoffs.
problem Lack of precise numerical forecasts in analyst reports.
method Empirical analysis of analyst reports to assess the predictive power of linguistic tone.
result The textual tone of analyst reports has predictive power for forecast errors and subsequent revisions, especially when language is vague and uncertainty is high.
We introduce a Bayesian solution for the problem in forensic speaker recognition, where there may be very little background material for estimating score calibration parameters. We work within the Bayesian paradigm of evidence reporting and develop a principled probabilistic treatment of the problem, which results in a…
Model identifies urgent radiology reports with high accuracy.
problem Lack of annotated training data for text analysis.
method Self-supervised contextual language representation using BERT.
result Model achieved 97.0% precision, 93.3% recall, and 95.1% F-measure.
In this paper, we consider the joint design of data compression and 802.15.4-based medium access control (MAC) protocol for smartgrids with renewable energy. We study the setting where a number of nodes, each of which comprises electricity load and/or renewable sources, report periodically their injected powers to a da…
Automatic understanding of domain specific texts in order to extract useful relationships for later use is a non-trivial task. One such relationship would be between railroad accidents' causes and their correspondent descriptions in reports. From 2001 to 2016 rail accidents in the U.S. cost more than $4.6B. Railroads i…
Study clusters Kenyan medical insurance companies based on financial performance and reporting consistency.
problem Identifying financial health and reporting consistency in Kenyan medical insurance companies.
method Advanced clustering techniques (KMeans, DTW) on financial ratios and time series data.
result Four distinct clusters identified, each representing different financial performance and reporting consistency combinations.
Paper proposes transparent reporting of algorithmic energy usage to promote environmental sustainability.
problem Need for transparent reporting of algorithmic energy usage for environmental sustainability.
method Developed a Python package to make analyses of energy usage accessible to individual researchers, localized to specific power grids, and compared with global benchmarks.
result Demonstrated the use of automatically-generated Energy Usage Reports in model-choice for machine learning.
In this paper, we present a data science automation system called Prediction Factory. The system uses several key automation algorithms to enable data scientists to rapidly develop predictive models and share them with domain experts. To assess the system's impact, we implemented 3 different interfaces for creating pre…
SusGen-GPT improves financial NLP and ESG report generation.
problem Lack of advanced NLP tools for finance and ESG domains.
method Developed SusGen-30K dataset and SusGen-GPT models.
result Achieved state-of-the-art performance in financial NLP tasks.
Deep learning improves cancer report classification accuracy.
problem Automatically assigning ICD-O3 codes to cancer reports.
method State-of-the-art deep learning techniques, including hierarchical and flat models, with attention mechanisms.
result Best model achieves 90.3% accuracy on topography site assignment and 84.8% on morphology type assignment.
Data labeling is currently a time-consuming task that often requires expert knowledge. In research settings, the availability of correctly labeled data is crucial to ensure that model predictions are accurate and useful. We propose relatively simple machine learning-based models that achieve high performance metrics in…
Model uses unsupervised learning to classify medical reports with less labeled data.
problem Lack of labeled data for fine-grained disease classification in medical reports.
method Developed a pipeline combining an unsupervised encoder-language model and a supervised classifier model.
result Improved classification accuracy with less labeled data compared to previous methods.
Model estimates non-reported GHG emissions for companies using machine learning.
problem Incomplete GHG emissions reporting by companies.
method Interpretable machine learning model tailored for non-reporting companies.
result Model accurately estimates emissions for diverse company groups.
Study reveals LLM personality patterns but lacks behavioral consistency.
problem Understanding and validating personality traits in LLMs.
method Characterized LLM personality across three dimensions: training dynamics, self-report validity, and intervention effects.
result Self-reported traits do not reliably predict behavior, and instructional alignment affects trait expression but not behavior.
Paper fine-tunes a language model to predict long-term stock buy signals.
problem Predicting long-term stock price movements with narrative text.
method Fine-tuning a small language model on 10-K reports for buy/sell decisions.
result Buy signals generated from 10-K text are most precise at 6 and 9 months, providing 4.8-9% improvement over random selection.
Gaussian process improves nowcasting of COVID-19 deaths.
problem Correcting for delays in reporting daily deaths.
method Flexible Gaussian process model with latent variables.
result Gaussian process nowcasts outperform other methods.
New method uses LLMs to extract financial insights from Q&A sections of reports.
problem Scalability and accuracy issues in extracting valuable insights from financial report Q&A sections.
method Combines retrieval-augmented generation technique with metadata.
result Empirically demonstrates superior performance of the proposed method.
Study shows how missing data from certain groups can unfairly bias risk models.
problem Data missingness without indicators of missingness can unfairly bias risk models.
method Developed an analytically tractable model of differential feature under-reporting and proposed new methods to mitigate bias.
result Under-reporting typically leads to increasing disparities in risk models.
Generative AI boosts analyst reports but increases forecast errors.
problem Improving financial analyst reports with AI.
method Natural experiment using FactSet's AI platform.
result AI-assisted reports are more comprehensive but lead to higher forecast errors.
Replication confirms CGD's effectiveness in competitive games.
problem Reproducibility of a novel Nash equilibrium algorithm.
method Replicated experiments and provided Python implementation.
result CGD avoids oscillatory and divergent behaviours.
Bayesian model averaging under predictor redundancy
problem Reporting Bayesian model averaging posterior without changing the Bayesian target
method Using hard or soft regions of support space
result Region reports often give shorter and clearer summaries while preserving the main posterior information
Study uses social media to analyze COVID-19 impact.
problem Understanding the global impact of COVID-19.
method Machine learning and linguistic tools to analyze social media posts.
result Automatic detection of positive reports of COVID-19.
Study analyzes carbon footprint of 1,417 ML models on Hugging Face.
problem Scarce knowledge on measuring and reporting carbon footprint of ML models.
method Repository mining study on Hugging Face Hub API.
result Stalled carbon emissions-reporting models, slight decrease in carbon footprint over 2 years.
New method decomposes profits and losses continuously, avoiding discrete reporting issues.
problem Analyzing profits and losses at discrete dates ignores detailed paths.
method Constructs a large class of continuous-time decompositions using extended Itô's formula.
result Identifies a preferred decomposition from exactness, symmetry, and normalization axioms.