Predicts academic risk in college students using interpretable machine learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper presents deep learning and ML for automated student performance estimation.
Study uses ML and causal analysis to predict student performance factors.
Paper discusses how financial institutions' model risk management can benefit academic research.
Study evaluates the impact of academic support center's face-to-face assistance on student performance.
The study uses Hidden Markov Models to analyze student enrollment patterns and academic performance.
System identifies high impact research from academic papers.
Study predicts academic achievement using students' support networks.
KT models struggle with student concept drift, but BKT remains the most stable.
Machine learning methods tend to outperform traditional statistical models at prediction. In the prediction of academic achievement, ML models have not shown substantial improvement over logistic regression. So far, these results have almost entirely focused on college achievement, due to the availability of administra…
A new method for name disambiguation in academic networks using multi-view attention and recurrent neural networks.
Study on academic finance evolution over 30 years.
Tool to estimate research impact for low-resource institutions.
Reasoning models outperform LLMs on CFA exams.
The failure of landing a job for college students could cause serious social consequences such as drunkenness and suicide. In addition to academic performance, unconscious biases can become one key obstacle for hunting jobs for graduating students. Thus, it is necessary to understand these unconscious biases so that we…
Efficiently trains BERT on academic GPUs in 12 days.
Kaggle competitions offer valuable insights for business forecasting.
Social media enhances or diminishes scientific status, depending on usage.
This is a brief survey of the research performed by Grandata Labs in collaboration with numerous academic groups around the world on the topic of human mobility. A driving theme in these projects is to use and improve Data Science techniques to understand mobility, as it can be observed through the lens of mobile phone…
Paper introduces a new algorithm to detect LLM-generated text.
We develop an algorithm that can detect pneumonia from chest X-rays at a level exceeding practicing radiologists. Our algorithm, CheXNet, is a 121-layer convolutional neural network trained on ChestX-ray14, currently the largest publicly available chest X-ray dataset, containing over 100,000 frontal-view X-ray images w…
We develop dependent hierarchical normalized random measures and apply them to dynamic topic modeling. The dependency arises via superposition, subsampling and point transition on the underlying Poisson processes of these measures. The measures used include normalised generalised Gamma processes that demonstrate power …
As recent literature has demonstrated how classifiers often carry unintended biases toward some subgroups, deploying machine learned models to users demands careful consideration of the social consequences. How should we address this problem in a real-world system? How should we balance core performance and fairness me…
AI predicts dyslexia and dysgraphia in children.
Variable Annuity (VA) products expose insurance companies to considerable risk because of the guarantees they provide to buyers of these products. Managing and hedging these risks requires insurers to find the value of key risk metrics for a large portfolio of VA products. In practice, many companies rely on nested Mon…
Over the last 23 years, the U.S. Securities and Exchange Commission has required over 34,000 companies to file over 165,000 annual reports. These reports, the so-called "Form 10-Ks," contain a characterization of a company's financial performance and its risks, including the regulatory environment in which a company op…
The Wallenius distribution is a generalisation of the Hypergeometric distribution where weights are assigned to balls of different colours. This naturally defines a model for ranking categories which can be used for classification purposes. Since, in general, the resulting likelihood is not analytically available, we a…
Estimates citation impact to recommend best publication venue.
Study evaluates digital transformation impact on financial performance using LLMs.
Randomized control methods improve asset pricing and performance analysis.
Bitcoin price prediction models fail to outperform a simple 'today's price' baseline, especially at longer horizons.
Recently, due to the booming influence of online social networks, detecting fake news is drawing significant attention from both academic communities and general public. In this paper, we consider the existence of confounding variables in the features of fake news and use Propensity Score Matching (PSM) to select gener…
Machine learning algorithms designed to characterize, monitor, and intervene on human health (ML4H) are expected to perform safely and reliably when operating at scale, potentially outside strict human supervision. This requirement warrants a stricter attention to issues of reproducibility than other fields of machine …
Shifu2 discovers advisor-advisee relationships in collaboration networks.
This report was originally written as an industry white paper on Hedge Funds. This paper gives an overview to Hedge Funds, with a focus on risk management issues. We define and explain the general characteristics of Hedge Funds, their main investment strategies and the risk models employed. We address the problems in H…
The DoD needs a robust process to evaluate AI/ML model performance and robustness.
One of the important measures of quality of education is the performance of students in the academic settings. Nowadays, abundant data is stored in educational institutions about students which can help to discover insight on how students are learning and how to improve their performance ahead of time using data mining…
This reviews the econophysics activities in Belgium from my admittedly biased point of view. Unknown historical notes or facts are presented for the first time explaining the aims, whence evolution of the research papers and friendly connections with colleagues. Comments on endeavors are also provided. The lack of offi…
Computer generated academic papers have been used to expose a lack of thorough human review at several computer science conferences. We assess the problem of classifying such documents. After identifying and evaluating several quantifiable features of academic papers, we apply methods from machine learning to build a b…
This paper uses information theory to improve risk modeling in big data.
Enhances insurance loss models using InsurTech data and machine learning.
Autonomous driving is getting a lot of attention in the last decade and will be the hot topic at least until the first successful certification of a car with Level 5 autonomy. There are many public datasets in the academic community. However, they are far away from what a robust industrial production system needs. Ther…
Peer-reviewed research and mined data predict stock returns similarly.
The professional services sector is at a turning point, with some industries showing growth opportunities.
Replicated and validated Rank-N-Contrast for robust regression.
Improved graph matching using covariates for network data integration.
Lecture notes introduce differential geometry using sheaves and differential operators.
The paper proposes to analyze a data set of Finnish ranks of academic publication channels with Extreme Learning Machine (ELM). The purpose is to introduce and test recently proposed ELM-based mislabel detection approach with a rich set of features characterizing a publication channel. We will compare the architecture,…