Study predicts academic achievement using students' support networks.
problem Predicting academic achievement in college students.
method Decision tree and random forest algorithms applied to Ties data.
result Different types of support are important for different demographics and genders.
Study uses ML and causal analysis to predict student performance factors.
problem Understanding socio-academic and economic factors affecting student performance.
method Employed machine learning techniques and causal analysis on 1,050 student profiles.
result Ridge Regression achieved robust predictions with MAE of 0.12 and MSE of 0.024.
Machine learning methods tend to outperform traditional statistical models at prediction. In the prediction of academic achievement, ML models have not shown substantial improvement over logistic regression. So far, these results have almost entirely focused on college achievement, due to the availability of administra…
Predicts academic risk in college students using interpretable machine learning.
problem Predicting academic risk from high-dimensional, unbalanced student data.
method Binary classification task using LightGBM model and Shapley value.
result 8 predictors for academic risk identified, including quality of academic partners and dormitory study atmosphere.
Paper discusses how financial institutions' model risk management can benefit academic research.
problem Improving academic research process and mitigating limitations.
method Adopting financial institutions' model risk management practices.
result Lessons from financial institutions can enhance academic research reliability.
Reasoning models outperform LLMs on CFA exams.
problem Previous research showed LLMs failing CFA exams; reasoning models show promise.
method Evaluated state-of-the-art reasoning models on CFA exams using pass/fail criteria.
result Most reasoning models pass all three CFA levels; Gemini 3.0 Pro achieves highest scores.
Paper presents deep learning and ML for automated student performance estimation.
problem Evaluation of students' performance during the pandemic.
method In-depth analysis of deep learning and machine learning approaches.
result Better performance across different prediction tasks with fully data-driven approach.
Study on academic finance evolution over 30 years.
problem Understanding changes in academic finance research over time.
method Analysis of 32 finance journals from 1992 to 2021 using Structural Topic Model.
result Most journals have become more generalist over time, covering more topics.
System identifies high impact research from academic papers.
problem Identifying high impact research amidst a growing volume of papers.
method Combines visual and text classifiers on a large dataset of PDFs and citation counts.
result Improved accuracy in predicting high impact research.
Tool to estimate research impact for low-resource institutions.
problem Costly databases limit access for third-world institutions.
method Machine Learning for data analysis and panel regression.
result Approximation of SCOPUS Impact Factor for free.
Study evaluates the impact of academic support center's face-to-face assistance on student performance.
problem Underestimation of Academic Support Center's true impact due to group bias.
method Applied causal inference theory and T-learner to evaluate conditional average treatment effect (CATE) of F2F personal assistance.
result Developed a new CATE function that depends on the number of F2F sessions, predicting improved CATE performance.
Paper introduces a new algorithm to detect LLM-generated text.
problem Detecting LLM-generated text to prevent misinformation.
method Adaptively learns the distance between original and rewritten text.
result Empirically, the new algorithm outperforms existing methods in most scenarios.
KT models struggle with student concept drift, but BKT remains the most stable.
problem Impact of student concept drift on KT models.
method Applied four KT models to five academic years of data.
result KT models generally degrade in performance with concept drift, BKT remains stable.
AI predicts dyslexia and dysgraphia in children.
problem Early detection and assessment of dyslexia and dysgraphia.
method Machine learning on datasets of children's handwriting and audio recordings.
result Preliminary model shows high performance in classifying dyslexic and dysgraphic children.
Social media enhances or diminishes scientific status, depending on usage.
problem Impact of social media on scientific stratification and mobility.
method Logistic Attribution Analysis combining statistical and machine learning methods.
result Social media promotes stratification and mobility, but beyond a threshold, it negatively impacts status.
Improved graph matching using covariates for network data integration.
problem Matching networks without unique identifiers.
method Two novel covariate-assisted seeded graph matching methods.
result Improved alignment accuracy through covariate information.
We develop an algorithm that can detect pneumonia from chest X-rays at a level exceeding practicing radiologists. Our algorithm, CheXNet, is a 121-layer convolutional neural network trained on ChestX-ray14, currently the largest publicly available chest X-ray dataset, containing over 100,000 frontal-view X-ray images w…
The study uses Hidden Markov Models to analyze student enrollment patterns and academic performance.
problem Limited understanding of how enrollment patterns affect academic performance.
method Applied Hidden Markov Models to categorize enrollment strategies and compare academic outcomes.
result Mixed enrollment strategies lead to better academic performance, especially during part-time semesters.
As recent literature has demonstrated how classifiers often carry unintended biases toward some subgroups, deploying machine learned models to users demands careful consideration of the social consequences. How should we address this problem in a real-world system? How should we balance core performance and fairness me…
Deep learning predicts S&P 500 index direction.
problem Accurate stock price prediction remains challenging.
method Convolutional neural network model for S&P 500 index forecasting.
result Model achieves over 55% accuracy in predicting index direction.
Most of the current game-theoretic demand-side management methods focus primarily on the scheduling of home appliances, and the related numerical experiments are analyzed under various scenarios to achieve the corresponding Nash-equilibrium (NE) and optimal results. However, not much work is conducted for academic or c…
A new method for name disambiguation in academic networks using multi-view attention and recurrent neural networks.
problem Disambiguating authors with the same name in large-scale academic networks.
method Multi-view Attention-based Pairwise Recurrent Neural Network (MA-PairRNN) that divides papers into blocks based on author attributes and merges blocks of the same author.
result MA-PairRNN significantly improves name disambiguation performance on real-world datasets.
Estimates citation impact to recommend best publication venue.
problem Choosing optimal publication venue for academic papers.
method Treatment effect estimation and bias correction method.
result Effective recommendation of publication venues based on citation potential.
State-of-the-art approaches for Knowledge Base Completion (KBC) exploit deep neural networks trained with both false and true assertions: positive assertions are explicitly taken from the knowledge base, whereas negative ones are generated by random sampling of entities. In this paper, we argue that random sampling is …
As a powerful tool of asynchronous event sequence analysis, point processes have been studied for a long time and achieved numerous successes in different fields. Among various point process models, Hawkes process and its variants attract many researchers in statistics and computer science these years because they capt…
Bitcoin price prediction models fail to outperform a simple 'today's price' baseline, especially at longer horizons.
problem Lack of robust models that consistently outperform a naive price predictor at various horizons.
method Surveyed peer-reviewed papers, categorized by evaluation methodology, contrasted with social media discourse, and proposed methodological standards.
result No peer-reviewed study has shown robust superiority over the naive baseline across multiple market regimes at short-to-medium horizons.
Efficiently trains BERT on academic GPUs in 12 days.
problem Training large-scale BERT models is expensive and time-consuming.
method Optimizes training on multiple GPUs and nodes, reducing costs.
result Trains BERT on academic GPUs in 12 days, not requiring expensive hardware.
Develops NFCF to reduce gender bias in social media recommendation systems.
problem Reduces gender bias in collaborative filtering systems on social media data.
method Pre-training and fine-tuning neural collaborative filtering with bias correction techniques.
result Achieves better performance and fairness in gender de-biased recommendations.
Shifu2 discovers advisor-advisee relationships in collaboration networks.
problem Discovering hidden advisor-advisee relationships in scientific collaboration networks.
method Network Representation Learning (NRL) model, considering both network structure and node/edge semantics.
result Improved stability and effectiveness compared to state-of-the-art methods.
Kaggle competitions offer valuable insights for business forecasting.
problem Lack of attention to Kaggle competitions in academic forecasting studies.
method Review of results from six Kaggle competitions featuring real-life business forecasting tasks.
result Global ensemble models outperform local single models in Kaggle competitions.
This report was originally written as an industry white paper on Hedge Funds. This paper gives an overview to Hedge Funds, with a focus on risk management issues. We define and explain the general characteristics of Hedge Funds, their main investment strategies and the risk models employed. We address the problems in H…
This reviews the econophysics activities in Belgium from my admittedly biased point of view. Unknown historical notes or facts are presented for the first time explaining the aims, whence evolution of the research papers and friendly connections with colleagues. Comments on endeavors are also provided. The lack of offi…
The failure of landing a job for college students could cause serious social consequences such as drunkenness and suicide. In addition to academic performance, unconscious biases can become one key obstacle for hunting jobs for graduating students. Thus, it is necessary to understand these unconscious biases so that we…
Computer generated academic papers have been used to expose a lack of thorough human review at several computer science conferences. We assess the problem of classifying such documents. After identifying and evaluating several quantifiable features of academic papers, we apply methods from machine learning to build a b…
Large speech dataset for commercial use with 9.98% word error rate.
problem Creating a diverse speech recognition dataset for commercial purposes.
method Internet search for licensed audio data with transcriptions, training model on the dataset.
result Model trained on dataset achieves 9.98% word error rate on Librispeech's test-clean test set.
Bayesian Causal Forests model assesses part-time work's impact on student growth.
problem Estimating causal effects of part-time work on student growth in mathematics achievement.
method Longitudinal Bayesian Causal Forests model combining non-parametric and difference-in-differences methods.
result Negative impact of part-time work for most students, potential benefits for those with low school belonging, widening achievement gap identified.
Enhances insurance loss models using InsurTech data and machine learning.
problem Traditional insurance loss models lack predictive accuracy due to limited data sources.
method Combining proprietary claims data with InsurTech data and applying machine learning techniques.
result Improved predictive accuracy of the loss model through machine learning.
Autonomous driving is getting a lot of attention in the last decade and will be the hot topic at least until the first successful certification of a car with Level 5 autonomy. There are many public datasets in the academic community. However, they are far away from what a robust industrial production system needs. Ther…
The professional services sector is at a turning point, with some industries showing growth opportunities.
problem Identifying growth opportunities in the professional services sector after decades of growth.
method A simple framework applied to the US economic context to diagnose growth opportunities.
result The professional services sector is expected to stall at a national level, but some industries still offer growth opportunities.
Accurate time series prediction over long future horizons is challenging and of great interest to both practitioners and academics. As a well-known intelligent algorithm, the standard formulation of Support Vector Regression (SVR) could be taken for multi-step-ahead time series prediction, only relying either on iterat…
Replicated and validated Rank-N-Contrast for robust regression.
problem Deep regression models struggle with continuous sample orders.
method Contrastive learning of continuous representations by ranking samples.
result Improved performance and robustness of RNC framework.
Lecture notes introduce differential geometry using sheaves and differential operators.
problem Exploring differential geometry concepts.
method Using sheaves, differential operators, and horizontal subbundles.
result Presented an approach to fundamental differential geometry structures.
The paper explores neural networks for improving delta hedging in financial markets.
problem Real-world financial markets do not perfectly match the assumptions of the Black-Scholes model.
method The authors test various neural architectures (RNN, TCN, Attention, MLP) for delta hedging and combine them with traditional models.
result NNHedge framework provides a pipeline for model development and assessment.
The paper proposes to analyze a data set of Finnish ranks of academic publication channels with Extreme Learning Machine (ELM). The purpose is to introduce and test recently proposed ELM-based mislabel detection approach with a rich set of features characterizing a publication channel. We will compare the architecture,…
New algorithm for mean estimation in add-remove model achieves optimal error.
problem Mean estimation in add-remove model of differential privacy.
method Proposed new algorithm achieving min-max optimality.
result Achieves best possible constant in mean squared error for all ε.
Paper uses ML to predict SME defaults with interpretability.
problem Lack of interpretability in ML models for SME default prediction.
method Model-agnostic approach using Accumulated Local Effects and Shapley values.
result eXtreme Gradient Boosting algorithm provides highest classification power with interpretability.
This paper develops models for cryptocurrency trading and evaluates Bitcoin options.
problem Evaluating European options on Bitcoin with realistic models.
method Jump-diffusion models for cryptocurrency dynamics.
result Models accurately predict Bitcoin option prices.
Recent academic work has developed a method to determine, in real time, if a given stock is exhibiting a price bubble. Currently there is speculation in the financial press concerning the existence of a price bubble in the aftermath of the recent IPO of LinkedIn. We analyze stock price tick data from the short lifetime…