R2DE assesses new exam questions quickly and accurately.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
GPT-4 assesses its confidence in answering USMLE questions with and without feedback.
EduQG generates better educational questions by pre-training on scientific text.
PerSense assesses personality traits from text for commonsense reasoning.
AI tested on 10 math questions from research.
New method approximates CV for model assessment and selection.
Shai is a 10B model for asset management tasks, outperforming baselines.
Testing symmetry of a probability distribution is a common question arising from applications in several fields. Particularly, in the study of observables used in the analysis of stock market index variations, the question of symmetry has not been fully investigated by means of statistical procedures. In this work a di…
Study evaluates neural networks for corporate credit rating assessment.
This paper assesses Gaussian and Exponential mechanisms for certifying adversarial robustness.
When assessing group solvency, an important question is to what extent intragroup transfers may be considered, as this determines to which extent diversification can be achieved. We suggest a framework to describe the families of admissible transfers that range from the free movement of capital to excluding any transac…
Although compelling assessments have been examined in recent years, more studies are required to yield a better understanding of the several methods where assessment techniques significantly affect student learning process. Most of the educational research in this area does not consider demographics data, differing met…
Automatically assesses the quality of online health articles.
A problem faced by many instructors is that of designing exams that accurately assess the abilities of the students. Typically these exams are prepared several days in advance, and generic question scores are used based on rough approximation of the question difficulty and length. For example, for a recent class taught…
Simplifies F-measure for better interpretability.
This paper addresses clustering with missing data using Rubin's rules.
Active Learning (AL) methods seek to improve classifier performance when labels are expensive or scarce. We consider two central questions: Where does AL work? How much does it help? To address these questions, a comprehensive experimental simulation study of Active Learning is presented. We consider a variety of tasks…
FactTest assesses LLM factuality with Type I error control.
As deep learning applications are becoming more and more pervasive in robotics, the question of evaluating the reliability of inferences becomes a central question in the robotics community. This domain, known as predictive uncertainty, has come under the scrutiny of research groups developing Bayesian approaches adapt…
New method assesses individual training points' privacy risk without retraining.
New ensemble models classify mouse movement trajectories to assess survey question difficulty.
New methods assess rankability of datasets for meaningful item rankings.
Automated scoring engines are increasingly being used to score the free-form text responses that students give to questions. Such engines are not designed to appropriately deal with responses that a human reader would find alarming such as those that indicate an intention to self-harm or harm others, responses that all…
The objective of ordinal embedding is to find a Euclidean representation of a set of abstract items, using only answers to triplet comparisons of the form "Is item closer to the item or item ?". In recent years, numerous algorithms have been proposed to solve this problem. However, there does not exist a fai…
LLMs show surprising confidence in their answers, beyond just tokens.
New framework assesses graph-learning datasets for better evaluation.
Community-based Question Answering (CQA) sites play an important role in addressing health information needs. However, a significant number of posted questions remain unanswered. Automatically answering the posted questions can provide a useful source of information for online health communities. In this study, we deve…
Most work in machine reading focuses on question answering problems where the answer is directly expressed in the text to read. However, many real-world question answering problems require the reading of text not because it contains the literal answer, but because it contains a recipe to derive an answer together with …
Develops method to assess feature importance in black-box models for unconditional distribution.
Optimal allocation of human effort to correct AI assessments in decision-making.
Study evaluates bias mitigation methods in deep learning, finds they often exploit hidden biases.
Survey data imputation methods impact feature selection and importance assessment.
We propose SPARFA-Trace, a new machine learning-based framework for time-varying learning and content analytics for education applications. We develop a novel message passing-based, blind, approximate Kalman filter for sparse factor analysis (SPARFA), that jointly (i) traces learner concept knowledge over time, (ii) an…
Using sequence to sequence algorithms for query expansion has not been explored yet in Information Retrieval literature nor in Question-Answering's. We tried to fill this gap in the literature with a custom Query Expansion engine trained and tested on open datasets. Starting from open datasets, we built a Query Expansi…
A framework isolates VQA reasoning from perception for better model evaluation.
Bayesian test assesses dependence between mixed data types.
One question central to Reinforcement Learning is how to learn a feature representation that supports algorithm scaling and re-use of learned information from different tasks. Successor Features approach this problem by learning a feature representation that satisfies a temporal constraint. We present an implementation…
This study introduces balanced DRPS and OrderedLogitNN for better QDE of discrete-level questions.
This paper investigates uncertainty calibration in multimodal large language models.
Framework assesses autograders' reliability and biases.
A cluster tree provides a highly-interpretable summary of a density function by representing the hierarchy of its high-density clusters. It is estimated using the empirical tree, which is the cluster tree constructed from a density estimator. This paper addresses the basic question of quantifying our uncertainty by ass…
Benchmark study evaluates 8 clustering methods on 99 UCR time series datasets.
Measures three types of noise in LLM evaluations.
In recent years, many non-traditional classification methods, such as Random Forest, Boosting, and neural network, have been widely used in applications. Their performance is typically measured in terms of classification accuracy. While the classification error rate and the like are important, they do not address a fun…
Modern machine learning methods are critical to the development of large-scale personalized learning systems that cater directly to the needs of individual learners. The recently developed SPARse Factor Analysis (SPARFA) framework provides a new statistical model and algorithms for machine learning-based learning analy…
The present research aims to highlight the main factors influencing the development of entrepreneurial innovation in a rural environment and to perform an empirical study with the purpose of assessing the main problems in rural development. The research performed is mostly of a quantitative nature, being based on the u…
Many questions in Data Science are fundamentally causal in that our objective is to learn the effect of some exposure, randomized or not, on an outcome interest. Even studies that are seemingly non-causal, such as those with the goal of prediction or prevalence estimation, have causal elements, including differential c…
Digital personas improve survey results for stable attributes but fail for subjective responses.