LaRT models LLMs' response accuracy and CoT length to evaluate reasoning ability and speed.
problem Valid evaluation of Large Language Models (LLMs) via response accuracy and chain-of-thought length.
method Introduces Latency-Response Theory (LaRT) to jointly model response accuracy and CoT length using latent ability and latent speed.
result LaRT yields higher estimation accuracy and shorter confidence intervals for latent traits compared to IRT.
The paper proposes a new method to evaluate LLM agent responses using ECDF clustering.
problem The standard evaluation of LLM agent responses via majority voting obscures response quality and distribution.
method The paper introduces a novel evaluation framework based on ECDF of cosine similarities and clustering of ECDFs using distances and k-medoids algorithm. result ECDF clustering reveals interpretable group structures in LLM responses, offering insights into agent settings.
Study evaluates different price response definitions for NASDAQ stocks.
problem Understanding the long-lasting effects of trading activity on stock prices.
method Examined two different price response implementations for NASDAQ Trades and Quotes (TAQ) data.
result Results are qualitatively the same for two different time scale definitions, but response can vary by up to a factor of two.
There are non-vanishing price responses across different stocks in correlated financial markets. We further study this issue by performing different averages, which identify active and passive cross-responses. The two average cross-responses show different characteristic dependences on the time lag. The passive cross-r…
IRT improves algorithm evaluation across datasets.
problem Evaluating the performance of algorithm portfolios.
method Modified IRT framework for evaluating algorithm portfolios across datasets.
result Richer characteristics of algorithm performance are revealed.
When response variables are nominal and populations are cross-classified with respect to multiple polytomies, questions often arise about the degree of association of the responses with explanatory variables. When populations are known, we introduce a nominal association vector and matrix to evaluate the dependence of …
Meta-Router optimizes LLM selection using gold-standard and preference-based data.
problem Training a high-quality LLM router with combined data sources is challenging due to bias and scarcity.
method Developed an integrative causal router training framework to correct bias and improve routing accuracy.
result Our approach delivers more accurate routing and improves the trade-off between cost and quality.
Bayesian optimization (BO) aims to minimize a given blackbox function using a model that is updated whenever new evidence about the function becomes available. Here, we address the problem of BO under partially right-censored response data, where in some evaluations we only obtain a lower bound on the function value. T…
Study optimizes classifiers for credit card mail campaigns and default prediction.
problem Optimizing classifiers for credit card mail campaigns and default prediction.
method Three distinct models: response, risk, and response-risk. Optimized various performance metrics.
result Random Forest classifier achieves highest accuracy (83.2%) in multi-class response-risk model.
ChatGPT's medical response accuracy is 56%, but studies vary widely.
problem Lack of standard guidelines for evaluating ChatGPT's performance in medicine.
method Systematic review and meta-analysis of 17 studies.
result ChatGPT's overall integrated accuracy in medical queries is 56%.
Paper introduces metrics for evaluating multi-agent policies using best response dynamics.
problem Evaluation and ranking of multi-agent policies in reinforcement learning.
method Adopting strict best response dynamics (SBRD) to model selfish behaviors, proposing perturbed SBRD for dynamic and non-stationary settings.
result Proposed perturbed SBRD can observe policies with maximum metrics and differ from optimal by any given tolerance.
New method for explaining dialogue response generation models.
problem Interpreting sequence generation models, especially dialogue response generation.
method Local Explanation of Response Generation (LERG) method.
result LERG improves dialogue response generation explanations compared to existing methods.
Neural network-based Open-ended conversational agents automatically generate responses based on predictive models learned from a large number of pairs of utterances. The generated responses are typically acceptable as a sentence but are often dull, generic, and certainly devoid of any emotion. In this paper, we present…
ISMCTS-BR learns best responses in large games, approximating worst-case performance.
problem Learning robustness to worst-case outcomes in large games.
method ISMCTS-BR, a scalable search-based algorithm for deep reinforcement learning.
result ISMCTS-BR approximates worst-case performance in large games.
In this study we present a kernel based convolution model to characterize neural responses to natural sounds by decoding their time-varying acoustic features. The model allows to decode natural sounds from high-dimensional neural recordings, such as magnetoencephalography (MEG), that track timing and location of human …
Inference methods are often formulated as variational approximations: these approximations allow easy evaluation of statistics by marginalization or linear response, but these estimates can be inconsistent. We show that by introducing constraints on covariance, one can ensure consistency of linear response with the var…
Paper estimates AI hallucinations in conditional generation tasks.
problem Estimating the frequency of AI-generated incorrect responses.
method Developed a method to estimate hallucination probability from generated responses and log probabilities.
result Method accurately estimates hallucination rate in natural language and synthetic tasks.
Scorio.jl ranks systems from repeated tasks using various methods.
problem Evaluating and ranking systems from repeated responses to shared tasks.
method Common tensor-based interface for multiple ranking methods.
result Pilot experiments show stability and runtime scaling.
Proposes an IRT-based ensemble method to improve machine learning accuracy.
problem Improving the accuracy of machine learning models, especially for hard-to-classify instances.
method Introduces Item Response Theory (IRT) to evaluate sample difficulty and classifier ability, creating three models with different assumptions.
result The proposed IRT ensemble model outperforms other methods on 19 datasets.
We propose an adversarial learning approach for generating multi-turn dialogue responses. Our proposed framework, hredGAN, is based on conditional generative adversarial networks (GANs). The GAN's generator is a modified hierarchical recurrent encoder-decoder network (HRED) and the discriminator is a word-level bidirec…
Active learning suffers from biased non-response, which this paper addresses.
problem Active learning's effectiveness is compromised by biased non-response in real-world contexts.
method Proposes a cost-based correction to the sampling strategy, UCB-EU, to mitigate the impact of biased non-response.
result UCB-EU successfully reduces the harm from labelling non-response in many settings.
Automated dialogue quality evaluation using user satisfaction estimates across multiple domains.
problem Lack of automated and domain-independent dialogue quality evaluation metrics.
method Created a new Response Quality annotation scheme, introduced five domain-independent feature sets, and experimented with six machine learning models.
result Gradient Boosting Regression model achieved best prediction performance, with a 16% relative improvement in binary satisfaction class prediction accuracy.
Demand response is designed to motivate electricity customers to modify their loads at critical time periods. The accurate estimation of impact of demand response signals to customers' consumption is central to any successful program. In practice, learning these response is nontrivial because operators can only send a …
Framework assesses autograders' reliability and biases.
problem Mixed reliability and biases in autograders for LLM evaluation.
method Bayesian GLMs to model evaluation outcomes.
result Explicit quantification of scoring differences and biases.
Detects physiological patterns to hemodynamic stress using unsupervised deep learning.
problem Identify and characterize physiological responses to hemorrhage in raw vital sign data.
method Transform vital sign time series into latent space using unsupervised deep learning, identify clusters, and evaluate latent embeddings.
result Clusters in latent embeddings correspond to physiological response patterns matching physicians' intuition.
Objective: Predict individual septic children's personalized physiologic responses to vasoactive titrations by training a Recurrent Neural Network (RNN) using EMR data. Materials and Methods: This study retrospectively analyzed EMR of patients admitted to a pediatric ICU from 2009 to 2017. Data included charted time se…
When simulating a complex stochastic system, the behavior of output response depends on input parameters estimated from finite real-world data, and the finiteness of data brings input uncertainty into the system. The quantification of the impact of input uncertainty on output response has been extensively studied. Most…
A new framework evaluates large language models efficiently and accurately.
problem Evaluation of large language models is challenging due to stochasticity and heterogeneity of benchmarks.
method Interpretable and scalable framework based on Item Response Theory (IRT) and majorization-minimization principle.
result Our method achieves superior scalability and interpretability compared to existing approaches.
We introduce the multiresolution recurrent neural network, which extends the sequence-to-sequence framework to model natural language generation as two parallel discrete stochastic processes: a sequence of high-level coarse tokens, and a sequence of natural language tokens. There are many ways to estimate or learn the …
The study optimizes sampling in complex systems with probabilistic response distributions.
problem Calibrating and optimizing complex systems with probabilistic response distributions.
method Non-parametric Bayesian approach to modeling spatial fields of probability distributions, introducing adaptive sampling strategies.
result Adaptive sampling strategies improve system evaluations by guiding focus towards key features.
AI predicts dementia onset from emotional face evaluations.
problem Early detection of dementia in aging societies.
method Behavioral responses analysis and AI regression.
result Encouraging AI-based prediction results for MoCA scores.
Identity-link IRT improves TVD-MI scores without curvature violations.
problem Preserving additivity in TVD-MI scores for efficient LLM evaluation.
method Derives clipped-linear model from Gini entropy maximization, using identity link.
result Identity-link yields lower curvature violations (median curl 0.080-0.150) compared to probit/logit.
Paper proposes using generalized lambda distributions for stochastic simulators.
problem Uncertainty quantification with complex stochastic models is computationally challenging.
method Flexible generalized lambda distribution approximates response PDF, parameters are sparse polynomial chaos expansions.
result Local inference of response PDF at each point of experimental design using replicated model evaluations.
Study off-policy evaluation and learning in dynamic pricing with context.
problem Dynamic personalized pricing and operations management problems with high-dimensional user types.
method Formalize causal structure, leverage single time-step evaluation, estimate marginal MDP.
result Improved out-of-sample policy performance in dynamic and capacitated pricing.
New IRT method identifies useful datasets for ML classifier evaluation.
problem Lack of standard evaluation strategy for ML benchmarks.
method Applied Item Response Theory (IRT) to OpenML-CC18 benchmark.
result Not all datasets are useful for evaluating classifiers.
Efficient methods estimate concordance probability for big data.
problem Efficiently calculating concordance probability in large datasets.
method Proposes two estimation methods for discrete and continuous settings.
result Estimators are accurate and computationally efficient.
Item Response Theory (IRT) aims to assess latent abilities of respondents based on the correctness of their answers in aptitude test items with different difficulty levels. In this paper, we propose the β3-IRT model, which models continuous responses and can generate a much enriched family of Item Characteristic Cur…
Study presents MMC model for better fitting multiple choice data.
problem Improving accuracy of latent trait estimates in IRT models.
method Fit autoencoders to MMC model, demonstrating better fit than nominal response model.
result MMC model outperforms traditional IRT models in fit.
New kernel methods estimate complex causal relationships.
problem Estimating nonparametric causal functions like dose-response curves.
method Kernel ridge regression with decomposition property.
result Uniform consistency with finite sample rates proved.
A hybrid ML method improves ship response predictions across different sea conditions.
problem Improving accuracy and generalizability of ML methods for ship response predictions.
method A hybrid machine learning method that corrects forces in a low-fidelity equation of motion.
result The hybrid method offers improved prediction accuracy and generalizability compared to benchmarks.
In this paper, we introduce a novel method to interpret recurrent neural networks (RNNs), particularly long short-term memory networks (LSTMs) at the cellular level. We propose a systematic pipeline for interpreting individual hidden state dynamics within the network using response characterization methods. The ranked …
An automated metric to evaluate dialogue quality is vital for optimizing data driven dialogue management. The common approach of relying on explicit user feedback during a conversation is intrusive and sparse. Current models to estimate user satisfaction use limited feature sets and rely on annotation schemes with low …
We exploit a continuous time random walk description of stock prices to obtain a fast and accurate evaluation of their volatility from intraday data. We show that financial markets are usefully described as open physical systems. Indeed we find that the process determining market volatility is not stationary while the …
MSRL learns a representation maximizing mutual info with response variables.
problem Learning sufficient representations for complex, multi-dimensional data.
method Variational mutual information, deep neural networks, generalized Dudley's inequality.
result MSRL achieves consistent and accurate representation learning.
Study develops efficient algorithm for probabilistic penetration response of composite plates.
problem Probabilistic modeling of discrete structural response, focusing on binary events like buckling.
method Adaptive domain-based decomposition, sparse grid sampling, assumption of monotonic behavior.
result Efficient computational framework for probabilistic penetration response of composite plates.
Study finds no significant short-term impact on liquidity supply after protocol fees were reduced.
problem Liquidity provider welfare is affected by protocol fees, but the impact on liquidity supply is unclear.
method Used a matched-overlap event-study difference-in-differences design to estimate the liquidity-supply response to take-rate cuts.
result No significant short-term impact on active liquidity or local depth; no change in LP participation or composition.
Incorporating nonlinearity is paramount to predicting the future states of a dynamical system, its response to shocks, and its underlying causal network. However, most existing methods for causality detection and impulse response, such as Vector Autoregression (VAR), assume linearity and are thus unable to capture the …
Paper develops methods to estimate derivative of dose-response curve for continuous treatments.
problem Estimating the derivative of the dose-response curve for continuous treatments.
method Doubly robust (DR) inference method using kernel smoothing, bias-corrected IPW and DR estimators.
result Proposes novel bias-corrected IPW and DR estimators for continuous treatments.