New method evaluates language model forecasters by checking consistency of predictions.
problem Evaluating the performance of language model forecasters is difficult due to lack of ground truth.
method Developed a consistency check framework based on arbitrage to evaluate forecasters.
result Consistency metrics correlate with ground truth performance of LLM forecasters.
Time-aware fact-checking improves veracity predictions for time-sensitive claims.
problem Fact-checking decisions should consider temporal information of claims and evidence.
method Investigated four temporal ranking methods to optimize evidence ranking for fact-checking models.
result Time-aware evidence ranking surpasses relevance assumptions and improves veracity predictions for time-sensitive claims.
Research aims to make fact-checking models more transparent.
problem Making fact-checking models explainable in a complex field.
method Combines fact-checking methods with explainable AI techniques.
result Developed initial solutions for explainable fact-checking.
We present FAKTA which is a unified framework that integrates various components of a fact checking process: document retrieval from media sources with various types of reliability, stance detection of documents with respect to given claims, evidence extraction, and linguistic analysis. FAKTA predicts the factuality of…
Social networks are getting closer to our real physical world. People share the exact location and time of their check-ins and are influenced by their friends. Modeling the spatio-temporal behavior of users in social networks is of great importance for predicting the future behavior of users, controlling the users' mov…
New research shows calibration error is flawed when dealing with model uncertainty.
problem Current model evaluation techniques conflate model uncertainty with aleatoric uncertainty.
method Posterior predictive checks to evaluate deep learning models.
result Calibration error and variants are incorrect when model uncertainty is present.
New method improves spatial prediction validation accuracy.
problem Validation methods fail for spatial prediction tasks due to mismatch between validation and test locations.
method Proposes a new validation method that adapts existing covariate-shift ideas to spatial settings.
result Proves and demonstrates the new method's superiority in spatial prediction validation.
The chapter improves deep learning models by interpreting and improving their performance.
problem Deep learning models often lack interpretability, leading to poor understanding of their predictions.
method The approach involves attributing importance to features and feature groups, including interactions, to improve model performance.
result The proposed attributions provide insights across various domains and can be used to improve model generalization.
SOAK assesses data subset similarity for better model training.
problem Estimating similarity between data subsets for accurate predictions.
method Same/Other/All K-fold cross-validation method.
result SOAK estimates similarity of learnable/predictable patterns in data subsets.
The choice of model class is fundamental in statistical learning and system identification, no matter whether the class is derived from physical principles or is a generic black-box. We develop a method to evaluate the specified model class by assessing its capability of reproducing data that is similar to the observed…
We contribute the largest publicly available dataset of naturally occurring factual claims for the purpose of automatic claim verification. It is collected from 26 fact checking websites in English, paired with textual sources and rich metadata, and labelled for veracity by human expert journalists. We present an in-de…
With the availability of vast amounts of user visitation history on location-based social networks (LBSN), the problem of Point-of-Interest (POI) prediction has been extensively studied. However, much of the research has been conducted solely on voluntary checkin datasets collected from social apps such as Foursquare o…
Developed a machine-checked Itô calculus for Brownian motion.
problem Formal verification of Itô calculus for Brownian motion.
method Machine-checked formalization in Lean over Mathlib.
result First machine-checked constructions of the Itô integral and Itô's formula.
We give a definition of an integer-valued function ∑iαixi∗ derived from arrow diagrams for the ambient isotopy classes of oriented spherical curves. Then, we introduce certain elements of the free Z-module generated by the arrow diagrams with at most l arrows, called relators of Type~($\check{…
We propose a procedure for assigning a relevance measure to each explanatory variable in a complex predictive model. We assume that we have a training set to fit the model and a test set to check the out of sample performance. First, the individual relevance of each variable is computed by comparing the predictions in …
A machine-checked Itô calculus for Brownian motion on [0,T]
problem Developing an L2 Itô calculus for Brownian motion method Formalized in Lean 4 on top of Mathlib and the BrownianMotion package
result First machine-checked proof of Itô's formula and construction of Itô integral as martingale-valued process
LogGENE uses log-cosh loss for deep learning in gene expression datasets, improving accuracy and interpretability.
problem Mining large gene expression datasets for reliable deep learning predictions.
method Develops a smooth alternative to check loss (log-cosh) for quantile regression in gene expression datasets.
result Achieves state-of-the-art performance in accuracy and provides robust uncertainty estimates.
This paper deforms complex tori and their mirrors using gerbes.
problem Deforming complex tori and their mirror partners.
method Using flat gerbes to deform complex tori and their mirrors, constructing holomorphic line bundles over deformed objects.
result Deformed complex tori and their mirrors can be studied using flat gerbes.
We present SemEval-2019 Task 8 on Fact Checking in Community Question Answering Forums, which features two subtasks. Subtask A is about deciding whether a question asks for factual information vs. an opinion/advice vs. just socializing. Subtask B asks to predict whether an answer to a factual question is true, false or…
PCS-UQ framework improves uncertainty quantification for machine learning models.
problem Ensuring trustworthy uncertainty quantification for machine learning models in high-stakes domains.
method PCS-UQ framework based on Predictability, Computability, and Stability principles, integrating prediction-checking, bootstrap samples, and multiplicative calibration.
result PCS-UQ maintains target coverage while outperforming or matching conformal methods in interval width and subgroup coverage.
Statistical model checking for PCTL on MDPs using reinforcement learning.
problem Model checking PCTL specifications on MDPs with statistical methods.
method Reinforcement learning for policy search, statistical model checking with UCB-based Q-learning.
result Provably guaranteed statistical model checking method for PCTL specifications on MDPs.
We conjecture a relation between the sl(N) knot homology, recently introduced by Khovanov and Rozansky, and the spectrum of BPS states captured by open topological strings. This conjecture leads to new regularities among the sl(N) knot homology groups and suggests that they can be interpreted directly in topological st…
DC-Check helps guide ML development by considering data-centric aspects.
problem Lack of standardized framework for data-centric considerations in ML.
method DC-Check is a checklist-style framework for data-centric AI at ML pipeline stages.
result Promotes thoughtfulness and transparency in ML development.
New model corrects bias in crowdsourced ratings for diverse items.
problem Bias and noise in crowdsourced ratings for training data.
method Bayesian rating model with item-level effects for difficulty, discriminativeness, and guessability.
result New model avoids bias in training data, improving model goodness of fit.
We check the claims that data from Google Trends contain enough data to predict future financial index returns. We first discuss the many subtle (and less subtle) biases that may affect the backtest of a trading strategy, particularly when based on such data. Expectedly, the choice of keywords is crucial: by using an i…
Paper presents a universal baseline for binary prediction models.
problem Need a robust baseline to evaluate model performance.
method Dutch Draw (DD) baseline method for binary classification models.
result Reduces to almost always predicting zero or one in most situations.
Proves SYZ mirror symmetry for del Pezzo and rational elliptic surfaces.
problem Proving mirror symmetry for specific Calabi-Yau surfaces.
method Adapting Hein's work, constructing asymptotically semi-flat Calabi-Yau metrics, and defining a mirror map.
result Existence and uniqueness of Calabi-Yau metrics on Y∖D. Proposes a copula-based filter for diabetes risk prediction.
problem Feature selection for robust and interpretable predictive modeling in medicine, especially for extreme patient strata.
method Copula-based supervised filter using Gumbel-copula implied upper-tail concordance score (lambda U).
result The proposed filter outperforms standard filters and provides clinically coherent predictors.
This work improves motion planning for quadcopters by learning and reasoning about controller performance.
problem Improving motion planning for quadcopters with safety margins and execution reliability.
method Introspective learning and reasoning to correct execution bias and improve collision checking.
result Substantial reduction in safety margins for motion actions, leading to safer execution.
We describe the infinitesimal moduli space of pairs (Y,V) where Y is a manifold with G2 holonomy, and V is a vector bundle on Y with an instanton connection. These structures arise in connection to the moduli space of heterotic string compactifications on compact and non-compact seven dimensional spaces, e.…
By the SYZ construction, a mirror pair (X,Xˇ) of a complex torus X and a mirror partner Xˇ of the complex torus X is described as the special Lagrangian torus fibrations X→B and Xˇ→B on the same base space B. Then, by the SYZ transform, we can construct a simpl…
We prove the following result announced in Todorov and Valov: Any homogeneous, metric ANR-continuum is a VGn-continuum provided dimGX=n≥1 and Hˇn(X;G)=0, where G is a principal ideal domain. This implies that any homogeneous n-dimensional metric ANR-continuum with $\check{H}^n(X;G)\neq…
We specify a result of Yokoi \cite{yo} by proving that if G is an abelian group and X is a homogeneous metric ANR compactum with dimGX=n and Hˇn(X;G)=0, then X is an (n,G)-bubble. This implies that any such space X has the following properties: Hˇn−1(A;G)=0 for every closed…
Paper checks SSC for matrix factorizations using Gurobi.
problem Checking the SSC for various matrix factorizations.
method Formulated as a non-convex quadratic optimization problem over a bounded set, solved with Gurobi.
result SSC can be checked in reasonable time for realistic scenarios.
New diagnostic method detects misspecified models in inverse PDE problems.
problem Misleading residual-norm diagnostics in inverse PDE problems.
method Structure-sensitive sequential diagnostic using e-processes.
result Rejects fitted models that produce biased predictions.
A variation of the Minority Game has been applied to study the timing of promotional actions at retailers in the fast moving consumer goods market. The underlying hypotheses for this work are that price promotions are more effective when fewer than average competitors do a promotion, and that a promotion strategy can b…
Study how past radiation determines present matter in Penrose's cyclic cosmology.
problem Determining matter content in the present eon from past radiation in Penrose's cyclic cosmology.
method Solve Einstein's equations for a spherical wave in the past eon, then apply reciprocity to find the present eon's matter content.
result The present eon is filled with three types of radiation: a damped wave, an in-going wave, and randomly scattered waves.
Next point-of-interest (POI) recommendation aims to offer suggestions on which POI to visit next, given a user's POI visit history. This problem has a wide application in the tourism industry, and it is gaining an increasing interest as more POI check-in data become available. The problem is often modeled as a sequenti…
Paper improves ETF tail-risk monitoring reliability.
problem Unreliable ETF risk monitoring under degraded data.
method Combines quality checks, prediction, scoring, and adjustment.
result Improves tail-risk monitoring, especially during stressed periods.
New method simplifies checking consistency of differentiable loss functions.
problem Verifying consistency of differentiable loss functions is difficult.
method Developed a new approach called strong indirect elicitation (strong IE) to simplify checking consistency.
result Strong IE is equivalent to calibration for strongly convex, differentiable surrogates.
This study compares deep learning and statistical models for stock price forecasting.
problem Accurate stock price prediction is challenging due to market volatility.
method Used deep learning (LSTM, RNN, CNN, FULL CNN) and statistical models (ARIMA, Moving Averages) on S&P 500 data.
result LSTM model showed the lowest Mean Absolute Error (MAE), indicating highest accuracy.
Mathematical framework for brane quantization using SYZ mirror symmetry.
problem Developing a mathematical framework for brane quantization.
method Applying SYZ mirror symmetry to construct and analyze branes.
result Established a mathematical definition of endomorphism algebras and their isomorphisms.
Proposes flexible spatial models for better understanding spatial heterogeneity.
problem Poor characterisation of spatial heterogeneity in conventional models.
method Spatial Bayesian Neural Networks (SBNNs) incorporating a spatial embedding layer and possibly spatially-varying parameters.
result SBNNs better match the finite-dimensional distribution of target spatial processes.
Q-learning with function approximation is one of the most popular methods in reinforcement learning. Though the idea of using function approximation was proposed at least 60 years ago, even in the simplest setup, i.e, approximating Q-functions with linear functions, it is still an open problem on how to design a pr…
We study the Yamabe invariants of cylindrical manifolds and compact orbifolds with a finite number of singularities, by means of conformal geometry and the Atiyah-Patodi-Singer L2-index theory. For an n-orbifold M with singularities ΣΓ={(pˇ1,Γ1),...,(pˇs,Γs)} (where each group $Γ_j<O…
For an immersed Lagrangian submanifold, let Aˇ be the Lagrangian trace-free second fundamental form. In this note we consider the equation ∇∗T=0 on Lagrangian surfaces immersed in C2, where T=−2∇∗(Aˇ┘ω), and we prove a gap theorem for the Whitney sphere as a solution …
Paper verifies RNNs using automata learning and model checking.
problem Verifying the correctness of RNNs is challenging.
method Learn a deterministic finite automaton from RNN, use model checking for verification.
result Can discover and generalize counterexamples to faulty flows.
A simple method treats heteroscedastic variance variatively, improving model calibration and sample quality.
problem Brittle optimization impacts model likelihoods for mean and variance estimation.
method Proposes a variational approach to heteroscedastic variance, improving predictive mean and variance calibration.
result The proposed method significantly improves parameter calibration and sample quality for regression and VAEs.