Crowd-powered system flags misinformation for fact checking.
problem Reduce the spread of fake news and misinformation on social media.
method Flexible temporal point process representation and scalable online algorithm Curb for optimal fact checking selection.
result Our scalable algorithm Curb can effectively reduce the spread of fake news and misinformation.
FAKTA automates fact checking across media sources.
problem Automating fact checking across diverse media sources.
method Unified framework integrating document retrieval, stance detection, evidence extraction, and linguistic analysis.
result FAKTA predicts factuality and provides evidence for claims.
Research aims to make fact-checking models more transparent.
problem Making fact-checking models explainable in a complex field.
method Combines fact-checking methods with explainable AI techniques.
result Developed initial solutions for explainable fact-checking.
Time-aware fact-checking improves veracity predictions for time-sensitive claims.
problem Fact-checking decisions should consider temporal information of claims and evidence.
method Investigated four temporal ranking methods to optimize evidence ranking for fact-checking models.
result Time-aware evidence ranking surpasses relevance assumptions and improves veracity predictions for time-sensitive claims.
Ranked second in fact-checking task, using DRR NN with embeddings.
problem Fact-checking questions in community forums.
method Deeply Regularized Residual Neural Network (DRR NN) with Universal Sentence Encoder embeddings, ensemble methods.
result Ranked second in fact-checking task.
Task focuses on fact checking in Q&A forums, improving over baseline systems.
problem Fact checking in community Q&A forums to distinguish factual from opinion.
method Two subtasks: distinguishing factual vs. opinion/advice/socializing, predicting answer truthfulness.
result Improved over baseline systems for both subtasks, but not for Subtask B.
We create a large dataset for fact checking claims and improve prediction accuracy.
problem Fact checking claims from multiple sources is challenging.
method We created a comprehensive dataset and developed a novel method for automatic veracity prediction.
result Our model achieves a Macro F1 of 49.2%, showing significant performance improvements.
A novel fact-checking method using debate dynamics on knowledge graphs.
problem Fact-checking on knowledge graphs with user comprehension and interactive reasoning.
method Reinforcement learning agents debate on paths in the graph to classify facts as true or false.
result Interactive reasoning and user understanding of AI decisions on knowledge graphs.
Paper develops a model for verifying facts in tables without pre-retrieved evidence.
problem Verification of factual claims in structured data, especially in open-domain settings.
method Joint reranking-and-verification model that fuses evidence documents.
result Model achieves comparable performance to closed-domain state-of-the-art on TabFact dataset.
Deep learning model detects and corrects outliers in crowd-sourced weather data.
problem Data quality issues in crowd-sourced weather data.
method Bayesian deep learning approach with Gaussian-uniform mixture density network.
result Automated outlier detection in spatio-temporal environmental modeling.
Paper tackles stance detection across domains using adversarial domain adaptation.
problem Stance detection in different domains is costly and tedious.
method Adversarial domain adaptation for stance detection.
result Model effectively transfers knowledge for accurate stance detection across domains.
Formalizes interpreting natural language rules for answering questions, collecting 32k task instances.
problem Interpreting regulations and answering 'Can I...?' or 'Do I have to...?' questions.
method Formalization of task, crowd-sourcing strategy to collect 32k instances, analysis of challenges, evaluation of performance.
result Promising results when no background knowledge is needed, substantial room for improvement when background knowledge is needed.
Analyzed Indian stock market data to find stylized facts with deviations.
problem Identifying stylized facts in the Indian stock market.
method Historical daily data analysis of NIFTY index stocks over 11 years.
result Significant deviations in leverage, asymmetry, and autocorrelation observed.
Developed a neural topic model for classifying COVID-19 disinformation.
problem Tackles the challenge of disinformation during the COVID-19 pandemic.
method Classification-aware neural topic model (CANTM) for COVID-19 disinformation.
result Demonstrated the effectiveness of CANTM in classifying COVID-19 disinformation.
We describe the infinitesimal moduli space of pairs (Y,V) where Y is a manifold with G2 holonomy, and V is a vector bundle on Y with an instanton connection. These structures arise in connection to the moduli space of heterotic string compactifications on compact and non-compact seven dimensional spaces, e.…
By exploiting standard facts about N=1 and N=2 supersymmetric Yang-Mills theory, the Donaldson invariants of four-manifolds that admit a Kahler metric can be computed. The results are in agreement with available mathematical computations, and provide a powerful check on the standard claims about supersymmetric Yang…
Paper optimizes summarization of multiple document groups for better distinction.
problem Comparative document summarization to select representative documents from multiple groups.
method Formulated new objective functions based on binary classification and maximum mean discrepancy, using gradient-based optimization.
result Gradient-based optimization outperforms other methods in automatic and crowd-sourced evaluations.
Combines low-fidelity and high-fidelity labels using Gaussian process co-kriging.
problem Classification with variable fidelity labels.
method Gaussian process co-kriging for latent functions, extended Laplace inference for multi-fidelity data.
result More resistant to labeling discrepancy than other fusion methods.
The paper examines circle graphs of Gauss diagrams and finds counterexamples to previous descriptions.
problem Problems with previous descriptions of realizable Gauss diagrams.
method Experimental checking and formulation of new descriptions of realizable circle graphs.
result New descriptions of realizable circle graphs and an algorithm for checking realizability.
Bayesian model improves truth inference from highly redundant crowd annotations.
problem Inferring true annotations from highly redundant crowd annotations.
method Bayesian graphical model with conjugate priors and iterative expectation-maximisation inference.
result Our technique significantly outperforms majority vote heuristic at one-sided level 0.025.
Crowd-sourced mosquito audio dataset for malaria research.
problem Understanding mosquito locations for malaria reduction.
method Release of a large mosquito audio dataset with labels from contributors.
result Demonstrated the feasibility of training a CNN on mosquito audio data.
We investigate the rigidity and asymptotic properties of quantum SU(2) representations of mapping class groups. In the spherical braid group case the trivial representation is not isolated in the family of quantum SU(2) representations. In particular, they may be used to give an explicit check that spherical braid grou…
The topological underpinnings are presented for a new algorithm which answers the question: `Is a given knot the unknot?' The algorithm uses the braid foliation technology of Bennequin and of Birman and Menasco. The approach is to consider the knot as a closed braid, and to use the fact that a knot is unknotted if and …
BUDS balances privacy and utility by shuffling data, achieving strong privacy with minimal loss.
problem Balancing privacy and utility in crowd-sourced statistical databases.
method One-hot encoding, iterative shuffling, loss estimation, risk minimization.
result Achieves ε=0.02 for privacy, maintaining a privacy bound of ε=ln[t/((n1−1)S)]. The paper tackles ranking experts based on their answers to questions, considering statistical and computational challenges.
problem Ranking experts based on their answers to questions, considering isotonic constraints.
method Investigates the existence of statistically optimal and computationally efficient procedures for ranking experts under isotonic constraints.
result Disproves the existence of computational-statistical gaps for the problem.
Activates speech DNNs to generate understandable examples.
problem Difficulty in understanding DNN classifications for speech.
method Activation maximization to generate speech samples.
result Activation maximization can generate understandable speech samples.
Improves neural program synthesis by addressing aliasing and syntax issues.
problem Ignoring program aliasing and syntax in neural program synthesis.
method Reinforcement learning and direct syntax maximization training.
result Improved accuracy, especially with limited training data.
Paper identifies food types from Yelp photos using machine learning.
problem Ineffective labeling of food photos on Yelp.
method Image pre-processing, CNN feature extraction, and classification algorithms.
result Identifies up to 10 food types from raw photos with high accuracy.
When modelling stock market dynamics, the price formation is often based on an equilbrium mechanism. In real stock exchanges, however, the price formation is goverend by the order book. It is thus interesting to check if the resulting stylized facts of a model with equilibrium pricing change, remain the same or, more g…
We study coordinate-invariance of some asymptotic invariants such as the ADM mass or the Chruściel-Herzlich momentum, given by an integral over a "boundary at infinity". When changing the coordinates at infinity, some terms in the change of integrand do not decay fast enough to have a vanishing integral at infinity; bu…
Gradient descent solves sparse skill estimation in crowdsourcing.
problem Crowd-sourced worker skill estimation with sparse and irregular assignments.
method Rank-one matrix completion and projected gradient descent.
result Skill estimates converge to global optima for specific sampling matrices.
Algorithm learns nearest neighbor graph from noisy distance queries.
problem Learning nearest neighbor graph from noisy distance samples.
method Active algorithm to find graph with high probability, analyzing query complexity.
result Empirically and theoretically efficient, needing only O(n log(n)Delta^-2) queries.
Deep learning predicts bridge load capacity from images.
problem Lack of data on aging bridges in post-disaster zones.
method Crowd-sourced images trained on a new CNN for multiclass classification.
result Improved prediction accuracy and practical optimisation.
We introduce a new model in order to describe the fluctuation of tick-by-tick financial time series. Our model, based on marked point process, allows us to incorporate in a unique process the duration of the transaction and the corresponding volume of orders. The model is motivated by the fact that the "excitation" of …
Model predicts political ideology using context vectors to mitigate bias and scarcity.
problem Scarcity and selection bias in political ideology prediction.
method Proposes a statistical model decomposing embeddings into context and position vectors, training an end-to-end model for deployment.
result Model can predict ideological labels even with minimal biased data, outperforming state-of-the-art methods.
BiLA uses variational Bayesian inference to aggregate noisy labels online.
problem Aggregating noisy labels from crowd workers in real-time.
method Variational Bayesian inference and stochastic optimization.
result BiLA reduces label error by at least 10-1.5% points.
New algorithms find all ε-good arms in stochastic bandits.
problem Finding all arms with means above a specified threshold in stochastic bandits.
method Two algorithms introduced to identify all ε-good arms.
result Demonstrated great empirical performance on large datasets.
Multifractal analysis and extensive statistical tests are performed upon intraday minutely data within individual trading days for four stock market indexes (including HSI, SZSC, S&P500, and NASDAQ) to check whether the indexes (instead of the returns) possess multifractality. We find that the mass exponent τ(q) is l…
Predicts lead contamination in Flint's water system based on home attributes.
problem Understanding and predicting lead contamination in Flint's water system.
method Data science approach using a large dataset of water tests and a crowd-sourced prediction challenge.
result Elevated lead risks can be weakly predicted from observable home attributes.
New functions derived from arrow diagrams for spherical curves, invariant under certain deformations.
problem Defining and analyzing integer-valued functions on spherical curves.
method Introducing new functions and relators to study spherical curves and their isotopy classes.
result Functions derived from arrow diagrams are invariant under specific deformations.
Wide-AdGraph detects ads and trackers using a graph of resource requests.
problem Detecting and blocking ad trackers to protect user privacy.
method Combining a large-scale graph of resource requests from multiple websites to train a machine learning algorithm.
result High accuracy (96.1% biased, 90.9% unbiased) in detecting ads and trackers.
Two geometric tests for forward-flatness are shown to be dual.
problem Checking forward-flatness in discrete-time systems.
method Two geometric tests based on involutive distributions and integrable codistributions.
result The two tests are dual to each other.
This is the first of three articles on the Fibered Isomorphism Conjecture of Farrell and Jones for L-theory. We apply the general techniques developed in [15] and [16] to the L-theory case of the conjecture and prove several results. Here we prove the conjecture, after inverting 2, for poly-free groups. In particular, …
Python models predict stock sentiment for market-beating returns.
problem Predicting public sentiment for stock trading.
method Crowd-sourced labeled data, trained and evaluated various models.
result Best models predict market-beating returns from public sentiment.
This paper deforms complex tori and their mirrors using gerbes.
problem Deforming complex tori and their mirror partners.
method Using flat gerbes to deform complex tori and their mirrors, constructing holomorphic line bundles over deformed objects.
result Deformed complex tori and their mirrors can be studied using flat gerbes.
Statistical model checking for PCTL on MDPs using reinforcement learning.
problem Model checking PCTL specifications on MDPs with statistical methods.
method Reinforcement learning for policy search, statistical model checking with UCB-based Q-learning.
result Provably guaranteed statistical model checking method for PCTL specifications on MDPs.
Community moderation drifts towards majority, study finds.
problem How to ensure crowd-sourced moderation systems trust and reward accurate evaluations.
method Consensus-based auditing with a two-stage algorithm that weights contributors by the stability of their past residuals.
result Minority contributors' evaluations drift towards the majority, and their participation share falls on controversial topics.
LLMs will inevitably hallucinate due to their mathematical structure.
problem The inherent limitations of Large Language Models (LLMs).
method Analysis of LLMs using computational theory and Godel's Incompleteness Theorem.
result Hallucinations in LLMs are an inevitable feature, not just errors.