Deep learning model detects and corrects outliers in crowd-sourced weather data.
problem Data quality issues in crowd-sourced weather data.
method Bayesian deep learning approach with Gaussian-uniform mixture density network.
result Automated outlier detection in spatio-temporal environmental modeling.
Bayesian model improves truth inference from highly redundant crowd annotations.
problem Inferring true annotations from highly redundant crowd annotations.
method Bayesian graphical model with conjugate priors and iterative expectation-maximisation inference.
result Our technique significantly outperforms majority vote heuristic at one-sided level 0.025.
It is common for CCTV operators to overlook inter- esting events taking place within the crowd due to large number of people in the crowded scene (i.e. marathon, rally). Thus, there is a dire need to automate the detection of salient crowd regions acquiring immediate attention for a more effective and proactive surveil…
Integrates multiple datasets to solve open set crowdsourcing problems.
problem Crowdsourcing with unknown label space and unfamiliar tasks.
method Integrates multiple crowdsourced datasets, weights them based on category correlation, and uses open set transfer learning.
result Proves OSCrowd solves open set crowdsourcing problems and outperforms related solutions.
Combines low-fidelity and high-fidelity labels using Gaussian process co-kriging.
problem Classification with variable fidelity labels.
method Gaussian process co-kriging for latent functions, extended Laplace inference for multi-fidelity data.
result More resistant to labeling discrepancy than other fusion methods.
Enhances crowd safety through AI and data-driven models.
problem Improving crowd safety during events.
method Innovative data collection, AI, and machine learning.
result Accurate multi-day forecasts for event planning.
Wisdom of the crowd, the collective intelligence derived from responses of multiple human or machine individuals to the same questions, can be more accurate than each individual, and improve social decision-making and prediction accuracy. This can also integrate multiple programs or datasets, each as an individual, for…
Paper optimizes summarization of multiple document groups for better distinction.
problem Comparative document summarization to select representative documents from multiple groups.
method Formulated new objective functions based on binary classification and maximum mean discrepancy, using gradient-based optimization.
result Gradient-based optimization outperforms other methods in automatic and crowd-sourced evaluations.
BiLA uses variational Bayesian inference to aggregate noisy labels online.
problem Aggregating noisy labels from crowd workers in real-time.
method Variational Bayesian inference and stochastic optimization.
result BiLA reduces label error by at least 10-1.5% points.
BUDS balances privacy and utility by shuffling data, achieving strong privacy with minimal loss.
problem Balancing privacy and utility in crowd-sourced statistical databases.
method One-hot encoding, iterative shuffling, loss estimation, risk minimization.
result Achieves ε=0.02 for privacy, maintaining a privacy bound of ε=ln[t/((n1−1)S)]. Crowd-sourced mosquito audio dataset for malaria research.
problem Understanding mosquito locations for malaria reduction.
method Release of a large mosquito audio dataset with labels from contributors.
result Demonstrated the feasibility of training a CNN on mosquito audio data.
The clustering ensemble technique aims to combine multiple clusterings into a probably better and more robust clustering and has been receiving an increasing attention in recent years. There are mainly two aspects of limitations in the existing clustering ensemble approaches. Firstly, many approaches lack the ability t…
The paper tackles ranking experts based on their answers to questions, considering statistical and computational challenges.
problem Ranking experts based on their answers to questions, considering isotonic constraints.
method Investigates the existence of statistically optimal and computationally efficient procedures for ranking experts under isotonic constraints.
result Disproves the existence of computational-statistical gaps for the problem.
Activates speech DNNs to generate understandable examples.
problem Difficulty in understanding DNN classifications for speech.
method Activation maximization to generate speech samples.
result Activation maximization can generate understandable speech samples.
Formalizes interpreting natural language rules for answering questions, collecting 32k task instances.
problem Interpreting regulations and answering 'Can I...?' or 'Do I have to...?' questions.
method Formalization of task, crowd-sourcing strategy to collect 32k instances, analysis of challenges, evaluation of performance.
result Promising results when no background knowledge is needed, substantial room for improvement when background knowledge is needed.
Paper identifies food types from Yelp photos using machine learning.
problem Ineffective labeling of food photos on Yelp.
method Image pre-processing, CNN feature extraction, and classification algorithms.
result Identifies up to 10 food types from raw photos with high accuracy.
MRCNet tackles crowd counting and density mapping in aerial imagery.
problem Accurate crowd counting and density estimation in aerial imagery.
method MRCNet is a novel encoder-decoder CNN that combines VGG-16 with FPN-inspired lateral connections.
result MRCNet outperforms state-of-the-art methods in aerial and CCTV-based crowd counting.
Wide-AdGraph detects ads and trackers using a graph of resource requests.
problem Detecting and blocking ad trackers to protect user privacy.
method Combining a large-scale graph of resource requests from multiple websites to train a machine learning algorithm.
result High accuracy (96.1% biased, 90.9% unbiased) in detecting ads and trackers.
Algorithm learns nearest neighbor graph from noisy distance queries.
problem Learning nearest neighbor graph from noisy distance samples.
method Active algorithm to find graph with high probability, analyzing query complexity.
result Empirically and theoretically efficient, needing only O(n log(n)Delta^-2) queries.
Deep learning predicts bridge load capacity from images.
problem Lack of data on aging bridges in post-disaster zones.
method Crowd-sourced images trained on a new CNN for multiclass classification.
result Improved prediction accuracy and practical optimisation.
SafeRNet uses IoT and cloud computing to provide real-time safe routes.
problem High traffic fatality rates despite advanced technology.
method Bayesian network for safe route modeling.
result Demonstrated effectiveness with real traffic data.
New algorithms find all ε-good arms in stochastic bandits.
problem Finding all arms with means above a specified threshold in stochastic bandits.
method Two algorithms introduced to identify all ε-good arms.
result Demonstrated great empirical performance on large datasets.
Max-MIG tackles crowdsourced label learning without knowing crowd information structure.
problem Learning from crowds without knowing the information structure among crowds.
method Max-MIG is an information theoretic approach that simultaneously aggregates crowdsourced labels and learns a data classifier.
result Max-MIG achieves state-of-the-art results in most settings, including real-world data.
We found that factors decay over time, with momentum fitting best.
problem Understanding how factors decay over time and their impact on performance.
method Derived a hyperbolic decay model for factors, tested against linear and exponential alternatives.
result Momentum exhibits hyperbolic decay, outperforming linear and exponential models.
Study shows how 'crowding' in equity trading affects performance and costs.
problem Deterioration of strategy performance, increased trading costs, and systemic risk due to equity factor crowding.
method Direct metrics of crowding based on imbalances of trades executed on the market, analyzing U.S. equity market data.
result Significant signs of crowding in well-known equity signals, especially Momentum, affecting order flow and portfolio rebalancing.
MTCNet uses MTL to estimate crowd density and count.
problem Crowd count estimation challenges due to scale variations and perspective.
method MTL deep neural network architecture with two tasks: density estimation and count classification.
result Achieves lower MAE than state-of-the-art methods on multiple datasets.
Crowd opinions in microblogs can predict event outcomes, matching with expert opinions.
problem Utilizing crowd wisdom for event outcome prediction in microblogs.
method Multi-label sentiment classification of tweets to gauge crowd opinion and compare with expert predictions.
result Crowd opinions in microblogs often match with expert opinions, especially in non-debate events.
ACFM predicts crowd flow adaptively integrating various factors.
problem Adaptive integration of factors affecting crowd flow changes.
method Unified neural network module with attention mechanism.
result Significant improvements over state-of-the-art methods.
New method improves crowd counting accuracy using inverse k-NN maps and multiscale upsampling.
problem Improving accuracy of crowd density maps for high-density gatherings.
method Developed MUD-ikNN architecture using inverse k-NN maps and multiscale upsampling. result New network architecture outperforms state-of-the-art crowd counting.
Crowded trades cluster investors, affecting stock price stability.
problem Crowded trades lead to price instability and systemic risk.
method Market clustering measure using granular trading data.
result Market clustering has a causal effect on stock return distribution tails, especially positive tail.
System detects social interactions in crowds using mobile phone sensors.
problem Detecting social interactions in crowded settings.
method Multi-modal mobile sensing (BLE, accelerometer, gyroscope) and machine learning.
result 77.8% precision and 86.5% recall for predicting social interactions.
Python models predict stock sentiment for market-beating returns.
problem Predicting public sentiment for stock trading.
method Crowd-sourced labeled data, trained and evaluated various models.
result Best models predict market-beating returns from public sentiment.
Machine learning outperforms crowd investors in predicting loan defaults and investment returns.
problem Determining if machine learning can outperform human decision-making in crowd lending.
method Using data from Prosper.com, a sophisticated ML algorithm was trained to predict loan defaults and investment returns.
result The ML algorithm outperforms crowd investors in predicting loan defaults and investment returns, especially for risky loans.
The paper interprets financial markets as crowds during booms and busts.
problem Understanding market irrationality during booms and busts.
method Integrates crowd psychology into behavioural finance.
result Markets behave like psychological crowds during booms and busts.
A model-free hedging method using stock crowding scores.
problem Designing costless portfolio strategies to hedge market risk.
method Network analysis of fund holdings to compute crowding scores, constructing long-short portfolios without numerical optimization.
result Long-short portfolios provide protection against both small and large market price fluctuations.
Crowdsourced wisdom improves causal learning.
problem Improving causal learning through collective intelligence.
method Crowdsourcing, expert knowledge elicitation, aggregation techniques, and LLMs.
result Collective contributions enhance global causal structure.
Community moderation drifts towards majority, study finds.
problem How to ensure crowd-sourced moderation systems trust and reward accurate evaluations.
method Consensus-based auditing with a two-stage algorithm that weights contributors by the stability of their past residuals.
result Minority contributors' evaluations drift towards the majority, and their participation share falls on controversial topics.
Investment herding can reduce household consumption, a phenomenon called crowding-out effect.
problem Investment herding's impact on household consumption.
method Optimal control theory to model and solve for household investment and consumption decisions.
result Existence of crowding-out effect due to investment herding.
The average portfolio structure of institutional investors is shown to have properties which account for transaction costs in an optimal way. This implies that financial institutions unknowingly display collective rationality, or Wisdom of the Crowd. Individual deviations from the rational benchmark are ample, which il…
This paper proposes a general model for synchronized crowding behavior. An order parameter is introduced to quantify the level of synchronization which is shown a function of percentage of agents in reactive state. Further, synchronization is shown to be driven by the most active agents with the highest volatility. A t…
Crowdsourced predictions from microservices improve supply chain efficiency.
problem Improving supply chain efficiency through high-quality predictions.
method Trials of a multi-agent system with microservices and economic incentives.
result Empirical lessons suggest potential for a Prediction Web.
Combines foundation models with weak supervision to improve NLP and video tasks.
problem Leveraging weak supervision with foundation models without labeled data.
method Liger, a combination of foundation model embeddings and weak supervision techniques.
result Liger outperforms existing weak supervision methods by 14.1 points on benchmark NLP and video tasks.
New method addresses crowding in high-dimensional data visualization.
problem Crowding issue in visualizing high-dimensional data.
method Adjusting capacity of high-dimensional balls and estimating correlation dimension.
result Mitigates crowding in various distance metrics.
Crowdsourcing can improve scientific investigation by enabling reproducibility and transparency.
problem Current research methods lack reproducibility and transparency, leading to unreliable decisions.
method Next-generation investigative approach leveraging human diversity, micro-specialized crowds, and computer-assisted control methods.
result The Theory of Enablers provides specific cognitive and non-cognitive enablers for crowd-based scientific investigation.
Low dimensional embeddings that capture the main variations of interest in collections of data are important for many applications. One way to construct these embeddings is to acquire estimates of similarity from the crowd. However, similarity is a multi-dimensional concept that varies from individual to individual. Ex…
System accurately identifies birds in real-world settings.
problem Identifying birds in diverse, realistic environments.
method Trained kNN and SVM classifiers on crowd-sourced audio data.
result Both classifiers perform similarly, with kNN offering flexibility.
Study improves forecasting of ED crowding using advanced ML models.
problem Improving forecasting of emergency department crowding.
method Advanced machine learning models (N-BEATS, LightGBM, DeepAR) using multivariable input data.
result N-BEATS and LightGBM outperform benchmarks in forecasting ED occupancy.
Cost forecasting helps crowdsourcers manage growing task sets efficiently.
problem Crowdsourcers face resource overwhelm with growing task sets.
method Cost forecasting to decide between new tasks or existing ones based on cost efficiency.
result Cost forecasting improves accuracy and efficiency in crowdsourcing.