Data science teams collaborate extensively, using various tools and stakeholders.
problem Lack of understanding in how data science workers collaborate in practice.
method Conducted an online survey with 183 data science workers.
result Data science teams are highly collaborative and use multiple tools and stakeholders.
Crowdsourced investigation shows differing results for technical analysis strategies.
problem Differing results in technical analysis strategies due to lack of method.
method Collaborative scientific computational framework using Monte Carlo simulations and historical back testing.
result Results are not repeatable by other researchers, highlighting the need for transparency and robustness.
Study designs statistical inference for collaborative science teams.
problem Maintaining scientific rigor in distributed, collaborative research.
method Analyzes hypothesis testing with strategic agents and principals.
result Principal can design policies to control posterior probability of null.
FRACTI framework supports large-scale collaboration and transparent investigation in finance.
problem Complex, multidisciplinary research in finance requires robust support systems.
method Defines scientific support systems, shares contributions, classifies facets, and outlines a meta-model.
result FRACTI enables provenance tracking and large-scale investigation in computational finance.
Machine learning applied to particle physics, focusing on future research and development.
problem Improving particle physics analysis and event identification.
method Roadmap for machine learning applications, resource requirements, and collaborative initiatives.
result Identification of resource needs and areas for external collaboration.
Recommender systems leverage product and community information to target products to consumers. Researchers have developed collaborative recommenders, content-based recommenders, and (largely ad-hoc) hybrid systems. We propose a unified probabilistic framework for merging collaborative and content-based recommendations…
Federated learning collaborates clients to train models without sharing data.
problem Privacy and data sharing in machine learning.
method Central server orchestrates collaborative training of models on decentralized data.
result Recent advances and open problems in FL.
The authors seek financial datasets to benchmark feature engineering methods on US market data.
problem Improving predictive models for financial data science competitions.
method Feature engineering methods applied to multivariate time-series data from the US market.
result Predictive power of models tested against Numerai-Signals targets.
Rubin LSST DESC uses AI/ML for dark energy research.
problem Challenges in uncertainty quantification and model robustness for AI/ML in DESC.
method Bayesian inference, physics-informed methods, validation frameworks, active learning.
result AI/ML methods are essential but require rigorous evaluation and governance.
'Ergodicity economics' is criticized as pseudoscience.
problem Flawed conceptual basis of mainstream economic theory.
method Claims 'ergodicity economics' is more parsimonious and clearer.
result Peters' approach has not produced falsifiable implications.
OpenML-Python API simplifies access to OpenML for Python users.
problem Limited access to OpenML for Python users.
method Developed a Python API (OpenML-Python) to integrate OpenML with Python-based tools.
result Facilitates easy access to OpenML's datasets, tasks, and experiments.
Kaggle chronicles 15 years of competitions, innovation, and data science.
problem Exploring 15 years of data science competitions and innovations.
method Longitudinal trend analysis and exploratory data analysis of millions of kernels and discussion threads.
result Kaggle is a growing platform with diverse use cases and adaptable Kagglers.
Sigma simplifies collaboration in economics with a streamlined computational representation.
problem Lack of effective collaboration tools in economics for large-scale projects.
method Introduces Sigma, a domain-specific computational representation for economics based on facets, contributions, and constraints of data.
result Sigma enables sharing and formalizing domain-specific concepts in economics for crowd-based scientific investigations.
A measure of relative importance of variables is often desired by researchers when the explanatory aspects of econometric methods are of interest. To this end, the author briefly reviews the limitations of conventional econometrics in constructing a reliable measure of variable importance. The author highlights the rel…
In the same way as the Hilbert Program was a response to the foundational crisis of mathematics, this article tries to formulate a research program for the socio-economic sciences. The aim of this contribution is to stimulate research in order to close serious knowledge gaps in mainstream economics that the recent fina…
By mapping the most advanced elements of the contemporary social interactions, the world scientific collaboration network develops an extremely involved and heterogeneous organization. Selected characteristics of this heterogeneity are studied here and identified by focusing on the scientific collaboration community of…
Survey of mobility studies using mobile phone data.
problem Understanding human mobility patterns.
method Data Science techniques applied to mobile phone datasets.
result Applications in urban planning, data traffic prediction, etc.
This paper analyzes machine learning workflows in climate modeling.
problem Challenges in integrating machine learning with climate modeling.
method Analysis of case studies focusing on design patterns and workflow structure.
result Synthesis of workflow design patterns across diverse projects in ML-enabled climate modeling.
Enhances insurance loss models using InsurTech data and machine learning.
problem Traditional insurance loss models lack predictive accuracy due to limited data sources.
method Combining proprietary claims data with InsurTech data and applying machine learning techniques.
result Improved predictive accuracy of the loss model through machine learning.
Prediction Factory automates predictive model development and evaluation.
problem Rapidly developing and sharing predictive models with domain experts.
method Data science automation system with three interfaces: baseline, full, and optional automation.
result Full automation interface generated reports funded 57.5% of the time, compared to 42.5% for baseline.
New algorithms reduce communication costs in collaborative learning.
problem Reducing communication costs in collaborative learning.
method Distributed boosting and adaptation to classification noise.
result Communication-efficient algorithms for collaborative PAC learning robust to noise.
Geotechnics adopts data-driven methods from materials informatics.
problem Soil complexity and lack of comprehensive data.
method Leveraging deep learning and transfer learning for feature extraction.
result Revolutionary potential of advanced computational tools in geotechnics.
Supply Chain Management often requires independent organizations to work together to achieve shared objectives. This collaboration is necessary when coordinated actions benefit the group more than the uncoordinated efforts of individual firms. Despite the commonly reported benefits that can be gained in close relations…
This paper considers the problem of high dimensional signal detection in a large distributed network whose nodes can collaborate with their one-hop neighboring nodes (spatial collaboration). We assume that only a small subset of nodes communicate with the Fusion Center (FC). We design optimal collaboration strategies w…
Proposes MC-AE for better unsupervised clustering of unlabeled data.
problem Lack of consideration for multi-local collaborative relationships in autoencoders.
method Integrates LSH for multi-local cross blocks, mcrRBM and mcrGRBM models.
result MC-AE improves unsupervised clustering performance.
Paper tackles unobserved confounding in human-AI collaborations.
problem Unobserved confounding undermines human-AI collaboration effectiveness.
method Combines sensitivity analysis from causal inference with AI-driven statistical modeling.
result Enhances robustness and reliability of collaborative outcomes.
Networks are a fundamental model of complex systems throughout the sciences, and network datasets are typically analyzed through lower-order connectivity patterns described at the level of individual nodes and edges. However, higher-order connectivity patterns captured by small subgraphs, also called network motifs, de…
The paper identifies collaborations in codebases using commit activity and language usage.
problem Identifying organic team interactions and collaborations in large codebases.
method Embedding and clustering commit activity, language usage, and code identifier topics.
result Restores engineering organization and reveals hidden collaborations.
Firms' collaboration networks can decline but remain resilient.
problem Resilience of firms' collaboration networks during decline.
method Analysis of 21,500 R&D collaborations over 25 years, simulating drop-out cascades.
result Firms' collaboration networks can adapt to mitigate decline and recover.
A new model VCM improves collaborative filtering by synchronously linking two VAEs.
problem Cold start and data sparsity issues in CF-based recommender systems.
method Proposes a variational collaborative model (VCM) that synchronously links two VAEs.
result VCM outperforms state-of-the-art methods on real-life datasets.
Study collaborative learning among multi-agents in multi-armed bandits.
problem Minimizing group cumulative regret in a heterogeneous multi-agent setting.
method Developed decentralized algorithms for collaboration between N agents learning M stochastic multi-armed bandits. result Proved near-optimal behavior of proposed algorithms for group regret.
DEMVC improves multi-view clustering with collaborative training and deep autoencoders.
problem Existing multi-view clustering methods have high computation and space complexities or lack representation capability.
method DEMVC learns embedded representations of multiple views individually using deep autoencoders and collaboratively trains all views.
result DEMVC achieves significant improvements over state-of-the-art methods on multi-view datasets.
Advances in collaborative filtering and ranking methods.
problem Improving recommendation systems efficiency and accuracy.
method Graph information encoding, pairwise and listwise approaches, regularization techniques, personalization.
result New methods significantly improve recommendation system performance.
PCL tackles collaborative learning for diverse agents, reducing sample complexity.
problem Balancing collaborative speedup with personalization for heterogeneous agents.
method AffPCL, with bias and importance correction mechanisms.
result AffPCL reduces sample complexity by a factor of max{n−1,δ}, where n is the number of agents and δ∈[0,1] measures heterogeneity. In this paper we examine the effect of applying ensemble learning to the performance of collaborative filtering methods. We present several systematic approaches for generating an ensemble of collaborative filtering models based on a single collaborative filtering algorithm (single-model or homogeneous ensemble). We pr…
Collaborative recommendation is an information-filtering technique that attempts to present information items (movies, music, books, news, images, Web pages, etc.) that are likely of interest to the Internet user. Traditionally, collaborative systems deal with situations with two types of variables, users and items. In…
Meta clustering categorizes learners for collaborative learning.
problem Filtering out unqualified collaborators in collaborative learning.
method Select-Exchange-Cluster (SEC) method to classify learners by their supervised functions.
result SEC can cluster learners into accurate collaboration sets and enhance single-learner performance.
Proposes a deep latent factor model for better recommendation systems.
problem Improving collaborative filtering in recommendation systems.
method Introduces a deeper latent factor model using deep learning.
result Significantly outperforms state-of-the-art techniques in experiments.
Improved subgradient method tackles ill-conditioned composite optimization problems.
problem Slow convergence of subgradient method for composite optimization problems.
method Preconditioned subgradient method with Levenberg-Marquardt approach.
result Linear convergence rate for composite optimization problems under mild conditions.
Optimal algorithm found for collaborative learning in bandits with optimal regret bounds.
problem Minimizing regret in collaborative multi-agent bandit problems.
method Proposed an algorithm with optimal regret bounds for collaborative multi-agent multi-armed bandit model.
result First algorithm with order optimal regret bounds for collaborative bandit model.
A collaborative machine teaching method that improves learner performance with privacy and efficiency.
problem Improving learner performance with distributed teachers while maintaining privacy and scalability.
method Formulates collaborative teaching as a consensus and privacy-preserving optimization process to minimize teaching risk.
result The proposed method delivers significantly more accurate teaching results with high speed compared to non-collaborative MINLP-based super teaching.
Partner-aware algorithms improve AI collaboration in multi-agent settings.
problem Improving AI cooperation in teams with shared rewards.
method Proposed Partner-Aware strategy extending Upper Confidence Bound for decentralized MAB.
result Achieves logarithmic regret in collaborative decision-making.
A new privacy-preserving deep learning scheme for asymmetrically collaborative machine learning.
problem Privacy and efficiency in collaborative machine learning across different data owners.
method Decomposes neural network steps for privacy-preserving training; novel protocol for information leakage.
result Efficient training with stable performance and significant speedup.
A new method for online personalized learning reduces gradient variance by dynamically selecting peers.
problem Online personalized decentralized learning with statistically heterogeneous clients.
method Gradient-based collaboration criterion allowing clients to dynamically select peers with similar gradients.
result The method acts as a variance reduction method, achieving optimal performance in certain conditions.
New algorithms for collaborative learning in uncertain, decentralized environments.
problem Byzantine collaborative learning in uncertain, decentralized environments.
method Two asynchronous solutions to averaging agreement, each optimal according to some dimension.
result New algorithms achieve optimal Byzantine resilience and collaborative learning.
New algorithms balance collaboration and adversarial behavior in linear bandits.
problem Minimizing regret in a collaborative linear bandit problem with adversarial agents.
method Robust collaborative phased elimination algorithm with tight analyses.
result Achieves near-optimal regret bounds of $O\left(α+ 1/\sqrt{M}
ight) \sqrt{dT}$ for good agents.
Game theory models incentivizes honesty in collaborative learning among competitors.
problem Incentivizing honest updates among competitors in collaborative learning schemes.
method Formulated a game to model interactions, studied two learning tasks, proposed mechanisms to incentivize honest communication.
result Rational clients are incentivized to manipulate their updates, preventing learning; proposed mechanisms ensure comparable learning quality to full cooperation.
Collaborative filtering is used to recommend items to a user without requiring a knowledge of the item itself and tends to outperform other techniques. However, collaborative filtering suffers from the cold-start problem, which occurs when an item has not yet been rated or a user has not rated any items. Incorporating …