Crowdsourcing can improve scientific investigation by enabling reproducibility and transparency.
problem Current research methods lack reproducibility and transparency, leading to unreliable decisions.
method Next-generation investigative approach leveraging human diversity, micro-specialized crowds, and computer-assisted control methods.
result The Theory of Enablers provides specific cognitive and non-cognitive enablers for crowd-based scientific investigation.
xVal tokenizes numbers continuously for better scientific model training.
problem Lack of continuous numerical tokenization for scientific datasets in LLMs.
method xVal: Continuous numerical tokenization strategy.
result xVal outperforms other numerical tokenization methods on scientific datasets.
Differentiable programming aids in solving differential equations and their sensitivities.
problem Computing gradients of numerical solutions of differential equations.
method Review of existing techniques and mathematical foundations.
result Established a coherent framework for combining differential equations with data-driven approaches.
New AI approach improves quantum device calibration by leveraging prior scientific discoveries.
problem Lack of abundant data in scientific disciplines hinders model generalizability.
method Introduces a new machine learning approach that combines prior scientific knowledge with data.
result Accuracy in predicting quantum device energy spectrum surpasses current state-of-the-art by over 20%.
This study compares Matlab and OpenCV for machine learning algorithms.
problem Comparing execution speeds of Matlab and OpenCV for machine learning.
method 20 real datasets, 20 different machine learning algorithms.
result OpenCV is significantly faster than Matlab in execution.
Machine learning impacts computational math, offering new functions approximations.
problem Machine learning's black box nature hinders further progress in computational math.
method Analyzes machine learning's impact on computational math and vice versa.
result Integrating computational math with machine learning can enhance both fields.
Fast emulators built with neural search accelerate expensive scientific simulations.
problem Slow execution of accurate simulations limits scientific discovery.
method Neural architecture search to build accurate emulators with limited data.
result Simulations accelerated by up to 2 billion times in various scientific fields.
Study examines scientific research on Bitcoin across various disciplines.
problem Limited time but large number of papers on Bitcoin.
method Bibliometric analysis of papers indexed in Web of Science Core Collection.
result Observes research clusters, emerging topics, and leading scholars.
FRACTI framework supports large-scale collaboration and transparent investigation in finance.
problem Complex, multidisciplinary research in finance requires robust support systems.
method Defines scientific support systems, shares contributions, classifies facets, and outlines a meta-model.
result FRACTI enables provenance tracking and large-scale investigation in computational finance.
Quantum computing promises new financial modeling.
problem Traditional financial modeling limitations.
method Overview of quantum computing applications in finance.
result Quantum computing can enhance financial modeling.
New deep learning method outperforms random training data.
problem Improving accuracy of deep learning algorithms in high dimensions.
method Training with low-discrepancy sequences instead of random data.
result Significantly outperforms standard deep learning algorithms.
New approach uses interpolation models and error bounds for verifiable scientific machine learning.
problem Challenges in verifying and validating modern scientific machine learning workflows.
method Combines multiple standard interpolation techniques with error bounds for efficient computation and comparative performance analysis.
result Error bounds for interpolation techniques can be computed or estimated efficiently, aiding in validation goals.
Etalumis bridges scientific simulators and probabilistic programming.
problem Infeasibility of rewriting scientific simulators for Bayesian inference.
method Cross-platform probabilistic execution protocol, MCMC and IC engines, distributed training of 3DCNN-LSTM.
result Achieved largest-scale posterior inference in a Turing-complete PPL for LHC use-case.
Develops a new dataset and model for summarizing scientific papers.
problem Lack of large datasets for summarizing scientific papers.
method Exploits author-provided summaries, uses neural sentence encoding and summarisation features.
result Models that encode sentences and their context perform best, significantly outperforming baselines.
A new MCMC method combines low and high-fidelity models to reduce computation.
problem Inefficient computation of expensive target densities in scientific applications.
method Pseudo-marginal MCMC approach using a telescoping series of low-fidelity models.
result Asymptotically exact multi-fidelity MCMC algorithms for reduced computational cost.
Model learns code representations from comments for data analysis tasks.
problem Lack of descriptive labels for analyzing large code corpora.
method Weakly supervised transformer architecture for joint code and comment representation.
result Model achieves 38% accuracy increase over expert-supplied heuristics.
New sampling strategy preserves relationships in multivariate scientific data.
problem Reducing storage and enabling efficient multivariate analyses on large scientific data.
method Uses principal component analysis for multivariate data and combines with existing univariate sampling algorithms.
result Efficacy demonstrated on real-world data sets, showing data reduction and multivariate analysis ease.
SVGP KAN integrates uncertainty quantification into Kolmogorov-Arnold networks.
problem Uncertainty quantification in scientific machine learning models.
method Sparse variational Gaussian process inference with Kolmogorov-Arnold topology.
result Demonstrated ability to distinguish aleatoric and epistemic uncertainty in various scientific applications.
Instrumented data enables causal scientific machine learning
problem Insufficient data for causal scientific machine learning
method Instrumented data with explicit model, uncertainty, and counterfactuals
result Supports causal interventions through Pearl's do-operator
Scalable solution for interpreting complex data-driven models.
problem Interpreting black box models and handling large datasets.
method Streaming neighborhood graph construction, topology computation, and data aggregation.
result Interactive exploration of high-dimensional data.
This paper explores probabilistic numerical methods for integrating statistical computations.
problem Handling numerical error as epistemic uncertainty in statistical computation.
method Probabilistic integrators that model numerical error as a distribution.
result Probabilistic integrators can achieve posterior contraction rates similar to Monte Carlo methods.
Efficient deep learning on exascale supercomputers solves materials imaging inverse problems.
problem Solving scientific inverse problems in materials imaging using deep learning.
method Novel communication strategies in synchronous distributed deep learning, including decentralized gradient reduction and computational graph-aware grouping.
result Achieved near-linear scaling of distributed training up to 27,600 GPUs on Summit, reaching 2.15(4) EFLOPS16. Quantum model discovery uses DQCs to solve equations from data.
problem Discovering differential equations from data using quantum computing.
method Differentiable quantum circuits (DQCs) to solve parameterized equations, regression on data and equations.
result Successful parameter inference and equation discovery on various systems.
Bayesian emulator tackles models with discontinuities efficiently.
problem Uncertainty quantification of models with multiple discontinuities.
method TENSE framework, designed correlation structures, single emulator object.
result Efficient emulation of complex models with multiple discontinuities.
Survey of Gaussian process constraints for modeling expensive data.
problem Modeling expensive data with physical constraints.
method Overview of various Gaussian process constraints and their implementation.
result Discussion of computational challenges introduced by constraints.
In recent years, ideas from statistics and scientific computing have begun to interact in increasingly sophisticated and fruitful ways with ideas from computer science and the theory of algorithms to aid in the development of improved worst-case algorithms that are useful for large-scale scientific and Internet data an…
The paper clarifies conditions for using benchmark scores in machine learning.
problem Using benchmark scores to draw scientific inferences about learning problems.
method Developing conditions of construct validity inspired by psychological measurement theory.
result Clarifies conditions under which benchmark scores support diverse scientific claims.
Paper introduces a method for operator learning using random features.
problem Estimating maps between infinite-dimensional spaces using input-output pairs.
method Function-valued random features method, building a linear combination of random operators.
result The method provides convergence guarantees and error bounds for nonlinear problems.
TGDS integrates scientific theory into data science for better model interpretation and discovery.
problem Limited applicability of data science models in scientific problems involving complex phenomena.
method Integrating scientific theory into data science models to improve model effectiveness and interpretability.
result TGDS aims to advance scientific understanding by discovering novel insights.
New method speeds up model selection for complex scientific tasks.
problem Exhaustive model selection is computationally infeasible for large model spaces.
method Branch-and-bound algorithm with non-monotonic criteria.
result Guaranteed identification of optimal models with significant computational speedups.
This work uses scientific constraints to validate neural network predictions in fusion physics.
problem Verifying the scientific plausibility of neural network predictions in fusion physics.
method Using known scientific constraints as a validation tool.
result Validated neural network predictions in fusion physics using scientific constraints.
Survey of de Casteljau's algorithm's applications in geometric data analysis.
problem No specific problem stated; focuses on algorithm applications.
method Constructive approach to generalize parametric smooth curves to manifolds.
result Algorithm provides principled way to analyze geometric data.
HollowFlow speeds up likelihood evaluation for large-scale models.
problem Prohibitive scaling of sample likelihood computations in flow-based models.
method Introduces HollowFlow, a flow-based generative model using a NoBGNN with a block-diagonal Jacobian structure.
result Achieves up to O(n^2) speed-up in likelihood evaluation for large systems.
ParaMonte simplifies Monte Carlo simulations for various scientific fields.
problem Efficiently performing Monte Carlo simulations for complex models.
method Unified, high-performance, parallelized library for C, C++, Fortran.
result Automates and streamlines Monte Carlo sampling for arbitrary-dimensional functions.
RPCholesky approximates kernel matrices with few evaluations.
problem Approximating kernel matrices efficiently.
method Randomly pivoted partial Cholesky factorization.
result RPCholesky provides nearly optimal low-rank approximations.
AWD distills neural network info into interpretable wavelets.
problem Imbalanced interpretability and efficiency in deep learning models.
method Adaptive wavelet distillation (AWD) penalizes neural network attributions in wavelet domain.
result AWD yields a concise, efficient, and interpretable model.
Revisits orbital minimization for neural operator decomposition.
problem Training neural networks to approximate eigenfunctions of operators.
method Adapts orbital minimization method (OMM) for neural networks.
result Justifies broader applicability of OMM in modern learning pipelines.
DLHub enables sharing and serving of scientific ML models.
problem Lack of specialized ML systems for scientific applications.
method Self-service model repository and scalable serving capabilities.
result DLHub offers better performance and more capabilities than existing systems.
Galactica learns from scientific literature to help researchers.
problem Information overload in scientific literature makes it hard to find useful insights.
method Trained on a large corpus of scientific papers, reference material, and knowledge bases.
result Outperforms existing models on various scientific tasks, including LaTeX equations and mathematical reasoning.
Efficiently projects vectors onto top PCA components without explicit PCA.
problem Efficiently project vectors onto top principal components of a matrix.
method Iterative algorithm using ridge regression and polynomial approximation.
result First runtime improvement for principal component regression.
Paper learns to generate scientific posters from papers.
problem Generating readable, informative, and visually aesthetic scientific posters is challenging.
method Data-driven framework using graphical models to learn poster elements.
result Model effectively synthesizes graphical elements for posters.
RAD estimates gradients with less memory, faster than small batch sizes.
problem Training deep models with stochastic gradient descent requires exact gradients, but they are not needed.
method Developed a framework for randomized automatic differentiation (RAD) to compute unbiased gradient estimates with reduced memory.
result RAD converges in fewer iterations than using a small batch size for feedforward networks and similar number for recurrent networks.
Celeste learns astronomical catalogs from large datasets.
problem Inferring astronomical catalogs from large-scale datasets.
method Scalable Bayesian inference with Julia, parallel optimization.
result Learned catalogs from modern large-scale astronomical datasets.
Social media enhances or diminishes scientific status, depending on usage.
problem Impact of social media on scientific stratification and mobility.
method Logistic Attribution Analysis combining statistical and machine learning methods.
result Social media promotes stratification and mobility, but beyond a threshold, it negatively impacts status.
A new framework for private Bayesian tests maintains interpretability and computational efficiency.
problem Lack of interpretability and inability to quantify evidence in confidential data.
method Differentially private Bayesian tests based on test statistics.
result Established results on Bayes factor consistency under the proposed framework.
Generative models improve for multiscale scientific data with new noise and interpolation techniques.
problem Numerical challenges in generating high-fidelity samples for multiscale scientific data.
method Design of noise distributions and interpolation schedules in function space to ensure Lipschitz regularity and finite noise roughness.
result Scale-adaptive noise and interpolation schedules improve numerical efficiency and fidelity of generated samples.
Extract keyphrases and relations from scientific documents.
problem Understanding which publications describe which processes, tasks, and materials.
method Evaluated 26 submissions across 3 scenarios.
result Task and findings relevant for researchers and information extraction communities.
Machine learning aids scientific discoveries by explaining complex data.
problem Extracting scientific insights from complex data.
method Combining machine learning with domain knowledge for transparency, interpretability, and explainability.
result Enhanced scientific consistency through machine learning and domain knowledge integration.