The study identifies influential bioinformatics algorithms for scalable computing.
problem Data deluge in bioinformatics and need for scalable computing solutions.
method Identifying and analyzing influential data mining and machine learning algorithms.
result Guiding scalable computing experts to focus on specific bioinformatics algorithms.
A scalable parallel BO method for asynchronous settings.
problem Expensive-to-evaluate problems in machine learning.
method Simple and scalable Bayesian optimization method for asynchronous parallel settings.
result Demonstrated promising performance on benchmark functions and hyperparameter optimization.
New scalable algorithm improves linear regression accuracy and efficiency.
problem Improving linear regression accuracy and efficiency for large datasets.
method Developed new theory to model significance and multicollinearity as constraints, scaling with n in the 10,000s. result Significantly improves accuracy, reduces false detection rate, and speeds up computation.
The article uses PageRank and persistent homology for scalable graph comparison.
problem Comparing the similarities between complex networks.
method Combines PageRank and persistent homology to compute a scalable graph descriptor.
result Shows the effectiveness of the method on shape mesh datasets.
New method for scalable barycenter computation using Wasserstein gradient flows.
problem Scalability and integration of label information in barycenter computation.
method Gradient flows in Wasserstein space, time discretization, mini-batch optimal transport, modular regularization, task-aware functions, supervised information integration.
result Empirically validated new state-of-the-art barycenter solver with labeled barycenters outperforming unlabeled ones.
Scalable3-BO tackles scalability issues in Bayesian optimization for big data and high dimensions.
problem Bayesian optimization scalability issues in big data and high dimensions.
method Sparse Gaussian process, random embedding, asynchronous parallelization.
result Scalable3-BO framework optimizes high-dimensional problems with 1 million data points and 10,000 dimensions.
Unified framework for scalable black-box optimization.
problem Expensive black-box evaluations in scientific and engineering domains.
method Integrates active learning, multi-armed bandits, and distributed computing.
result Consistently outperforms state-of-the-art black-box optimizers.
Scalable TS using sparse GPs improves efficiency without sacrificing performance.
problem Efficiently applying TS to complex, multi-modal problems.
method Sparse Gaussian Process models for scalable TS.
result Theoretical and empirical validation of scalable TS's effectiveness.
New method speeds up analysis of computer experiments.
problem Computational infeasibility of direct GP inference for large datasets.
method Adapted Vecchia's ordered conditional approximation to scaled input space.
result Significant performance improvement over existing methods.
This dissertation advances scalable Gaussian processes using iterative methods and pathwise conditioning.
problem The classical Gaussian process formulation is not scalable for large datasets and modern hardware.
method Combining iterative methods and pathwise conditioning to improve scalability.
result Significantly reduced memory requirements and facilitated application to larger datasets.
NodeSig efficiently computes binary node embeddings for scalable graph analysis.
problem Scalability issues in graph representation learning models.
method NodeSig uses random walk diffusion probabilities and stable random projections to compute binary node embeddings efficiently.
result NodeSig achieves a good balance between accuracy and efficiency on node classification and link prediction tasks.
Bayesian non-linear matrix completion tackles large, sparse data.
problem Predict missing elements in large, sparsely observed matrices.
method Bayesian Gaussian process latent variable models with data-parallel distributed computation.
result Scalable Bayesian non-linear matrix completion outperforms linear methods.
New heuristics for parallel Bayesian optimization.
problem Challenges in parallel and scalable Bayesian optimization.
method Review and propose practical heuristic algorithms.
result Simple, heuristic algorithms for parallel Bayesian optimization.
Meta-learning interpretable decision trees with synthetic data.
problem Lack of efficient, scalable methods for generating synthetic data for decision tree meta-learning.
method Synthetic generation of near-optimal decision trees using the MetaTree transformer architecture.
result Meta-learning of decision trees achieves performance comparable to real-world data or optimal decision trees, with significant computational cost reduction.
A scalable Gaussian process model using a mixture-of-experts approach.
problem Training Gaussian process models is computationally expensive and not scalable.
method A mixture-of-experts model with low-dimensional matrix inversions and importance sampling.
result The model offers comparable performance to Gaussian process regression at a lower computational cost.
This work improves scalability of Wasserstein distances in high dimensions.
problem Scalability issues in computing Wasserstein distances in high dimensions.
method Empirical convergence rates, robustness to data contamination, and computational methods.
result Established fast rates and robust estimation risks for sliced Wasserstein distances.
New method for scalable DPMM estimation in distributed data.
problem Efficiently handling new components in distributed Dirichlet Process Mixture Models.
method Locally creating new components, probabilistically consolidating them, and maintaining consistency with low communication cost.
result High scalability and consistent estimation in distributed and asynchronous environments.
TensorHyper-VQC improves VQC scalability and robustness.
problem Scalability and noise sensitivity in VQC.
method Tensor-train-guided hypernetwork framework.
result TensorHyper-VQC achieves superior performance and robust noise tolerance.
Paper combines scalable BMF algorithms for web-scale datasets.
problem High computational cost of Bayesian Matrix Factorization.
method Combines Posterior Propagation and asynchronous distributed implementation.
result Substantial improvements in scalability on web-scale datasets.
Sparse Vision MoE matches dense networks in image recognition while using less compute.
problem Scaling vision models efficiently in computer vision.
method Vision MoE (V-MoE) - a sparse version of Vision Transformer.
result V-MoE matches state-of-the-art dense networks in image recognition with half the compute.
Explosive growth in data and availability of cheap computing resources have sparked increasing interest in Big learning, an emerging subfield that studies scalable machine learning algorithms, systems, and applications with Big Data. Bayesian methods represent one important class of statistic methods for machine learni…
Efficiently computes tree-Wasserstein barycenter for large-scale multilevel clustering and scalable Bayes.
problem Large-scale multilevel clustering and scalable Bayes problems.
method Proposes an efficient algorithm for tree-Wasserstein barycenter and variants.
result Significantly improves efficiency in computation and memory usage for large-scale applications.
Bayesian deep learning methods improve uncertainty estimation in computer vision.
problem Estimating uncertainty in deep learning models for robust computer vision.
method Comprehensive evaluation framework for scalable epistemic uncertainty estimation methods.
result Ensembling provides more reliable uncertainty estimates than MC-dropout.
GNet uses Gaussian processes for scalable, flexible neural networks.
problem Large-scale predictive modeling with high computational and storage costs.
method GNet employs Gaussian processes with nonparametric activation functions and a fast algorithm for training and predictions.
result GNet achieves competitive performance across various test problems, including nonlinear function prediction and real-world data regression.
GNet uses Gaussian processes for scalable, flexible neural networks.
problem Large-scale predictive modeling with high computational and storage costs.
method GNet employs Gaussian processes with nonparametric activation functions and a fast algorithm for efficient training and predictions.
result GNet achieves competitive performance across various test problems, including nonlinear function prediction and real-world data regression.
Estimates complex dependency structures in multi-omics data.
problem Graphical model estimation from multi-omics data with scalability and consistency.
method Pseudolikelihood-based graphical model framework with ℓ1-penalized empirical risk. result Estimates partial correlation network from dual-omic liver cancer data.
This paper introduces a neural sampler for scalable sampling from complex distributions.
problem Efficiently sampling from high-dimensional un-normalized distributions.
method Neural implicit sampler trained with KL and Fisher divergence methods.
result The neural sampler generates large batches of samples with low computational costs.
A new method for scalable spectral clustering using random binning features.
problem Scalability issues in spectral clustering for large-scale problems.
method Random Binning features to accelerate similarity graph construction and eigendecomposition.
result Achieves similar accuracy to standard spectral clustering but with linear computational cost.
The paper discusses scalable learning for wireless data-driven systems.
problem Expanding data volume and model complexity limit centralized learning solutions.
method Discusses scalable architecture and local learning strategies.
result Promising research directions in scalable data-driven wireless communications.
This work connects BNNs to GPs, providing scalable inference and identifying key properties.
problem Scaling and inference challenges in Bayesian neural networks.
method General convergence from BNNs to GPs, new covariance function, and scalable Nyström approximation.
result Established a scalable maximum a posterior (MAP) training and prediction procedure.
XGBoost accelerates machine learning on GPUs.
problem Training large datasets efficiently on GPUs.
method Multi-GPU gradient boosting with data compression and end-to-end GPU parallelism.
result Processed 115 million instances in 3 minutes.
Celeste learns astronomical catalogs from large datasets.
problem Inferring astronomical catalogs from large-scale datasets.
method Scalable Bayesian inference with Julia, parallel optimization.
result Learned catalogs from modern large-scale astronomical datasets.
New scalable GP approximation using Fourier series decomposition.
problem Scalability and accuracy in Gaussian process approximations.
method Harmonic kernel decomposition (HKD) to decompose kernels orthogonally.
result Significantly outperforms standard variational methods in scalability and accuracy.
How should statistical procedures be designed so as to be scalable computationally to the massive datasets that are increasingly the norm? When coupled with the requirement that an answer to an inferential question be delivered within a certain time budget, this question has significant repercussions for the field of s…
COPML framework securely trains models across multiple data owners without revealing individual data.
problem Privacy-preserving collaborative machine learning with multiple data owners.
method Securely encodes data, distributes computation, performs distributed training.
result Achieves up to 16x speedup in training time while maintaining strong privacy.
Scalable verifier for recurrent neural networks using polyhedral abstractions.
problem Certifying the correctness of recurrent neural networks.
method Combining sampling, optimization, and Fermat's theorem for polyhedral abstractions; gradient descent for refinement.
result Successfully verified challenging recurrent models in various domains.
New scalable methods for log determinant computations speed up Gaussian process kernel learning.
problem Prohibitive computational cost of log determinant calculations for Gaussian process kernel learning.
method Stochastic approximations based on Chebyshev, Lanczos, and surrogate models.
result Lanczos method is superior for kernel learning, and surrogate models are highly efficient and accurate.
FIRAL is a scalable active learning algorithm for multiclass classification.
problem Scalability issues with FIRAL in large datasets.
method Proposed an approximate algorithm with reduced storage and computational complexity.
result Demonstrated strong scalability and accuracy on large datasets.
A new framework SPOT efficiently solves large scale optimal transport problems.
problem Heavy computational burden in optimal transport limits its use.
method Implicit generative learning framework (SPOT) approximates optimal transport plan and solves it using stochastic gradient algorithms.
result SPOT efficiently solves optimal transport problems and can recover the density of the plan.
New method improves SBI efficiency and scalability.
problem Scalability issues in SBI methods for large datasets.
method Langevin dynamics with score matching, exploiting likelihood structure.
result Structured score network enhances statistical efficiency and scalability.
New scalable Lipschitz bounds improve neural network robustness analysis.
problem Computing tight Lipschitz bounds for deep neural networks is challenging and computationally expensive.
method Derived new closed-form Lipschitz bounds using more general feasible points of LipSDP, avoiding SDP solvers.
result Improved scalability and precision of Lipschitz estimation for large neural networks.
This paper improves SVD for recommender systems using block-based matrix factorization.
problem Scalability and performance issues in recommender systems.
method Block-based Singular Value Decomposition (BMF) for matrix factorization.
result BMF paired with SVD enhances performance and scalability.
Improves scalability of Gaussian process for large datasets.
problem Scalability issues in Gaussian process for large datasets.
method Developed variational sparse inference algorithms (VSHGP, SVSHGP, DVSHGP) and distributed computing techniques.
result Demonstrated superior performance on various datasets compared to existing methods.
A scalable GPVAE method using local adjacencies to approximate GP inference.
problem Scalability issues in exact GP inference for large-scale GPVAEs.
method Neighbour-driven approximation strategy that confines computations to nearest neighbours.
result Outperforms other GPVAE variants in predictive performance and computational efficiency.
This work proposes a scalable framework for trusted multi-party computations using blockchain.
problem Ensuring trust in results from multi-agent computational experiments.
method Combining distributed validation and blockchain for immutable audits, reducing storage and communication costs.
result Guaranteed verifiability and validity of local computations in a scalable multi-agent environment.
Framework for applying GPs to real-world data with scalability guidelines.
problem Deployment of Gaussian Processes (GPs) is hindered by computational costs and lack of guidelines.
method Proposed a framework for identifying GP suitability and setting up robust models, formalizing decisions of experienced practitioners.
result More accurate results at test time for glacier elevation change case study.
SVB method provides scalable Bayesian proportional hazards model for high-dimensional gene expression data.
problem Bayesian methods for high-dimensional sparse survival data often sacrifice uncertainty quantification or computational scalability.
method Mean-field variational approximation for scalable Bayesian proportional hazards model.
result SVB method offers posterior distribution for parameters and variable selection via posterior inclusion probabilities.
New Krylov subspace methods speed up mixed-effects models with crossed random effects.
problem Slow computations for high-dimensional crossed random effects in mixed-effects models.
method Krylov subspace-based methods for generalized mixed-effects models with cross effects.
result Speedups by factors of up to 10,000 in computations for mixed-effects models.