Two-block ADMM outperformed multi-block ADMM in multi-task learning experiments.
problem Comparing two-block vs. multi-block ADMM in optimization performance.
method Compared two-block and multi-block ADMM on multi-task learning problems.
result Multi-block ADMM consistently outperformed two-block ADMM in optimization and prediction performance.
New method finds linear relationships across multiple data blocks using proximal gradient descent with ℓ1 constraint.
problem Finding leading generalized eigenvectors for multi-block CCA.
method Proximal gradient descent with ℓ1 constraint. result Rate-optimal solution under suitable assumptions.
A method reduces dimensionality for multi-block data, enhancing feature extraction and classification accuracy.
problem Tractable feature extraction from large-scale, multi-dimensional data.
method Common and individual feature extraction from multi-block data structures using tensor decompositions.
result Significant reduction in dimensionality and enhanced accuracy in feature extraction and classification.
Proposes a method for variable selection and imputation in incomplete data.
problem Missing data in high-dimensional supervised learning settings.
method Multi-block Data-Driven sparse PLS (mdd-sPLS) with Koh-Lanta algorithm for imputation and prediction.
result Shows lowest prediction error for response variables in simulations.
Paper tackles multi-block min-max optimization with applications in deep AUC maximization.
problem Multi-block min-max bilevel optimization with non-convex strongly-concave upper level and strongly convex lower level.
method Single-loop randomized stochastic algorithm for constant number of blocks per iteration.
result Sample complexity of O(1/ε^4) for finding ε-stationary point, matching optimal complexity.
TGCCA analyzes higher-order tensors using orthogonal rank-R CP decomposition.
problem Handling higher-order structures in multi-block data analysis.
method Tensor Generalized Canonical Correlation Analysis (TGCCA) with orthogonal rank-R CP decomposition.
result TGCCA outperforms state-of-the-art methods on simulated and real data.
Paper uses non-Euclidean analysis to classify brain structure variations.
problem Classifying joint variations in multi-object brain structures.
method Combines non-Euclidean statistics and non-parametric integrative analysis.
result Effective, robust, and interpretable joint structure found.
New algorithms tackle complex multi-block optimization problems in machine learning.
problem Non-convex multi-block bilevel optimization with hierarchical sampling challenges.
method Blockwise stochastic variance-reduced methods with parallel speedup.
result Achieves matching complexity to single-block problems with parallel speedup.
New method speeds up solving machine learning problems by splitting variables randomly.
problem Solving large-scale machine learning and signal processing problems efficiently.
method Randomized Multi-Block ADMM (RAC-MBADMM) for convex and nonconvex quadratic optimization.
result RAC-MBADMM converges linearly and outperforms other algorithms in solution time and quality.
In this paper we propose a randomized primal-dual proximal block coordinate updating framework for a general multi-block convex optimization model with coupled objective function and linear constraints. Assuming mere convexity, we establish its O(1/t) convergence rate in terms of the objective value and feasibility m…
Many problems in machine learning and other fields can be (re)for-mulated as linearly constrained separable convex programs. In most of the cases, there are multiple blocks of variables. However, the traditional alternating direction method (ADM) and its linearized version (LADM, obtained by linearizing the quadratic p…
A novel method designs multi-block neural networks using Q-learning.
problem Designing optimal neural network architectures efficiently.
method Reinforcement learning (Q-learning) to sequentially pick different types of blocks.
result Effective in creating multi-block neural networks with comparable or better performance.
Proposes a method to predict content preferences for mobile users in decentralized caching networks.
problem Determining caching schemes for decentralized caching networks with mobile traffic.
method Formulates content preference learning as a DRMTL problem, integrates mobility prediction, and uses ADMM for optimization.
result Mobility-aware content preference learning provides more accurate predictions and improved hit ratios.
New method clusters neurons with similar connectivity profiles.
problem Accurately determining which neurons have similar neurological tasks.
method Proposes clustered Gaussian graphical model and symmetric convex clustering penalty.
result Demonstrates effectiveness of the approach on synthetic and real-world data.
This paper examines MEV attacks in dynamic AMMs and proposes new protections.
problem Dynamic AMMs introduce new MEV attack vectors due to inter-block weight changes.
method Analyzed inter-block weight changes as analogous to trades, conducted simulations.
result New inter-block protections are required to guard against multi-block MEV attacks.
Paper improves Heavy-ball method convergence in convex settings.
problem Convergence analysis of Heavy-ball method in convex optimization.
method Improved convergence complexity results for Heavy-ball method with constant step size.
result First non-ergodic O(1/k) rate result for coercive objective functions.
Recent several years have witnessed the surge of asynchronous (async-) parallel computing methods due to the extremely big data involved in many modern applications and also the advancement of multi-core machines and computer clusters. In optimization, most works about async-parallel methods are on unconstrained proble…
iGecco+ integrates multi-view data for better clustering.
problem Discovering common group structure in mixed multi-view data.
method Integrative Generalized Convex Clustering Optimization (iGecco) with adaptive feature selection.
result iGecco+ achieves superior clustering performance on high-dimensional mixed multi-view data.
Paper analyzes complexity of proximal inertial gradient descent.
problem Computational complexity of proximal inertial gradient descent.
method Analyzed convergence rates and proved various rates under different conditions.
result Proved non-ergodic O(1/k) rate for coercive objective functions.
Unified framework for graph coarsening using node features and graph matrices.
problem Dimensionality reduction of large graphs while preserving node features.
method Optimization-based framework that unifies graph learning and dimensionality reduction.
result The learned coarsened graph is ε-similar to the original graph, where ε is a small positive number.
Recent years have witnessed the rapid development of block coordinate update (BCU) methods, which are particularly suitable for problems involving large-sized data and/or variables. In optimization, BCU first appears as the coordinate descent method that works well for smooth problems or those with separable nonsmooth …
Block Coordinate Update (BCU) methods enjoy low per-update computational complexity because every time only one or a few block variables would need to be updated among possibly a large number of blocks. They are also easily parallelized and thus have been particularly popular for solving problems involving large-scale …
The paper develops algorithms for solving complex optimization problems over Riemannian manifolds.
problem Nonconvex and nonsmooth multi-block optimization over Riemannian manifolds with coupled constraints.
method Develops an ADMM-like primal-dual approach with decoupled solvable subroutines.
result The algorithms achieve an iteration complexity of O(1/ε^2) to reach an ε-stationary solution.
Proposes RSP model for efficient big data analysis.
problem Efficiently partitioning big data sets for analysis.
method Random sample partition (RSP) data model and block-level sampling.
result RSP data blocks can estimate statistics and build models equivalent to whole data set.
Data preprocessing improves data quality for robust data mining.
problem Noisy and incomplete data hinders data mining models.
method Overview of data cleaning, transformation, and preprocessing methods.
result Preprocessing significantly affects data mining model performance.
A new method for handling imbalanced big data using ensembles and smart data.
problem Imbalanced data distribution in big data scenarios.
method Smart Data driven Decision Trees Ensemble (SD_DeTE) methodology.
result SD_DeTE outperforms Random Forest in handling imbalanced binary classification problems in big data.
Prevents sensitive data generation in diffusion models using labeled and unlabeled data.
problem Generating sensitive data in diffusion models using unlabeled data.
method Positive-Unlabeled Diffusion Models, approximating ELBO with labeled and unlabeled data.
result Prevents the generation of sensitive data without compromising image quality.
Study reveals Data Shapley's inconsistent performance in data selection tasks.
problem Inconsistency of Data Shapley's performance in data selection across different settings.
method Hypothesis testing framework and identification of utility functions.
result Data Shapley's performance is no better than random selection without specific constraints.
Survey on data collection challenges in machine learning.
problem Data scarcity and need for labeled data in machine learning.
method Comprehensive study of data acquisition, labeling, and improvement techniques.
result Identification of research challenges in data collection.
PRRO generates synthetic tabular data that improves SL performance and class distribution.
problem Low SL utility of synthetic data due to class imbalance and overlooked data relationships.
method Data pruning and column reordering to optimize SL utility.
result Synthetic data generated with PRRO enhances predictive performance and class distribution.
Defines data science as a natural ecosystem with challenges and missions.
problem Challenges and missions in data science due to 5D complexities and data life cycle phases.
method Systemic and data-centric view of data science as a fusion of data universe and its challenges, formalizing a general-purpose architecture.
result Essential data science as a natural ecosystem integrating specific disciplines and high-impact applications.
Synthetic data enhances analytics but requires careful volume management.
problem Accuracy of statistical methods on synthetic data vs. raw data.
method Synthetic Data Generation for Analytics framework using tabular diffusion models.
result Error rate decreases with more synthetic data but may stabilize or increase.
Data science redefines causal inference from observational data, classifying tasks into description, prediction, and counterfactual prediction.
problem Widespread misunderstandings about data science's role in causal inference from observational data.
method Organizing data science tasks into three classes: Description, prediction, and counterfactual prediction (including causal inference).
result The necessity of subject-matter expert knowledge for causal analyses in data science.
This paper evaluates how dirty data affects data mining and machine learning results.
problem Negative impacts of dirty data on data mining and machine learning results.
method Experimental comparison of missing, inconsistent, and conflicting data on classification and clustering algorithms.
result Guidelines for algorithm selection and data cleaning based on experimental findings.
DPASF stream preprocesses Big Data streams efficiently.
problem Efficient preprocessing of streaming Big Data.
method Implemented six preprocessing algorithms in Apache Flink.
result Preprocessing improves data accuracy in streaming Big Data.
This paper introduces C-DSL to improve data mining outcomes by considering context.
problem Data collection ambiguities, data imbalance, hidden biases, lack of domain info, and data incompleteness.
method Developed Context-Driven Data Science Lifecycle (C-DSL) to address data quality issues.
result Tangible improvements to data mining outcomes were achieved through C-DSL.
Proposes using probabilistic models for privacy-preserving synthetic data.
problem Designing high-quality synthetic data for privacy preservation.
method Formulate the problem through probabilistic modelling, choosing a model for the data.
result Statistical discoveries can be reliably reproduced from synthetic data.
Unlabeled data helps stop active learning better than labeled data.
problem Reducing the need for manual annotation in text classification.
method Compared stopping methods based on labeled, unlabeled, and training data.
result Stopping methods using unlabeled data are more effective.
New test ensures quality of shared data in machine learning.
problem Ensuring quality of external data in machine learning tasks.
method Distribution-free two-sample testing procedures grounded in conformal outlier detection.
result Identifies valuable external data agents for model personalization.
Paper creates fair synthetic data ensuring equal predictions across sensitive attributes.
problem Ensuring fair predictions across sensitive attributes in synthetic data.
method Equalizing target probability distributions across sensitive attributes in synthetic data generation.
result Synthetic data provides strong fair predictions, equal across all thresholds.
A new method classifies multiple correlated data streams simultaneously.
problem Classifying multiple correlated data streams in practical scenarios.
method Double-Coupling Support Vector Machines (DC-SVM) considers both internal and external correlations.
result The proposed method outperforms traditional methods on artificial and real-world data streams.
This paper improves neural machine translation training by selecting and denoising data.
problem Reduces negative impact of noisy data on neural machine translation training.
method Measures and selects domain data, applies denoising curriculum using online data selection.
result Significant effectiveness for training on noisy data.
DPA preserves data distribution in reduced dimensions.
problem Loss of data distribution in dimension reduction.
method DPA combines encoder and decoder to match data distribution.
result DPA successfully reconstructs data distribution.
Efficient synthetic data generation improves model performance on tabular data.
problem Improving model robustness and performance with scarce or low-quality data.
method Hardness characterization to identify high-value training points, generating synthetic data only from these points.
result Synthetic data generated from hardest points outperforms non-targeted methods on tabular datasets.
For most problems in science and engineering we can obtain data sets that describe the observed system from various perspectives and record the behavior of its individual components. Heterogeneous data sets can be collectively mined by data fusion. Fusion can focus on a specific target relation and exploit directly ass…
DAERNN models censored data using neural networks with data augmentation.
problem Handling censored data in expectile regression.
method Data augmentation based Expectile Regression Neural Networks (ERNNs).
result DAERNN outperforms existing censored ERNNs methods and achieves comparable predictive performance to fully observed data.
This paper quantifies uncertainty in Data Shapley using statistical inference.
problem Uncertainty in data valuation due to dynamic data distribution.
method Established relationship with U-statistics and quantified uncertainty using statistical inference.
result Confidence intervals for Data Shapley estimations are provided.
Generative Adversarial Networks create time series data from images.
problem Generating realistic time series data from images.
method Wasserstein GANs with gradient penalty for stability, synthesizing sinusoidal, PPG, and ECG data.
result Successfully generated time series data using image-based GANs.