Estimates tangent space variation on Riemannian submanifolds using feature size.
problem Estimating tangent space variation on Riemannian submanifolds.
method Using structural properties of local feature size and elementary Euclidean geometry.
result Asymptotically tight estimates of tangent space variation.
Deep Local-Global network improves lung nodule malignancy prediction.
problem Challenging task of classifying lung nodules as benign or malignant.
method Proposes a novel method combining local and global feature extraction.
result Achieved state-of-the-art results with AUC=95.62%.
Adaptive Siamese network improves local feature descriptor learning efficiency.
problem Estimating the size of neural networks for local feature descriptors.
method Adaptive pruning Siamese architecture based on neuron activation.
result Learned local feature descriptors outperform state-of-the-art methods in patch matching.
New model learns graph features for classification.
problem Graph classification with structural information loss.
method Transform graphs into vertex grids, apply vertex convolution.
result Model preserves structural information on local vertices.
Paper uses deep learning to improve thermal-hydraulic simulations.
problem Limited credibility of thermal-hydraulic codes in real plant conditions.
method Feature Similarity Measurement (FSM) and deep learning.
result Deep learning constructs relationships between local physical features and simulation errors.
Analyzes error sources in global feature effect estimation methods.
problem Unexplored error sources in global feature effect estimation methods.
method Systematic, estimator-level analysis of bias and variance.
result Holdout data is theoretically cleanest, but estimation variance depends on sample size and model characteristics.
Method learns feature maps from deep CNN layers for weakly supervised chest pathology localization.
problem Localization of chest pathologies in X-ray images is challenging due to varying sizes and appearances.
method Class-aware deep multiscale feature learning using intermediate feature maps from CNN layers.
result Improves localization performance of small pathologies like nodules and masses.
New molecular descriptors improve machine learning for small and large molecules.
problem Improving machine learning models for molecular properties.
method Developed constant-size molecular descriptors combining connectivity counts and encoded distances.
result Models using these descriptors perform comparably to or better than state-of-the-art models.
MR-GNN predicts interactions between structured entities using multi-resolution and dual graph neural networks.
problem Predicting interactions between structured entities, especially considering features in substructures of different sizes and interactions between entities.
method MR-GNN uses a multi-resolution architecture and dual graph-state L-STMs to extract features from different neighborhoods and pairwise graphs, respectively.
result MR-GNN improves prediction accuracy compared to state-of-the-art methods.
A new framework SIMBA improves graph classification performance on size-imbalanced datasets.
problem Size imbalance in graph classification leads to poor model performance.
method Energy-guided structural smoothing between head and tail graphs, re-weighting based on energy propagation.
result SIMBA outperforms existing methods in size-imbalanced graph classification tasks.
Framework learns asymmetric and local features in multi-dimensional data.
problem Learning features in multi-dimensional data, especially images.
method Bayesian hierarchical modeling with recursive wavelet transforms.
result Framework achieves high computational scalability and adaptivity.
Locally sparse neural networks improve interpretability for biomedical tabular data.
problem Overfitting and lack of interpretability in neural networks for tabular biomedical data.
method Locally sparse neural network with a gating network to select relevant features.
result The method outperforms state-of-the-art models in synthetic and real-world biomedical datasets.
Losaw improves FI scores by decorrelating features in ML models.
problem Feature correlation distorts feature importance scores in ML models.
method Losaw uses local sample weighting to decorrelate features.
result Losaw consistently improves feature importance scores and prediction accuracy.
Local convolutions bias neural networks towards high-frequency adversarial examples.
problem High-frequency adversarial examples in neural networks.
method Analysis of different linear and nonlinear architectures, focusing on the impact of local convolution operations.
result Local convolutions induce an implicit bias towards high frequency features, leading to high-frequency adversarial examples.
PFBP algorithm speeds up feature selection in big data.
problem Feature selection in high-dimensional and/or large sample size data.
method PFBP algorithm partitions data and uses local computations with early decisions.
result Asymptotic optimality for causal networks, super-linear speedup, linear scalability.
QuPWM detects epileptic spikes in MEG signals with high accuracy.
problem Manual detection of epileptic spikes in MEG signals is time-consuming and subjective.
method QuPWM combines PWM and SVM for feature extraction and classification.
result Average accuracy of 98% achieved on a balanced dataset of 3104 samples.
New method uncovers global topology through local interactions, reducing algorithm complexity.
problem Global interaction is necessary for forming feature maps that preserve global topology.
method Competing agents engage in local interactions to form feature maps without global interaction.
result Local interactions can uncover global topology, leading to consistent map quality across diverse datasets.
New method improves Gaussian kernel approximations for high-frequency data.
problem Limited scalability of kernel-based models to large data sets.
method Local random feature approximations using Maclaurin expansions and polynomial sketches.
result Significant improvement in kernel approximations and downstream performance for high-frequency data.
GOTabPFN improves tabular model performance with compact tokenization for HDLSS data.
problem Making tabular models effective for high-dimensional, low-sample size data without retraining.
method Introducing Graph-guided Ordering with Local Refinement (GO-LR) and Neuro-Inspired Subunit Compression (NSC) to create compact meta-features.
result GOTabPFN improves stability and accuracy in tabular benchmarks with compact tokenization.
New method explains deep neural networks by ranking feature importance.
problem Limited ability to explain deep neural networks.
method Proposes a novel approach to global feature ranking in DNNs, leveraging partial covariance structures and variable dependence.
result Demonstrates improved feature ranking and interpretation in various domains.
A new test for conditional independence adapts to nonlinear dependencies efficiently.
problem Testing conditional independence in nonlinear and high-dimensional data.
method Nearest-neighbor estimator of conditional mutual information combined with local permutation scheme.
result The test reliably simulates null distribution and is better calibrated for non-smooth densities.
Adaptive batch sizes improve local gradient methods in distributed training.
problem Communication bottlenecks in distributed deep learning.
method Adaptive batch size strategies for local gradient methods.
result Adaptive batch sizes reduce minibatch gradient variance and improve training efficiency.
Paper develops robust SVM classifiers for uncertain data.
problem Sensitivity of SVM classifiers to data uncertainty.
method Two probabilistic approaches: Single Perturbation and Extreme Empirical Loss.
result Both methods reduce data uncertainty effects efficiently.
A framework for collaborative learning reduces communication rounds.
problem Collaborative learning with distributed features and privacy concerns.
method Federated Stochastic Block Coordinate Descent (FedBCD) algorithm.
result The algorithm achieves O ( T ) O(\sqrt{T}) O ( T ) communication rounds and O ( 1 / T ) O(1/\sqrt{T}) O ( 1/ T ) accuracy. New semimetrics maximize test power for comparing distributions.
problem Comparing and understanding differences between distributions.
method Features optimized to maximize test power, converging with sample size.
result Linear-time tests achieve comparable performance to state-of-the-art methods.
Two graph auto-encoders decouple feature propagation from graph convolution layers.
problem Designing efficient graph auto-encoders with fixed receptive fields.
method L-GAE and L-VGAE using linear matrix computation before auto-encoder input.
result Comparable performance to VGAEs with smaller, simpler networks.
An algorithm reduces breast cancer detection data complexity using effect sizes.
problem Improving accuracy in breast cancer detection.
method Statistical feature selection and SVM classifier with linear kernel.
result SVM classifier achieved over 90% accuracy.
Value selection reduces model size while maintaining accuracy.
problem Space efficiency in model size reduction.
method Two probabilistic methods based on information theory's metric: PVS and P + VS.
result Value selection achieves balance between accuracy and model size reduction.
Flexible outlier detection using graph communities for robust performance.
problem Outlier detection in small sample size unbalanced problems.
method Local measure of label heterogeneity in a weighted graph topology.
result Overall outperforms local and global strategies in multi and single view settings.
Study uses machine learning to detect early COVID-19 from CT images.
problem Early detection of COVID-19 from CT images.
method Machine learning methods applied to patches of CT images, feature extraction (GLCM, LDP, GLRLM, GLSZM, DWT), SVM classification.
result Best classification accuracy of 99.68% with 10-fold cross-validation and GLSZM feature extraction.
New graph kernel scales well with graph size and number, achieving state-of-the-art performance.
problem Graph kernels lose structure information when representing graphs.
method Proposes a positive-definite global alignment graph kernel using random features and random graph embeddings.
result Achieves quasi-linear scalability with respect to graph size and number.
Algorithm segments glandular structures in colon histology images for cancer grading.
problem Manual gland segmentation is time-consuming and risky for patients.
method Local intensity and texture features, Random Forest classifier, multilevel approach.
result Fast, accurate automatic gland segmentation for clinical use.
NN-Stacking improves predictive power of regression models by adjusting stacking coefficients with features.
problem Low predictive power of linear stacking methods.
method NN-Stacking uses neural networks to estimate adaptive stacking coefficients.
result NN-Stacking leads to better predictive power, especially in large datasets.
The paper adapts step sizes in TD learning to identify relevant features.
problem Identifying which features are relevant for temporal-difference learning.
method Adapting step sizes in stochastic gradient descent for feature relevance in TD learning.
result TD IDBD effectively distinguishes relevant features in gridworld and robotic tasks.
Flexible classifier using Mahalanobis distances for non-elliptical distributions.
problem Classifying non-elliptical and multimodal distributions.
method Semiparametric classifier based on Mahalanobis distances and generalized additive models.
result The proposed classifiers outperform traditional methods in high-dimensional, low-sample-size scenarios.
Study shows DNNs often extract redundant features, influenced by network size and activation function.
problem Redundancy in deep neural network features.
method Hierarchical clustering of features based on cosine distances, varying network sizes and activation functions.
result Network size and activation function are key factors in DNN redundancy.
Embed nodes with multi-scale attributes for robust network analysis.
problem Capturing complex node attributes across different scales.
method Multi-scale attributed node embedding (AE & MUSAE) using Skip-gram approach.
result Proves node-feature mutual information is implicitly factorized by embeddings.
The paper investigates why GNNs struggle to generalize from small to large graphs.
problem Challenges in graph neural networks' ability to generalize across different graph sizes.
method Identified and studied the effect of local structure on size generalization; proposed a novel SSL task.
result GNNs can converge to non-generalizing solutions when there is a discrepancy in local structure.
Improved variational inequality algorithms using adaptive step sizes.
problem Solving monotone variational inequalities and convex-concave min-max problems efficiently.
method Adaptive step sizes that eliminate hyperparameters and global Lipschitz continuity requirements.
result Eliminated the need for the golden ratio in the algorithm and improved complexity bounds.
Study ridge ensembles in proportional feature-to-sample size regime, proving risk equivalence and GCV consistency.
problem Characterizing and optimizing ridge ensembles in proportional feature-to-sample size regimes.
method Proportional asymptotics analysis, GCV for tuning, proving risk equivalence.
result Risk of optimal full ridgeless ensemble matches optimal ridge predictor's risk.
AutoStep MCMC adapts step size locally for better sampling efficiency.
problem Challenging step size selection for complex, multiscale targets.
method AutoStep MCMC uses a locally adaptive step size for involutive proposals.
result AutoStep MCMC is π-invariant, irreducible, and aperiodic.
Ensembles of random-feature models can't outperform a single large model.
problem Finding the optimal balance between model size and ensemble size.
method Deterministic equivalent risk estimates and scaling laws analysis.
result Ensembles of random-feature models achieve near-optimal performance only under specific conditions.
Localized Lasso improves interpretable models for high-dimensional data.
problem High-dimensional regression with small sample size and interpretability.
method Sample-wise network regularization and exclusive group sparsity.
result Localized Lasso outperforms alternatives in simulated and genomic data.
Geometric step decay schedules improve stochastic algorithms' convergence on sharp nonconvex problems.
problem Convergence of stochastic algorithms on sharp nonconvex problems.
method Geometric step decay schedule applied to stochastic algorithms.
result Geometric step decay schedules lead to local linear convergence rates for sharp nonconvex problems.
NIS learns optimal embedding sizes for recommendation models.
problem Finding optimal embedding sizes for large-scale recommendation models.
method Neural Input Search (NIS) uses reinforcement learning to automatically find optimal vocabulary sizes and embedding dimensions.
result NIS improves prediction accuracy by 6.8% on Recall@1 and 1.8% on ROC-AUC.
Adapts DR objectives for both sample and feature size reduction.
problem Simultaneously reduce sample and feature sizes.
method Semi-relaxed Gromov-Wasserstein optimal transport.
result OT plan delivers competitive hard clustering.
New method learns domain-invariant local feature patterns for unsupervised domain adaptation.
problem Performance degradation due to domain-shift in unsupervised domain adaptation.
method Jointly learns domain-invariant local feature patterns and holistic feature distributions.
result Superior performance on benchmark datasets compared to state-of-the-art methods.
Large SGD step sizes lead to sparse feature learning in neural networks.
problem Sparse feature learning in neural networks with large step sizes.
method Empirical observations and theoretical analysis of SGD dynamics.
result Large step sizes induce implicit regularization leading to sparse predictors.