Survival analysis models research reproducibility, offering new insights.
problem Reproducibility crisis in machine learning research.
method Survival analysis to model reproducibility as a continuous process.
result Survival analysis provides deeper insights into research reproducibility.
The reproducibility of scientific research has become a point of critical concern. We argue that openness and transparency are critical for reproducibility, and we outline an ecosystem for open and transparent science that has emerged within the human neuroimaging community. We discuss the range of open data sharing re…
Study on reproducibility in optimization with bounds on limits.
problem Limits of reproducibility in noisy or error-prone optimization procedures.
method Defined a quantitative measure of reproducibility and analyzed convex optimization settings.
result Revealed a fundamental trade-off between computation and reproducibility.
ERICA assesses reproducibility in cluster analysis.
problem Lack of a unified framework for evaluating cluster analysis replicability.
method ERICA (iterative clustering assignments) method to quantify replicability.
result Demonstrates ERICA's ability to identify reproducible cluster structure.
Paper introduces RKHM and KME for richer data analysis.
problem Lack of rich data structures in kernel methods.
method Proposes RKHM and KME for functional data analysis.
result RKHM captures structural properties in functional data.
In this paper, we discuss the approaches we took and trade-offs involved in making a paper on a conceptual topic in pattern recognition research fully reproducible. We discuss our definition of reproducibility, the tools used, how the analysis was set up, show some examples of alternative analyses the code enables and …
Paper introduces RKHM for more explicit variable structures analysis.
problem Explicitly analyzing structures among variables.
method Orthonormal systems in Hilbert C∗-modules, RKHM. result Theoretical and practical procedures for RKHM orthonormalization.
Evaluating the computational reproducibility of data analysis pipelines has become a critical issue. It is, however, a cumbersome process for analyses that involve data from large populations of subjects, due to their computational and storage requirements. We present a method to predict the computational reproducibili…
Study assesses consistency and reproducibility of LLMs in finance and accounting tasks.
problem Consistency and reproducibility of LLM outputs in finance and accounting research.
method Extensive experimentation with 50 independent runs across 5 tasks using 3 OpenAI models.
result Task-specific patterns of consistency and reproducibility, with binary classification and sentiment analysis achieving near-perfect reproducibility.
Teaches reproducible research to medical students and postgrads.
problem Lack of reproducibility in medical research practices.
method Designed and delivered a lecture series on reproducible research.
result Encountered practical obstacles in reproducing a published analysis.
Study of regularized least squares in RKKS with indefinite kernels.
problem Asymptotic properties of regularized least squares with indefinite kernels in RKKS.
method Introducing a bounded hyper-sphere constraint, theoretical demonstration of globally optimal solution, modified error decomposition techniques, matrix perturbation theory.
result Derivation of learning rates in RKKS, same as RKHS under certain conditions.
What makes a paper independently reproducible? Debates on reproducibility center around intuition or assumptions but lack empirical results. Our field focuses on releasing code, which is important, but is not sufficient for determining reproducibility. We take the first step toward a quantifiable answer by manually att…
This study connects Gaussian processes and RKHS, bridging two machine learning communities.
problem Understanding the relationship between Gaussian processes and RKHS.
method Examining connections and equivalences in regression, interpolation, and other topics.
result Established the equivalence between Gaussian Hilbert space and RKHS.
The study uses reproducing kernels to model bond discount curves.
problem Estimating bond discount curves under no-arbitrage conditions.
method Introduced reproducing kernels as a regression basis for estimating bond discount curves.
result Reproducing kernels provide a tractable solution for calibrating models to market data.
Competency questions help experts select best clustering for energy data.
problem Ad hoc and subjective selection of clustering structures by domain experts.
method Formalize expert knowledge and requirements with competency questions.
result Competency questions improve reproducibility and evaluation of clustering applications.
A typical approach in estimating the learning rate of a regularized learning scheme is to bound the approximation error by the sum of the sampling error, the hypothesis error and the regularization error. Using a reproducing kernel space that satisfies the linear representer theorem brings the advantage of discarding t…
The paper uses Banach spaces to analyze neural networks.
problem Understanding the function spaces of neural networks.
method Theory of reproducing kernel Banach spaces.
result Representer theorem for wide class of Banach spaces.
Study improves LLMs for PPI analysis by addressing uncertainty.
problem Uncertainty in LLM predictions for PPIs.
method Fine-tuned LLaMA-3 and BioMedGPT models, LoRA ensembles, Bayesian LoRA for UQ.
result Competitive PPI identification performance across diverse disease contexts.
Survey on reproducibility and distortion issues in text clustering and topic modeling.
problem Reproducibility and misleading cluster geometry in unsupervised learning for text categorization.
method Systematic literature review of text clustering and topic modeling from 2011-2022.
result Outliers and initialization issues are significant factors in text clustering and topic modeling.
Kernel methods are studied in a mean field limit for high-dimensional data.
problem Analyzing kernel methods in high-dimensional data with many variables.
method Investigation of kernel methods in the mean field limit of interacting particle systems.
result Rigorous mean field limit of kernels and detailed analysis of the limiting reproducing kernel Hilbert space.
NetML provides datasets and challenges for network traffic analysis.
problem Lack of representative datasets and reproducibility issues in network traffic analysis.
method Released three open datasets with flow features and raw packets, implemented machine learning methods.
result NetML datasets will serve as a common platform for AI-driven research.
Unsupervised clustering can reproduce categorization systems if features and metrics are correctly selected.
problem Reproducing expert-provided categorization systems using unsupervised clustering.
method Investigated using toy datasets and real-world fund categorization. Used appropriate feature selection and a supervised Random Forest-based distance metric.
result Unsupervised clustering can reproduce ground truth classes if features and metrics are correctly selected.
Python package for SPD matrix distances, reproducible and extensible.
problem Computing distances between SPD matrices for various applications.
method Unified, extensible framework supporting multiple SPD metrics.
result Reproducible and accessible SPD matrix comparison tool.
Study uses SGD to learn operators in Hilbert spaces with convergence analysis.
problem Learning operators in general Hilbert spaces with SGD.
method Proposes weak and strong regularity conditions for convergence analysis.
result SGD converges to best linear approximation of nonlinear operators.
Recently, there has been emerging interest in constructing reproducing kernel Banach spaces (RKBS) for applied and theoretical purposes such as machine learning, sampling reconstruction, sparse approximation and functional analysis. Existing constructions include the reflexive RKBS via a bilinear form, the semi-inner-p…
Targeted Learning uses robust statistics for reproducible research.
problem Improving reproducibility and rigor in statistical analyses.
method Principled standard for statistical estimation and inference, minimizing assumptions.
result Enhances reliability of statistical conclusions.
Develops exact and invariant study-based decompositions for network meta-analysis.
problem Lack of exact contribution decompositions in network meta-analysis.
method Contrast-space projection formulation of NMA, study-based definition of direct and indirect evidence.
result Exact covariance-aware decompositions of NMA estimator into direct and indirect contributions.
Paper proposes a method for early stopping in regression using reproducing kernels.
problem Early stopping for iterative learning algorithms in nonparametric regression.
method Data-driven rule based on minimum discrepancy principle, validated by fixed-point analysis of localized Rademacher complexities.
result The proposed rule is minimax-optimal and performs comparably to cross-validation.
Paper extends RPD for better handling multiple modalities and non-convexity.
problem Handling multiple modalities and non-convexity in data clouds.
method Computes RPD in a reproducing kernel Hilbert space using kernel principal component analysis.
result The method outperforms RPD and is comparable to other models on benchmark datasets.
This study assesses the reproducibility of 1H-MRS scans across different vendors and sessions.
problem Lack of harmonization in magnetic resonance spectroscopy protocols among vendors.
method Analysis of CV and ICC for within- and between-sessions, and correlation coefficients for across machines.
result Metabolite concentrations are highly reproducible across different vendors and sessions.
Tract-specific diffusion measures, as derived from brain diffusion MRI, have been linked to white matter tract structural integrity and neurodegeneration. As a consequence, there is a large interest in the automatic segmentation of white matter tract in diffusion tensor MRI data. Methods based on the tractography are p…
Develops vector-valued RKBS for neural networks and operators.
problem Understanding function spaces of Rd-valued neural networks and neural operators. method Defines and constructs vector-valued RKBS (vv-RKBS) without restrictive assumptions.
result Establishes Representer Theorem for neural architectures.
The paper advocates for interpretable, accountable, reproducible machine learning in medicine.
problem Black box models in medicine lack transparency and regulatory approval.
method Intrinsically interpretable modeling approaches and collaborative learning paradigms.
result Interpretable machine learning models can support clinical decisions and gain regulatory approval.
Proposes incorporating noise sources in machine learning evaluation for more reliable conclusions.
problem Inadequate handling of nondeterminism in machine learning research leads to unreliable results.
method Uses linear mixed effects models (LMEMs) and generalized likelihood ratio tests (GLRT) to analyze performance evaluation scores and assess performance differences.
result Demonstrates how to incorporate various sources of noise and data properties into statistical significance testing and reliability analysis.
Eluder dimension and information gain are equivalent for reproducing kernel Hilbert spaces.
problem Complexity measures in bandit and reinforcement learning.
method Equivalence of eluder dimension and information gain for reproducing kernel Hilbert spaces.
result Eluder dimension and information gain are equivalent for reproducing kernel Hilbert spaces.
New approach to supervised learning in RKHS and vvRKHS using C∗-algebras.
problem Traditional supervised learning in RKHS and vvRKHS.
method Generalizing supervised learning to RKHM using C∗-algebras. result Constructing RKHMs with enhanced representation power.
Two new methods for analyzing repeated measures data using embeddings into Reproducing Kernel Hilbert Spaces.
problem Analyzing complex data structures with multiple features over time.
method Two generalizations of canonical correlation analysis for repeated measures data using embeddings into Reproducing Kernel Hilbert Spaces.
result Consistency rates for transformation and correlation estimators, relaxing common assumptions.
Study finds long-range dependence in financial markets, but deep generative models struggle to replicate it.
problem Long-range dependence in financial markets and challenges of deep generative models.
method Empirical analysis of financial data from three sectors, including LRD through various statistical methods and deep learning models.
result Deep generative models can reproduce stylized features but fail to capture long-range dependence structures.
Much of human knowledge sits in large databases of unstructured text. Leveraging this knowledge requires algorithms that extract and record metadata on unstructured text documents. Assigning topics to documents will enable intelligent search, statistical characterization, and meaningful classification. Latent Dirichlet…
Estimates kernel eigenvalues for compositional dot-product kernels.
problem Improving estimates for kernel eigenvalues.
method Eigenvalue decay estimates of integral operators associated with dot-product kernels.
result Improved estimates for kernel volumes in reproducing kernel Hilbert spaces.
New theoretical tools simplify kernel-based tests analysis.
problem Asymptotic behavior of kernel-based tests in various scenarios.
method Avoids complex expansions and limit theorems, works directly with Hilbert spaces random functionals.
result Framework leads to simpler analysis with minimal regularity conditions.
This paper presents a stochastic behavior analysis of a kernel-based stochastic restricted-gradient descent method. The restricted gradient gives a steepest ascent direction within the so-called dictionary subspace. The analysis provides the transient and steady state performance in the mean squared error criterion. It…
The paper verifies deep neural networks' ability to approximate functions on spheres.
problem Theoretical verification of deep neural networks' performance on spherical functions.
method Spherical analysis using reproducing kernels and convolutional factorizations.
result Rates of uniform approximation for functions in Sobolev spaces and additive ridge forms.
Kernelized cumulants improve statistical analysis in high-dimensional spaces.
problem Statistical analysis in high-dimensional spaces with low variance estimators.
method Extending cumulants to RKHS using tensor algebra and kernel trick.
result Kernelized cumulants provide new all-purpose statistics with computational tractability.
New rates for GLD and SGLD in infinite-dimensional spaces without dimensionality issues.
problem Gradient Langevin dynamics and SGLD convergence rates in high-dimensional spaces.
method Analysis of GLD and SGLD in infinite-dimensional Hilbert spaces, using stochastic differential equations and Markov chains.
result Derivation of dimension-free convergence rates for GLD and SGLD.
We derive a system of stochastic differential equations simulating the dynamics of the three agent groups with herding interaction. Proposed approach can be valuable in the modeling of the complex socio-economic systems with similar composition of the agents. We demonstrate how the sophisticated statistical features of…
Develops robust persistence diagrams using kernel methods.
problem Persistence diagrams are sensitive to data perturbations.
method Constructs robust persistence diagrams from superlevel filtrations of robust density estimators using reproducing kernels.
result Robust persistence diagrams are consistent estimators in bottleneck distance.
The paper studies convergence of kernel autocovariance operators for stationary processes.
problem Estimating autocovariance operators of stationary processes on Polish spaces.
method Investigates convergence of empirical estimates of autocovariance operators under various conditions.
result Provides consistency results for kernel PCA and spectral analysis methods.