Twinning splits data into fast, statistically similar sets.
problem Creating statistically similar data splits for Big Data.
method Twinning is a method based on SPlit for fast, model-independent dataset splitting.
result Twinning is orders of magnitude faster than SPlit.
Proposes using Monte Carlo Dropout in Autoencoder and VAE for synthetic data generation.
problem Handling large amounts of data in costly or difficult-to-collect scenarios.
method Incorporates Monte Carlo Dropout within Autoencoder and Variational Autoencoder.
result Generated data sets are statistically and predictively similar to actual data.
New similarity index avoids limitations of CCA in neural networks.
problem Limitations of existing methods in measuring neural network representation similarity.
method Introducing a similarity index based on centered kernel alignment (CKA) to measure representational similarity matrices.
result CKA reliably identifies correspondences between representations in networks trained from different initializations.
A new test assesses text similarity between two groups of documents.
problem Comparing similarity between two groups of documents.
method Neural network-based language models estimate entropy, and a test statistic derived from an estimation-and-inference framework is used.
result The proposed test maintains the nominal Type one error rate while offering greater power compared to existing methods.
Study on communication limits for distributed convex learning and optimization.
problem Identifying communication efficiency limits for distributed convex learning and optimization.
method Analyzing different assumptions and types of functions under varying information and computational power.
result Many communication rounds may be required without similarity between local objective functions.
Estimates set overlap and similarity using random samples.
problem Estimating set overlap and similarity with limited data.
method Binomial model for predicting set overlap, comparing to previous methods.
result Binomial model provides better estimates with small sample sizes.
Neural networks auto-denoise similar inputs, enabling new statistical analysis.
problem Estimating similarity of inputs for neural networks.
method Define and quantify similarity from neural network perspective, using parameter variation impact on outputs.
result Estimate sample density and quantify denoising effect without true labels.
Unified approach to private statistics from empirical to population data.
problem Divided focus on empirical vs population statistics in private statistics.
method Unified methods for both types of statistics.
result Methods for empirical statistics can be applied to population statistics.
A test ranks generative models based on their similarity to real-world data.
problem Challenges in model selection for generative models due to lack of likelihoods.
method Statistical test of relative similarity using maximum mean discrepancies (MMDs).
result The test provides a meaningful ranking of model performance.
This paper improves prediction in small data sets by eliciting expert knowledge about feature similarities.
problem Improving predictive models from small high-dimensional data sets.
method Eliciting expert knowledge about pairwise feature similarities and using sequential decision making techniques.
result Improvement in predictive performance on both simulated and real data.
CrowdMI uses crowdsourcing to impute missing data, achieving results similar to complex models.
problem Missing data imputation using human crowdsourcing.
method Replicating a multiple imputation framework with multiple crowdworkers completing a survey.
result Valid imputations for both qualitative and quantitative missing data comparable to complex models.
Paper develops efficient statistical estimators for distributed data.
problem Communication and privacy issues in distributed statistical inference.
method Iterative algorithms for distributed optimization, adapting to loss function similarity.
result CEASE estimators achieve statistical efficiency in finite steps.
ProbMinHash improves Jaccard similarity hashing for big data applications.
problem Efficiently estimating set similarities in big data with weighted elements.
method Locality-sensitive hash algorithms that calculate signatures collectively.
result Significantly faster than the original approach, with improved estimation error.
Machine learning and statistical modeling complement each other in healthcare analytics.
problem Choosing between machine learning and statistical modeling for analytics challenges.
method Choosing based on problem, data, and desired outcomes.
result Machine learning and statistical modeling are complementary, using similar principles but different tools.
The paper optimizes similarity learning for better machine learning performance.
problem Improving machine learning performance through better similarity measures.
method Probabilistic framework for pairwise bipartite ranking, focusing on pointwise ROC optimization.
result Universal and faster learning rates derived for the optimization problem.
Method measures weight similarity in neural networks using normalization and statistical inference.
problem Quantifying weight similarity in non-convex neural networks.
method Chain normalization rule and hypothesis-training-testing statistical inference.
result Weights of identical neural networks converge to similar local solutions.
Develops a deep learning approach for statistical arbitrage.
problem Temporal price differences between similar assets.
method Constructs arbitrage portfolios using latent asset pricing factors and a convolutional transformer for time series signals.
result High risk-adjusted returns and Sharpe ratios with optimal trading policy.
Proposes a method to compare noisy high-dimensional datasets with low-dimensional manifolds.
problem Comparing distributions on manifolds in noisy high-dimensional datasets.
method Linking low-rank structure to manifold geometry, developing a scale-invariant distance measure.
result Superior robustness and statistical power compared to existing methods.
Similarity encoding improves learning from messy categorical data.
problem Learning from categorical variables with high cardinality and redundancy.
method Similarity encoding, a generalization of one-hot encoding that uses similarities between categories.
result Similarity encoding significantly outperforms traditional encoding methods in prediction accuracy.
Similar models predict similarly, reducing overfitting risk.
problem Excessive reuse of test data in machine learning.
method Proved model similarity mitigates overfitting and provided a generalization bound.
result Model similarity reduces the risk of overfitting, even when accuracy levels suggest otherwise.
Robust method estimates self-similarity for mammogram images, improving cancer detection.
problem Statistical assessment of self-similarity in real data with large mean level shifts.
method Theil-type weighted regression for wavelet-based estimation, compared to OLS and AV.
result Robust approach shows nearly 68% accuracy in cancer vs non-cancer classification.
Proposes RSP model for efficient big data analysis.
problem Efficiently partitioning big data sets for analysis.
method Random sample partition (RSP) data model and block-level sampling.
result RSP data blocks can estimate statistics and build models equivalent to whole data set.
Detects change-points in similarity networks to identify anomalous nodes.
problem Detecting changes in network structure that affect node similarity.
method Sequential node-wise average similarity measures for change detection; community detection for anomaly isolation.
result Simple sequential procedure effectively identifies change-points and anomalous nodes.
The paper examines how to test if two learning algorithms produce similar outcomes.
problem Testing if two learning algorithms produce similar outcomes when trained on different data sets.
method Using Total Variation (TV) distance to measure similarity of posterior distributions.
result TV indistinguishable learning rules are equivalent to existing stability notions and can be statistically amplified.
Method embeds numeric tabular datasets into a shared vector space for similarity and retrieval.
problem Lack of meaningful representation for numeric tabular datasets in large language models.
method Structured exploratory data analysis descriptors, sentence transformer embedding, CCA for cross-dataset alignment.
result Total P@1 score of 0.9 across 15 datasets, robust nearest-neighbor retrieval and cluster structure.
We introduce a new discrepancy score between two distributions that gives an indication on their similarity. While much research has been done to determine if two samples come from exactly the same distribution, much less research considered the problem of determining if two finite samples come from similar distributio…
Paper proposes an anomaly detection system for DBMS diagnosis.
problem Difficulty in detecting anomalies in DBMS due to increasing metrics.
method Uses deep autoencoder and statistical process control for anomaly detection, and time series similarity for event finding.
result Demonstrates effectiveness of the proposed model in detecting anomalies and finding related events.
Proposes dynamic borrowing method for historical data in clinical trials.
problem Insufficient statistical power in rare and pediatric disease clinical trials.
method Dynamic borrowing method based on frequentist approach using similarity measures.
result Demonstrates usefulness of dynamic borrowing in reanalyzing clinical trial data.
New procedures identify market graph from sign similarity networks.
problem Identifying market graph from sign similarity networks.
method Introducing new statistical procedures for market graph identification in sign similarity networks.
result Optimal procedures in sign similarity networks are less sensitive to stock attribute distributions.
Defines a similarity measure for classification distributions.
problem Measuring similarity between classification distributions.
method Proposes task similarity, a novel measure quantifying performance of source distributions on target distributions.
result Empirical task similarity correlates with transfer efficiency and semantic similarity of source distributions.
The paper reports the construction of artificial stock market that emerges the similar statistical facts with real data in Indonesian stock market. We use the individual but dominant data, i.e.: PT TELKOM in hourly interval. The artificial stock market shows standard statistical facts, e.g.: volatility clustering, the …
Paper proposes a method to validate statistical models based on data consistency.
problem Validation of statistical modeling assumptions in scientific inference problems.
method Automatic evaluation of model consistency with observed data.
result The proposed criterion assesses models' ability to generate similar data.
We investigate scaling and memory effects in return intervals between price volatilities above a certain threshold q for the Japanese stock market using daily and intraday data sets. We find that the distribution of return intervals can be approximated by a scaling function that depends only on the ratio between the …
The study improves the assessment of fairness in face recognition using ROC curves and statistical guarantees.
problem Improving the assessment of fairness in face recognition systems.
method Proves asymptotic guarantees for empirical ROC curves and fairness metrics, and introduces a recentering technique to avoid bootstrap pitfalls.
result Demonstrates the practical relevance of the methods for assessing fairness in face recognition systems.
There are plenty of problems where the data available is scarce and expensive. We propose a generator of semi-artificial data with similar properties to the original data which enables development and testing of different data mining algorithms and optimization of their parameters. The generated data allow a large scal…
BHLR predicts hyperlink weights from data vectors using symmetric similarity functions and Bregman divergence.
problem Predicting hyperlink weights from data vectors in a general framework.
method BHLR learns a symmetric similarity function to minimize Bregman-divergence between hyperlink weights and estimated similarities.
result BHLR is statistically consistent and computationally tractable, providing theoretical guarantees for various methods.
Bayesian neural network improves feature selection and prediction.
problem Improving feature selection and prediction accuracy in neural networks.
method BNN-ARD with l2-norm feature importance measure.
result Improves variable selection and predictive performance on real-world data.
Develops statistical guarantees for neural networks with regularization.
problem Lack of comprehensive mathematical theories for neural networks.
method General statistical guarantee for least-squares with regularizers.
result Prediction error increases sub-linearly in layers, logarithmically in parameters.
New statistical theory explains contrastive learning effectiveness.
problem Understanding why contrastive learning works well for representation extraction.
method Developed a new theoretical framework based on approximate sufficient statistics.
result Near-sufficient encoders derived from contrastive learning can be adapted for downstream tasks.
Proposes rpf-kernel for clustering via random projection forests.
problem Clustering similar data points while distinguishing them from dissimilar ones.
method Random projection forests to learn a similarity kernel.
result rpf-kernel effectively clusters data with competitive performance.
Pairwise quantile regression tackles similarity scoring in biometric systems.
problem Analyzing errors in similarity scoring for facial recognition.
method Established theoretical guarantees for pairwise quantile regression solutions, leveraging sharp concentration results for U-processes. result Proved generalization bounds and identified conditions for fast learning rates.
Flexible multi-task learning framework using summary statistics.
problem Data-sharing constraints in healthcare settings.
method Proposes a flexible multi-task learning framework utilizing summary statistics and adaptive parameter selection.
result Systematic non-asymptotic analysis and simulations demonstrate the method's performance.
Paper introduces new method for statistical inference with stochastic gradients.
problem Uncertainty quantification for solutions from iterative optimization methods.
method Moment-adjusted stochastic gradient descent.
result Established non-asymptotic theory for statistical inference.
Adapts auxiliary losses using gradient similarity to improve neural network performance.
problem Statistical inefficiency in neural networks and difficulty in selecting helpful auxiliary tasks.
method Uses cosine similarity between gradients of tasks to adaptively weight auxiliary losses.
result Guaranteed convergence to critical points of the main task and practical usefulness across domains.
Stock markets are complex systems exhibiting collective phenomena and particular features such as synchronization, fluctuations distributed as power-laws, non-random structures and similarity to neural networks. Such specific properties suggest that markets operate at a very special point. Financial markets are believe…
The study improves firm size data normality for statistical analysis.
problem Firm size data often do not follow a normal distribution.
method Applied Box-Cox transformation to improve normality.
result Transformed firm size data show strong linearity.
We propose a route for the evaluation of risk based on a transformation of the covariance matrix. The approach uses a `potential' or `objective' function. This allows us to rescale data from different assets (or sources) such that each data set then has similar statistical properties in terms of their probability distr…
Estimates shared parameters across related learning problems using robust statistics and LASSO.
problem Simultaneously learning related but heterogeneous problems like store demand or patient risk.
method Two-stage multitask learning estimator combining robust statistics and LASSO regression.
result Improved sample complexity bounds for multitask learning, especially beneficial for 'data-poor' instances.