Algorithm optimizes measurement sequence to minimize data acquisition.
problem Efficiently measure high-dimensional data with minimal measurements.
method Active sequential inference using variational autoencoder (VAE) latent space.
result Optimal measurement sequences chosen to recover high-dimensional data.
Robust variable selection for high-dimensional data with missing and measurement errors.
problem Missing data and measurement errors confound data distribution.
method Exponential loss function with inverse probability weighting and additive error models.
result The Atan punishment method improves robust variable selection.
Paper proposes data quality measures for large-scale high-dimensional data.
problem Lack of practical data quality measures for large-scale high-dimensional data.
method Proposes two data quality measures: class separability and in-class variability. Efficient algorithms based on random projections and bootstrapping are provided.
result Efficient algorithms for computing data quality measures on large-scale high-dimensional data.
This research evaluates measures of dependence for financial time-series data.
problem Accurately preparing time series data and selecting an appropriate measure of dependence is challenging.
method Review and establishment of a comprehensive analysis framework for shaping time-series data and evaluating measures of dependence.
result A method, framework, and example for selecting and evaluating a suitable measure of dependence are presented.
Improves comparison of F-measures for imbalanced datasets.
problem Comparing F-measures for classification algorithms on imbalanced data.
method Two improvements to existing F-measure comparison methods.
result Enhanced accuracy in comparing F-measures for classification.
New stability measures for similar features improve feature selection accuracy.
problem Existing stability measures fail to distinguish similar features in highly correlated datasets.
method Introduce new adjusted stability measures that consider feature similarities.
result One new stability measure considers highly similar features as interchangeable.
In this paper, we propose a new method of Bayesian measurement for spectral deconvolution, which regresses spectral data into the sum of unimodal basis function such as Gaussian or Lorentzian functions. Bayesian measurement is a framework for considering not only the target physical model but also the measurement model…
A new measure of dependence for various data types.
problem Measuring dependence in multivariate, functional, and structured data.
method Combines local normalization with RKHS flexibility.
result Validates the measure's properties and competitive performance.
The study introduces measures of collective mobility from aggregated OD data.
problem Understanding large-scale mobility patterns from aggregated data.
method Developed a framework using synthetic and real data to interpret network-level mobility.
result Aggregated mobility measures reveal network structure and flow constraints.
Logistic regression is a widely used method in several fields. When applying logistic regression to imbalanced data, for which majority classes dominate over minority classes, all class labels are estimated as `majority class.' In this article, we use an F-measure optimization method to improve the performance of logis…
Paper estimates spectral risk measures for insurance data with truncated and censored data.
problem Estimating spectral risk measures for insurance data with left truncation and right censoring.
method Proposes a non-parametric estimator using product limit estimator and establishes asymptotic normality.
result Proposed estimator outperforms existing methods for small k and small sample sizes.
Defining similarity measures is a requirement for some machine learning methods. One such method is case-based reasoning (CBR) where the similarity measure is used to retrieve the stored case or set of cases most similar to the query case. Describing a similarity measure analytically is challenging, even for domain exp…
The paper analyzes Laplace learning for Gaussian measure data in infinite dimensions, proving convergence.
problem Analyzing Laplace learning for infinite-dimensional Gaussian measure data.
method Minimizes Dirichlet energy on a graph constructed from the full dataset.
result Proves pointwise convergence of the graph Dirichlet energy for Gaussian measure data.
Paper infers intrinsic dimension from quasi-convex measurements.
problem Inferring intrinsic dimension from measurements by quasi-convex functions.
method Developed a method using filtration of Dowker complexes based on discrete data of point orderings.
result Correct intrinsic dimension can be inferred in the limit of large data under generic assumptions.
New measures detect HFT activity, revealing its impact on stock prices.
problem Lack of public data on HFT activity.
method Developed machine learning models to predict HFT activity using proprietary and public data.
result Measures outperform conventional proxies and reveal HFT's impact on price discovery.
Community recovery is a central problem that arises in a wide variety of applications such as network clustering, motion segmentation, face clustering and protein complex detection. The objective of the problem is to cluster data points into distinct communities based on a set of measurements, each of which is associat…
Confidence intervals improve evaluation of binary prediction rules in data mining.
problem Uncertainty in performance measures estimation from finite datasets.
method Asymptotic normal approximations for confidence intervals, with a blurring correction.
result Improved finite sample coverage probabilities and general performance measures inference.
The paper deals with the adaptation of a new measure for the unsupervised feature selection problems. The proposed measure is based on space filling concept and is called the coverage measure. This measure was used for judging the quality of an experimental space filling design. In the present work, the coverage measur…
A new measure identifies clusters without assuming data distribution.
problem Identifying the correct number of clusters in data without distribution assumptions.
method Nonparametric interpoint distance-based approach.
result Superior to existing clustering measures, validated on synthetic and real data.
Hybrid clustering combines partitional and hierarchical clustering for computational effectiveness and versatility in cluster shape. In such clustering, a dissimilarity measure plays a crucial role in the hierarchical merging. The dissimilarity measure has great impact on the final clustering, and data-independent prop…
The paper proposes a framework for information-theoretic predictive uncertainty measures.
problem The need for reliable estimation of predictive uncertainty in machine learning.
method Revisiting core concepts, categorizing predictive uncertainty measures based on model and approximation of true distribution.
result Identification of conditions under which certain predictive uncertainty measures excel.
Interestingness measures provide information that can be used to prune or select association rules. A given value of an interestingness measure is often interpreted relative to the overall range of the values that the interestingness measure can take. However, properties of individual association rules restrict the val…
kdiff measures distances for time series and structured data.
problem Estimating distances between time series and structured data.
method kdiff uses non-linear kernel distances based on matching overlapping distributions.
result kdiff is more robust to noise and partial occlusions.
The paper tackles binary classification with measure data using topological descriptors.
problem Binary classification with measure data.
method Develops classifiers for measure data using topological descriptors (persistence diagrams).
result Upper and lower bounds on the Rademacher complexity of classifiers on measures.
Paper discusses methods to measure privacy in synthetic tabular data.
problem Lack of standard methods to quantify privacy in synthetic data.
method Discusses proposed quantification approaches for synthetic data privacy.
result Contributes to SD privacy standards and stimulates discussion.
Measures time-delay embedding for noisy, sparse data.
problem Applying Takens' embedding theorem to real-world, noisy data.
method Formulated a measure-theoretic generalization of the embedding theorem, using optimal transport.
result Reconstructed full state of dynamical systems from time-lagged partial observations robust to noise and sparsity.
Mining association rules is an important technique for discovering meaningful patterns in transaction databases. Many different measures of interestingness have been proposed for association rules. However, these measures fail to take the probabilistic properties of the mined data into account. In this paper, we start …
In data science, it is often required to estimate dependencies between different data sources. These dependencies are typically calculated using Pearson's correlation, distance correlation, and/or mutual information. However, none of these measures satisfy all the Granger's axioms for an "ideal measure". One such ideal…
A good measure of similarity between data points is crucial to many tasks in machine learning. Similarity and metric learning methods learn such measures automatically from data, but they do not scale well respect to the dimensionality of the data. In this paper, we propose a method that can learn efficiently similarit…
Diffusion Maps framework is a kernel based method for manifold learning and data analysis that defines diffusion similarities by imposing a Markovian process on the given dataset. Analysis by this process uncovers the intrinsic geometric structures in the data. Recently, it was suggested to replace the standard kernel …
New RF dissimilarity measures improve multi-view learning accuracy.
problem Improving multi-view learning accuracy in HDLSS problems.
method Modified Random Forest proximity measures for HDLSS multi-view classification.
result Second method significantly more accurate than other state-of-the-art methods.
This paper tackles deep clustering evaluation challenges in high-dimensional data.
problem Evaluation of deep clustering methods is problematic due to the curse of dimensionality and variations in embedding spaces.
method Develops a theoretical framework to highlight the ineffectiveness of internal validation measures and proposes a systematic approach to applying clustering validity indices in deep learning.
result The proposed framework reduces misguidance from improper use of clustering validity indices in deep learning.
The paper assesses quality measures for machine learning models using cross-validation.
problem Evaluating the accuracy and robustness of quality measures for machine learning models.
method Cross-validation approach to estimate prediction error and quantify explained variation. Confidence bounds and local quality measures derived from residuals.
result The reliability and robustness of quality measures are assessed through numerical examples and confidence bounds.
Fourier representation improves KSD for infinite-dimensional data.
problem Applying KSD to infinite-dimensional data.
method Combining measure equations with kernel methods for a Fourier representation of KSD.
result KSD can separate measures in infinite-dimensional Hilbert spaces.
Causal discovery algorithms infer causal relations from data based on several assumptions, including notably the absence of measurement error. However, this assumption is most likely violated in practical applications, which may result in erroneous, irreproducible results. In this work we show how to obtain an upper bo…
A new measure DCSI quantifies separability for density-based clustering.
problem Quantifying meaningful clusters in data sets.
method Developed a new separability measure DCSI based on separation and connectedness.
result Correctly identifies touching or overlapping classes that do not correspond to meaningful density-based clusters.
Bounds on factual and counterfactual distributions under measurement error in discrete models.
problem Measurement errors in discrete data and their impact on inference.
method Expressing modeling assumptions as linear constraints and using linear programming to derive bounds.
result Sharp bounds on factual and counterfactual distributions for various models, including instrumental variable scenarios.
New measures capture tail dependence and non-exchangeability in financial data.
problem Underestimation of tail dependence and inability to capture non-exchangeable tail dependence.
method Tail copulas and novel tail dependence measures (MTCM, ATCM) are proposed.
result Captures non-exchangeable tail dependence and provides analytical forms for various copulas.
The paper analyzes skewness and kurtosis measures for skew-elliptical distributions.
problem Examining skewness and kurtosis measures for skew-elliptical distributions.
method Deriving exact expressions for skewness and kurtosis measures for skew-elliptical distributions, constructing test statistics, and comparing measures through simulations and real data analysis.
result Exact expressions and test statistics for skewness and kurtosis measures for various skew-elliptical distributions.
Machine learning models for repeated measurements are limited. Using topological data analysis (TDA), we present a classifier for repeated measurements which samples from the data space and builds a network graph based on the data topology. When applying this to two case studies, accuracy exceeds alternative models wit…
Spectral Clustering(SC) is a prominent data clustering technique of recent times which has attracted much attention from researchers. It is a highly data-driven method and makes no strict assumptions on the structure of the data to be clustered. One of the central pieces of spectral clustering is the construction of an…
We develop correlated random measures, random measures where the atom weights can exhibit a flexible pattern of dependence, and use them to develop powerful hierarchical Bayesian nonparametric models. Hierarchical Bayesian nonparametric models are usually built from completely random measures, a Poisson-process based c…
Solves Ricci flow on Riemann surfaces with measure initial data.
problem Existence and smoothness of Ricci flow on Riemann surfaces.
method Formulation and solution of existence problem using Ricci flow.
result New examples of nongradient expanding Ricci solitons.
This paper improves the robustness of risk estimation for financial positions.
problem Ensuring robustness of risk measures in the presence of data noise.
method Proposes a quantitative approach using the Fortet-Mourier metric to quantify the variation of true probability measures.
result Derives explicit error bounds for discrepancies between laws of estimators based on true and perturbed data.
Measurement error in the observed values of the variables can greatly change the output of various causal discovery methods. This problem has received much attention in multiple fields, but it is not clear to what extent the causal model for the measurement-error-free variables can be identified in the presence of meas…
Causal inference is central to many areas of artificial intelligence, including complex reasoning, planning, knowledge-base construction, robotics, explanation, and fairness. An active community of researchers develops and enhances algorithms that learn causal models from data, and this work has produced a series of im…
The paper introduces a statistical test to assess and rank distance measures.
problem Assessing the relative information retained by different distance measures.
method Developed a statistical test to compare distance measures.
result Identifies the most informative distance measure among candidates.
Private method measures nonlinear correlations between data hosted across two entities.
problem Measuring nonlinear correlations between sensitive data hosted across multiple parties while preserving privacy.
method Differentially private estimator of distance correlation.
result First private estimator of nonlinear correlations in a multi-party setup.