Study optimizes scoring rules for incentivizing agent's information gathering in online settings.
problem Optimizing incentives for agents to acquire information in online settings.
method Designing a sample-efficient algorithm that tailors the UCB algorithm to the strategic agent's model.
result Achieves sublinear T2/3-regret after T iterations, independent of the number of states. The paper introduces a method for detecting principal communities and embedding vertices.
problem Detecting and embedding vertices in graphs with community structure.
method Principal graph encoder embedding method that detects principal communities and produces vertex embeddings.
result The method successfully detects principal communities and produces accurate vertex embeddings.
New methods improve anomaly detection from large streaming data.
problem Anomaly detection in large streaming data fails with current methods.
method Two novel randomized algorithms (rPS and gPS) for better detection of correlated anomalies.
result High and balanced recall and estimated accuracy for anomaly detection.
Develops a method for causal inference with noisy confounders.
problem Noisy measurements of confounders in treatment effects models.
method Local principal subspace approximation combining K-nearest neighbors matching and PCA.
result Estimators of treatment effects and counterfactual distributions are constructed.
sPCA models may not have orthogonal scores and loadings, complicating interpretation.
problem sPCA scores and loadings may not be orthogonal.
method Illustrated and numerically demonstrated the implications of sPCA on scores, residuals, and variance explained.
result sPCA approaches perform poorly on noise-free, sparse data.
Directly compute classification by learning features with class scores.
problem Classification efficiency and accuracy on various datasets.
method PCA for feature encoding, supervised learning model with encoder-decoder structure.
result Effective classification performance on multiple datasets.
DPA autoencoders learn data distribution and intrinsic dimensionality with guarantees.
problem Learning data distribution and intrinsic dimensionality in unsupervised learning.
method Combines distributionally correct reconstruction with principal-component-like interpretability.
result Exact theoretical guarantees on disentangling factors of variation and intrinsic dimensionality.
With the development of high-throughput technologies, principal component analysis (PCA) in the high-dimensional regime is of great interest. Most of the existing theoretical and methodological results for high-dimensional PCA are based on the spiked population model in which all the population eigenvalues are equal ex…
A novel outlier detection method for high-dimensional data.
problem Challenges in outlier detection in high-dimensional data.
method Principal component analysis and kernel density estimation.
result The proposed method outperforms benchmark methods in F1-score and execution time. Develops an ℓ_p theory for PCA and spectral clustering.
problem Lack of precise characterizations of PCA scores for low-dimensional embedding.
method An ℓ_p perturbation theory for PCA in Hilbert spaces, analyzing eigenvectors and Gram matrix.
result Optimal recovery results for Gaussian mixture and stochastic block models.
Non-parametric method predicts multi-stream longitudinal data evolution.
problem Predicting the evolution of multi-stream longitudinal data for an in-service unit.
method Decomposes each stream into eigenfunctions and FPC scores, uses Gaussian process prior and empirical Bayesian updating.
result Framework outperforms state-of-the-art approaches and achieves high predictive accuracy.
P-OCS detects OOD samples in a low-dimensional subspace, outperforming existing methods.
problem Efficient OOD detection for deep learning models in open-world environments.
method P-OCS operates in the orthogonal complement of the principal subspace, applying a single projected perturbation.
result P-OCS achieves state-of-the-art OOD detection with negligible computational cost and without requiring model retraining.
Explains various PCA and SPCA methods with theory and applications.
problem No specific problem stated; focuses on explaining methods.
method Explains PCA, SPCA, kernel PCA, and kernel SPCA methods with theory and applications.
result Comprehensive coverage of PCA and SPCA methods with theory and applications.
A new stable similarity measure for time series using persistent homology.
problem Constructing a robust measure of time series similarity.
method Persistent homology for stability, bi-conditional periodicity score for similarity.
result Stability of the bi-conditional periodicity score under perturbations and dimension reduction.
Optimizes vessel hull forms using PCA and DNN.
problem Designing optimal hull forms for vessel performances.
method PCA compresses hull forms, DNN predicts performances.
result DNN accurately predicts hull form performances.
Principal component regression (PCR) is a widely used two-stage procedure: principal component analysis (PCA), followed by regression in which the selected principal components are regarded as new explanatory variables in the model. Note that PCA is based only on the explanatory variables, so the principal components a…
Proposes a method to detect anomalies in financial time series using PCA and neural networks.
problem Anomalies in financial time series lead to miscalibrated risk models.
method Extract features using PCA, define anomaly score with neural network, calibrate cutoff value.
result The proposed PCA NN approach outperforms other anomaly detection methods.
Study on reliability of latent reuse in diffusion models under distribution shift.
problem When can latent spaces from a source dataset be reused for a target dataset with different distributions?
method Considered a source-target setting with approximately low-dimensional datasets near different subspaces. Analyzed the target-domain score error due to principal-angle misalignment and target ambient noise.
result Latent reuse is reliable only if the source and target subspaces are close and the target ambient noise is not too amplified.
Improved histogram-based anomaly detector using extended principal component features.
problem Challenges in detecting anomalies in large, correlated datasets.
method Extended histogram-based anomaly detection using principal components.
result Significant improvement in anomaly detection accuracy with no significant increase in runtime.
A novel online framework for analyzing multidimensional functional data.
problem Analysis of multidimensional functional data streams poses significant challenges.
method Online functional principal component analysis using tensor product splines on a Stiefel manifold with Riemannian stochastic gradient descent.
result Efficient and scalable modeling of multidimensional functional data.
Improved MUSE boosts performance and reduces error in Bayesian inference.
problem Hierarchical Bayesian inference problems
method Implicit differentiation applied to MUSE algorithm
result Significant speedup and improved accuracy compared to Hamiltonian Monte Carlo
Principal component analysis (PCA) is very popular to perform dimension reduction. The selection of the number of significant components is essential but often based on some practical heuristics depending on the application. Only few works have proposed a probabilistic approach able to infer the number of significant c…
Unified kernel-based methods improve nonlinear causal discovery.
problem Identifying nonlinear causal relationships between time series variables.
method Unified Kernel Principal Component Regression (KPCR) and Gaussian Process score-based model with Smooth Information Criterion.
result Improved performance in time series nonlinear causal discovery.
Personalized treatment of patients based on tissue-specific cancer subtypes has strongly increased the efficacy of the chosen therapies. Even though the amount of data measured for cancer patients has increased over the last years, most cancer subtypes are still diagnosed based on individual data sources (e.g. gene exp…
Principal component analysis (PCA) for binary data, known as logistic PCA, has become a popular alternative to dimensionality reduction of binary data. It is motivated as an extension of ordinary PCA by means of a matrix factorization, akin to the singular value decomposition, that maximizes the Bernoulli log-likelihoo…
No policy can simultaneously be fully autonomous, optimally calibrated, and helpful, proving a trilemma.
problem Proving impossibility of a policy achieving maximum helpfulness, optimal calibration, and full autonomy.
method Geometric proof showing that adding any non-affine autonomy incentive to a strictly proper scoring rule destroys strict properness.
result The Behavioral Credibility Trilemma: no policy can achieve all three goals simultaneously.
New algorithm balances spatial data approximation and prediction accuracy.
problem Lack of methods considering spatial correlation and downstream modeling in dimension reduction.
method Formalizes approximation and modeling utility as metrics, proposes a balanced algorithm.
result Optimal trade-off between approximation accuracy and downstream modeling utility.
This study examines whether PCA can effectively identify nitrogen pollution sources in rivers.
problem Identifying pollution sources in rivers for effective environmental management.
method Principal Component Analysis and its modifications, along with Independent Component Analysis and Factor Analysis, are applied to nitrogen pollution source identification.
result PCA and related techniques can be powerful tools for uncovering nitrogen pollution sources in rivers.
Performance of nuclear threat detection systems based on gamma-ray spectrometry often strongly depends on the ability to identify the part of measured signal that can be attributed to background radiation. We have successfully applied a method based on Principal Component Analysis (PCA) to obtain a compact null-space m…
The paper proposes a new method for clustering survival data using smoothed log-hazard trajectories.
problem Clustering survival data based on instantaneous risk dynamics.
method Functional Principal Component Analysis applied to B-spline smoothed log-hazard trajectories.
result The proposed method provides an interpretable representation of relative temporal risk dynamics.
New method quantifies uncertainty in denoising models.
problem Uncertainty quantification in denoising models.
method Derives a relation between posterior moments and derivatives, uses it for efficient uncertainty quantification.
result Efficient computation of principal components and full marginal distributions of the posterior.
This paper introduces a new unsupervised method for dimensionality reduction via regression (DRR). The algorithm belongs to the family of invertible transforms that generalize Principal Component Analysis (PCA) by using curvilinear instead of linear features. DRR identifies the nonlinear features through multivariate r…
A new model integrates covariates with grade of membership analysis for better latent structure recovery.
problem Improving latent structure recovery in multivariate categorical data analysis.
method Covariate-assisted grade of membership model exploiting shared low-rank simplex geometry.
result Auxiliary covariates can provably improve latent structure recovery, leading to faster convergence rates.
This research creates efficient models for cyclo-stationary systems using generative methods.
problem Efficiently modeling systems with periodic forcing.
method Score-based generative modeling for reduced-order models.
result Accurately reproduces statistical properties and temporal correlations of cyclo-stationary time series.
In 2018, at the World Economic Forum in Davos it was presented a new countries' economic performance metric named the Inclusive Development Index (IDI) composed of 12 indicators. The new metric implies that countries might need to realize structural reforms for improving both economic expansion and social inclusion per…
Proposes a new method to explain model predictions for consumer recourse.
problem Current explanation methods fail to provide meaningful recourse to decision subjects.
method Develops feature responsiveness scores to highlight actionable features.
result Standard practices can undermine decision subjects by highlighting unresponsive features.
Rotates MFVI for better Gaussian approximations.
problem Improving variational approximations for complex distributions.
method Rotated coordinate system, PCA-based rotation, iterative Gaussianization.
result Significantly more accurate approximations with lower computational cost.
As a popular tool for producing meaningful and interpretable models, large-scale sparse learning works efficiently when the underlying structures are indeed or close to sparse. However, naively applying the existing regularization methods can result in misleading outcomes due to model misspecification. In particular, t…
LLmFPCA-detect detects anomalies in sparse longitudinal text data using LLMs and mFPCA.
problem Challenges in detecting patterns and anomalies in sparse longitudinal textual data.
method Pairs LLM-based text embeddings with mFPCA to detect clusters and anomalies.
result LLmFPCA-detect outperforms state-of-the-art baselines on Amazon and Wikipedia datasets.
Analysis of an organization's computer network activity is a key component of early detection and mitigation of insider threat, a growing concern for many organizations. Raw system logs are a prototypical example of streaming data that can quickly scale beyond the cognitive power of a human analyst. As a prospective fi…
The main purpose of this work is to examine the behavior of the implied volatility smiles around jumps, contributing to the literature with a high-frequency analysis of the smile dynamics based on intra-day option data. From our high-frequency SPX S\&P500 index option dataset, we utilize the first three principal compo…
Study evaluates information leakage in Polymarket markets, finding limited applicability and resolution ambiguity.
problem Limited applicability of ILS-dl framework across Polymarket markets.
method Scaling from single-case to population-scale evaluation using ILS-dl framework.
result Only 0.7% of candidate markets yield computable ILS-dl values, and resolution semantics are the main obstacle.
ACA identifies and explains anomalies in data.
problem Explaining anomalies in non-supervised data analysis.
method Abnormal Component Analysis (ACA) using data depth.
result ACA provides a linear explanation for anomalies.
A novel unsupervised outlier detection method using Randomized PCA Forest.
problem Unsupervised outlier detection in datasets.
method Randomized Principal Component Analysis (RPCA) Forest for deriving an outlier score.
result Superior performance compared to classical and state-of-the-art methods.
Introduces generalized principal bundles and connections, linking them to standard gauge theories.
problem Generalized principal bundles and connections in field theories.
method Local coordinate transformation laws and horizontal lifts.
result Generalized principal connections are associated to Lie group fiber bundle connections.
A new method uses string method to explore diffusion models.
problem Understanding the geometry of learned distributions in diffusion models.
method String method to compute continuous paths between samples.
result The string method identifies realistic morphing sequences and transition pathways.
Introduces principal fairness for fair decision-making.
problem Discrimination among similarly affected individuals.
method Uses principal stratification from causal inference.
result Explicitly accounts for decision impacts, not just protected attributes.
Study of semi-principal bundles using group actions and wreath products.
problem Understanding bundles with fibers as free G-spaces. method Defining semi-principal bundles, bases, and frame bundles; using wreath products and functors.
result Semi-principal bundles can be retracted to principal bundles, preserving parallel transport.