Hybrid LLM generates synthetic data preserving causal parameters.
problem Synthetic data fails to accurately estimate causal effects.
method Combines model-based covariate synthesis with separately learned propensity and outcome models.
result Hybrid framework ensures causal structure in synthetic data.
A novel approach using graph learning and synthetic long positions for statistical arbitrage in options markets.
problem Exploiting statistical arbitrage opportunities in options markets using machine learning.
method Two-stage graph learning approach: first stage defines a novel prediction target isolating pure arbitrages via synthetic bonds; second stage proposes SLSA positions.
result Statistically significant outperformance of GL baselines and consistent positive returns with an average P&L-contract information ratio of 0.1627.
Synthetic experiments are crucial for assessing causal machine learning methods.
problem Current empirical evaluations of causal machine learning methods are insufficient and unreliable.
method Propose principles for conducting rigorous empirical analyses with synthetic data.
result Rigorous synthetic experiments are essential for building trust in causal machine learning methods.
The paper corrects bias in synthetic data for imbalanced learning.
problem Challenges in balancing false positive and negative rates in imbalanced data.
method Proposes a bias correction procedure to generate synthetic data for minority groups.
result Enhances prediction accuracy while avoiding overfitting.
In-context learning solves PU classification without iterative optimization.
problem Binary classification with only labeled positives and unlabeled samples.
method Pretrained transformer (PUICL) that learns from synthetic PU datasets.
result Outperforms four standard PU learning baselines on 20 benchmarks.
Paper proposes a method to estimate true positive proportion without knowing it.
problem Bias in binary classifier performance due to different positive item proportions.
method Maximum likelihood estimator for true proportion of positives.
result Method accurately estimates true positive proportion in data sets.
New synthetic data analysis reveals high type 1 error rates.
problem Analyzing synthetic data for inference raises significant methodological challenges.
method Developed statistical inference tools and conducted a simulation study.
result Type 1 error rates are unacceptably high in synthetic data analysis.
Proposes efficient calibration for indoor localization models.
problem Calibration data scarcity in wireless indoor localization.
method Uses synthetic labels and prediction sets to fine-tune a predictor and estimate bias.
result Yields rigorous coverage guarantees for prediction sets.
New k-means method clusters radar image sequences using SPD matrices.
problem Clustering radar image sequences efficiently.
method Developed k-means on SPD matrices for non-Euclidean data. result Effective clustering of radar image sequences via SPD matrices.
A framework evaluates synthetic tabular data quality objectively.
problem Lack of an objective interpretation of tabular data metrics.
method Proposes a single mathematical objective for synthetic tabular data distribution, structurally decomposes it, and unifies existing metrics.
result Synthesizers that represent tabular structure outperform other methods, especially on smaller datasets.
SynthBH uses synthetic data to control FDR in multiple testing.
problem Controlling false discovery rate in multiple hypothesis testing.
method SynthBH, a synthetic-powered multiple testing procedure.
result SynthBH guarantees FDR control with synthetic data.
Generative synthetic data can preserve predictive accuracy but distort causal inference.
problem Distortion of average treatment effect estimates in synthetic data.
method Hybrid synthetic-data framework that generates covariates while modeling treatment and outcome mechanisms separately.
result Hybrid synthesis improves causal fidelity compared to fully generative baselines.
The Penrose theorem and Hawking's topology theorem are extended to weighted spacetimes.
problem Extending Penrose's singularity theorem and Hawking's topology theorem to weighted spacetimes.
method Using weighted null energy condition and synthetic dimension to generalize the theorems.
result Generalized versions of the Penrose and Hawking theorems hold under a weighted null energy condition.
New mass inequalities and proofs for causal variational principles.
problem Proving new mass inequalities for causal variational principles.
method Proved a new inequality for minimizers of causal variational principles and applied it to prove the positive mass theorem.
result Introduced a positive quasilocal mass and proved new mass inequalities.
SHAP Distance assesses semantic fidelity of synthetic tabular data.
problem Semantic fidelity of synthetic tabular data is not well evaluated.
method SHAP Distance, defined as cosine distance between global SHAP attribution vectors.
result SHAP Distance detects semantic discrepancies overlooked by standard measures.
The rapid digital transformation without security considerations has resulted in the rise of global-scale cyberattacks. The first line of defense against these attacks are Network Intrusion Detection Systems (NIDS). Once deployed, however, these systems work as blackboxes with a high rate of false positives with no mea…
We develop Square Root Graphical Models (SQR), a novel class of parametric graphical models that provides multivariate generalizations of univariate exponential family distributions. Previous multivariate graphical models [Yang et al. 2015] did not allow positive dependencies for the exponential and Poisson generalizat…
EmDT generates synthetic fraud data to improve detection accuracy.
problem Imbalanced datasets in fraud detection lead to poor performance on rare fraudulent transactions.
method EmDT uses UMAP clustering to identify fraudulent patterns and a Transformer denoising network to generate synthetic data.
result EmDT significantly improves classification performance compared to existing methods.
Latent Noise Injection improves synthetic data generation for privacy and statistical alignment.
problem Slow convergence of generative models in high-dimensional settings.
method Latent Noise Injection using Masked Autoregressive Flows (MAF).
result Synthetic data closely reflects the underlying distribution, especially in high-dimensional settings.
NP-PROV separates mean and variance spaces to improve function uncertainty.
problem Neural Processes fail on out-of-domain tasks due to shared latent space uncertainty.
method Separates mean and variance into function-value-related and position-related latent spaces.
result NP-PROV achieves state-of-the-art likelihood with bounded variance in drifts.
New inequality for eigenfunctions on curved spaces.
problem Eigenfunctions on non-smooth spaces with Ricci curvature.
method Sharp reverse-Hölder inequality for Dirichlet Laplacian eigenfunctions.
result Generalizes classical comparison theorem to curved spaces.
Improved method for computing Fréchet means on SPD matrices.
problem Computing Fréchet means on the manifold of SPD matrices.
method Random matrix theory-based approach for estimating Fréchet means.
result Significantly outperforms state-of-the-art methods in experiments.
Study dynamic assortment and positioning of products with varying display effects.
problem Dynamic assortment and positioning of products with varying display effects.
method Design round-based learning algorithms for both multiplicative and general position effects models, and develop efficient subroutines for optimization.
result First regret-optimal characterization for both models, with matching upper and lower bounds.
Positive-confidence (Pconf) classification [Ishida et al., 2018] is a promising weakly-supervised learning method which trains a binary classifier only from positive data equipped with confidence. However, in practice, the confidence may be skewed by bias arising in an annotation process. The Pconf classifier cannot be…
Gen-LRA attacks synthetic data leakage without model knowledge.
problem Auditing synthetic data privacy leakage.
method Generative Likelihood Ratio Attack (Gen-LRA).
result Gen-LRA outperforms other attacks across metrics.
Syntax designs adaptive trials for subpopulations with potential benefits.
problem Identifying subpopulations with positive treatment effects in diverse patient populations.
method Adaptive patient recruitment and synthetic control estimation.
result Syntax outperforms conventional trial designs in identifying beneficial subpopulations.
CTSyn generates high-quality synthetic tabular data.
problem Challenges in generating high-quality synthetic tabular data.
method Diffusion-based generative foundation model with autoencoder and conditional latent diffusion.
result CTSyn outperforms existing table synthesizers on standard benchmarks.
We study the deformation of spherical conical metrics with at least some of the cone angles larger than 2π. We show in this note via synthetic geometry that for one family of such metrics, there is local rigidity in the choice of cone positions if angles are fixed. This gives an evidence of the analytic obstruction c…
New method debiases selection bias in PU classification with exposure data.
problem Binary classification from positive and unlabeled data with selection bias.
method Automatic Debiased PUE (ADPUE) learning method.
result ADPUE outperforms traditional PU learning methods on various datasets.
A novel method for learning DAGs from positive-valued data.
problem Causal discovery from observational data of positive-valued variables.
method Hybrid Moment-Ratio Scoring (H-MRS) algorithm combining moment-based scoring and log-scale regression.
result H-MRS integrates log-scale Ridge regression for moment-ratio estimation with a greedy ordering procedure based on raw-scale moment ratios, followed by Elastic Net-based parent selection.
New method classifies manifold-valued data using Riemannian geometry.
problem Classifying data on curved Riemannian manifolds.
method Probabilistic Learning Vector Quantization on Symmetric Positive Definite Matrices.
result The method outperforms traditional Euclidean methods on manifold-valued data.
The study improves harmonic map theory for metric spaces with curvature bounds.
problem Harmonic maps between specific metric spaces with curvature constraints.
method Synthetic geometry, Optimal Transport, Heat Flow, viscosity theory.
result Established Lipschitz continuity and Bochner-Eells-Sampson inequality.
In this work, we consider the task of classifying binary positive-unlabeled (PU) data. The existing discriminative learning based PU models attempt to seek an optimal reweighting strategy for U data, so that a decent decision boundary can be found. However, given limited P data, the conventional PU models tend to suffe…
New ranking algorithms improve online content delivery by learning from click data.
problem Bias in ranking systems due to production system biases.
method Proposed novel extensions of LinUCB and Linear Thompson Sampling algorithms to handle position-based click model.
result Validated the proposed algorithms through offline and online experiments.
New MCMC methods improve efficiency for large network inference.
problem Efficiency of Metropolis within Gibbs for large networks.
method Combination of split Hamiltonian Monte Carlo and Firefly Monte Carlo.
result New methods outperform Metropolis within Gibbs on synthetic and real networks.
Meta-learning framework for few-shot one-class classification using order-equivariant networks.
problem Few labeled examples for positive class in one-class classification tasks.
method Order-equivariant networks for meta-learning a binary classifier conditioned on positive examples.
result Meta-learning framework outperforms baselines on unseen synthetic streams.
The paper tackles sparse graph learning under Laplacian-related constraints, improving upon existing methods.
problem Learning a sparse undirected graph from multivariate data under Laplacian-related constraints.
method Modifications to penalized log-likelihood approaches to enforce total positivity and lasso/adaptive lasso penalties using ADMM.
result The proposed constrained adaptive lasso approach significantly outperforms existing Laplacian-based approaches.
In many applications, different populations are compared using data that are sampled in a biased manner. Under sampling biases, standard methods that estimate the difference between the population means yield unreliable inferences. Here we develop an inference method that is resilient to sampling biases and is able to …
Serial crystallography is the field of science that studies the structure and properties of crystals via diffraction patterns. In this paper, we introduce a new serial crystallography dataset comprised of real and synthetic images; the synthetic images are generated through the use of a simulator that is both scalable …
Adaptive source selection for positive transfer in linear models improves target dataset performance.
problem Limited task-specific labeled data in business settings.
method Greedily decides from which sources and how many samples to incorporate into the target dataset using an accept/reject rule based on a data-dependent estimate of the transfer gain.
result Consistent gains over classical and recent strong baselines while avoiding negative transfer.
Continuous metrics on R^3 with specific properties have non-negative harmonic mass.
problem Proving non-negativity of mass for continuous metrics.
method Defining harmonic mass and using properties of approximating smooth metrics.
result The harmonic mass of continuous metrics is non-negative.
Introduces a new model for mapping matrices to matrices, subsuming linear regression.
problem Learning matrix-to-matrix mappings from data.
method Partial trace regression model, leveraging quantum information theory.
result Relevance demonstrated in matrix-to-matrix regression and positive semidefinite matrix completion.
Kernel K-means clusters probability distributions.
problem Clustering a sample of probability distributions.
method Mapping distributions to kernel mean embeddings in RKHS, then applying K-means.
result Effective unsupervised classification of probability distributions.
New method predicts positive samples with missing labels.
problem Missing labels due to response-dependent factors.
method P(U)U-O-Mixture algorithm for joint estimation.
result Non-convex algorithm leads to optimal statistical error.
We propose a novel deep learning architecture suitable for the prediction of investor interest for a given asset in a given time frame. This architecture performs both investor clustering and modelling at the same time. We first verify its superior performance on a synthetic scenario inspired by real data and then appl…
We consider gradient descent with `momentum', a widely used method for loss function minimization in machine learning. This method is often used with `Nesterov acceleration', meaning that the gradient is evaluated not at the current position in parameter space, but at the estimated position after one step. In this work…
Measure contraction property is a synthetic Ricci curvature lower bound for metric measure spaces. We consider Sasakian manifolds with non-negative Tanaka-Webster Ricci curvature equipped with the metric measure space structure defined by the sub-Riemannian metric and the Popp measure. We show that these spaces satisfy…
Framework uses human feedback to safely set OOD detection thresholds, reducing false positives.
problem Challenges in setting OOD detection thresholds for safety-critical applications.
method Mathematically grounded framework leveraging expert feedback to dynamically update thresholds.
result Guaranteed to meet FPR constraint while minimizing human feedback, maintaining FPR at most 5%.