Improved nonparametric regression with debiasing for root-n consistency.
problem Challenges in achieving root-n consistency and normal distribution for nonparametric estimators.
method Debiasing technique by adding a correction term to nonparametric estimators.
result Achieves root-n consistency and asymptotic normality.
New method estimates mutual information using normalizing flows.
problem Mutual information estimation in high-dimensional data.
method Normalizing flows to map data to target distributions with known MI.
result Theoretical guarantees and practical advantages demonstrated.
We prove that the only compact, origin-symmetric, strictly convex ancient solutions of the planar p centro-affine normal flows are contracting origin-centered ellipses.
We consider partially observed multiscale diffusion models that are specified up to an unknown vector parameter. We establish for a very general class of test functions that the filter of the original model converges to a filter of reduced dimension. Then, this result is used to justify statistical estimation for the u…
New priors can update posteriors without re-estimating likelihoods.
problem Degradation of classification approaches when class priors change.
method Recompute posteriors using recovered likelihoods from original posteriors and new priors.
result Dynamic update of original posteriors is possible without re-estimating likelihoods.
New method efficiently interpolates nonparametric density estimators.
problem Efficient evaluation of nonparametric density estimators.
method Piecewise multivariate polynomial interpolation scheme.
result New estimator with low space requirements and efficient querying.
Study geodesic Lie groups' convergence to limits with quantitative estimates.
problem Quantifying convergence rates of geodesic Lie groups to their limits.
method Estimates on the difference between original metrics and asymptotic/tangent metrics.
result Sharpens existing bounds on convergence rates.
Improved MoM estimator enhances classical shadows protocol for quantum measurements.
problem Efficient estimation of expectation values with reduced measurement shots.
method Modified median-of-means estimator with optimal constants and U-statistics.
result Improved performance of modified estimator for Clifford measurements.
CDRE estimates density ratios in streaming data without historical samples.
problem Online learning with shifting data distributions.
method Iterative estimation of density ratios between initial and current distributions.
result CDRE outperforms standard DRE in estimating divergences between distributions.
Improved KernelSHAP via linear regression for ML model interpretation.
problem Efficiently estimating Shapley values in model-agnostic settings.
method Revisiting KernelSHAP via linear regression, developing techniques for convergence and uncertainty.
result Original KernelSHAP incurs negligible bias for significant variance reduction.
New nonconvex penalty smooths at origin for deep learning.
problem Improving variable selection and bias in high-dimensional statistical learning.
method Developed a new nonconvex penalty function smooth at origin.
result Asymptotic bias of new penalty function vanishes exponentially fast.
Javanmard and Montanari propose a debiased estimator for high-dimensional regression.
problem Bias in high-dimensional regression models.
method Debiased LASSO estimator.
result Debiased LASSO yields asymptotically normal estimators and valid hypothesis tests.
We investigated the topological properties of stock networks through a comparison of the original stock network with the estimated stock network from the correlation matrix created by the random matrix theory (RMT). We used individual stocks traded on the market indices of Korea, Japan, Canada, the USA, Italy, and the …
Paper solves long-standing Gaussian curvature conjecture for minimal graphs.
problem Gaussian curvature of minimal graphs over the unit disk.
method Complex-analytic methods, conformal harmonic parameterization.
result Sharp estimate for Gaussian curvature at the origin of minimal graphs.
Simplified argument for second order estimate in quaternionic Calabi-Yau problem.
problem Second order estimates for quaternionic Calabi-Yau problem on hyperkähler manifolds.
method Simplified argument to derive the estimate.
result Simplified derivation of second order estimate.
Calibrated Prediction-Powered Inference improves semisupervised mean estimation by calibrating prediction scores.
problem Semisupervised mean estimation with a small labeled sample and a large unlabeled sample, and miscalibrated prediction models.
method Calibrated Prediction-Powered Inference (Calibeating) post-hoc calibrates the prediction score on the labeled sample before using it for semisupervised estimation.
result Calibrated Prediction-Powered Inference can improve the original score both as a predictor of the outcome and as a regression adjustment for semisupervised inference.
δ-CLUE generates diverse explanations for model uncertainty.
problem Lack of constraints in generating explanations for uncertainty estimates.
method Augmenting CLUE approach to provide a set of plausible explanations.
result Returns a set of diverse inputs that yield confident predictions.
The paper analyzes the statistical properties of GANs using f-divergence.
problem Understanding the statistical behavior of GANs and comparing different f-divergences. method Asymptotic analysis of f-divergence GANs, including Kullback-Leibler divergence. result Asymptotically equivalent GANs with the same discriminator classes for correctly specified models.
Due to the limited resources and the scale of the graphs in modern datasets, we often get to observe a sampled subgraph of a larger original graph of interest, whether it is the worldwide web that has been crawled or social connections that have been surveyed. Inferring a global property of the original graph from such…
Replication study shows Deep-SE still not as effective as previously thought for agile effort estimation.
problem Improving accuracy in estimating agile software development effort.
method Close replication of Deep-SE using additional data and comparison with multiple baselines.
result Deep-SE outperforms only a few cases, suggesting more work is needed.
The paper proves the consistency and efficiency of a volatility estimator in noisy data.
problem Proving the consistency and efficiency of a volatility estimator in the presence of microstructure noise.
method Proves asymptotic normality using Central Limit Theorem for Fourier spot volatility estimator.
result Proves consistency and asymptotic efficiency of the Fourier spot volatility estimator in noisy data.
Proves Hessian estimates for special Lagrangian equation with new proofs.
problem Interior Hessian estimates for special Lagrangian equation
method Doubling proofs, higher codimension analogue of previous methods
result Higher codimension analogue of gradient estimate for minimal hypersurfaces
Post-estimation smoothing improves prediction accuracy with structural indices.
problem Using natural structural indices in machine learning without losing robustness.
method A post-estimation smoothing operator that separates from the original predictor.
result Post-estimation smoothing improves accuracy over original predictors under simple conditions.
Paper presents a new way to estimate model changes without full model evaluation.
problem Efficiently estimating changes in model parameters and outputs due to data point removal.
method Dual representation of influence functions for linearizable models, reducing computational complexity.
result The dual representation can be an efficient alternative to original influence functions, especially for large models.
New method improves PCA for high-dimensional data with n < p.
problem PCA struggles in high-dimensional settings with n < p.
method Pairwise differences covariance estimation with four regularized versions.
result Proposed methods outperform existing estimators in high-dimensional data settings.
We give a new and complete proof of Hamilton's injectivity radius estimate for sequences with bounded and almost nonnegative curvature operators, unbounded diameters, and bump-like origins. Such sequences arise in particular from dilations about a singularity of the Ricci flow on a 3-manifold.
Study on residual Monge-Ampère mass for symmetric plurisubharmonic functions.
problem Analyzing the residual Monge-Ampère mass of symmetric plurisubharmonic functions.
method Proved zero mass for functions with zero Lelong number at origin and S1-invariance. result Zero mass conjecture answered for symmetric functions.
Clinical models can be unstable, leading to unreliable predictions.
problem Stability of clinical prediction models developed using statistical or machine learning methods.
method Simulation and case studies of statistical and machine learning approaches to show instability in model predictions.
result Model instability often leads to miscalibration of predictions in new data.
Study improves curvature estimate for stable marginally outer trapped hypersurfaces with a free boundary.
problem Curvature estimate for stable marginally outer trapped hypersurfaces with a free boundary.
method Iteration argument based on uniform area bound.
result Improved curvature estimate for stable marginally outer trapped hypersurfaces.
We extend the randomized singular value decomposition (SVD) algorithm \citep{Halko2011finding} to estimate the SVD of a shifted data matrix without explicitly constructing the matrix in the memory. With no loss in the accuracy of the original algorithm, the extended algorithm provides for a more efficient way of matrix…
This paper considers the problem of estimating a high-dimensional vector of parameters θ∈Rn from a noisy observation. The noise vector is i.i.d. Gaussian with known variance. For a squared-error loss function, the James-Stein (JS) estimator is known to dominate the simple maximum-likelihood (…
In this paper, we propose a new threshold-kernel jump-detection method for jump-diffusion processes, which iteratively applies thresholding and kernel methods in an approximately optimal way to achieve improved finite-sample performance. We use the expected number of jump misclassifications as the objective function to…
This paper studies directed exploration for reinforcement learning agents by tracking uncertainty about the value of each available action. We identify two sources of uncertainty that are relevant for exploration. The first originates from limited data (parametric uncertainty), while the second originates from the dist…
New method reduces copyright risks in AI-generated images.
problem Copyright issues in AI-generated images.
method Genericization method using originality estimation and PREGen technique.
result PREGen reduces likelihood of generating copyrighted characters by over half.
Paper constructs non-symmetric collapsing spacetimes without symmetries.
problem Forming non-symmetric collapsing spacetimes in vacuum.
method Modified Christodoulou's a priori estimates and gluing construction.
result Past geodesic completeness and asymptotic Minkowski space.
Kuwert and Schätzle showed in 2001 that the Willmore flow converges to a standard round sphere, if the initial energy is small. In this situation, we prove stability estimates for the barycenter and the quadratic moment of the surface. Moreover, in codimension one we obtain stability bounds for the enclosed volume and …
Novel compression method preserves privacy while reducing communication costs.
problem Reducing communication costs in differential privacy mechanisms.
method Poisson private representation (PPR) for compressing and simulating local randomizers.
result Achieves compression within a logarithmic gap from theoretical lower bound.
This study proposes sparse estimation methods for the generalized linear models, which run one of least angle regression (LARS) and least absolute shrinkage and selection operator (LASSO) in the tangent space of the manifold of the statistical model. This study approximates the statistical model and subsequently uses e…
Proposes a method to stabilize treatment effect estimation with unbalanced data.
problem Unbalanced treatment assignment leading to unstable propensity score estimations.
method Undersamples data for propensity score modeling and calibrates scores to match original distribution.
result The estimator retains asymptotic properties of the DML estimator and improves finite sample performance.
The Lugannani-Rice formula is a saddlepoint approximation method for estimating the tail probability distribution function, which was originally studied for the sum of independent identically distributed random variables. Because of its tractability, the formula is now widely used in practical financial engineering as …
Robustly aligns datasets with partial GW distance to handle contamination.
problem Aligning contaminated datasets using Gromov-Wasserstein distances.
method Proposes a partial GW distance estimator to minimize distortion from outliers.
result The partial GW distance estimator is minimax optimal and near-optimal in finite samples.
As opposed to standard empirical risk minimization (ERM), distributionally robust optimization aims to minimize the worst-case risk over a larger ambiguity set containing the original empirical distribution of the training data. In this work, we describe a minimax framework for statistical learning with ambiguity sets …
Investigates numerical issues in GP interpolation parameter estimation.
problem Numerical issues in maximum likelihood parameter estimation for Gaussian process interpolation.
method Investigates and proposes strategies to improve open-source software implementations.
result Improves reliability and reproducibility of studies relying on GP implementations.
Proposes a new method for kernel density estimation using stagewise minimization and a simple dictionary.
problem Kernel density estimation with data-adaptive weighting parameters and sparse representation.
method Stagewise minimization algorithm based on U-divergence and a simple dictionary. result Develops non-asymptotic error bound for the proposed estimator.
Estimates input from output of nonlinear systems using ANN.
problem Estimating unknown compositional input from system output.
method Artificial Neural Networks (ANNs) for nonlinear system inversion.
result ANNs can compete with optimal bounds for linear systems and demonstrate promising results for nonlinear systems.
In this paper we propose using the principle of boosting to reduce the bias of a random forest prediction in the regression setting. From the original random forest fit we extract the residuals and then fit another random forest to these residuals. We call the sum of these two random forests a \textit{one-step boosted …
We analyze differences between two information-theoretically motivated approaches to statistical inference and model selection: the Minimum Description Length (MDL) principle, and the Minimum Message Length (MML) principle. Based on this analysis, we present two revised versions of MML: a pointwise estimator which give…
Digital technologies ignited a revolution in the agrifood domain known as precision agriculture: a main question for enabling precision agriculture at scale is if accurate product quality control can be made available at minimal cost, leveraging existing technologies and agronomists' skills. As a contribution along thi…