Model complexity is an important factor to consider when selecting among graphical models. When all variables are observed, the complexity of a model can be measured by its standard dimension, i.e. the number of independent parameters. When hidden variables are present, however, standard dimension might no longer be ap…
Study proposes local effective dimension to measure model capacity and generalization error.
problem Capturing the generalization power of machine learning models.
method Proposes local effective dimension as a capacity measure.
result Local effective dimension bounds the generalization error and correlates well with it.
A new measure of model complexity based on Fisher Information.
problem Model complexity measurement in statistical models.
method Effective dimension defined by the number of cubes needed to cover the model space.
result The effective dimension is scale-dependent and measures model complexity.
New method corrects Laplace/BIC errors in singular models, revealing effective dimension.
problem Laplace/BIC errors in singular models due to incorrect effective dimension assumption.
method RLCT (real log canonical threshold) to correct effective dimension in linear models.
result Correct evidence slope and effective dimension estimation in linear settings.
Study shows how to effectively predict functions on manifolds using kernel methods.
problem Regression on manifolds with limited data.
method Reproducing kernel Hilbert space methods, Weyl law, effective dimension.
result Kernel regression estimator yields minimax-optimal error bounds controlled by effective dimension.
New algorithm reduces sketching dimension to effective problem size.
problem Solving L2-regularized least-squares problems efficiently.
method Randomized algorithm using Gaussian and SRHT embeddings.
result Preserves convergence guarantees with reduced embedding dimension.
NGD models have higher effective dimension than SGD models.
problem Measuring model complexity accurately.
method Comparison of NGD and SGD models using effective dimension measures.
result NGD models have a higher effective dimension than SGD models.
Sliced inverse regression (SIR) is a pioneer tool for supervised dimension reduction. It identifies the effective dimension reduction space, the subspace of significant factors with intrinsic lower dimensionality. In this paper, we propose to refine the SIR algorithm through an overlapping slicing scheme. The new algor…
Paper analyzes ensemble Kalman updates for effective dimension and localization.
problem Why small ensemble sizes work well in inverse problems and data assimilation.
method Non-asymptotic analysis of ensemble Kalman updates, focusing on effective dimension and localization.
result Rigorously explains why a small ensemble size is sufficient when prior covariance has moderate effective dimension.
Estimates mean dimension of neural networks to reveal interaction effects.
problem Understanding interaction effects in neural networks.
method Estimation procedure for mean dimension from datasets, analyzing layer-by-layer evolution and impact of activation functions.
result Mean dimension reveals differences in interaction magnitude across neural network architectures.
We obtain an Einstein metric of constant negative curvature given an arbitrary boundary metric in three dimensions, and a conformally flat one given an arbitrary conformally flat boundary metric in other dimensions. In order to compute the on-shell value of the gravitational action for these solutions, we propose to in…
Paper improves learning efficiency by focusing on effective dimensionality.
problem Dimensionality bottleneck in modern learning tasks.
method Developed tools to reduce dimensional costs using effective dimensionality.
result Uniform concentration bounds involving effective dimensionality, improving over existing results.
New theory shows deep networks adapt to data's intrinsic dimensionality even when data isn't on a low-dimensional manifold.
problem Existing theories on deep nonparametric regression assume data lie on a low-dimensional manifold, which is often not the case in real-world applications.
method Introduces effective Minkowski dimension to characterize the intrinsic dimension of data subsets and proves sample complexity depends on this new complexity notation.
result Deep neural networks can adapt to the effective Minkowski dimension of data, circumventing the curse of dimensionality for moderate sample sizes.
Study extends compactness theorems to weighted manifolds with integral curvature bounds.
problem Estimating diameter of weighted manifolds under curvature constraints.
method Extended Sprouse's compactness theorems to weighted manifolds with integral curvature bounds. Used ε-range to handle specific cases. Extended segment inequality to weighted manifolds.
result Proved theorems for weighted manifolds with effective dimension ≤ 1 and ≥ dimension.
Generalization in nonlinear least squares can be studied via algorithmic stability and effective dimension.
problem Generalization in nonlinear least squares models
method Deriving error bounds for local minimizers using algorithmic stability and effective dimension
result Bounds depend on learned geometry rather than parameter count
SignSGD analysis quantifies its effects in high dimensions.
problem Understanding signSGD's effects in high-dimensional settings.
method High-dimensional analysis of signSGD, deriving SDE and ODE for risk.
result Quantification of signSGD's effects: effective learning rate, noise compression, diagonal preconditioning, gradient noise reshaping.
New stability analysis improves generalization of multipass SGD.
problem Improper preconditioning affects generalization in multipass SGD.
method Developed on-average stability analysis for multipass SGD.
result Proper preconditioning yields optimal effective dimension dependence.
SGD generalizes well in high dimensions without regularization.
problem Generalization of overparameterized models in high dimensions.
method Stochastic Gradient Descent (SGD) for convex and locally convex loss functions.
result Generalization error is independent of the ambient dimension p under certain conditions. Study bandit problems under censorship, estimating performance loss.
problem Estimating performance loss in bandit problems with censored feedback.
method Introduced a broad class of censorship models and analyzed their effective dimension.
result Effective dimension naturally leads to results analogous to uncensored settings.
SCBMs model causal effects using low-dimensional bottlenecks.
problem Causal effect estimation in high-dimensional systems.
method Structural causal models with low-dimensional summary statistics.
result SCBMs provide a flexible framework for task-specific dimension reduction.
Estimates manifold dimension from random samples.
problem Estimating the dimension of a manifold from random samples.
method Explicit theoretical and heuristic bounds for data set size.
result Data set needs to be sufficiently large for accurate dimension estimation.
A faster method for estimating effects in large data using fixed-point trees.
problem Estimating heterogeneous effects in large dimensions with computational efficiency.
method Fixed-point approximation to eliminate Jacobian estimation and speed up GRFs.
result Significant computational efficiency improvement without sacrificing statistical accuracy.
We prove the existence of Sasakian-Einstein metrics on infinitely many rational homology spheres in all odd dimensions greater than 3. In dimension 5 we obain somewhat sharper results. There are examples where the number of effective parameters in the Einstein metric grows exponentially with dimension.
This paper proposes a novel kernel approach to linear dimension reduction for supervised learning. The purpose of the dimension reduction is to find directions in the input space to explain the output as effectively as possible. The proposed method uses an estimator for the gradient of regression function, based on the…
Introduces relative information gain for improving Gaussian process regression rates.
problem Improving the sample complexity of estimating or maximizing unknown functions.
method Introduces relative information gain, interpolates between effective dimension and information gain, and proves PAC-Bayesian bounds.
result Obtains minimax-optimal rates of convergence through the relative information gain.
A mathematical model describes deforming manifolds with precise vectors and fields.
problem Modeling and describing the deformation of complex manifolds in practical applications.
method Proposes a modified differential dynamic model with constraints on spatial and temporal continuity, presenting deforming vector and field.
result Demonstrates the effectiveness of an autonomous deforming field in data dimension reduction tasks.
New method improves counterfactual distribution learning for high-dimensional outcomes.
problem Counterfactual distribution learning for high-dimensional outcomes with concentrated structure.
method Geometry-adaptive diffusion-guided smoothing estimators combining causal nuisance adjustment and local outcome geometry.
result Geometry-adaptive methods show steeper error decay in semi-synthetic experiments.
We present MDP Playground, a testbed for Reinforcement Learning (RL) agents with dimensions of hardness that can be controlled independently to challenge agents in different ways and obtain varying degrees of hardness in toy and complex RL environments. We consider and allow control over a wide variety of dimensions, i…
The paper improves confidence set construction for statistical inference.
problem Constructing reliable confidence sets in statistical inference.
method Establishes a finite-sample bound using effective dimension and generalized self-concordance.
result Developed a confidence set adapted to optimization landscapes.
The current study proposes a dimension reduction method, stepwise support vector machine (SVM), to reduce the dimensions of large p small n datasets. The proposed method is compared with other dimension reduction methods, namely, the Pearson product difference correlation coefficient (PCCs), recursive feature eliminati…
New method corrects missing data bias in dimension reduction.
problem Missing data complicates high-dimensional data analysis.
method Developed a bias-corrected Gram matrix for heterogeneous missingness.
result Proposed method improves dimension reduction techniques significantly.
Our goal in this paper is to develop an effective estimator of fractal dimension. We survey existing ideas in dimension estimation, with a focus on the currently popular method of Grassberger and Procaccia for the estimation of correlation dimension. There are two major difficulties in estimation based on this method. …
Study on VC dimension of GCNNs with input resolution effects.
problem Understanding the generalization capabilities of GCNNs.
method Derived upper and lower bounds for VC dimension, analyzed factors affecting it.
result Extended previous results on VC dimension of GCNNs, providing insights into input resolution dependence.
SQUEAK reduces space complexity for Nystrom approximations in KRR.
problem Large datasets in KRR require impractical storage space.
method SQUEAK uses unnormalized ridge leverage scores for incremental updates.
result Space complexity improved with constant factor worse than exact RLS.
In this text we give a decomposition result on polynomial poly-vector fields generalizing a result on the decomposition of homogeneous Poisson structures. We discuss consequences of this decomposition result in particular for low dimensions and low degrees. We provide the tools to calculate simple cubic Poisson structu…
Proposes a method for evaluating multiple dimensions of organizational effectiveness using DEA.
problem Evaluating multiple dimensions of organizational effectiveness in large data sets.
method Introduces two regularized DEA models (SBM and GP-SBM) to estimate both dimension-specific and aggregate efficiency scores.
result Demonstrates improved efficiency and validity compared to conventional methods.
Recently we introduced T-duality in the study of topological insulators, and used it to show that T-duality trivialises the bulk-boundary correspondence in 2 dimensions. In this paper, we partially generalise these results to higher dimensions and briefly discuss the 4D quantum Hall effect.
This paper explores saturation effects in spectral algorithms over large dimensions.
problem Saturation effects in spectral algorithms over large dimensions.
method Improved minimax lower bound and gradient flow with early stopping strategy.
result Exact convergence rates of spectral algorithms in large dimensional settings.
The correspondence between Riemann-Finsler geometries and effective field theories with spin-independent Lorentz violation is explored. We obtain the general quadratic action for effective scalar field theories in any spacetime dimension with Lorentz-violating operators of arbitrary mass dimension. Classical relativist…
Proposes a deep learning method for effective data representation.
problem Constructing effective data representations for prediction.
method A deep dimension reduction approach to learning representations with sufficiency, low dimensionality, and disentanglement.
result The proposed deep nonparametric representation is consistent and performs better than existing methods.
The aim of this paper is to give an upper bound for the dimension of a torus T which acts on a GKM manifold M effectively. In order to do that, we introduce a free abelian group of finite rank, denoted by A(Γ,α,∇), from an (abstract) (m,n)-type GKM graph (Γ,α,∇). Here, an (m,n)-type GKM …
Kernel balancing weights are generalized as KRRR, providing better confidence intervals for treatment effects.
problem Lack of generalization error, correct feature specification, and limited to average effects.
method Interpreting kernel balancing weights as KRRR, relaxing feature specification, and extending Gaussian approximation.
result KRRR provides strong generalization properties and justifies confidence sets for causal functions.
We study the relationship between national culture and the disposition effect by investigating international differences in the degree of investors' disposition effect. We utilize brokerage data of 387,993 traders from 83 countries and find great variation in the degree of the disposition effect across the world. We fi…
Variable selection and dimension reduction are two commonly adopted approaches for high-dimensional data analysis, but have traditionally been treated separately. Here we propose an integrated approach, called sparse gradient learning (SGL), for variable selection and dimension reduction via learning the gradients of t…
Characterizes learnability of forgiving 0-1 loss functions in multiclass settings.
problem Understanding when multiclass learning with forgiving 0-1 loss functions is possible.
method Introduces a new combinatorial dimension based on Natarajan Dimension to determine learnability.
result A hypothesis class is learnable if and only if the Generalized Natarajan Dimension is finite.
Augmented KRnet improves flow-based generative modeling by maintaining exact invertibility.
problem Maintaining exact invertibility in flow-based generative models.
method Integrates augmented dimensions into KRnet to achieve full nonlinear updates in two iterations, keeping exact invertibility.
result Augmented KRnet achieves full nonlinear updates in two iterations, maintaining exact invertibility.
Functional determinant for mixed signature sphere products depends on sphere dimensions and parity.
problem Determining the functional determinant for scalar fields on mixed signature sphere products.
method Analyzing the GJMS operator on SqimesSp to derive the functional determinant. result The functional determinant depends only on the total dimension and parity of the sphere dimensions.
Paper proposes adaptive parameter selection for KGD algorithms.
problem Improving parameter selection for kernel-based gradient descent.
method Integrates bias-variance analysis with splitting method, introduces empirical effective dimension.
result Adaptive parameter selection strategy achieves optimal generalization error bound.