Develops graphical calculus for monoidal categories with twisted pivotal structures.
problem Constructing modules for surfaces with Morse functions or foliations.
method Graphical calculus and string nets for monoidal categories with twisted pivotal structures.
result Twisted string net modules assemble in an oriented categorified 2-TQFT.
New construction of Turaev-Viro invariants invariant under Morita equivalence.
problem Constructing Turaev-Viro invariants invariant under Morita equivalence.
method Pivotal bicategory construction of spherical module categories.
result The invariant recovers the standard Turaev-Viro invariant and is independent of the skeleton.
A string-net model associates a vector space to a surface in terms of graphs decorated by objects and morphisms of a pivotal fusion category modulo local relations. String-net models are usually considered for spherical fusion categories, and in this case the vector spaces agree with the state spaces of the correspondi…
Adapts pivoting technique to circle homeomorphisms for proofs.
problem Probabilistic Tits alternative and exponential synchronization.
method Adapts Gou{ë}zel's pivoting technique.
result Different proofs of probabilistic Tits alternative and exponential synchronization.
In many real-world scenarios, an autonomous agent often encounters various tasks within a single complex environment. We propose to build a graph abstraction over the environment structure to accelerate the learning of these tasks. Here, nodes are important points of interest (pivotal states) and edges represent feasib…
New pivoting strategy improves trace norm contraction in low-rank approximation.
problem Finding good low-rank approximations of symmetric, positive-definite matrices.
method Choosing rows with likelihood proportional to Aii2 for randomly pivoted partial Cholesky algorithm. result Same trace norm contraction result in Frobenius norm for improved pivoting strategy.
New quadrature method using randomly pivoted Cholesky outperforms existing techniques.
problem Efficiently approximating integrals of functions in reproducing kernel Hilbert spaces.
method Nodes drawn by randomly pivoted Cholesky algorithm.
result Randomly pivoted Cholesky quadrature is fast and achieves comparable accuracy to more computationally intensive methods.
Structural correspondence learning (SCL) is an effective method for cross-lingual sentiment classification. This approach uses unlabeled documents along with a word translation oracle to automatically induce task specific, cross-lingual correspondences. It transfers knowledge through identifying important features, i.e…
Paper reinterprets ARP algorithm and improves its analysis and speed.
problem Column subset selection in data analysis.
method Volume sampling and active learning connections, new analysis, rejection sampling.
result Faster implementations and new analysis for ARP algorithm.
Estimates proportions of LLM-generated text in mixed documents.
problem Estimating the proportion of text generated by a pre-specified LLM in mixed documents.
method Developed estimators for two observation regimes: full observation and pivotal reduction, and established sample complexity bounds.
result Full observation estimators require fewer samples than pivotal reduction estimators.
Study on estimating Gumbel--Max watermark proportions in edited documents.
problem Estimating the proportion of a document generated from a watermarked LLM.
method Comparison of full observation and pivotal reduction observation regimes; development of estimators and information-theoretic lower bounds.
result Full observation yields a substantially smaller sample complexity compared to pivotal reduction.
The paper constructs semistrict monoidal 2-categories from foam evaluations.
problem Creating examples of semistrict monoidal 2-categories.
method Using a closed foam evaluation formula as input, the paper rigorously constructs semistrict monoidal 2-categories.
result The constructed monoidal 2-categories are semistrict, have duals and adjoints, and carry a spatial duality structure.
New algorithm improves plant breeding by clustering soybean genotypes more accurately and efficiently.
problem Low accuracy and high computational complexity in clustering plant genotypes.
method Spectral Clustering with Pivotal Sampling for phenotypic data.
result Our algorithm achieves substantially more accuracy than existing methods.
We extend the notion of an ambidextrous trace on an ideal (developed by the first two authors) to the setting of a pivotal category. We show that under some conditions, these traces lead to invariants of colored spherical graphs (and so to modified 6j-symbols).
PiVoT improves real-time multi-object detection and tracking in clutter.
problem Challenges in multi-object detection and tracking from noisy point clouds.
method Variational inference for fast, clutter-resilient multi-object tracking.
result Substantial performance improvement over existing Bayesian trackers.
Develops an empirical likelihood framework for random forests and ensembles.
problem Quantifying the statistical uncertainty of random forests and ensembles.
method Empirical likelihood framework exploiting the incomplete U-statistic structure of ensemble predictions. result Modified empirical likelihood statistic achieves accurate coverage and practical reliability.
In high dimensional sparse regression, pivotal estimators are estimators for which the optimal regularization parameter is independent of the noise level. The canonical pivotal estimator is the square-root Lasso, formulated along with its derivatives as a "non-smooth + non-smooth" optimization problem. Modern technique…
Methods for prediction and tolerance intervals in non-normal models.
problem Constructing prediction and tolerance intervals for non-normal data.
method Two approaches: pivotal quantity approximation and confidence interval for mean.
result Intuitive, simple, efficient methods with proper operating characteristics.
Exact selective inference with randomization for Gaussian regression models.
problem Exact selective inference in Gaussian regression models.
method Introduces a pivot for exact selective inference with randomization, reducing the problem to a bivariate truncated Gaussian distribution.
result Our pivot leads to exact inference and produces narrower confidence intervals than related methods.
Generalizes string-net modular functors to non-spherical categories.
problem Extending string-net models to non-spherical categories.
method Using non-semisimple string-nets and Drinfeld centers.
result Equivalence between string-net and Lyubashenko modular functors.
We study unsupervised multilingual alignment, the problem of finding word-to-word translations between multiple languages without using any parallel data. One popular strategy is to reduce multilingual alignment to the much simplified bilingual setting, by picking one of the input languages as the pivot language that w…
A modular object in a symmetric monoidal bicategory is a Frobenius algebra object whose product and coproduct are biadjoint, equipped with a braided structure and a compatible twist, satisfying rigidity, ribbon, pivotality, and modularity conditions. We prove that the oriented 3-dimensional bordism bicategory of 1-, 2-…
Estimates watermarked content proportions in mixed-source texts.
problem Optimally estimating the proportion of watermarked content in texts with mixed sources.
method Casting the problem as estimating a proportion parameter in a mixture model based on pivotal statistics.
result Proposes efficient estimators for watermark proportion and shows their accuracy through evaluations.
Cross-domain sentiment classification (CDSC) is an importance task in domain adaptation and sentiment classification. Due to the domain discrepancy, a sentiment classifier trained on source domain data may not works well on target domain data. In recent years, many researchers have used deep neural network models for c…
Several techniques for domain adaptation have been proposed to account for differences in the distribution of the data used for training and testing. The majority of this work focuses on a binary domain label. Similar problems occur in a scientific context where there may be a continuous family of plausible data genera…
Study on knotting in very long polymer chains, finding Poisson distribution for prime knot types.
problem Understanding knotting in very long polymer chains.
method Generated and analyzed 243−k polygons of size n=2k using tree data structure and pivot algorithm. Used new knot diagram simplification and invariant-free classification. result Number of prime summands of knot type K in a random n-gon is well described by a Poisson distribution. New 4-manifold invariant defined from trisection diagrams.
problem Defining a new 4-manifold invariant from trisection diagrams.
method Algebraic data from bimodule categories and spherical fusion categories, described diagrammatically.
result Includes Hopf algebraic invariants and modular fusion category invariants.
PIVOT bridges Black-Scholes price and implied volatility spaces via a differentiable layer.
problem Lack of a differentiable interface between price and implied volatility spaces.
method Develops PIVOT, a differentiable layer that preserves LBR's forward pass and avoids backpropagation through branch logic, addressing singularity issues.
result PIVOT achieves high performance and accuracy, reducing price and implied volatility errors by up to 43.4% and 21.3% respectively.
Accelerated RPCholesky speeds up kernel matrix approximations.
problem Efficiently approximating large kernel matrices.
method Accelerated randomly pivoted Cholesky (RPCholesky) with block matrix computations and rejection sampling.
result Approximates kernel matrices up to 40 times faster.
In the setting of high-dimensional linear regression models, we propose two frameworks for constructing pointwise and group confidence sets for penalized estimators which incorporate prior knowledge about the organization of the non-zero coefficients. This is done by desparsifying the estimator as in van de Geer et al.…
Novel framework uses synthetic data to quantify uncertainty in complex data.
problem Uncertainty quantification in complex, unstructured data.
method Perturbation-Assisted Sample Synthesis (PASS) and Perturbation-Assisted Inference (PAI) framework.
result Statistically guaranteed validity in inference, enhancing reliability of synthetic data.
PANDA improves linear discriminant analysis in high dimensions with minimal tuning.
problem Linear discriminant analysis in high-dimensional settings.
method PANDA: a tuning-insensitive method for linear discriminant analysis.
result PANDA achieves optimal convergence rates in estimation error and misclassification rate.
Random walk speed on Teichmüller space is a proper function.
problem Understanding the speed of random walks on Teichmüller space.
method Adaptation of Gouëzel's pivoting techniques to Teichmüller space.
result Speed of random walk is a proper function on Teichmüller space.
BiHRNN predicts inflation by leveraging hierarchical structure and bidirectional RNNs.
problem Accurate inflation forecasting is challenging due to dynamic factors and the layered structure of the Consumer Price Index.
method Bi-directional Hierarchical Recurrent Neural Network (BiHRNN) model that uses bidirectional information flow between levels and informative constraints on RNN parameters.
result BiHRNN significantly outperforms traditional RNN models in forecasting accuracy.
RPCholesky approximates kernel matrices with few evaluations.
problem Approximating kernel matrices efficiently.
method Randomly pivoted partial Cholesky factorization.
result RPCholesky provides nearly optimal low-rank approximations.
Proposes a method to identify causal relationships using background knowledge.
problem Identifying causal relationships in the presence of background knowledge.
method Learning local structure using all types of causal background knowledge (direct, non-ancestral, ancestral). Criteria for identifying causal relationships based on local structure.
result Effective and efficient method for local structure learning and causal relationship identification.
Introduces admissible skein modules for non-semisimple categories.
problem No specific problem stated; generalization of Kauffman skein algebra.
method Introduces admissible skein modules associated to ideals in pivotal categories.
result These modules generalize Kauffman skein algebra and relate to quantum invariants.
Paper shows linear convergence of ISTA and FISTA for ill-conditioned images.
problem Solving linear inverse problems with sparse representation in signal and image processing.
method Revisits iterative shrinkage-thresholding algorithms (ISTA) and improves their convergence properties.
result Linear convergence of ISTA and FISTA for strongly convex smooth parts, even in ill-conditioned cases.
This paper analyzes Local SGD for federated learning, achieving both statistical and communication efficiency.
problem Statistical estimation and inference in federated learning with decentralized data.
method Local SGD, a multi-round estimation procedure using intermittent communication.
result Local SGD achieves both statistical efficiency and communication efficiency.
We investigate the relationship between the algebra of tensor categories and the topology of framed 3-manifolds. On the one hand, tensor categories with certain algebraic properties determine topological invariants. We prove that fusion categories of nonzero global dimension are 3-dualizable, and therefore provide 3-di…
We announce a Godbillon-Vey index formula for longitudinal Dirac operators on a foliated bundle $(X,\F)$ with boundary; in particular, we define a Godbillon-Vey eta invariant on the boundary foliation, that is, a secondary invariant for longitudinal Dirac operators on type III foliations. Our theorem generalizes the cl…
Stochastic Gradient Descent (SGD) is a workhorse in machine learning, yet its slow convergence can be a computational bottleneck. Variance reduction techniques such as SAG, SVRG and SAGA have been proposed to overcome this weakness, achieving linear convergence. However, these methods are either based on computations o…
Study on statistical inference for nonlinear stochastic approximation with Markovian data.
problem Statistical inference for nonlinear stochastic approximation algorithms with Markovian data.
method Established a functional central limit theorem for the partial-sum process of the target parameter estimate, providing asymptotic pivotal statistics for constructing confidence intervals.
result Valid and efficient asymptotic inference method for nonlinear stochastic approximation algorithms with Markovian data.
Generative AI improves stock selection by synthesizing features from diverse data sources.
problem Automating feature discovery in stock market data.
method Used large language models with retrieval-augmented generation and structured prompting to synthesize features from various data sources.
result AI-generated features consistently outperform baselines, with Sharpe improvements ranging from 14% to 91%.
Estimates time-series drifts from i.i.d. data using a direct Nadaraya-Watson plug-in method.
problem Nonparametric estimation of Schrödinger bridge drifts from single time interval data.
method Direct Nadaraya-Watson plug-in estimator based on kernelized numerator and denominator terms.
result Uniform non-asymptotic bound, CLT under undersmoothing, and adaptive bandwidth selector.
Low-rank modeling plays a pivotal role in signal processing and machine learning, with applications ranging from collaborative filtering, video surveillance, medical imaging, to dimensionality reduction and adaptive filtering. Many modern high-dimensional data and interactions thereof can be modeled as lying approximat…
This paper explores how MoE layers improve deep learning performance.
problem Understanding the Mixture-of-Experts (MoE) layer in deep learning.
method Formal study of MoE layer's effectiveness and mechanism.
result MoE layer improves performance by leveraging cluster structure and non-linearity.
We consider multi-task regression models where observations are assumed to be a linear combination of several latent node and weight functions, all drawn from Gaussian process (GP) priors that allow nonzero covariance between grouped latent functions. We show that when these grouped functions are conditionally independ…