Transformers can predict pseudo-random sequences from LCGs with unseen parameters and moduli.
problem Learning pseudo-random number sequences from linear congruential generators with unknown parameters and moduli.
method Investigated the ability of Transformers to learn LCG sequences with varying complexity and moduli. Analyzed embedding layers and attention patterns.
result Transformers can predict pseudo-random sequences from LCGs with unseen parameters and moduli, up to mexttest=216, using a two-step strategy. Random number generators (RNGs) that are crucial for cryptographic applications have been the subject of adversarial attacks. These attacks exploit environmental information to predict generated random numbers that are supposed to be truly random and unpredictable. Though quantum random number generators (QRNGs) are ba…
Analyzes generalization error in generalized linear models, explaining double descent phenomenon.
problem Understanding generalization of machine learning models in high dimensions.
method Develops a framework to characterize asymptotic generalization error for generalized linear models.
result Rigorously explains the double descent phenomenon in generalized linear models.
Generative LLE modifies LLE to generate stochastic embeddings.
problem Nonlinear dimensionality reduction and manifold learning.
method Generative LLE modifies LLE by using stochastic linear reconstruction.
result Generative LLE can generate various LLE embeddings stochastically.
Let F_n denote the free group generated by n letters. The purpose of this article is to show that Hol(F_2), the holomorph of the free group on two generators, is linear. Consequently, any split group extension of F_2 by a linear group H is linear. This result gives a large linear subgroup of Aut(F_3). A second applicat…
Study one-dimensional topological theories with linear generating functions.
problem Understanding one-dimensional topological theories with defects.
method Construct bases of hom spaces for decorated unoriented one-dimensional cobordisms.
result Gram determinant and linear generating functions constructed.
Linear F-manifolds are studied with connections and dual spaces.
problem Understanding linear F-manifolds and their dual spaces.
method Developed systematic treatment and defined duality using connections.
result Defined compatibility conditions between linear F-manifolds and generalized tangent bundle.
New bounds for nearly-linear networks without training.
problem Generalization of neural networks close to linearity.
method Perturbation of linear networks to derive bounds.
result First non-vacuous bounds for neural nets.
Investigates linearity of group amalgams and new examples of non-linear groups.
problem Linearity of group amalgams and examples of non-linear groups.
method Investigates linearity of amalgams of subgroups of algebraic groups.
result Establishes linearity of certain 'doubles' of linear groups and finds new non-linear examples.
Paper introduces structured sparsity estimators for Generalized Linear Models.
problem Estimating structured sparsity in GLMs with debiased estimators.
method Extends Stucky and van de Geer's results to GLMs with structured sparsity.
result Proves oracle inequalities for structured sparsity estimators in GLMs.
The paper identifies generators of linear SDEs with noise types.
problem Identifying the generator of linear SDEs from their solution distribution.
method Deriving sufficient and necessary conditions for additive noise, and sufficient conditions for multiplicative noise.
result Generic conditions for identifying the generator of linear SDEs with both types of noise.
Extends linear MDP to handle nonlinear rewards.
problem Restrictive linear MDP assumption limits real-world applicability.
method Proposes Generalized Linear MDP (GLMDP) with GLMs for rewards.
result Develops offline RL algorithms achieving suboptimality guarantees.
Study on RCD(0,N) spaces with small linear diameter growth.
problem Understanding structure properties of RCD(0,N) spaces.
method Analyzing the (revised) fundamental group of RCD(0,N) spaces.
result Proved that the revised fundamental group is finitely generated for RCD(0,N) spaces with small linear diameter growth.
Neural networks with DAGs show linearity as width increases.
problem Understanding linearity in neural networks with arbitrary DAG structures.
method Analyzing the transition to linearity in networks with arbitrary DAGs, characterizing width by minimum in-degree.
result General neural networks with DAGs exhibit linearity as width approaches infinity.
New linear flows using exponential of linear transformations improve generative models.
problem Improving generative models in machine learning.
method Developed convolution exponentials and generalized Sylvester Flows using the exponential of linear transformations.
result Convolution exponentials and Convolutional Sylvester Flows outperform other models in log-likelihood.
StarNet trains deep models without gradients using linear equations.
problem Training deep generative models with gradients.
method Solving determined systems of linear equations.
result Least-square bounds for latent codes and model parameters.
In recent work on both generative and discriminative score to log-likelihood-ratio calibration, it was shown that linear transforms give good accuracy only for a limited range of operating points. Moreover, these methods required tailoring of the calibration training objective functions in order to target the desired r…
Paper analyzes agnostic learning of mixed linear regression without generative models.
problem Learning mixed linear regression without assuming stochastic generation.
method Expectation Maximization (EM) and Alternating Minimization (AM) algorithms.
result AM and EM algorithms converge to population loss minimizers under standard conditions.
Develops estimators for near-optimal linear regression under distribution shift.
problem Linear regression under distribution shift with scarce target domain data.
method Minimax linear risk estimators covering various transfer learning settings.
result Achieves near-optimal risk for linear regression problems under distribution shift.
New insights into optimization and generalization for linear models.
problem Understanding the implicit regularization of optimization methods for linear models.
method Investigating the norms minimized by interpolating solutions and using projections to move between solutions.
result Proving that for over-parameterized linear classification, projections onto the data-span enable the use of under-parameterized techniques.
New bounds for transfer learning in linear models, improving generalization.
problem Understanding when auxiliary data helps in improving generalization in linear models.
method Derivation of exact error bounds and optimal task weights for linear regression and linear neural networks.
result First non-vacuous sufficient conditions for beneficial auxiliary learning in linear neural networks.
Observing a linear superposition principle, a family of new minimal hypersurfaces in Euclidean space is found, as well as that linear combinations of generalized helicoids induce new algebraic minimal cones of arbitrarily high degree.
Study Loday algebroids, prove splitting theorem, and linearize problems.
problem Splitting and linearization of Loday algebroids.
method Local splitting-type results, Euler-like derivations.
result Established a general linearization principle.
The linear transports along paths in vector bundles introduced in Ref. [1] are applied to the special case of tensor bundles over a given differentiable manifold. Links with the transports along paths generated by derivations of tensor algebras are investigated. A possible generalization of the theory of geodesics is p…
Study on neural scaling laws for solving linear systems in-context.
problem Theoretical guarantees for solving linear systems using a linear transformer architecture.
method Neural scaling laws and task diversity for in-domain and out-of-domain generalization.
result Novel notion of task diversity for necessary and sufficient condition of generalization under task shifts.
NGSLL combines DNN accuracy with linear model interpretability.
problem Combining high accuracy of DNNs with interpretability of linear models.
method Neural generators of sparse local linear models (NGSLL) using DNNs to approximate non-linear functions.
result Effective in real-world datasets, achieving high predictive performance and interpretability.
Paper proposes linear transformers for efficient in-context learning without context length limitations.
problem Quadratic complexity of softmax transformers limits data processing speed.
method Investigates linear transformers under domain generalization, showing they learn mappings from context distributions to response functions.
result Linear transformers achieve in-context learning with a linear complexity in context length, offering a dimension-independent convergence rate.
BELIEF framework interprets GLMs using binary linear models.
problem Understanding and interpreting generalized linear models (GLMs) with binary outcomes.
method Developed a framework called binary expansion linear effect (BELIEF) to interpret GLMs through transparent linear models.
result BELIEF framework reveals perfect predictors in complete separation scenarios.
The paper characterizes compatible linear connections on 3D Finsler manifolds.
problem Characterizing compatible linear connections on Finsler manifolds of dimension three.
method Intrinsic method to characterize compatible linear connections, focusing on indicatrices and Euclidean symmetries.
result If a compatible linear connection is not unique, indicatrices must be Euclidean surfaces of revolution.
We construct analogues of FI-modules where the role of the symmetric group is played by the general linear groups and the symplectic groups over finite rings and prove basic structural properties such as Noetherianity. Applications include a proof of the Lannes--Schwartz Artinian conjecture in the generic representatio…
Linear models can grok without understanding, improving generalization.
problem Understanding the phenomenon of grokking in linear models.
method Analytical and numerical derivation of training and generalization dynamics in linear networks.
result Grokking can occur in linear networks without reaching understanding, and its timing depends on various parameters.
LOL method simplifies forming linear combinations of latent variables.
problem Lack of general-purpose methods for manipulating latent variables.
method Latent Optimal Linear combinations (LOL) method.
result LOL simplifies creation of expressive low-dimensional representations.
We prove that the semistability growth of hyperbolic groups is linear, which implies that hyperbolic groups which are sci (simply connected at infinity) have linear sci growth. Based on the linearity of the end-depth of finitely presented groups we show that the linear sci is preserved under amalgamated products over f…
This paper tackles efficient federated learning for generalized linear bandits.
problem Limited communication efficiency restricts existing federated learning solutions to linear models.
method Proposes a communication-efficient solution framework using online and offline regression.
result Proves sub-linear regret and communication cost for generalized linear bandits.
The paper provides a non-asymptotic error bound for linear system identification under nonlinear policies.
problem System identification for linear systems with nonlinear and/or time-varying policies under i.i.d. random excitation noises.
method Least square estimation with non-asymptotic error bound for bounded state and action trajectories.
result The error bound is consistent with linear policies and generalizes existing guarantees.
Motivated by drug design, we consider the best-arm identification problem in generalized linear bandits. More specifically, we assume each arm has a vector of covariates, there is an unknown vector of parameters that is common across the arms, and a generalized linear model captures the dependence of rewards on the cov…
We show that classical Wilczynski--Se-ashi invariants of linear systems of ordinary differential equations are generalized in a natural way to contact invariants of non-linear ODEs. We explore geometric structures associated with equations that have vanishing generalized Wilczynski invariants and establish relationship…
We show how the tangent functor extends from ordinary smooth maps to "microformal morphisms" (also called "thick morphisms") of supermanifolds. Microformal morphisms generalize ordinary maps and correspond to formal canonical relations between the cotangent bundles specified by generating functions depending on positio…
Proposes a method for differentially private linear regression and synthetic data generation.
problem Lack of valid inference and synthetic data generation methods for small-scale datasets in privacy-aware settings.
method Gaussian differentially private linear regression with bias-corrected estimator and SDG procedure.
result Improves accuracy and provides valid confidence intervals for downstream tasks.
The study classifies and characterizes totally symmetric sets in the general linear group.
problem Understanding the structure and properties of totally symmetric sets in the general linear group.
method Formulated a notion of irreducibility for totally symmetric sets in the general linear group and classified them.
result Classification of irreducible totally symmetric sets and those of maximal cardinality.
Holistic GLMs add constraints for better model quality.
problem Improving classical linear regression models.
method Sparsity-inducing, sign-coherence, and linear constraints.
result Holistic GLMs reliably solve GLMs for various responses.
We assume a vector bundle p:E→M with a general linear connection K and a classical linear connection $\Lam$ on M. We prove that all classical linear connections on the total space E naturally given by $(\Lam, K)$ form a 15-parameter family. Further we prove that all connections on J1E naturally given by…
In this paper we give an example of a linear group such that its tensor square is not linear. Also, we formulate some sufficient conditions for the linearity of non-abelian tensor products G⊗H and tensor squares G⊗G. Using these results we prove that tensor squares of some groups with one relation a…
Study on online regression with noise, achieving near-optimal regret bounds.
problem Online generalized linear regression with stochastic noise.
method Sharp analysis of FTRL algorithm for stochastic label noise.
result Achieved near-optimal regret bounds for O(σ2dlogT)+o(logT). Motivated by Pan-Yang [PY] and Ma-Cheng [MC], we study a general linear nonlocal curvature flow for convex closed plane curves and discuss the short time existence and asymptotic convergence behavior of the flow. Due to the linear structure of the flow, this partial differential equation problem can be resolved using a…
Characterizes problems solvable via linear convergence algorithms.
problem Optimization problems solvable with linear convergence.
method Riemannian gradient descent.
result Characterized problems solvable via linear convergence.
The Chern-Simons forms for R-linear connections on Lie algebroids are considered. A generalized Chern-Simons formula for such R-linear connections is obtained. We it apply to define Chern character and secondary characteristic classes for R-linear connections of Lie algebroids.
New classifier combines locally linear kernels for fast and accurate non-linear classification.
problem Developing a fast and accurate non-linear classifier.
method Combines locally linear classifiers using a ℓ1 Multiple Kernel Learning (MKL) problem with scalable MKL training for streaming kernels. result The resulting classifier achieves high accuracy with fast inference time.