A new method compresses deep neural networks by predicting and quantizing weights between layers.
problem Resource constraints in deep neural networks.
method Inter-Layer Weight Prediction (ILWP) and quantization based on Smoothly Varying Weight Hypothesis (SVWH).
result The method achieves higher weight compression rates at the same accuracy level.
Study finds varying market efficiency in prewar and wartime Japanese stock market.
problem Measuring market efficiency in prewar and wartime Japanese stock market.
method Using a new market capitalization-weighted stock price index, the study examines market efficiency over time and historical events.
result The adaptive market hypothesis is supported in the prewar and wartime Japanese stock market, with efficiency varying over time and with historical events.
Suppose some classifiers are selected from a set of hypothesis classifiers to form an equally-weighted ensemble that selects a member classifier at random for each input example. Then the ensemble has an error bound consisting of the average error bound for the member classifiers, a term for selectivity that varies fro…
A new algorithm for non-stationary linear bandits with improved regret bound.
problem Non-stationary linear bandit problem with time-varying rewards.
method D-LinUCB, a discounted linear regression algorithm with exponential weights.
result Upper bound on dynamic regret of order d^{2/3} B_T^{1/3}T^{2/3}, optimal in slowly-varying and abruptly-changing environments.
We propose a novel stacked generalization (stacking) method as a dynamic ensemble technique using a pool of heterogeneous classifiers for node label classification on networks. The proposed method assigns component models a set of functional coefficients, which can vary smoothly with certain topological features of a n…
Paper develops sparse learning for heavy-tailed time series with locally stationary dynamics.
problem Sparse learning for high-dimensional heavy-tailed locally stationary time series.
method Additive modeling with kernel smoothing, sparsity-inducing penalized estimation.
result Prediction-error bounds and convergence rates for different sparsity structures.
ACORE improves hypothesis testing and confidence sets in likelihood-free inference.
problem Constructing hypothesis tests and confidence sets in likelihood-free inference settings.
method Formulates classical LRT as a classification problem, uses machine learning to improve estimates.
result Demonstrates improved accuracy in hypothesis testing and confidence sets.
New insights into why sparse networks perform well, including Supermasks.
problem Understanding why sparse networks trained from scratch perform better than non-sparse models.
method Analyzed three critical components of the Lottery Ticket algorithm: zeroing weights, signs, and masking.
result Discovered Supermasks that can improve performance of untrained networks.
Method measures weight similarity in neural networks using normalization and statistical inference.
problem Quantifying weight similarity in non-convex neural networks.
method Chain normalization rule and hypothesis-training-testing statistical inference.
result Weights of identical neural networks converge to similar local solutions.
New model predicts energy prices volatility by smoothing time variation and persistence.
problem Separate study of volatility's time variation and persistence.
method Dynamic persistence model that allows shocks with heterogeneous persistence to vary smoothly over time.
result Significantly improves volatility forecasts over state-of-the-art models.
In this paper, we study the benefits of using polyharmonic splines and node layouts with smoothly varying density for developing robust and efficient radial basis function generated finite difference (RBF-FD) methods for pricing of financial derivatives. We present a significantly improved RBF-FD scheme and successfull…
We construct a smooth codimension-one foliation on the five-sphere in which every leaf is a symplectic four-manifold and such that the symplectic structure varies smoothly. Our construction implies the existence of a complete regular Poisson structure on the five-sphere.
Scaling laws found for reinforcement learning performance with model size and compute.
problem Challenges in extending generative modeling scaling laws to reinforcement learning.
method Introduced intrinsic performance as a monotonic function of mean episode return.
result Intrinsic performance scales as a power law in model size and environment interactions.
Using an obstruction based on Donaldson's theorem, we derive strong restrictions on when a Seifert fibered space Y=F(e;q1p1,…,qkpk) over an orientable base surface F can smoothly embed in S4. This allows us to classify precisely when Y smoothly embeds provided e>k/2, where $…
Estimates RL data for dynamic treatment effects using GMM.
problem Estimating dynamic treatment effects from RL data with nonstationary behavior policies.
method Weighted GMM approach to stabilize variance in adaptive RL settings.
result Valid hypothesis testing and confidence regions for dynamic treatment effects.
Study shows cryptocurrency market efficiency changes over time.
problem Measuring cryptocurrency market efficiency over time.
method Used a generalized least squares-based time-varying model to measure efficiency without sample size dependence.
result Bitcoin's market efficiency is higher than Ethereum's over most periods.
SEF-M improves spiking neural network classification accuracy by 14%.
problem Improving spiking neural network classification accuracy.
method Meta-neuron based learning algorithm with time-varying weight model.
result Time-varying weight model improves classification accuracy by 14%.
New framework detects time-varying economic persistence.
problem Time-varying persistence in economic shocks.
method Localized regression techniques to identify evolving heterogeneity.
result Substantial persistence variations align with macroeconomic events.
Paper tackles hypothesis transfer learning for black-box models.
problem Difficult to build universal machine learning models across different institutions.
method Dynamic Knowledge Distillation (dkdHTL) with instance-wise weighting.
result Empirical results show the effectiveness of dkdHTL.
This study examines the adaptive market hypothesis (AMH) in Japanese stock markets (TOPIX and TSE2). In particular, we measure the degree of market efficiency by using a time-varying model approach. The empirical results show that (1) the degree of market efficiency changes over time in the two markets, (2) the level o…
Smoothly conjugate Anosov flows on 3D manifolds are actually smoothly conjugate.
problem Smoothly conjugate 3D Anosov flows are not always smoothly conjugate.
method Proved smooth rigidity for volume preserving Anosov flows on 3-manifolds.
result Smooth conjugacy implies smooth conjugacy for volume preserving Anosov flows.
We study deformations of free boundary constant mean curvature (CMC) hypersurfaces whose Jacobi operator is degenerate due to symmetries of the ambient space. The value of the mean curvature and the ambient metric are allowed to vary simultaneously, provided that the infinitesimal ambient symmetries change smoothly. We…
Study excess capacity in neural networks using Rademacher complexity.
problem Understanding how much capacity deep networks have beyond what's needed for classification.
method Unified Rademacher complexity bounds for function composition and convolutional layers, considering Lipschitz constants and initialization norms.
result There is substantial excess capacity per task, and capacity can be kept similar across different tasks.
Extends double linear policy with time-varying weights and proves robust positive expectation.
problem Ensuring robustness in policy optimization with time-varying parameters.
method Employed a novel elementary symmetric polynomials characterization approach to prove robust positive expectation (RPE). Derived explicit expressions for expected cumulative gain-loss and variance.
result Proved the robust positive expectation property holds for the extended double linear policy.
We investigate the ramifications of the Legendrian satellite construction on the relation of Lagrangian cobordism between Legendrian knots. Under a simple hypothesis, we construct a Lagrangian concordance between two Legendrian satellites by stacking up a sequence of elementary cobordisms. This construction narrows the…
In this article, the long-term behavior of the stock market index of the New York Stock Exchange is studied, for the period 1950 to 2013. Specifically, the CRSP Value-Weighted and CRSP Equal-Weighted index are analyzed in terms of market efficiency, using the standard ratio variance test, considering over 1600 one week…
Paper tests for time-varying entropy in stock prices, finding periods of inefficiency.
problem Testing for time-varying entropy in stock price dynamics.
method Unbiased approximation of Shannon entropy variance, optimal rolling window selection, hypothesis testing.
result Existence of periods of market inefficiency for meme stocks.
Motivated by the definition of the smooth manifold structure on a suitable mapping space, we consider the general problem of how to transfer local properties from a smooth space to an associated mapping space. This leads to the notion of smoothly local properties. In realising the definition of a local property at a pa…
In this work, we consider hypothesis testing and anomaly detection on datasets where each observation is a weighted network. Examples of such data include brain connectivity networks from fMRI flow data, or word co-occurrence counts for populations of individuals. Current approaches to hypothesis testing for weighted n…
Restricted Boltzmann Machines are described by the Gibbs measure of a bipartite spin glass, which in turn corresponds to the one of a generalised Hopfield network. This equivalence allows us to characterise the state of these systems in terms of retrieval capabilities, both at low and high load. We study the paramagnet…
SATL adapts to varying smoothness in hypothesis transfer learning.
problem Fixed kernel regularization fails in varying smoothness settings.
method Proposes SATL, a two-phase KRR algorithm with adaptive Gaussian kernels.
result SATL achieves minimax optimality with matching upper and lower bounds.
Improved hypothesis testing and change-point detection using diffusion-based methods.
problem Limited power of score-based hypothesis tests and change-point detection.
method Extending score-based Fisher divergence to diffusion-divergence by multiplying score functions with a matrix-valued function or weight matrix.
result Theoretical quantification and demonstration of optimal performance of diffusion-based algorithms.
Randomly initialized networks contain subnetworks that perform similarly to target networks.
problem Proving the lottery ticket hypothesis for neural networks.
method Pruning over-parameterized neural networks to find subnetworks.
result Randomly initialized networks contain subnetworks with similar performance to target networks without additional training.
Paper proposes an EKF for estimating time-varying market efficiency.
problem Estimating time-varying market efficiency under nonlinear dynamics.
method Extended Kalman Filter (EKF) for time-varying autoregressive models.
result U.S. market generally remained weak-form efficient since mid-1946.
Residuals improve deep neural networks without increasing hypothesis complexity.
problem Understanding how residual connections affect hypothesis complexity and generalization.
method Analyzing the covering number of the hypothesis space and deriving a margin-based generalization bound.
result Residual connections do not increase the hypothesis complexity of neural networks.
A new approach optimizes weights in DLP for better risk-adjusted performance.
problem Optimizing time-varying weights in Double Linear Policy (DLP) for better risk-adjusted performance.
method Stochastic Model Predictive Control (SMPC) framework to maximize risk-adjusted returns while enforcing constraints.
result Empirical results show improved risk-adjusted performance and drawdown control.
Develops a Bayesian non-parametric approach for signal separation with varying components.
problem Signal separation with varying components across different input locations.
method Augments Gaussian Process Latent Variable Models with weighted sums of pure component signals and incorporates priors for linear weights.
result Framework allows for non-linear variations in signals and incorporates useful priors for linear weights.
Proposes DSW for unbiased ITE estimation with dynamic confounders.
problem Estimating ITE from dynamic observational data with time-varying confounders.
method Deep Sequential Weighting (DSW) infers hidden confounders using current treatment assignments and historical information.
result DSW generates unbiased and accurate treatment effects.
Novel method models dynamic brain graphs from time series data.
problem Generating hypotheses for dynamic brain states.
method Conditionally weighted superposition of static graphs.
result Improves f1-scores by 22-28% on average over baselines.
Geodesic flows between hypersurfaces in Euclidean spaces using Lorentzian geometry.
problem Interpolation between hypersurfaces in Euclidean spaces.
method Lorentzian geodesic flow between tangent spaces of hypersurfaces.
result Geodesic flow is preserved by rigid transformations and homotheties.
A new method trains deep networks by separating weight locations from values.
problem Training deep networks efficiently and effectively.
method Lookahead Permutation (LaPerm) to train DNNs by reconnecting weights.
result LaPerm can train DNNs with random and dense, sparse, or single-valued initial weights.
The lottery ticket hypothesis finds multiple winning sub-networks in neural networks.
problem Finding a single winning sub-network in neural networks.
method Analyzing neural networks trained in isolation and on different tasks.
result Neural networks contain multiple sub-networks that match the accuracy of the original network, not just one.
The paper refutes the manifold hypothesis for image data and proposes the union of manifolds hypothesis.
problem The manifold hypothesis fails to capture the structure of image data.
method Empirical verification of the union of manifolds hypothesis on image datasets.
result Image data lies on a disconnected set with varying intrinsic dimensions.
The paper explores how the Gauss curvature of Riemannian surfaces can be represented as the divergence of a vector field.
problem Representing the Gauss curvature of Riemannian surfaces as the divergence of a vector field.
method Investigates the existence of a metric linear connection of zero curvature and its role in differential geometry.
result Provides conditions under which a Riemannian surface can be considered a generalized Berwald surface.
The paper tests if LLMs' capabilities are executed by small subnetworks (circuits).
problem Understanding how LLMs execute their capabilities.
method Formalized criteria for circuits, developed hypothesis tests, applied to six circuits.
result Synthetic circuits align with idealized properties, while Transformer circuits vary in their alignment.
Analyzes the complexity of linear hypothesis sets using Rademacher complexity.
problem Understanding the complexity of linear hypothesis sets for various norms.
method Tight analysis of empirical Rademacher complexity for linear hypothesis classes with bounded weights.
result Improved bounds on Rademacher complexity for linear hypothesis sets, matching or improving existing results.
New method relaxes spatial invariance in locally connected layers, improving accuracy.
problem Improving classification accuracy with locally connected layers.
method Designing a low-rank locally connected layer with varying spatially varying combining weights.
result Relaxing spatial invariance improves classification accuracy over convolution and locally connected layers.
In this paper we address the problem of pool based active learning, and provide an algorithm, called UPAL, that works by minimizing the unbiased estimator of the risk of a hypothesis in a given hypothesis space. For the space of linear classifiers and the squared loss we show that UPAL is equivalent to an exponentially…