BiTAT improves neural network quantization for edge devices by focusing on weight dependencies and disentangling them.
problem Performance degradation of compact neural networks under extreme quantization.
method Task-dependent Aggregated Transformation (BiTAT) method that orthonormalizes weights and progressively quantizes them.
result BiTAT effectively preserves model performance on ImageNet and CIFAR-100 with compact backbones.
New method estimates spatial weights matrix for lattice data, improving prediction accuracy.
problem Estimating spatial dependence structure for regular lattice data.
method Adaptive lasso with cross-sectional resampling to estimate sparse spatial weights matrix.
result Improves prediction accuracy of nitrogen dioxide concentrations.
New method clusters hypergraphs using weighted random walks and Laplacians.
problem Clustering hypergraph data with edge-dependent weights.
method Random walks with edge-dependent vertex weights, constructing hypergraph Laplacians for clustering.
result Proposed methods outperform existing hypergraph clustering algorithms.
Study proves no minimal surfaces can be contained in certain half-spaces or cones.
problem Prohibiting minimal surfaces from certain geometric configurations.
method Analyzes weighted minimal surfaces in R3 with height-dependent weights. result No proper surfaces can be contained in specific half-spaces or cones.
Reweighting improves risk bounds in certain data regions.
problem Improving risk bounds in classification and heteroscedastic regression.
method Weighted empirical risk minimization with a data-dependent weight function.
result A weighted ERM estimator can achieve superior performance in specific sub-regions.
CDST improves ensemble prediction by adjusting model weights based on covariates.
problem Improving ensemble prediction accuracy in complex scenarios.
method Covariate-dependent stacking (CDST) with flexible model weights estimated via cross-validation.
result CDST consistently outperforms conventional model averaging methods in complex datasets.
Robust feature-weighted jump models for time-dependent clustering
problem Temporal clustering
method Robust feature-weighted jump model
result Accurate recovery of true cluster sequence and feature identification
We improve prediction set coverage by assigning weights to individual sets.
problem Aggregating multiple prediction sets weakens overall coverage guarantee.
method Propose a framework for weighted aggregation of prediction sets.
result Achieve tighter coverage bounds that interpolate between 1−2α and 1−α guarantees. Marginal structural models (MSMs) estimate the causal effect of a time-varying treatment in the presence of time-dependent confounding via weighted regression. The standard approach of using inverse probability of treatment weighting (IPTW) can lead to high-variance estimates due to extreme weights and be sensitive to …
The paper calculates sensitivities for financial derivatives using path weighting methods.
problem Computing sensitivities for path-dependent financial derivatives with high variance and degeneracy issues.
method Proposes explicit path weighting formula, variance reduction adjustment, and covariance inflation technique.
result Effective methods to address high variance and degeneracy in sensitivities computation.
The paper analyzes how re-weighting helps in reducing variance in high-dimensional kernel methods under covariate shifts.
problem The challenge of high-dimensional kernel methods under covariate shifts and the role of re-weighting.
method Derives asymptotic expansion of high-dimensional kernels under covariate shifts, analyzes bias-variance decomposition, and characterizes the regularized kernel.
result Re-weighting helps in decreasing variance and can be seen as a data-dependent regularization.
The study analyzes how neural reward models learn features for policy optimization in a Gaussian single-index model.
problem Reward modeling in policy optimization and its impact on downstream value.
method Two-stage neural reward model: first learns hidden direction, then fits readout layer.
result For any feature-learning temperature above a dimension-free threshold, a constant fraction of neurons recover the hidden direction.
Hypergraphs are used in machine learning to model higher-order relationships in data. While spectral methods for graphs are well-established, spectral theory for hypergraphs remains an active area of research. In this paper, we use random walks to develop a spectral theory for hypergraphs with edge-dependent vertex wei…
We empirically investigate the best trade-off between sparse and uniformly-weighted multiple kernel learning (MKL) using the elastic-net regularization on real and simulated datasets. We find that the best trade-off parameter depends not only on the sparsity of the true kernel-weight spectrum but also on the linear dep…
Improved GRU model with weighted time-delay feedback for long-term dependencies.
problem Modeling long-term dependencies in sequential data.
method Introducing a gated recurrent unit (GRU) with a weighted time-delay feedback mechanism.
result τ-GRU outperforms state-of-the-art models on various tasks.
This paper introduces hierarchical Gaussian process priors for neural networks to capture weight correlations and inductive biases.
problem Capturing weight correlations and inductive biases in neural networks.
method Hierarchical Gaussian process priors with unit embeddings and input-dependent kernels.
result Hierarchical Gaussian process priors provide competitive predictive performance and desirable uncertainty estimates.
We prove structure theorems for complete manifolds satisfying both the Ricci curvature lower bound and the weighted Poincaré inequality. In the process, a sharp decay estimate for the minimal positive Green's function is obtained. This estimate only depends on the weight function of the Poincaré inequality, and yields …
Test-asset construction affects factor model performance.
problem How test assets are constructed impacts factor model performance.
method Forming characteristic-unsorted random portfolios and varying stock selection, initial weighting, holding, and rebalancing.
result Test-asset construction shifts factor model rankings materially.
This paper presents a general framework for norm-based capacity control for Lp,q weight normalized deep neural networks. We establish the upper bound on the Rademacher complexities of this family. With an Lp,q normalization where q≤p∗, and 1/p+1/p∗=1, we discuss properties of a width-independent ca…
The study examines how weight sharing, equivariance, and locality affect the sample complexity of neural networks.
problem Understanding the impact of design choices on the generalization error of neural networks.
method Statistical learning theory applied to single hidden layer networks with weight sharing, equivariance, and locality.
result Lower and upper bounds for sample complexity are derived, showing that locality has benefits but comes with a trade-off.
New knot invariants from biquandle arrow weights.
problem Creating nontrivial knot invariants.
method Introducing biquandle arrow weights for knot diagrams.
result Enhancements are nontrivial and not determined by homset cardinality.
Bayesian neural networks with dependent weights converge to Gaussian mixtures.
problem Limitations of standard Gaussian priors in neural networks.
method Posterior analysis with Gaussian likelihood for networks with dependent weights.
result Posterior distribution identified in the wide-width limit, ensuring invertibility of random covariance matrix.
Corrects distribution shift in target shift scenarios using importance weighting.
problem Analyzes importance weighting for correcting distribution shift under target shift.
method Analyzed importance-weighted kernel ridge regression under target shift.
result Shows that importance weighting corrects the train-test mismatch without altering input-space complexity.
Paper explores exact recovery of communities in weighted graphs using Gaussian and exponential distributions.
problem Exact recovery of communities in weighted graphs with Gaussian and exponential distributions.
method Introduces a new semi-metric to describe conditions for exact recovery and analyzes conditions for both complete and incomplete graphs.
result Necessary and sufficient conditions for exact recovery are asymptotically tight and applicable to both complete and incomplete graphs.
Investigates portfolio selection for rank-dependent utilities in incomplete markets.
problem Portfolio selection for agents with rank-dependent utility in incomplete financial markets.
method Characterizes deterministic strict equilibrium strategies for constant-coefficient and time-invariant probability weighting functions. Addresses the issue of selecting an optimal strategy from multiple equilibrium strategies for time-variant probability weighting functions.
result Characterizes deterministic strict equilibrium strategies and identifies optimal strategies from multiple equilibrium strategies.
AGBoost uses attention weights to improve GBM for regression problems.
problem Improving gradient boosting machine for regression tasks.
method Attention-based modification of GBM with trainable attention weights.
result AGBoost achieves better performance on regression datasets.
This paper improves risk control for financial markets by calibrating VaR forecasts using conformal methods.
problem Nonstationary and regime-dependent losses in financial markets.
method Regime-weighted conformal risk control (RWC) for VaR forecasting.
result RWC improves regime-conditional stability in some settings with modest conservativeness changes.
In this paper we study convex stochastic search problems where a noisy objective function value is observed after a decision is made. There are many stochastic search problems whose behavior depends on an exogenous state variable which affects the shape of the objective function. Currently, there is no general purpose …
Enhanced Neural ODEs outperform traditional models in image classification and video prediction.
problem Efficiently modeling time-varying dynamics in neural networks.
method Proposed a novel family of non-autonomous Neural ODEs with time-varying weights.
result Outperformed previous Neural ODE variants in speed and representational capacity.
Improves Gower's similarity for mixed-type variables with automatic weighting.
problem Handling missing values and unbalanced variable contributions in Gower's similarity for mixed-type data.
method Automatic weighting scheme minimizing differences in correlation between contributing dissimilarities and weighted Gower's dissimilarity.
result Improved performance in classification and imputation of missing values.
The paper provides an efficient method to price path-dependent derivatives using multiscale stochastic volatility models.
problem Pricing path-dependent derivatives under multiscale stochastic volatility models.
method Derives a Malliavin representation for the first-order approximation of the price of path-dependent derivatives.
result An efficient Monte Carlo approximation for pricing path-dependent derivatives is derived.
Unified framework for measuring concentration in weighted networks considering both weight distributions and network structure.
problem Traditional indices neglect the topology of relationships among network elements.
method Develops a family of topology-aware concentration indices that jointly account for weight distributions and network structure.
result The proposed indices preserve key properties and allow concentration to be evaluated across different dimensions of dependence.
In many problem settings, parameter vectors are not merely sparse but dependent in such a way that non-zero coefficients tend to cluster together. We refer to this form of dependency as "region sparsity." Classical sparse regression methods, such as the lasso and automatic relevance determination (ARD), which model par…
Study of deep neural networks with dependent weights leading to new model limits and properties.
problem Characterizing deep neural networks with dependent weights in the infinite-width limit.
method Modeling weights as a mixture of Gaussian distributions and analyzing the infinite-width limit.
result Characterization of neural network layers by scalar parameters and Lévy measures, leading to new model limits.
Unexpectedly, weighted Pareto variables are stochastically dominant.
problem Understanding stochastic dominance in Pareto distributions.
method Analyzing weighted averages of Pareto random variables with infinite mean.
result The weighted average of Pareto variables is stochastically dominant.
Long short-term memory (LSTM) is normally used in recurrent neural network (RNN) as basic recurrent unit. However,conventional LSTM assumes that the state at current time step depends on previous time step. This assumption constraints the time dependency modeling capability. In this study, we propose a new variation of…
We propose an iterative scheme for feature-based positioning using a new weighted dissimilarity measure with the goal of reducing the impact of large errors among the measured or modeled features. The weights are computed from the location-dependent standard deviations of the features and stored as part of the referenc…
Investment strategies for rank-dependent utility agents are derived in a continuous-time market.
problem Time inconsistency in rank-dependent utility models.
method Study of consistent planners seeking intra-personal equilibrium strategies.
result Explicit final wealth profile replicating equilibrium strategies, with scaling function derived.
A new method for deep learning under distribution shift by iteratively refining importance weighting.
problem Handling distribution shift in deep learning models when training and test data distributions differ.
method Dynamic Importance Weighting (dynamic IW) that iterates between weight estimation and weighted classification, using a pre-trained feature extractor and stochastic optimization.
result Dynamic IW outperforms state-of-the-art methods in experiments with various types of distribution shift on multiple datasets.
T-BFA targets and misleads specific DNN inputs to a chosen output.
problem Targeted attack on DNN weight parameters to hijack function.
method Identifies critical weight bits, ranks them by class dependence, and flips them to mislead inputs.
result Successfully misclassifies images from 'Hen' to 'Goose' class with 100% success rate, maintaining 59.35% validation accuracy.
Recurrent neural networks (RNNs) are notoriously difficult to train. When the eigenvalues of the hidden to hidden weight matrix deviate from absolute value 1, optimization becomes difficult due to the well studied issue of vanishing and exploding gradients, especially when trying to learn long-term dependencies. To cir…
The paper analyzes methods for estimating linear functionals from observational data, proving upper bounds and showing optimal procedures.
problem Estimating linear functionals from observational data in causal inference and bandit literature.
method Two-stage procedures that first estimate treatment effect function, then use it to estimate the linear functional.
result Proves non-asymptotic upper bounds on mean-squared error for two-stage procedures and shows instance-dependent optimality.
Differentially private method for estimating individualized treatment rules.
problem Estimating individualized treatment rules while preserving privacy.
method Differentially private two-stage empirical risk minimization (DP-2ERM).
result Improved privacy-utility trade-off demonstrated through simulations and applications.
Neural networks with integer weights approximate continuous functions efficiently.
problem Approximating continuous functions using neural networks with integer weights.
method Integrates superexpressive activation functions and integer weights.
result Convergence rate of order n2β+d−2βlog2n for neural network regression. New method clusters matrix-valued data by latent variables.
problem Clustering matrix-valued data with hidden structure.
method Latent variable model with hierarchical clustering.
result Algorithm attains clustering consistency in high dimensions.
Study on learning rates in neural networks of varying depth.
problem Dependence of maximal update learning rate on network depth.
method Analysis of random fully connected ReLU networks with mean-field weight initialization.
result Maximal update learning rate scales like L−3/2 with network depth. Sleep-based regularization stabilizes STDP in recurrent neural networks.
problem Pathological weight dynamics in recurrent SNNs.
method Periodic offline phases with stochastic decay and spontaneous activity.
result Sleep-based renormalization prevents weight saturation and preserves learned structure.
Proposes a weighted conformal approach for cluster label uncertainty.
problem Cluster label uncertainty in unlabeled data.
method Develops a conformal inference algorithm to correct label mismatch.
result Improves confidence set size in nonlinear and high-dimensional clustering.