Algorithm analyzes trading profits from size factor using equal-weighted portfolios.
problem Estimating trading profits from systematic rebalancing contributions.
method Uses INTECH's algorithm on equal-weighted portfolios combining size factor exposure.
result Natural test subject for Stochastic Portfolio Theory.
Dropout improves matrix factorization by controlling factor size.
problem Understanding regularization properties of dropout for matrix factorization.
method Theoretical analysis of dropout's equivalence to a deterministic model with adaptive dropout rates.
result Dropout's regularization effect is limited by the fixed dropout rate, suggesting adaptive rates.
Introduces human capital into asset pricing model for better return prediction.
problem Improving asset pricing models to better predict stock returns.
method Used OLS and IVGMM to estimate six-factor model parameters from four sets of portfolios.
result Human capital component shares predictive power with other factors in explaining stock returns.
Investment strategy depends on many factors for venture capital funds.
problem Finding the optimal portfolio size for venture capital funds.
method Analyzes various factors affecting fund returns and optimal portfolio size, starting with basic assumptions and increasing complexity.
result Investment strategy depends on many factors, not a one-size-fits-all formula.
Gradient descent with large steps leads to chaotic parameter space and unpredictable outcomes.
problem Understanding the behavior of gradient descent with large step sizes in matrix factorization.
method Analyzing the fractal structure of the parameter space and deriving critical step sizes for convergence.
result Gradient descent with large steps exhibits chaotic behavior and sensitivity to initialization, creating a fractal boundary between converging and diverging minimizers.
This paper investigates factors influencing SGD minima.
problem Understanding the factors that influence the minima found by SGD.
method Examined learning rate, batch size, Hessian, and gradient covariance; used stochastic differential equations to model SGD.
result The ratio of batch size to learning rate is a main factor in SGD dynamics.
We propose a fast algorithm for computing the expected tranche loss in the Gaussian factor model. We test it on portfolios ranging in size from 25 (the size of DJ iTraxx Australia) to 100 (the size of DJCDX.NA.HY) with a single factor Gaussian model and show that the algorithm gives accurate results. The algorithm prop…
We add size factor to CAPM and normalize residuals by Volatility Index.
problem Capturing the size effect in CAPM and making residuals Gaussian.
method Insert size effect, normalize residuals by Volatility Index, and fit model to real-world data.
result The new model shows long-term stability and connects to Stochastic Portfolio Theory.
SGD minima influenced by learning rate, batch size, and gradient covariance.
problem Characterizing the relation between learning rate, batch size, and the properties of SGD minima.
method Approximated SGD by SDE to investigate learning rate, batch size, and gradient covariance effects.
result The ratio of learning rate to batch size is a key determinant of SGD dynamics and minima width, leading to better generalization.
Study on adversarial examples from data size, task, and model factors.
problem Understanding adversarial examples from data size, task, and model perspectives.
method Systematic study on adversarial examples from three aspects: data size, task-dependent, and model-specific factors.
result Adversarial generalization requires more data than standard generalization.
We speed up marginal inference by ignoring factors that do not significantly contribute to overall accuracy. In order to pick a suitable subset of factors to ignore, we propose three schemes: minimizing the number of model factors under a bound on the KL divergence between pruned and full models; minimizing the KL dive…
Enhanced Dantzig selector reduces recovery error in high dimensions.
problem Improving recovery accuracy in ultra-high dimensional settings.
method Constrained Dantzig selector with sequential linear programming.
result Achieves convergence rates within a logarithmic factor of the sample size of oracle rates.
This paper diagnoses factor-model pricing errors using a new method.
problem Measuring pricing errors in factor models with general characteristic axes.
method Developed a method to measure factor-model pricing errors as bridge-alpha curves, using a predetermined characteristic order and prefix portfolios.
result Adding a counterpart factor flips the curve's sign on every axis, but only HML and CMA overcorrect enough to be rejected.
Paper examines M-SVM for multi-task learning, showing reliability and pre-convergence-rate factor improvements.
problem Whether MTL always provides reliable results and how MTL outperforms independent learning.
method Regularized multi-task learning (MTL) based on SVM models (M-SVM).
result M-SVM is Bayes risk consistent in large sample size, improving pre-convergence-rate factor (PCR) for small data.
EigenDamage reduces neural network size and FLOPs with structured pruning in the Kronecker-Factored Eigenbasis.
problem Reducing neural network size and FLOPs while maintaining accuracy for resource-constrained devices.
method Kronecker-Factored Eigenbasis reparameterization and Hessian-based structured pruning.
result Empirically validated improvements in model size and FLOPs with negligible accuracy loss.
Gradient descent on matrix factorization leads to minimum nuclear norm solution.
problem Optimizing underdetermined quadratic objectives over matrices.
method Gradient descent on a full dimensional factorization of the matrix.
result Gradient descent converges to the minimum nuclear norm solution.
Study finds Value Granger-causes Size during crisis regimes but not during normal times.
problem Understanding regime-dependent predictive relationships between equity factors.
method Used 35 years of Fama-French data and a Student-t Hidden Markov Model (HMM) to identify crisis regimes.
result Value Granger-causes Size during crisis regimes but not during normal times, validating across multiple historical events.
Deep learning speeds up IFA estimation for large datasets.
problem Slow MML estimation for large-scale IFA models.
method Importance-weighted autoencoder (IWAE) for fast VI.
result IWAE yields accurate estimates faster than MH-RM.
Derives a size premium from automated market makers in decentralized AI subnets.
problem Determining the profitability and risk of decentralized AI subnets.
method Analyzes daily data on 128 subnets, tests the size premium, and calculates transaction costs.
result The size premium is reduced by a halving of token emissions but remains profitable only below a certain asset threshold.
The paper examines how gradient descent stabilizes low-rank matrix factorization in noisy conditions.
problem Stability of low-rank implicit regularization in perturbed deep matrix factorization.
method Derives spectral conditions for gradient descent to exhibit a low-rank phase in noiseless settings and analyzes perturbed dynamics.
result Gradient descent converges to a low-rank solution under perturbation, with explicit dependence on perturbation size.
Study market-to-book ratios using Stochastic Portfolio Theory.
problem Identify the value factor in stock returns.
method Develop functionally generated portfolios using book values and analyze their relative returns.
result The value factor (market-to-book ratio) affects portfolio performance.
Analyzes factors affecting flow VI performance.
problem Consistent performance of flow VI across studies.
method Step-by-step analysis of capacity, objectives, batchsize, estimators, and step-sizes.
result Specific recommendations and a flow VI recipe.
The common wisdom argues that, in general, large trades cause large price changes, while small trades cause small price changes. However, for extremely large price changes, the trade size and news play a minor role, while the liquidity (especially price gaps on the limit order book) is a more influencing factor. Hence,…
We propose a 4-factor model for overnight returns and give explicit definitions of our 4 factors. Long horizon fundamental factors such as value and growth lack predictive power for overnight (or similar short horizon) returns and are not included. All 4 factors are constructed based on intraday price and volume data a…
Faster convergence and handling larger mini-batches for deep neural networks.
problem Generalization gap in large-scale distributed training of deep neural networks.
method Second-order optimization using Kronecker-factored approximate curvature.
result Achieved 75% Top-1 validation accuracy with mini-batch size of 131,072 in 978 iterations.
DeepThin compresses deep neural networks, improving performance and reducing resource usage.
problem Efficiently compressing large neural networks for mobile devices.
method Combining rank factorization with a reshaping process to add nonlinearity.
result DeepThin achieves significant improvements in word error rates and test loss compared to existing methods.
Efficient NTF algorithm for large sparse tensors.
problem Sparse multi-dimensional data and limitations of existing NTF algorithms.
method Saturating Coordinate Descent with element selection based on Lipschitz continuity.
result Proposes a scalable NTF algorithm for large tensors.
A new criterion HBIC improves model selection for factor analysis with missing data.
problem Model selection for factor analysis with incomplete data.
method Proposes a novel criterion HBIC that uses actual observed information in the penalty term.
result HBIC is more accurate than BIC when missing data rates are high.
A method to reduce knowledge graph embedding models by binarizing parameters.
problem Large memory requirements for tensor factorization models in knowledge graph completion.
method Introducing a quantization function to binarize parameters of CP tensor decomposition.
result Successfully reduced model size by more than an order of magnitude while maintaining task performance.
Pruning improves model generalization in over-parameterized models, contradicting traditional theories.
problem Pruning's effect on generalization in over-parameterized models.
method Empirical study on standard pruning algorithms and additional regularization effects.
result Pruning leads to better training and regularization, improving generalization.
The study addresses overlooked data-generating processes in time-series asset pricing.
problem The literature on time-series asset pricing overlooks the data-generating processes for factors expressed in return differences.
method The study proposes a new definition of returns and compound returns for factors, and uses OLS with net returns for single-index models.
result OLS with net returns for single-index models leads to inflated alphas, exaggerated t-values, and overestimated Sharpe ratios.
Audit fees change based on company and economic factors during auditor switching.
problem Understanding how audit fees change when auditors switch firms.
method Examined the impact of auditor switching on audit fees, considering company characteristics and economic data.
result The direction and magnitude of audit fee changes during switching depend on economic stability and company characteristics.
This paper compares two stock factor models in China's A-share market.
problem Contradicting results in existing research on stock factor models.
method Empirical analysis using China's A-share data from 2005-2020, orthogonalizing redundant factors, and 25-group portfolio returns calculation.
result The five-factor model outperforms the three-factor model in explaining excess return rates.
In the standard equilibrium and/or arbitrage pricing framework, the value of any asset is uniquely specified from the belief that only the systematic risks need to be remunerated by the market. Here, we show that, even for arbitrary large economies when the distribution of the capitalization of firms is sufficiently he…
Gradient descent balances layer magnitudes in deep neural networks without explicit regularization.
problem Balancing magnitudes across layers in deep neural networks.
method Gradient descent with infinitesimal step size enforces layer magnitude balance.
result Gradient descent automatically balances layer magnitudes without explicit regularization.
Motivated by an application in computational biology, we consider low-rank matrix factorization with {0,1}-constraints on one of the factors and optionally convex constraints on the second one. In addition to the non-convexity shared with other matrix factorization schemes, our problem is further complicated by a c…
Paper proposes a new deflation varimax method for vintage factor analysis.
problem Finding a scientifically meaningful low-dimensional representation of data.
method Deflation varimax procedure for orthogonal matrix rotation.
result The proposed method achieves minimax optimal factor loading estimation.
Graph clustering involves the task of dividing nodes into clusters, so that the edge density is higher within clusters as opposed to across clusters. A natural, classic and popular statistical setting for evaluating solutions to this problem is the stochastic block model, also referred to as the planted partition model…
We propose a fast algorithm for computing the expected tranche loss in the Gaussian factor model. We test it on a 125 name portfolio with a single factor Gaussian model and show that the algorithm gives accurate results. We choose a 125 name portfolio for our tests because this is the size of the standard DJCDX.NA.HY p…
AdaBatch dynamically adjusts batch size during training for deep learning models.
problem Choosing optimal batch size for deep neural networks.
method Adaptive batch size adjustment during training.
result Adaptive batch sizes improve performance by up to 6.25x on 4 GPUs with minimal accuracy loss.
Improved Naive Bayes for text classification with small datasets.
problem Poor performance of Naive Bayes in small training datasets.
method Introducing a correlation factor to Naive Bayes estimator.
result Our method achieves better accuracy than traditional Naive Bayes.
A new bootstrapping method reduces key sizes and runtime in FHE.
problem Large plaintext evaluation in FHE increases bootstrapping complexity.
method New polynomial vector representation and monic monomial permutation matrices.
result Polynomial factor improvement in key size and constant factor in runtime.
This paper tackles multi-asset market making by reducing dimensionality and considering different transaction sizes.
problem Optimizing bid and ask prices for multiple assets while managing inventory risk in volatile markets.
method Proposes a dimensionality reduction technique using a factor model and considers different transaction sizes.
result Generalizes existing market making models by incorporating different transaction sizes and prices.
Compression method reduces word embedding size for NLP models.
problem Memory constraints in deploying deep learning models for NLP tasks.
method Low rank matrix factorization during training to compress word embeddings.
result 90% compression with minimal accuracy loss for sentence classification tasks.
Improved robustness in optimization methods using second-order information.
problem Scalability and sensitivity to mini-batch size in optimization methods.
method Mini-Batch Stochastic Variance-Reduced Newton (extttMb−SVRN) algorithm incorporating partial second-order information. result Achieves a fast linear convergence rate independent of mini-batch size for large data sizes.
Paper explores subdifferential chain rules for matrix factorization and related machine learning models.
problem Clarke subdifferential chain rules for matrix factorization and factorization machines.
method Analyzes conditions for subdifferential chain rules to hold, especially for overparameterized models.
result Subdifferential chain rules hold for matrix factorization and factorization machines under certain conditions.
A new method decouples set representation learning from posterior modeling for efficient amortized inference.
problem Efficient inference for large sets of observations with shared factors.
method Train a mean-pool Deep Set on sets of size at most two, then finetune the inference head on pre-aggregated embeddings.
result Matches or outperforms standard baselines at a fraction of the compute cost for large N.
Introduces BMF for efficient matrix factorization of large data.
problem Efficiently factorizing large scale matrices with limited memory.
method Uses block matrix approach and factorization at a block level.
result Demonstrates faster convergence on large matrices.