Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

54107161214 · Jun 202019922001200920172026
48 results for popularity adjustment

New algorithms improve community detection in network data with strong consistency.

problem Challenges in effectively adapting spectral clustering techniques and achieving strong consistency in label recovery.
method Proposed Thresholded Cosine Spectral Clustering (TCSC) and one-step Refined TCSC algorithms, with strong consistency proofs.
result One-step Refined TCSC achieves strong consistency in community detection under PABM, correctly recovering all labels with high probability.

New algorithms improve community detection and parameter estimation for PABM.

problem Improving community detection and parameter estimation for PABM.
method Connecting PABM to GRDPG, constructing new algorithms, and deriving asymptotic properties.
result Absolute number of community detection errors tends to zero as graph vertices increase.

In the present paper we study a sparse stochastic network enabled with a block structure. The popular Stochastic Block Model (SBM) and the Degree Corrected Block Model (DCBM) address sparsity by placing an upper bound on the maximum probability of connections between any pair of nodes. As a result, sparsity describes o…

2019-10-03abs ↗pdf ↗

Adjusted for chance measures are widely used to compare partitions/clusterings of the same data set. In particular, the Adjusted Rand Index (ARI) based on pair-counting, and the Adjusted Mutual Information (AMI) based on Shannon information theory are very popular in the clustering community. Nonetheless it is an open …

2015-12-03abs ↗pdf ↗

Paper characterizes optimal graph clustering limits under a new model.

problem Graph clustering under varying edge density signals.
method Introduced Popularity-Adjusted Block Model (PABM) to address SBM and DCBM limitations.
result Cluster recovery possible even when edge density signals vanish, highlighting local connectivity differences.

The rise in popularity of major social media platforms have enabled people to share photos and textual information about their daily life. One of the popular topics about which information is shared is food. Since a lot of media about food are attributed to particular locations and restaurants, information like spatio-…

2019-06-26abs ↗pdf ↗

Metaheuristics optimize portfolios with pre-assignment and margin trading for better risk-adjusted returns.

problem Maximizing returns while minimizing risk in portfolio optimization.
method Incorporates pre-assignment constraints and margin trading strategies using Genetic Algorithms and Particle Swarm Optimization.
result Metaheuristic-based portfolio optimization yields superior risk-adjusted returns compared to traditional methods.

Does adding a theorem to a paper affect its chance of acceptance? Does labeling a post with the author's gender affect the post popularity? This paper develops a method to estimate such causal effects from observational text data, adjusting for confounding features of the text such as the subject or writing quality. We…

2019-05-29abs ↗pdf ↗

This paper identifies and analyzes biases in risk-adjusted index weighting methods, affecting social welfare and market fairness.

problem Biases in risk-adjusted index weighting methods lead to tracking errors and fraud in indices and ETFs.
method Characterizes and analyzes the biases and adverse effects of risk-adjusted index weighting methods.
result These biases reduce social welfare and can enable harmful arbitrage activities.

Modified cosine distance improves similarity performance in data with variance and correlation.

problem Limitations of traditional cosine similarity in random variable spaces with variance and correlation.
method Proposed a variance-adjusted cosine distance metric to overcome limitations of traditional cosine similarity.
result Modified cosine distance shows 100% test accuracy in KNN model on the Wisconsin Breast Cancer Dataset.

Lower bounds on MALA and HMC for well-conditioned distributions.

problem Understanding the performance limits of Metropolized sampling methods.
method Analyzing the Metropolis-adjusted Langevin algorithm (MALA) and multi-step Hamiltonian Monte Carlo (HMC) with a leapfrog integrator.
result Nearly-tight lower bound of Ω~(κd)\widetildeΩ(κd) on the mixing time of MALA from an exponentially warm start.

New method dynamically adjusts UTD ratio to balance under- and overfitting in RL.

problem Balancing under- and overfitting in world model learning for RL.
method Dynamic adjustment of UTD ratio based on validation performance on a small subset of experience data.
result Our method improves balance between under- and overfitting compared to default settings and competitive with extensive hyperparameter search.

Binscatter is a popular method for visualizing bivariate relationships and conducting informal specification testing. We study the properties of this method formally and develop enhanced visualization and econometric binscatter tools. These include estimating conditional means with optimal binning and quantifying uncer…

2019-02-25abs ↗pdf ↗

A new sampler for complex discrete distributions efficiently updates all variables in parallel.

problem Sampling complex high-dimensional discrete distributions efficiently and accurately.
method Discrete Langevin proposal (DLP) for parallel coordinate updates with controlled stepsize.
result DLP efficiently explores high-dimensional and strongly correlated variables with asymptotic bias of zero for log-quadratic distributions.

The Kelly Criterion is applied to prediction markets to analyze risk and return.

problem Mean beliefs in prediction markets often differ from actual prices.
method Logarithmic utility and Kullback-Leibler divergence are used to study risk and return adjustments.
result Misjudgment of bias and investment fraction affect portfolio growth rate.

New model corrects bias in crowdsourced ratings for diverse items.

problem Bias and noise in crowdsourced ratings for training data.
method Bayesian rating model with item-level effects for difficulty, discriminativeness, and guessability.
result New model avoids bias in training data, improving model goodness of fit.

Machine learning factors outperform traditional portfolio optimization methods.

problem Comparing machine learning and traditional portfolio optimization methods.
method Examined machine learning and factor-based portfolio optimization using autoencoder neural networks and dimensionality reduction techniques.
result Minimum-variance portfolios using latent factors derived from autoencoders and sparse methods outperform simpler benchmarks in risk minimization.

A new noise model for preferential Bayesian optimization using user anchors.

problem Inadequate assumption of homoscedastic noise in human-in-the-loop settings.
method Proposes a heteroscedastic noise model with anchors and a KDE uncertainty map.
result Risk-adjusted performance improvement and clarified anchor placement effects.

Exact second-order optimization for deep learning reduces computational cost and improves performance.

problem Inadequate use of second-order optimization methods in deep learning due to high computational cost and non-convexity.
method Developed an exact stochastic second-order Newton method that addresses the non-convexity issue and provides an expression for the stochastic Hessian.
result Exact second-order Newton direction formula and its application in deep learning datasets.

Confounding bias, missing data, and selection bias are three common obstacles to valid causal inference in the data sciences. Covariate adjustment is the most pervasive technique for recovering casual effects from confounding bias. In this paper, we introduce a covariate adjustment formulation for controlling confoundi…

2019-07-02abs ↗pdf ↗

Improved financial performance through better regime prediction.

problem Predicting financial market regimes for profitable trading.
method A novel method combining contrarian trading and frequent short positions.
result Significant performance improvements over four years across three asset classes.

This paper considers the problem of choosing a good classifier. For each problem there exist an optimal classifier, but none are optimal, regarding the error rate, in all cases. Because there exists a large number of classifiers, a user would rather prefer an all-purpose classifier that is easy to adjust, in the hope t…

2018-02-10abs ↗pdf ↗

Investigates adjustments on Lie group crossed modules for gauge theory.

problem Existence and classification of adjustments on crossed modules of Lie groups.
method Differentiation/integration correspondence with infinitesimal adjustments; Lie algebra techniques.
result Infinitesimal adjustments exist if and only if the Kassel-Loday class lies in the image of the Chern-Weil homomorphism.

Paper discusses optimal CP for second-order predictions.

problem How to incorporate second-order predictions into conformal prediction.
method Introduces Bernoulli prediction sets (BPS) for second-order predictions and applies conformal risk control for compromised validity.
result BPS provides the smallest prediction sets with conditional coverage.

Reducing communication in training large-scale machine learning applications on distributed platform is still a big challenge. To address this issue, we propose a distributed hierarchical averaging stochastic gradient descent (Hier-AVG) algorithm with infrequent global reduction by introducing local reduction. As a gen…

2019-03-12abs ↗pdf ↗

The role of uncertainty quantification (UQ) in deep learning has become crucial with growing use of predictive models in high-risk applications. Though a large class of methods exists for measuring deep uncertainties, in practice, the resulting estimates are found to be poorly calibrated, thus making it challenging to …

2019-10-30abs ↗pdf ↗

Cross-entropy loss together with softmax is arguably one of the most common used supervision components in convolutional neural networks (CNNs). Despite its simplicity, popularity and excellent performance, the component does not explicitly encourage discriminative learning of features. In this paper, we propose a gene…

2016-12-07abs ↗pdf ↗

We describe principal 3-bundles with adjusted connections using Lie algebras and groupoids.

problem Describing principal 3-bundles with adjusted connections.
method Derived explicit forms of adjustment data for 3-term LL_\infty-algebras, integrated action Lie 3-algebroids to Lie 3-groupoids, and used differential cohomology.
result Explicit description of principal 3-bundles with adjusted connections in terms of differential cohomology.

Deep metric learning (DML) is a popular approach for images retrieval, solving verification (same or not) problems and addressing open set classification. Arguably, the most common DML approach is with triplet loss, despite significant advances in the area of DML. Triplet loss suffers from several issues such as collap…

2019-11-28abs ↗pdf ↗

A novel graphical matching approach improves pairs trading by reducing portfolio variance and risk-adjusted returns.

problem Common pairs trading methods lead to high portfolio variance and low risk-adjusted returns due to focusing on highly cointegrated assets.
method Model all assets and their cointegration levels with a weighted graph. Select pairs as a maximum weighted matching to ensure no shared assets and lower portfolio variance.
result The matching-based strategy shows a significant improvement in risk-adjusted performance, with a gross Sharpe ratio of 1.23.

Efficient adjustment sets found for cost-minimized causal estimations.

problem Estimating interventional means with minimum cost in causal graphical models.
method Defined cost-adjustment sets, constructed flow networks, and used maximum flow algorithms.
result Minimum cost optimal adjustment sets exist and can be found efficiently.

The paper provides PAC bounds for estimating causal effects using covariate adjustment with a valid set.

problem Estimating causal effects in high-dimensional settings without randomized experiments.
method PAC learning perspective, valid adjustment set, $\eps$-Markov blanket, constraint-based algorithms.
result PAC-bounds the estimation error of covariate adjustment by a term exponential in the size of the adjustment set.

Study optimal adjustment sets for causal policies with hidden variables.

problem Estimating dynamic treatment regimes with hidden variables.
method Developed criteria for graphs without hidden variables to compare estimators, extended to dynamic policies and hidden variables.
result Existence and computation of optimal minimal and globally optimal adjustment sets.

Improved ARMA-GARCH model for illiquid assets like cryptocurrencies.

problem Inadequate modeling of illiquid assets, especially cryptocurrencies, with traditional ARMA-GARCH models.
method Introducing liquidity-adjusted liquidity jump and diffusion metrics into ARMA-GARCH framework.
result The liquidity-adjusted model improves model fit and volatility sensitivity for cryptocurrencies.