Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

3016029031,204 · Jun 202019922001200920172026
48 results for Alternative Data

Investigates optimal consumption and investment using alternative data sources.

problem Optimal consumption and investment decisions under hidden economic regimes.
method Develops a novel duality theory for a jump-diffusion process with alternative data.
result Provides conditions for using control approach based on dynamic programming.

Alternative app data improves credit scoring for underserved borrowers.

problem Improving credit scoring for low-wealth and young individuals.
method Use of alternative data from app-based marketplaces, validated with TreeSHAP method.
result Alternative data sources predict financial behavior better than traditional bureau data.

Introduces alternators for modeling sequences, outperforming baselines.

problem Modeling complex sequential data with stability and efficiency.
method Two neural networks (OTN and FTN) alternate between outputting samples in observation and feature spaces, learned via cross-entropy criterion.
result Alternators outperform strong baselines in various domains (Lorenz equations, Neuroscience, Climate Science).

Paper predicts market implied volatility using alternative data and machine learning.

problem Predicting market implied volatility using alternative data.
method Used Google News statistics and Wikipedia site traffic as alternative data sources, and applied Logistic Regression, Support Vector Machines, and AdaBoost as machine learning models.
result Movements in market implied volatility can be predicted using machine learning techniques.

Study optimizes sampling to avoid extreme tail risks in unknown heavy-tailed distributions.

problem Identify optimal alternative with minimal extreme tail risk from unknown heavy-tailed distributions.
method Data-driven sequential sampling policies to maximize likelihood of selecting the optimal alternative.
result Proposed methods outperform existing approaches in identifying the optimal alternative.

Predictive modeling applications increasingly use data representing people's behavior, opinions, and interactions. Fine-grained behavior data often has different structure from traditional data, being very high-dimensional and sparse. Models built from these data are quite difficult to interpret, since they contain man…

2016-07-21abs ↗pdf ↗

Model analyzes cooccurrence data for recommender systems and item relevance.

problem High-dimensional cooccurrence data from online platforms.
method Shared parameter Alternating Tweedie (SA-Tweedie) model with Fisher scoring and learning rate adjustment.
result SA-Tweedie model outperforms other methods in optimizing parameters.

The α\alpha-Alternator adapts to varying noise levels in sequences, improving robustness and performance.

problem Current models assume uniform noise levels, limiting performance on noisy temporal data.
method Introduces α\alpha-Alternator using Vendi Score to dynamically adjust noise sensitivity.
result Outperforms Alternators and state-of-the-art models in trajectory prediction, imputation, and forecasting.

We propose a method for finding alternate features missing in the Lasso optimal solution. In ordinary Lasso problem, one global optimum is obtained and the resulting features are interpreted as task-relevant features. However, this can overlook possibly relevant features not selected by the Lasso. With the proposed met…

2016-11-18abs ↗pdf ↗

Paper proposes an algorithm for PARAFAC2-based CMTF models with various constraints.

problem Jointly analyze matrices and tensors with irregular/ragged data.
method Alternating Optimization (AO) and ADMM for fitting PARAFAC2-based CMTF models with various constraints.
result Accurately recovers underlying patterns using various constraints and linear couplings.

Flexible framework for CMTF with ADMM for various constraints and couplings.

problem Challenges in data fusion from multiple sources with varying characteristics.
method Flexible algorithmic framework using AO and ADMM for various constraints, loss functions, and couplings.
result Accurate and computationally efficient results for various loss functions, including KL divergence.

The Volume conjecture claims that the hyperbolic Volume of a knot is determined by the colored Jones polynomial. The purpose of this article is to show a Volume-ish theorem for alternating knots in terms of the Jones polynomial, rather than the colored Jones polynomial: The ratio of the Volume and certain sums of coeff…

2004-03-25abs ↗pdf ↗

New method calculates knot and link properties using state codes.

problem Determining the unoriented genus and crosscap number of prime alternating knots and links.
method Encoding states as tuples and using them to compute genus and crosscap number.
result Computed values for all such links through 14 crossings and knots through 19 crossings, identifying patterns.

Alternating minimization represents a widely applicable and empirically successful approach for finding low-rank matrices that best fit the given data. For example, for the problem of low-rank matrix completion, this method is believed to be one of the most accurate and efficient, and formed a major component of the wi…

2012-12-03abs ↗pdf ↗

Data-driven predictive analytics are in use today across a number of industrial applications, but further integration is hindered by the requirement of similarity among model training and test data distributions. This paper addresses the need of learning from possibly nonstationary data streams, or under concept drift,…

2017-10-18abs ↗pdf ↗

Study tests uniformity of categorical data against missing-ball alternatives, finding chi-squared test outperforms.

problem Testing uniformity of categorical data against missing-ball alternatives.
method Characterizes minimax risk, uses collisions and chi-squared test, reduces to structured subset of alternatives.
result Minimax test outperforms chi-squared test under least favorable alternative.

We introduce a version of Khovanov homology for alternating links with marking data, ωω, inspired by instanton theory. We show that the analogue of the spectral sequence from Khovanov homology to singular instanton homology introduced in \cite{KM_unknot} for this marked Khovanov homology collapses on the E2E_2 page fo…

2018-01-08abs ↗pdf ↗

GT-PCA improves PCA for image and time series data.

problem Lack of robustness to transformations in PCA.
method GT-PCA is a neural network that estimates components invariant to specific transformations.
result GT-PCA outperforms alternative methods in synthetic and real data experiments.

Improves deep neural networks using soft labels through alternating minimization.

problem Improving deep neural networks training with soft labels.
method Co-Learns DNNs and soft labels via Alternating Minimization of two objectives.
result COLAM achieves improved performance on many tasks with better testing classification accuracy.

This paper proposes an alternating back-propagation algorithm for learning the generator network model. The model is a non-linear generalization of factor analysis. In this model, the mapping from the continuous latent factors to the observed signal is parametrized by a convolutional neural network. The alternating bac…

2016-06-28abs ↗pdf ↗

An alternating distance is a link invariant that measures how far away a link is from alternating. We study several alternating distances and demonstrate that there exist families of links for which the difference between certain alternating distances is arbitrarily large. We also show that two alternating distances, t…

2014-06-26abs ↗pdf ↗

New quasi-alternating links created from existing ones.

problem Creating new quasi-alternating links from existing ones.
method Extending the construction of quasi-alternating links by replacing a crossing with an alternating tangle of the same type.
result Jones polynomial of new quasi-alternating links has no gap if the original link has no gap.

AI improves MSME credit scoring using bank statement data.

problem Lack of access to financing for MSMEs due to traditional credit scoring methods.
method Developed a cash flow-based pipeline using bank statement data for machine learning credit scoring.
result Bank statement features significantly improve credit scoring models, achieving AUROC of 0.806.

This paper studies a stylized, yet natural, learning-to-rank problem and points out the critical incorrectness of a widely used nearest neighbor algorithm. We consider a model with nn agents (users) {xi}i[n]\{x_i\}_{i \in [n]} and mm alternatives (items) {yj}j[m]\{y_j\}_{j \in [m]}, each of which is associated with a latent feat…

2018-07-09abs ↗pdf ↗

A link is almost alternating if it is non-alternating and has a diagram that can be transformed into an alternating diagram via one crossing change. We give formulas for the first two and last two potential coefficients of the Jones polynomial of an almost alternating link. Using these formulas, we show that the Jones …

2017-07-18abs ↗pdf ↗

Bankwitz characterized an alternating diagram representing the trivial knot. A non-alternating diagram is called almost alternating if one crossing change makes the diagram alternating. We characterize an almost alternaing diagram representing the trivial knot. As a corollary we determine an unknotting number one alter…

2006-04-30abs ↗pdf ↗

The paper tackles learning true rankings from noisy, incomplete data.

problem Learning true rankings from incomplete and noisy data.
method Introduces a selective Mallows model for noisy rankings and derives upper and lower bounds on sample complexity.
result Strong asymptotically tight bounds on sample complexity for learning complete rankings and top-k rankings.

Data augmentation by mixing samples, such as Mixup, has widely been used typically for classification tasks. However, this strategy is not always effective due to the gap between augmented samples for training and original samples for testing. This gap may prevent a classifier from learning the optimal decision boundar…

2019-06-20abs ↗pdf ↗