Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

2855698541,138 · Jun 202019922001200920172026
48 results for Data Plagiarism Index

This paper introduces DPI, a new metric to assess data-copying risk in tabular data.

problem Measuring privacy risk of data-copying in tabular generative models.
method Proposes Data Plagiarism Index (DPI) for evaluating data-copying risk.
result DPI identifies data-copying threats to tabular data models, highlighting privacy and fairness issues.

Visual spoofing bypasses spam filters and plagiarism detection.

problem Vulnerability in spam filters that can be exploited by visually similar but differently encoded characters.
method Replaces characters with visually similar but differently encoded characters from a different alphabet.
result Spammers can create messages that bypass existing spam filters.

funcGNN uses graph neural networks to estimate program similarity efficiently.

problem Estimating accurate program similarity for software engineering tasks.
method funcGNN trains on labeled CFG pairs to predict GED between unseen programs using effective embedding vectors.
result funcGNN achieves lower error rate (0.00194) and is 23 times faster than traditional methods.

We introduce a Maximum Entropy model able to capture the statistics of melodies in music. The model can be used to generate new melodies that emulate the style of the musical corpus which was used to train it. Instead of using the nn-body interactions of (n1)(n-1)-order Markov models, traditionally used in automatic mus…

2016-10-11abs ↗pdf ↗

Machine learning conferences face ethical issues in review process.

problem Ethical issues in the review process of machine learning conferences.
method Study of recruitment issues, double-blind process infringements, fraudulent behaviors, biases, and appendix phenomenon.
result Highlighting the need for awareness in the machine learning community.

Index structures are important for efficient data access, which have been widely used to improve the performance in many in-memory systems. Due to high in-memory overheads, traditional index structures become difficult to process the explosive growth of data, let alone providing low latency and high throughput performa…

2019-05-08abs ↗pdf ↗

Improved stock index analysis using fuzzy parameters and machine learning.

problem Analyzing the S&P 500 stock index with long-term dependence.
method Combining fuzzy theory and machine learning to modify the Barndorff-Nielsen and Shephard model.
result The new model effectively captures the stochastic dynamics of the stock index time series.

DFR models dynamic distributional data with weighted Fréchet means.

problem Regression of distribution-valued responses over time.
method Dynamic Fréchet Regression (DFR) with index-aware weighting and feature selection.
result Improved predictive accuracy and feature recovery over existing methods.

Multifractal analysis and extensive statistical tests are performed upon intraday minutely data within individual trading days for four stock market indexes (including HSI, SZSC, S&P500, and NASDAQ) to check whether the indexes (instead of the returns) possess multifractality. We find that the mass exponent τ(q)τ(q) is l…

2007-06-14abs ↗pdf ↗

Pruning is an efficient model compression technique to remove redundancy in the connectivity of deep neural networks (DNNs). Computations using sparse matrices obtained by pruning parameters, however, exhibit vastly different parallelism depending on the index representation scheme. As a result, fine-grained pruning ha…

2019-05-14abs ↗pdf ↗

FunBaT extends Tucker decomposition to handle continuous-indexed tensor data.

problem Handling continuous-indexed tensor data that doesn't fit traditional Tucker decomposition.
method FunBaT treats continuous-indexed data as interactions between a core tensor and a group of latent functions modeled by Gaussian processes (GP). It converts each GP into a state-space prior and uses advanced message-passing techniques for scalable inference.
result FunBaT effectively handles real-world data with continuous indexes, demonstrating its advantage in synthetic and real-world applications.

We present a new model for credit index derivatives, in the top-down approach. This model has a dynamic loss intensity process with volatility and jumps and can include counterparty risk. It handles CDS, CDO tranches, Nth-to-default and index swaptions. Using properties of affine models, we derive closed formulas for t…

2009-11-09abs ↗pdf ↗

Quantum SVM improves financial data classification.

problem Classifying financial data using quantum machine learning.
method Application of quantum kernels to financial data, specifically DSEx Broad Index.
result Empirical quantum advantage demonstrated for financial data classification.

New method accurately reconstructs Russell 3000 index, revealing crowded portfolios.

problem Crowding in index portfolios during reconstitution events.
method Developed a Python package for accurate index reconstruction using CRSP US Stock data.
result Annual Russell 3000 portfolios are more crowded than quarterly ones, suggesting lower transaction costs.

This paper reviews and analyzes various modeling approaches for financial index tracking.

problem Efficient replication of market index performance in financial markets.
method Categorization into three frameworks: optimization, statistical, and machine learning; empirical study on S&P 500 dataset.
result Optimization-based models deliver the most precise index tracking, statistical-based models achieve the strongest return-risk balance, and data-driven models provide competitive performance.

The study examines the index of MOTS in Kerr-Newman-de Sitter spacetime and its relation to mass and charge.

problem Investigating the index of MOTS in Kerr-Newman-de Sitter spacetime.
method Analyzing the spatial cross section of the cosmological horizon in the Kerr-Newman-de Sitter spacetime, proving index bounds and establishing area-charge estimates.
result Established bounds on the index of MOTS and a connection between MOTS with index one and General Relativity.

Investment strategies involving cryptocurrencies and VIX INDEX show positive impact in market performance.

problem Investment strategies involving cryptocurrencies and VIX INDEX.
method Parameter estimation on raw data, comparison of two different portfolios, and analysis of different market conditions.
result VIX INDEX positively impacts the investment portfolio of cryptocurrencies in both standard and downward markets.

Develops a new cluster validity index to find multiple optimal cluster numbers.

problem Finding the optimal number of clusters in real-world data with varying densities, sizes, and shapes.
method A new correlation-based cluster validity index that yields multiple local peaks.
result The new index finds multiple optimal cluster numbers in various scenarios.

New validity index for fuzzy-possibilistic c-means clustering.

problem Conflicting results in determining the optimal number of clusters due to noisy data points and outliers.
method Introducing a new validity index (FP index) for fuzzy-possibilistic c-means clustering.
result FP index works well in datasets with varying cluster shapes and densities.

Bank transactions help predict macroeconomic indexes faster and more accurately.

problem Lag in macroeconomic index availability and autoregressive models' limitations in complex scenarios.
method Use financial transactions data to estimate macroeconomic indexes using neural networks and smart sampling.
result Neural network approach outperforms baseline methods on hand-crafted features based on transactions.

Training a neural network for a classification task typically assumes that the data to train are given from the beginning. However, in the real world, additional data accumulate gradually and the model requires additional training without accessing the old training data. This usually leads to the catastrophic forgettin…

2018-09-07abs ↗pdf ↗

Study creates a global living index to assess quality of life.

problem Long-term impacts of global economic changes on living conditions.
method Machine learning framework combining socio-economic factors.
result Developed a practical tool for policymakers to identify areas needing improvement.

Investment horizon approach has been used to analyze indexes of Polish stock market.Optimal time horizon for each return value is evaluated by fitting appropriate function form of the distribution. Strong asymmetry of gain-loss curves is observed for WIG index, whereas gain and loss curves look similar for WIG20 and fo…

2006-08-22abs ↗pdf ↗

Improved FDR control for sparse financial index tracking.

problem Maintaining FDR control in high-dimensional financial data with strong variable dependencies.
method Expanding T-Rex framework to handle overlapping groups of correlated variables with nearest neighbors penalization.
result Accurately tracks the S&P 500 index using only a small number of stocks.

A generalization of Callias' index theorem for self adjoint Dirac operators with skew adjoint potentials on asymptotically conic manifolds is presented in which the potential term may have constant rank nullspace at infinity. The index obtained depends on the choice of a family of Fredholm extensions, though as in the …

2012-10-11abs ↗pdf ↗

Paper models demand and solvency for index insurance, combining traditional and measurable index-based coverage.

problem Reducing protection gaps for emerging risks.
method Develops a model for demand and solvency conditions, combining traditional and index-based insurance.
result Deduces a product that benefits from both traditional and index-based insurance approaches.

SGD shows distinct phases in learning single-index models, achieving optimal sample complexity and regret.

problem Learning single-index models with SGD in adaptive data settings.
method Stochastic gradient descent (SGD) with an optimal learning rate schedule.
result SGD achieves near-optimal sample complexity and regret guarantees across both burn-in and learning phases.