Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

10.7%21.4%32.1%42.8% · Jun 202019922001200920172026
48 results for data alteration

Machine learning identifies types of alterations in historical manuscripts.

problem Understanding and categorizing alterations in historical manuscripts.
method Alteration Latent Dirichlet Allocation (alterLDA) model.
result High performance in recognizing alterations on labelled data, and interesting insights on unlabelled data.

Framework identifies brain connectivity alterations for MDD patients using limited rs-fMRI data.

problem Difficult to analyze brain connectivity alterations from limited rs-fMRI data.
method Proposed a multitask Gaussian Bayesian network (MTGBN) framework to learn individual disease-induced alterations.
result Framework efficiently learns Bayesian network structures from limited data, showing improved performance.

New algorithm detects tensor dependence structure alterations efficiently.

problem Detecting alterations in tensor dependence structures.
method Tensor-normal distributions, decorrelation, centralization, SERA (Sparsity-Exploited Reranking Algorithm).
result The proposed SERA algorithm controls false discovery rates effectively.

Collectives can manipulate learning platforms by coordinated data submission, requiring strategic assessments and algorithms.

problem Collectives can influence learning platforms by altering data, posing risks and requiring strategic planning.
method Developed a theoretical and algorithmic framework to understand and mitigate collective manipulation of learning platforms.
result Demonstrated the need for strategic assessments and implementable coordination algorithms to prevent collective manipulation.

Transformations of macroeconomic data affect machine learning forecasts, especially with regularization and nonlinearity.

problem The impact of data transformations on machine learning forecasts in macroeconomic contexts.
method Review and propose new data transformations, empirically evaluate their effects, and compare traditional and moving average rotations.
result Traditional factors should almost always be included as predictors, and moving average rotations can provide important gains.

Artificial intelligence, or AI, enhancements are increasingly shaping our daily lives. Financial decision-making is no exception to this. We introduce the notion of AI Alter Egos, which are shadow robo-investors, and use a unique data set covering brokerage accounts for a large cross-section of investors over a sample …

2019-07-08abs ↗pdf ↗

In the present work, we study the decompositions of codimension-one transitions that alter the singular set the of stable maps of S3S^3 into R3,\mathbb{R}^3, the topological behaviour of the singular set and the singularities in the branch set that involves cuspidal curves and swallowtails that alter the singular set. W…

2018-07-16abs ↗pdf ↗

A new mathematical approach detects frequency-based alterations in brain networks.

problem Understanding disease-relevant brain alterations through network analysis.
method Proposes a novel connectome harmonic analysis framework using common harmonic waves learned from Stiefel manifolds.
result Identifies more significant and reproducible network dysfunction patterns in Alzheimer's disease.

Foundation models alter medical data science workflow, challenging veridical data science principles.

problem Foundation models disrupt traditional data science practices in medicine.
method Critically examined the medical foundation model lifecycle and its deviation from veridical data science principles.
result Foundation models challenge veridical data science principles of predictability, computability, and stability.

Frequentist method estimates uncertainty in RNNs without altering architecture.

problem Uncertainty quantification in RNNs for decision-making.
method Jackknife resampling and influence functions to estimate variability.
result The method provides theoretical coverage guarantees on uncertainty intervals.

To characterize circumstellar systems in high contrast imaging, the fundamental step is to construct a best point spread function (PSF) template for the non-circumstellar signals (i.e., star light and speckles) and separate it from the observation. With existing PSF construction methods, the circumstellar signals (e.g.…

2020-01-02abs ↗pdf ↗

Model explanations based on pure observational data cannot compute the effects of features reliably, due to their inability to estimate how each factor alteration could affect the rest. We argue that explanations should be based on the causal model of the data and the derived intervened causal models, that represent th…

2019-09-19abs ↗pdf ↗

Proposes a method to model financial returns with extreme shocks using flexible tail transformations.

problem Capturing extreme shocks in financial return data.
method Introduces a transformation layer in normalizing flows to model heavy-tailed distributions.
result Trained models can generate synthetic sets of extreme returns.

New method creates universal perturbations to fool neural network interpretations.

problem Vulnerability of gradient-based saliency maps to adversarial perturbations.
method Gradient-based optimization and PCA-based approach to create UPI.
result Existence and successful application of Universal Perturbation for Interpretation (UPI).

Graph neural networks help assess how global changes affect plant-pollinator networks.

problem Interpreting GNN results to understand how global changes impact plant-pollinator networks.
method Simulation study and application on Spipoll dataset to assess effects of global changes on pollination networks.
result GNNs can detect interactive effects between covariates and plant genera on pollination network connectivity.

We introduce a new model of stochastic bandits with adversarial corruptions which aims to capture settings where most of the input follows a stochastic pattern but some fraction of it can be adversarially changed to trick the algorithm, e.g., click fraud, fake reviews and email spam. The goal of this model is to encour…

2018-03-25abs ↗pdf ↗

Humans increasingly interact with Artificial intelligence(AI) systems. AI systems are optimized for objectives such as minimum computation or minimum error rate in recognizing and interpreting inputs from humans. In contrast, inputs created by humans are often treated as a given. We investigate how inputs of humans can…

2019-12-08abs ↗pdf ↗

The paper analyzes how much data points can be altered to change their rank in nearest neighbor searches.

problem Vulnerability of nearest neighbor search in high-dimensional data.
method Statistical analysis of perturbation needed to change neighbor rank.
result Derived statistical distribution of perturbation needed to modify neighbor rank.

We demonstrate that the lowest possible price change (tick-size) has a large impact on the structure of financial return distributions. It induces a microstructure as well as it can alter the tail behavior. On small return intervals, the tick-size can distort the calculation of correlations. This especially occurs on s…

2010-01-28abs ↗pdf ↗

We consider data poisoning attacks, a class of adversarial attacks on machine learning where an adversary has the power to alter a small fraction of the training data in order to make the trained classifier satisfy certain objectives. While there has been much prior work on data poisoning, most of it is in the offline …

2018-08-27abs ↗pdf ↗

Combines global and local features for better social circle prediction in ego-networks.

problem Efficiently analyzing ego-networks with hidden local structures.
method Evolved deep learning techniques to capture both global and local network features.
result Social circle prediction benefits from a combination of global and local features.

A method to prevent image representation collapse through data-dependent augmentation.

problem Representation collapse due to image augmentations that damage information.
method Formalizing a stochastic encoding process with a tug-of-war between corruption and preserved information, using infoMax objective.
result Learning a data-dependent distribution of augmentations to avoid representation collapse.

Study examines how business units can benefit from group cohesion under regulatory constraints.

problem Regulatory constraints limit business units' ability to form a single cohesive group.
method Defined and analyzed cohesive risk measures to minimize capital costs.
result Cohesive risk measures allow groups to achieve minimal capital costs without altering individual liabilities.

'Big' high-dimensional data are commonly analyzed in low-dimensions, after performing a dimensionality-reduction step that inherently distorts the data structure. For the same purpose, clustering methods are also often used. These methods also introduce a bias, either by starting from the assumption of a particular geo…

2018-02-15abs ↗pdf ↗

Study shows adversarial attacks can fool speech-to-text models, and PCA is ineffective as a defense.

problem Adversarial attacks can mislead speech-to-text neural networks.
method Crafted adversarial waveforms, used PCA for defense, tested under black-box setting.
result PCA is ineffective as a defense mechanism against adversarial attacks in audio domain.

This study shows how trade policy uncertainty affects stock-T bill correlations.

problem The impact of trade policy uncertainty on stock-T bill relationships.
method Extended Dynamic Conditional Correlation (DCC) framework incorporating exogenous variables.
result Trade policy uncertainty significantly alters stock-T bill correlations, especially under specific political conditions.

Interestingness measures provide information that can be used to prune or select association rules. A given value of an interestingness measure is often interpreted relative to the overall range of the values that the interestingness measure can take. However, properties of individual association rules restrict the val…

2013-08-16abs ↗pdf ↗

Learning algorithms need bias to generalize and perform better than random guessing. We examine the flexibility (expressivity) of biased algorithms. An expressive algorithm can adapt to changing training data, altering its outcome based on changes in its input. We measure expressivity by using an information-theoretic …

2019-11-09abs ↗pdf ↗

We define the (total) center of mass for suitably asymptotically hyperbolic time-slices of asymptotically anti-de Sitter spacetimes in general relativity. We do so in analogy to the picture that has been consolidated for the (total) center of mass of suitably asymptotically Euclidean time-slices of asymptotically Minko…

2015-01-22abs ↗pdf ↗

We analyse the dependence of stock return cross-correlations on the sampling frequency of the data known as the Epps effect: For high resolution data the cross-correlations are significantly smaller than their asymptotic value as observed on daily data. The former description implies that changing trading frequency sho…

2007-04-09abs ↗pdf ↗