Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

66133199265 · Jun 202019922001200920182026
48 results for fundamental sources

The study identifies two sources of invariants in 2--nondegenerate CR geometries.

problem Characterizing fundamental invariants of 2--nondegenerate CR geometries.
method Analyzes the harmonic curvature and the difference in complex structures.
result Nontrivial examples of CR geometries can be obtained as deformations of models.

Study examines Bitcoin price formation beyond fundamental sources.

problem Understanding Bitcoin price formation beyond traditional factors.
method Bayesian quantile regression to analyze Bitcoin price and its determinants.
result Identified three groups of determinants influencing Bitcoin price under different market conditions.

Training a source model optimally for its own task is suboptimal for downstream transfer.

problem The optimality of a source model for its own task hinders downstream transfer performance.
method Analyzes L2-SP ridge regression, characterizes transfer-optimal source penalty, and identifies alignment-dependent effects.
result Transfer benefits from stronger source regularization when aligned imperfectly, and from weaker regularization when aligned perfectly.

In this paper, as a fundamental study on the theory of Morse functions and their higher dimensional versions or fold maps and applications to geometric theory of manifolds, which were started in 1950s by differential topologists such as Thom and Whitney and have been studied actively, we study algebraic and differentia…

2015-08-23abs ↗pdf ↗

Algorithm finds frequencies, amplitudes, and phases of sinusoids in noisy data.

problem Finding frequencies, amplitudes, and phases of sinusoids in noisy data.
method Maximum likelihood approach to estimate tone parameters from contaminated observations. Successively estimates frequencies and jointly optimizes amplitudes and phases.
result Near-linear computational complexity (O(N)) for estimating MM number of sinusoidal sources.

New method separates audio sources without needing known decompositions.

problem Difficulty in training source separation models on real-world mixtures.
method Generates estimated decompositions from stereo mixtures and trains a deep learning model.
result Trained model can separate single-channel audio sources effectively.

New method identifies diffusion sources on networks with statistical confidence.

problem Identifying sources of diffusion on networks without restrictive assumptions.
method Statistical framework and confidence set inference approach based on hypothesis testing.
result Efficiently produces a small subset of nodes covering the source node with any confidence level.

Paper proposes a technique to detect and predict sources of contaminants in complex systems.

problem Difficulty in identifying sources of contaminants in coupled natural and human systems.
method Developed a technique for simultaneous source detection and prediction.
result Outperforms other approaches in detecting potential groundwater contamination.

The Rashomon effect shows many models can perform similarly, explored in this paper.

problem Why do many models perform similarly in machine learning?
method Categorized causes into statistical, structural, and procedural sources.
result Structural multiplicity persists and cannot be resolved without additional assumptions.

Second-order methods fail to fully quantify epistemic uncertainty, leading to biased predictions.

problem Incomplete quantification of epistemic uncertainty in machine learning models.
method Analysis of existing second-order uncertainty estimation methods.
result Current methods overestimate aleatoric uncertainty and underestimate epistemic uncertainty, leading to biased predictions.

We give complete algorithms and source code for constructing (multilevel) statistical industry classifications, including methods for fixing the number of clusters at each level (and the number of levels). Under the hood there are clustering algorithms (e.g., k-means). However, what should we cluster? Correlations? Ret…

2016-07-17abs ↗pdf ↗

This paper develops a method to estimate the rate-distortion function for general data sources.

problem Estimating the rate-distortion function for general data sources.
method Develops an algorithm for sandwiching the R-D function of a general (not necessarily discrete) source using i.i.d. data samples.
result Estimates R-D sandwich bounds for various data sources, including natural images, indicating potential for improving compression methods.

This paper benchmarks FinGPT for financial datasets using open-source large language models.

problem Challenges in integrating GPT-based models with financial datasets.
method Instruction Tuning paradigm for open-source large language models adapted for financial contexts.
result Demonstrates the effectiveness and adaptability of FinGPT in financial tasks.

Paper studies clustering with transfer learning in high dimensions.

problem Improving clustering accuracy in high-dimensional settings with related datasets.
method Developed a minimax-optimal transfer-assisted clustering procedure.
result Characterized phase transitions for consistent target clustering.

Adding data can sometimes hurt model performance in multi-source healthcare tasks.

problem Identifying when adding more data helps or hinders model outcomes in multi-source healthcare tasks.
method Identified the Data Addition Dilemma, demonstrated empirically observed trade-offs, introduced distribution shift heuristics.
result Adding data can sometimes reduce model performance due to distribution shift.

A new method for binary ICA using non-stationary sources.

problem Independent component analysis of binary data.
method Linear mixing model in latent space, followed by binary observation model with non-stationary sources.
result Proves non-identifiability with few observed variables but identifies with more variables.

A framework isolates and learns approximately shared features for better domain adaptation.

problem Reducing performance degradation in unseen domains using machine learning models.
method Statistical framework distinguishing feature utilities based on correlation variance across domains. Learning approximately shared features from source tasks and fine-tuning on target tasks.
result Improved population risk compared to previous results on both source and target tasks, resolving the paradox of feature selection.

Non-negative blind source separation (BSS) has raised interest in various fields of research, as testified by the wide literature on the topic of non-negative matrix factorization (NMF). In this context, it is fundamental that the sources to be estimated present some diversity in order to be efficiently retrieved. Spar…

2013-08-26abs ↗pdf ↗

In this paper it is shown that a CR embedding from one strictly pseudoconvex hypersurface into another (of strictly larger dimension) sends chains on the source to chains on the target if and only if the embedding has a lift to a conformal isometry of the associated Fefferman bundles with vanishing pseudo-Riemannian se…

2010-10-17abs ↗pdf ↗

This study examines representation bias in open-source Qwen models for investment decisions.

problem Representation bias in financial applications of large language models.
method Balanced round-robin prompting over 150 U.S. equities, constrained decoding, token-logit aggregation.
result Firm size and valuation increase model confidence, while risk factors decrease it.

LEARNER improves low-rank matrix estimation using source population data.

problem Improving low-rank matrix estimation in target populations with diverse data sources.
method LEARNER uses similarity in latent spaces between source and target populations to enhance estimation.
result LEARNER often outperforms benchmark methods, especially with higher signal-to-noise ratios in the source population.

Paper models and compresses wideband CSI feedback in FDD MIMO systems.

problem Fundamental limits of channel state information (CSI) feedback in FDD massive MIMO systems.
method Modeling CSI as a Gaussian-mixture source with latent geometry states, proposing Gaussian-mixture transform coding (GMTC).
result Near-optimal CSI compression achieved through state-adaptive transform coding without large neural encoders.

Proposes a copula-based model for multi-view clustering with directional dependency.

problem Challenges in integrating multi-source datasets with directional dependency.
method Copula-based multi-view clustering model accounting for directional dependence.
result Ignoring directional dependence negatively impacts clustering performance.

Paper shows adversarial classification algorithms are inherently more sensitive to data manipulation.

problem Inadequate provable guarantees for machine learning performance, especially in unreliable data environments.
method Formal analysis of binary classification algorithms' sensitivity to adversarial manipulation.
result Fundamental tradeoff curve between accuracy and sensitivity is determined by data statistics, not algorithm tuning.

Study shows multi-distribution learning has slower rates than single-task learning.

problem Understanding the statistical complexity of learning from heterogeneous sources.
method Structured hypothesis-testing framework to capture the statistical cost of certifying near-optimality under bounded noise.
result Learning across multiple distributions incurs slow rates scaling with k/ε2k/ε^2, even under constant noise levels.

We present a simple agent-based model to study the development of a bubble and the consequential crash and investigate how their proximate triggering factor might relate to their fundamental mechanism, and vice versa. Our agents invest according to their opinion on future price movements, which is based on three source…

2008-06-18abs ↗pdf ↗

We study the limits and methods of training two-layer autoencoders.

problem Understanding the limits and methods of training two-layer autoencoders.
method Focus on non-linear two-layer autoencoders trained in the proportional regime, using gradient methods.
result Gradient methods achieve the minimizers of the population risk and reveal the structure of the features.

Modeling wildfire aerosols using satellite data to predict solar radiation reduction.

problem Accurately estimate and predict AOD propagation from wildfires using multi-source satellite data.
method Physics-informed statistical modeling integrating multi-source satellite data with an advection-diffusion equation.
result The proposed approach accurately predicts AOD propagation and demonstrates model interpretability.