Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

326395126 · Jun 202019922001200920182026
48 results for groundtruth collection

Satyam simplifies groundtruth collection for machine vision tasks.

problem Difficulty in collecting accurate groundtruth for machine vision systems.
method Crowdsourcing through Amazon Mechanical Turk, automation of task creation and quality control.
result Groundtruth collected by Satyam is comparable to that of trained experts and provides equivalent ML performance.

Self-supervised method estimates depth from monocular endoscopy videos.

problem Depth estimation from monocular endoscopy data without manual labeling.
method Convolutional neural networks trained with sparse supervision from stereo methods.
result Submillimeter mean residual error in cross-patient CT scans comparison.

iPrompt uses LLMs to generate natural-language explanations of data patterns.

problem Finding and explaining patterns in data using natural language.
method Interpretable autoprompting (iPrompt) that generates natural-language explanations based on LLMs.
result iPrompt can accurately find and explain data patterns, improving upon human-written prompts.

Paper studies nonnegative Tucker decomposition identifiability with sparsity conditions.

problem Identify nonnegative Tucker decomposition factors uniquely.
method Adapting NMF identifiability results, derive procedures using tensor unfoldings or slices.
result Nonnegative Tucker decomposition factors are identifiable under certain sparsity conditions.

This work provides guarantees for off-policy function estimation under realizability assumptions.

problem Estimating the value function of a policy under user-specified error-measuring distributions.
method The approach involves imposing a flexible regularization on the MIS objectives to account for an arbitrary user-specified distribution.
result Exact characterization of the optimal dual solution that determines the data-coverage assumption in the case of value-function learning.

Enhanced transformer converts whispered speech to natural speech.

problem Machine recognition of whispered speech is challenging.
method Proposes an enhanced transformer architecture trained end-to-end using supervised learning.
result Similar formant distributions of converted speech to groundtruth.

We apply variational inference to learn vehicle trajectory parameters from noisy data.

problem Learning parameters for vehicle trajectory estimation from noisy measurements.
method Gaussian variational inference with parameter learning in a motion and sensor model context.
result High-quality state estimates achieved even with outliers and false loop closures.

Study collective pricing and hedging with admissible risk exchanges forming a finitely generated convex cone.

problem Collective pricing and hedging with exchanges forming a finitely generated convex cone.
method Extend collective First Fundamental Theorem of Asset Pricing and pricing-hedging duality.
result No collective arbitrage implies the closedness of the aggregate feasibility cone.

Measures collectivity in financial covariances and correlations to reveal trends and precursors.

problem Capturing collective motion in financial markets to predict trends and precursors.
method Measures collectivity using the largest eigenvalue and average sector collectivity.
result Identifies collective signals around major financial events and captures trends in covariances and correlations.

Study finds strict collection policies improve portfolio quality of microfinance banks.

problem Improving portfolio quality of microfinance banks through better credit collection policies.
method Multi-stage sampling, regression analysis, descriptive statistics.
result Collection policy has a higher effect on portfolio quality.

The paper extends collective arbitrage concepts to multi-agent markets with cooperation.

problem Understanding collective market completeness and pricing in multi-agent systems.
method Develops new techniques and theorems to establish collective pricing-hedging duality and collective replication.
result Established a Second Fundamental Theorem of Asset Pricing in cooperative multi-agent settings.

Active data collection improves convergence rates in operator learning.

problem Improving convergence rates in operator learning with linear target and stochastic input.
method Active data collection strategies with mean-zero stochastic process and continuous covariance kernels.
result Achieves arbitrarily fast error convergence rates with eigenvalue decay of covariance kernels.

A universal collection of 4 invariants improves neural network accuracy for molecular dynamics.

problem Improving accuracy of neural networks in molecular dynamics.
method Developed a universal collection of 4 smooth scalar invariants on M(3) x M(3) and evaluated their effectiveness in a PONITA neural network architecture.
result Using a universal collection of invariants significantly improves neural network accuracy.

This study examines how learning algorithms affect collective action in machine learning.

problem The impact of collective action on machine learning is limited when not considering the choice of learning algorithms.
method Focuses on distributionally robust optimization and stochastic gradient descent, analyzing their effects on collective success.
result The choice of learning algorithm significantly impacts the effective size and success of a collective in machine learning.

This review explores the use of machine learning in discovering collective variables for biomolecular dynamics.

problem Understanding the conformational dynamics and molecular recognition in biomolecules.
method Statistical analysis of high-dimensional spatiotemporal data generated from molecular dynamics simulations.
result Machine learning algorithms can be used to discover abstract collective variables that describe biomolecular dynamics.

CUDC collects diverse data for offline RL by predicting future states.

problem Challenges in collecting task-agnostic data for offline RL.
method Adaptive temporal distances for curiosity-driven data collection.
result CUDC outperforms existing unsupervised methods in offline RL tasks.

Collectives can manipulate learning platforms by coordinated data submission, requiring strategic assessments and algorithms.

problem Collectives can influence learning platforms by altering data, posing risks and requiring strategic planning.
method Developed a theoretical and algorithmic framework to understand and mitigate collective manipulation of learning platforms.
result Demonstrated the need for strategic assessments and implementable coordination algorithms to prevent collective manipulation.

Post-ADC inference corrects bias in statistical inference after active data collection.

problem Bias in inference after active data collection.
method Post-ADC inference framework that corrects bias from both ADC process and data-driven target construction.
result Valid inference for data collected by SMBO methods like GP-UCB and TPE.

Study shows small groups can influence machine learning algorithms.

problem How small groups can influence machine learning algorithms deployed on digital platforms.
method Proposed a theoretical model and conducted experiments on a large-scale language model.
result Small groups can exert significant control over machine learning algorithms.

The study examines collective behavior in banking sectors across mature and emerging markets.

problem Understanding collective behavior in banking sectors across different market types.
method Applied Random Matrix Theory (RMT) to analyze the banking sectors of 4 world stock markets.
result Mature markets exhibit higher collective behavior compared to emerging markets.

Optimal online data collection for semiparametric inference reduces regret.

problem Sequential data collection decisions for efficient estimation under budget constraints.
method Online Moment Selection framework; Explore-then-Commit and Explore-then-Greedy policies.
result Online data collection policies achieve zero regret relative to an oracle policy.

We address the collective matrix completion problem of jointly recovering a collection of matrices with shared structure from partial (and potentially noisy) observations. To ensure well--posedness of the problem, we impose a joint low rank structure, wherein each component matrix is low rank and the latent space of th…

2014-12-05abs ↗pdf ↗

Proposes online debiasing to correct bias in adaptive data collection for high-dimensional linear regression.

problem Bias in adaptive data collection for high-dimensional linear regression.
method Online debiasing procedure for LASSO and other estimators.
result Optimal debiasing of LASSO estimator in specific sparsity regime.

New algorithm for collective Gaussian hidden Markov models inference.

problem Inference of collective Gaussian hidden Markov models from aggregate data.
method Collective Gaussian forward-backward algorithm, extending Sinkhorn belief propagation.
result Convergence guarantee and applicability to single individual Kalman filter.

Classifies collective motions in biological networks using graph dynamic mode decomposition.

problem Classifying complex collective motions in biological networks based on transient and complexly changing network properties.
method Data-driven spectral analysis (graph dynamic mode decomposition) to extract dynamical properties.
result Contextual node information and physical properties are crucial for classifying collective motions.

We determine the extent to which the collection of ΓΓ-Euler-Satake characteristics classify closed 2-orbifolds. In particular, we show that the closed, connected, effective, orientable 2-orbifolds are classified by the collection of ΓΓ-Euler-Satake characteristics corresponding to free or free abelian ΓΓ and are not…

2009-02-12abs ↗pdf ↗

Adapting policy learning for data collected from evolving systems.

problem Challenges in learning optimal policies from adaptively collected data.
method Proposes an algorithm based on generalized augmented inverse propensity weighted (AIPW) estimators to control worst-case estimation variance.
result Achieves minimax rate optimal regret guarantees even with diminishing exploration.

Improved disability insurance model with collective health claims.

problem Enhance disability insurance model with collective health claims.
method Expand classic semi-Markov model with collective health claims, solve many-body problem using mean-field approach.
result Mean-field approach simplifies complex model into a transparent pricing method.

We derive a class of macroscopic differential equations that describe collective adaptation, starting from a discrete-time stochastic microscopic model. The behavior of each agent is a dynamic balance between adaptation that locally achieves the best action and memory loss that leads to randomized behavior. We show tha…

2004-08-20abs ↗pdf ↗

A scalable topic model for large document collections using MapReduce.

problem Scalability issues in topic modeling for large document collections.
method Correlated Topic Model with variational Expectation-Maximization in MapReduce framework.
result Comparable topic coherences with LDA in MapReduce framework.

Let Ω=(ωj)jIΩ=(ω_{j})_{j\in I} be a collection of pairwise non-isotopic simple closed curves on the closed, orientable, genus gg surface SgS_{g}, such that ωiω_{i} and ωjω_{j} intersect exactly once for iji\neq j. It was recently demonstrated by Malestein, Rivin, and Theran that the cardinality of such a collection is no mo…

2012-10-10abs ↗pdf ↗

Develops efficient method to detect multiple collective anomalies in multivariate data streams.

problem Detecting anomalies in multivariate data streams, especially collective anomalies.
method MVCAPA: A method that efficiently detects multiple collective anomalies without approximations.
result MVCAPA consistently estimates the number and location of collective anomalies.