Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,878 papers · 148 categories

Trend · papers per month

2795588371,116 · Jun 202019922001200920172026
48 results for ancillary data

DART2 enhances multiple testing by leveraging ancillary information robustly.

problem Enhancing multiple testing power with uncertain ancillary information.
method Distance-assisted multiple testing procedure (DART2) that handles both helpful and misleading ancillary information.
result DART2 asymptotically controls FDR and improves power when ancillary information is helpful, maintaining FDR and power otherwise.

This paper improves bidding price prediction for ancillary services markets, boosting revenues.

problem Volatility in renewable energy sources affects grid stability and revenue optimization.
method Machine learning models (SVR, DT, k-NN) and offset adjustment for pay-as-bid markets.
result The proposed approach increases potential revenues by 27.43% to 37.31% compared to baseline models.

Simpler model outperforms state-of-the-art for disaggregating census data.

problem Disaggregating detailed census data into finer-grained, high-resolution mappings.
method Aggregate learning approach using ancillary data for an interpretable model.
result Simple model outperforms state-of-the-art on disaggregation metrics.

Ancillaries have become a major source of revenue and profitability in the travel industry. Yet, conventional pricing strategies are based on business rules that are poorly optimized and do not respond to changing market conditions. This paper describes the dynamic pricing model developed by Deepair solutions, an AI te…

2019-02-06abs ↗pdf ↗

Proposes a new GAN architecture for generating data conditioned on partial information.

problem Generating data conditioned on partial ancillary information.
method Introduces a new Adversarial Network architecture and training strategy.
result The proposed method outperforms standard Conditional GANs in generating data under partial conditioning.

Optimizes profit in targeted marketing across multiple markets with varying marketing expenditures.

problem Maximizing profit in a sequential marketing strategy with multiple markets and varying marketing costs.
method Near-optimal algorithms in an adversarial bandit setting, proving regret bounds for different demand curve types.
result Proved near-optimal regret bounds for the profit-maximization problem in targeted marketing.

We consider a natural model of random knotting- choose a knot diagram at random from the finite set of diagrams with n crossings. We tabulate diagrams with 10 and fewer crossings and classify the diagrams by knot type, allowing us to compute exact probabilities for knots in this model. As expected, most diagrams with 1…

2015-12-17abs ↗pdf ↗

In this paper, we propose a novel hash learning approach that has the following main distinguishing features, when compared to past frameworks. First, the codewords are utilized in the Hamming space as ancillary techniques to accomplish its hash learning task. These codewords, which are inferred from the data, attempt …

2019-02-22abs ↗pdf ↗

This paper strengthens the central limit theorem for order statistics using relative entropy.

problem Establishing a stronger mode of convergence for central limit behavior of order statistics.
method Using relative entropy to ensure a stronger mode of convergence for central limit behavior of order statistics.
result An order O(1/n)O(1/\sqrt{n}) rate of convergence is established under mild conditions.

Quantum-assisted Gaussian process speeds up data regression.

problem High computational complexity of Gaussian process regression for large datasets.
method Quantum-assisted sparse Gaussian process regression using random Fourier features.
result Achieves polynomial-order computational speedup compared to classical methods.

Multiple machine learning and prediction models are often used for the same prediction or recommendation task. In our recent work, where we develop and deploy airline ancillary pricing models in an online setting, we found that among multiple pricing models developed, no one model clearly dominates other models for all…

2019-05-21abs ↗pdf ↗

Auptimizer simplifies hyperparameter tuning for machine learning models.

problem Difficulty and time-consuming hyperparameter tuning for machine learning models.
method General HPO framework that distributes computing resources and integrates various HPO techniques.
result Simplified model tuning and bookkeeping for data scientists.

This paper improves hierarchical community detection efficiency using local structural properties.

problem Efficiency of hierarchical community detection methods in large networks.
method Use of local structural network properties as proxies to improve efficiency.
result Achieves competitive results in modularity with improved efficiency.

The general theory of boundary value problems for linear elliptic wedge operators (on smooth manifolds with boundary) leads naturally, even in the scalar case, to the need to consider vector bundles over the boundary together with general smooth fiberwise multiplicative group actions. These actions, essentially trivial…

2013-01-24abs ↗pdf ↗

Rigidity theorem for critical points of Allen-Cahn equation on S³.

problem Rigidity of critical points with low Morse index on S³.
method Analysis of nullity and symmetries of critical points, Frankel-type theorem for nodal sets.
result Critical points with index five are symmetric and vanish on a Clifford torus, realizing the fifth width of the min-max spectrum.

The sharp and recent increase in the availability of data captured by different sensors combined with their considerably heterogeneous natures poses a serious challenge for the effective and efficient processing of remotely sensed data. Such an increase in remote sensing and ancillary datasets, however, opens up the po…

2018-12-19abs ↗pdf ↗

TeLeS improves ASR confidence estimation by considering temporal alignment and lexical errors.

problem Inaccurate confidence scores from E2E ASR models, especially for overconfident predictions.
method Proposes TeLeS, a novel confidence score that considers temporal alignment and lexical errors, and uses shrinkage loss to handle data imbalance.
result TeLeS generalizes well across different languages and ASR models, leading to significant WER reduction.

This paper investigates the approximation power of three types of random neural networks: (a) infinite width networks, with weights following an arbitrary distribution; (b) finite width networks obtained by subsampling the preceding infinite width networks; (c) finite width networks obtained by starting with standard G…

2019-06-18abs ↗pdf ↗

Bayesian transfer learning improves predictive performance with limited source data.

problem Improving statistical procedures with limited target and source datasets.
method Total risk prior for joint parameter distribution, Bayesian Lasso, model averaging, Gibbs sampling.
result Superior predictive performance compared to frequentist baseline, especially with limited source data.

The problem of finding a reduced dimensionality representation of categorical variables while preserving their most relevant characteristics is fundamental for the analysis of complex data. Specifically, given a co-occurrence matrix of two variables, one often seeks a compact representation of one variable which preser…

2012-10-19abs ↗pdf ↗

Study uses SAR data to estimate forest vegetation indices, improving monitoring of temperate forests.

problem Limitations of optical satellite data in monitoring forest ecosystems, especially due to atmospheric effects.
method Estimating four vegetation indices (LAI, FAPAR, EVI, NDVI) using multitemporal Sentinel-1 SAR and ancillary data.
result Accurate estimation of forest vegetation indices using SAR data, achieving high R2R^2 and low MAE values.

This paper operationalizes the Exponential Mechanism using Normalizing Flows for private optimization.

problem Improving privacy in machine learning while maintaining accuracy and efficiency.
method Using Normalizing Flows to approximate sampling from the Exponential Mechanism for private optimization.
result ExpM+NF provides more privacy than non-private SGD but not as much as DPSGD.

Quantum algorithm solves financial option pricing using Hamiltonian simulation.

problem Efficiently solving the Black-Scholes equation for option pricing dynamics.
method Mapped Black-Scholes equation to Schrödinger equation, used efficient Hamiltonian simulation techniques.
result Quantum algorithm shows feasible approach for solving financial derivatives on a quantum computer.

Study examines barriers to grid-connected battery systems in Spain, finding high cycle cost remains main obstacle.

problem Barriers to grid-connected battery systems in Spain's deregulated electricity market.
method Utilization analysis and concept of 'potentially profitable utilization time' introduced.
result High cycle cost remains the main barrier for grid-connected battery systems in Spain.

New method uses rank-conditioned Horvitz-Thompson estimation for unbiased sample reuse in Plackett-Luce best-of-K objective.

problem Estimating the expected maximum reward in Plackett-Luce draws without replacement.
method Rank-conditioned Horvitz-Thompson estimation with joint-score REINFORCE for unbiased sample reuse.
result Unbiased estimation of the Plackett-Luce best-of-K objective with finite second moment guarantees.

Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…

2017-12-12abs ↗pdf ↗

Prevents sensitive data generation in diffusion models using labeled and unlabeled data.

problem Generating sensitive data in diffusion models using unlabeled data.
method Positive-Unlabeled Diffusion Models, approximating ELBO with labeled and unlabeled data.
result Prevents the generation of sensitive data without compromising image quality.

Study reveals Data Shapley's inconsistent performance in data selection tasks.

problem Inconsistency of Data Shapley's performance in data selection across different settings.
method Hypothesis testing framework and identification of utility functions.
result Data Shapley's performance is no better than random selection without specific constraints.

PRRO generates synthetic tabular data that improves SL performance and class distribution.

problem Low SL utility of synthetic data due to class imbalance and overlooked data relationships.
method Data pruning and column reordering to optimize SL utility.
result Synthetic data generated with PRRO enhances predictive performance and class distribution.

Defines data science as a natural ecosystem with challenges and missions.

problem Challenges and missions in data science due to 5D complexities and data life cycle phases.
method Systemic and data-centric view of data science as a fusion of data universe and its challenges, formalizing a general-purpose architecture.
result Essential data science as a natural ecosystem integrating specific disciplines and high-impact applications.

This paper introduces C-DSL to improve data mining outcomes by considering context.

problem Data collection ambiguities, data imbalance, hidden biases, lack of domain info, and data incompleteness.
method Developed Context-Driven Data Science Lifecycle (C-DSL) to address data quality issues.
result Tangible improvements to data mining outcomes were achieved through C-DSL.

Proposes using probabilistic models for privacy-preserving synthetic data.

problem Designing high-quality synthetic data for privacy preservation.
method Formulate the problem through probabilistic modelling, choosing a model for the data.
result Statistical discoveries can be reliably reproduced from synthetic data.

Paper creates fair synthetic data ensuring equal predictions across sensitive attributes.

problem Ensuring fair predictions across sensitive attributes in synthetic data.
method Equalizing target probability distributions across sensitive attributes in synthetic data generation.
result Synthetic data provides strong fair predictions, equal across all thresholds.