Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

2935868791,172 · Jun 202019922001200920182026
48 results for logged data

Dividing deep learning models for consistent anomaly detection in changing log data.

problem Anomaly detection methods fail when log data types change, leading to false negatives.
method Divide deep learning models based on log data correlation and extract correlations.
result Continues anomaly detection accuracy even when log data changes.

The paper proposes a SeqGAN model to generate balanced log messages for anomaly detection.

problem Imbalanced log data makes anomaly detection difficult.
method SeqGAN for generating balanced log messages, Autoencoder for feature extraction, GRU for anomaly detection.
result Oversampling and balancing data improves anomaly detection accuracy.

Logsy detects anomalies in logs using a novel classification-based approach.

problem Anomaly detection in unstructured logs is challenging due to limited model generalization.
method Logsy learns log representations by distinguishing normal and anomaly logs using a classification-based approach with an attention-based encoder and hyperspherical loss function.
result Logsy improves anomaly detection performance by 0.25 in F1 score compared to previous methods.

The study analyzes gaps in well logs to improve prediction models.

problem Improving prediction of missing data in well logs for oil exploration.
method Descriptive analysis, generation of artificial gaps, comparison of machine learning algorithms.
result Artificial Neural Networks, Random Forests, and Linear Regression algorithms were compared in predicting missing data.

Introduces CSLC models to bridge deep generative models and classical algorithms.

problem Mode collapse and memorization issues in deep generative models and restrictive assumptions in classical algorithms.
method Introduces conditionally strongly log-concave (CSLC) models, factorizing data distribution into strongly log-concave conditional distributions.
result Efficient parameter estimation and sampling algorithms with theoretical guarantees for non-log-concave data distributions.

We explain SSL objectives as log-likelihoods in a data curation model.

problem Lack of understanding of SSL objectives as log-likelihoods.
method Formulate SSL objectives as a log-likelihood in a generative model of data curation.
result SSL methods can be understood as lower-bounds on a principled log-likelihood.

We speed up inference in log-linear models by using random perturbations and pre-computed data structures.

problem Inference in log-linear models is slow and impractical for large output spaces.
method Gumbel random variable perturbations and Maximum Inner Product Search data structure.
result Sublinear amortized inference and learning in log-linear models with provable guarantees.

A novel method for learning DAGs from positive-valued data.

problem Causal discovery from observational data of positive-valued variables.
method Hybrid Moment-Ratio Scoring (H-MRS) algorithm combining moment-based scoring and log-scale regression.
result H-MRS integrates log-scale Ridge regression for moment-ratio estimation with a greedy ordering procedure based on raw-scale moment ratios, followed by Elastic Net-based parent selection.

The paper proposes a new method for clustering survival data using smoothed log-hazard trajectories.

problem Clustering survival data based on instantaneous risk dynamics.
method Functional Principal Component Analysis applied to B-spline smoothed log-hazard trajectories.
result The proposed method provides an interpretable representation of relative temporal risk dynamics.

Paper proposes robust estimators for heavy-tailed data with infinite variance.

problem Developing robust estimators for heavy-tailed data with infinite variance.
method Proposes two robust estimators: ridge log-truncated M-estimator and elastic net log-truncated M-estimator.
result Demonstrates robustness of log-truncated estimations over standard estimations through simulations and real data analysis.

Model non-stationary financial data using log-normal distributions and Langevin equations.

problem Modeling non-stationary volume-price distributions in finance.
method Model non-stationary volume-price distributions with a log-normal distribution. Derive Langevin equations from the series of log-normal parameters.
result Reconstructed statistics of volume-price distributions fit well empirical data.

The paper evaluates machine learning cyber defenses using log data against adversarial attacks.

problem Evaluating the robustness of machine learning cyber defenses against adversarial attacks.
method Developed a testing framework using deep reinforcement learning and adversarial natural language processing.
result Higher dropout levels increase robustness, with 90% dropout probability showing the highest robustness.

Paper tackles parameter learning for log-supermodular models, improving on existing bounds.

problem Parameter estimation for log-supermodular models with intractable log-partition functions.
method Use stochastic subgradient technique to maximize a lower-bound on the log-likelihood, extending to conditional maximum likelihood.
result Perturb-and-MAP bound outperforms separable optimization bound for parameter estimation.

New method detects anomalies in computing centers' logs.

problem Anomaly detection in continuously changing log data for predictive maintenance.
method Evolving granular classifiers using Fuzzy-set-Based evolving Modeling and evolving Granular Neural Network.
result Classification model prioritizes maintenance based on anomaly severity.

This paper addresses measurement errors in high-dimensional compositional data using a log-contrast model calibration approach.

problem Measurement errors in high-dimensional regression models involving compositional covariates.
method Calibration approach for the linear log-contrast model under lenient sparsity conditions.
result Established asymptotic normality of the estimator for inference.

Develops an anytime-valid framework for optimal policy identification from logged contextual bandit data.

problem Selecting the optimal policy from a candidate policy class while monitoring evidence continuously.
method Constructs a time-indexed set that retains the true optimal policy set uniformly over time.
result The procedure allows the analyst to monitor policy values, eliminate clearly suboptimal policies, and stop at data-dependent times without invalidating inference.

Semi-supervised GANs with log-signatures improve credit card fraud detection.

problem Detecting fraud in large, complex financial transaction data streams.
method Conditional GANs with Bayesian inference and log-signatures for robust feature encoding.
result Consistent improvements over benchmarks in global and domain-specific metrics.

Improved sampling for diffusion models and log-concave distributions.

problem Efficient sampling for diffusion models and log-concave distributions.
method Algorithms for sampling with δδ-error in polylog(1/δ)\mathrm{polylog}(1/δ) steps using accurate score estimates.
result Exponential improvement in complexity over previous results.

Improves SGM convergence bounds in W2-distance without strict assumptions.

problem Convergence bounds for SGMs in W2-distance require stringent assumptions.
method Novel framework using the OU process and PDE analysis.
result Log-concavity evolves from weak to strong over time.

Paper proposes a log-domain training method to reduce neural network complexity.

problem High computational complexity in training deep neural networks limits real-time training.
method End-to-end training and inference scheme using approximate logarithmic operations in the log-domain.
result 16-bit log-based training achieves within 1% accuracy of floating-point baselines.

Combines offline causal inference and online bandit learning for better decision-making.

problem Making adaptive decisions using both logged and streaming data to avoid user harm.
method Unified offline causal inference and online learning algorithms, deriving bounds on decision accuracy.
result First upper regret bound for forest-based online bandit algorithms.

A new method combines online and offline learning to tackle contextual bandits with missing action support.

problem Learning optimal policies with logged data when the logging policy has deficient support.
method Hybrid approach using online exploration to exploit supported actions and offline learning to avoid unnecessary explorations.
result Determines an optimal policy with theoretical guarantees using minimal online explorations.

We clarify the status of log-periodicity associated with speculative bubbles preceding financial crashes. In particular, we address Feigenbaum's [2001] criticism and show how it can be rebuked. Feigenbaum's main result is as follows: ``the hypothesis that the log-periodic component is present in the data cannot be reje…

2001-06-26abs ↗pdf ↗

Log-linear models are the popular workhorses of analyzing contingency tables. A log-linear parameterization of an interaction model can be more expressive than a direct parameterization based on probabilities, leading to a powerful way of defining restrictions derived from marginal, conditional and context-specific ind…

2014-09-09abs ↗pdf ↗