Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

306090120 · Jun 202019922001200920172026
48 results for Patent quality

In this work, we focus on fine-tuning an OpenAI GPT-2 pre-trained model for generating patent claims. GPT-2 has demonstrated impressive efficacy of pre-trained language models on various tasks, particularly coherent text generation. Patent claim language itself has rarely been explored in the past and poses a unique ch…

2019-07-01abs ↗pdf ↗

Paper introduces ML for rare-event prediction in patent quality estimation.

problem Lack of predictive modeling in econ, management, tech forecasting.
method Introduces ML approach for optimizing predictive performance.
result Demonstrates synergy between ML and inferential statistics.

Improved AI patent classifier measures U.S. and China's AI patenting.

problem Measuring AI patents with high precision and generalization.
method Fine-tuning PatentSBERTa on manually labeled data from USPTO's AI Patent Dataset.
result Rapid growth in AI patenting in both countries, but different organizational patterns.

Patent lawsuits are costly and time-consuming. An ability to forecast a patent litigation and time to litigation allows companies to better allocate budget and time in managing their patent portfolios. We develop predictive models for estimating the likelihood of litigation for patents and the expected time to litigati…

2016-03-23abs ↗pdf ↗

Estimates yearly improvement rates for nearly all technologies using US patent data.

problem Providing a comprehensive account of technological change rates.
method Mapping patents to technology domains, calculating average centrality, and estimating improvement rates.
result Variation in improvement rates from 1.9% to 228.8% per year, with software-based domains often the fastest.

Study predicts startup outcomes like funding, patenting, IPOs using machine learning.

problem Forecasting startup success metrics like funding, patenting, IPOs.
method Developed interpretable machine learning framework, used preprocessing, class imbalance handling, and compared multiple models.
result Achieved high AUROC values for patent, funding, and exit predictions.

Automatic measurement of semantic text similarity is an important task in natural language processing. In this paper, we evaluate the performance of different vector space models to perform this task. We address the real-world problem of modeling patent-to-patent similarity and compare TFIDF (and related extensions), t…

2018-09-24abs ↗pdf ↗

We describe a number of devices for pulling candy, called taffy pullers,that are related to pseudo-Anosov maps of punctured spheres. Though the mathematical connection has long been known for the two most common taffy puller models, we unearth a rich variety of early designs from the patent literature, and introduce a …

2016-07-30abs ↗pdf ↗

DREAM model improves computational efficiency for non-linear effects in relational event models.

problem Efficiently modeling non-linear effects in dynamic relational networks.
method Introduces Deep Relational Event Additive Model (DREAM) using Neural Additive Models.
result Demonstrates superior computational efficiency compared to traditional REM approaches.

We study the relationship between firms' performance and their technological portfolios using tools borrowed from the complexity science. In particular, we ask whether the accumulation of knowledge and capabilities related to a coherent set of technologies leads firms to experience advantages in terms of productive eff…

2017-07-07abs ↗pdf ↗

New algorithm classifies and generates genomic sequences using RG-flow categorifier.

problem Classifying and generating genomic sequences for disease prediction.
method RG-flow based categorifier combining quantum field theory, holographic duality, and neural ODEs.
result RG categorifier can classify and generate new sequences from genomic data.

The paper analyzes tech specialization and diversification at various scales.

problem Trade-offs between specialization and diversification in economic development.
method Patent data and Economic Complexity framework.
result Technological Coherence positively impacts growth at metropolitan areas but negatively at larger scales.

We perform an optimal localization of asymptotically flat initial data sets and construct data that have positive ADM mass but are exactly trivial outside a cone of arbitrarily small aperture. The gluing scheme that we develop allows to produce a new class of NN-body solutions for the Einstein equation, which patently…

2014-07-17abs ↗pdf ↗

Breakthrough discoveries and inventions involve unexpected combinations of contents including problems, methods, and natural entities, and also diverse contexts such as journals, subfields, and conferences. Drawing on data from tens of millions of research papers, patents, and researchers, we construct models that pred…

2019-10-18abs ↗pdf ↗

GSR optimizes tasks in scientific workflows, improving performance across diverse applications.

problem Uncertainty in task selection and evaluation in scientific workflow optimization.
method Generate-Select-Refine (GSR) framework that alternates between task generation and optimization.
result GSR outperforms existing LLM-based optimizers in various scientific applications.

Network embedding is a method to learn low-dimensional representation vectors for nodes in complex networks. In real networks, nodes may have multiple tags but existing methods ignore the abundant semantic and hierarchical information of tags. This information is useful to many network applications and usually very sta…

2019-04-19abs ↗pdf ↗

FinTech negatively impacts Chinese banks' financial sustainability.

problem Impact of FinTech on financial sustainability of Chinese commercial banks.
method Three-stage network DEA-Malmquist model and two-way fixed effects model.
result FinTech primarily undermines financial sustainability by eroding loan efficiency and profitability.

Technological change and innovation are vitally important, especially for high-tech companies. However, factors influencing their future research and development (R&D) trends are both complicated and various, leading it a quite difficult task to make technology tracing for high-tech companies. To this end, in this pape…

2020-01-02abs ↗pdf ↗

Unlike other industries in which intellectual property is patentable, the financial industry relies on trade secrecy to protect its business processes and methods, which can obscure critical financial risk exposures from regulators and the public. We develop methods for sharing and aggregating such risk exposures that …

2011-11-19abs ↗pdf ↗

Understanding cities is central to addressing major global challenges from climate and health to economic resilience. Although increasingly perceived as fundamental socio-economic units, the detailed fabric of urban economic activities is only now accessible to comprehensive analyses with the availability of large data…

2014-05-13abs ↗pdf ↗

The availability of large idea repositories (e.g., the U.S. patent database) could significantly accelerate innovation and discovery by providing people with inspiration from solutions to analogous problems. However, finding useful analogies in these large, messy, real-world repositories remains a persistent challenge …

2017-06-17abs ↗pdf ↗

AirRL uses RL to infer urban air quality from selected stations.

problem Inferring fine-grained urban air quality from limited monitoring stations.
method Reinforcement learning model with a dynamic station selector and air quality regressor.
result AirRL achieves highest performance in air quality inference experiments.

Study on sparse recovery with mixed-quality data, establishing sample-size conditions.

problem Sparse recovery with heterogeneous noise from high- and low-quality sources.
method Establishes linear trade-off for sufficient conditions, analyzes LASSO algorithm.
result Linear trade-off for sufficient conditions, robustness of LASSO to data heterogeneity.

Protein structure prediction has been a grand challenge problem in the structure biology over the last few decades. Protein quality assessment plays a very important role in protein structure prediction. In the paper, we propose a new protein quality assessment method which can predict both local and global quality of …

2016-02-13abs ↗pdf ↗

Framework improves ML performance by identifying high-quality data.

problem Poor data quality hampers ML performance.
method Intelligent data-centric evaluation framework combining quality measurements and unsupervised learning.
result Framework improves ML system performance in real-world use case.

This note investigates the causes of the quality anomaly, which is one of the strongest and most scalable anomalies in equity markets. We explore two potential explanations. The "risk view", whereby investing in high quality firms is somehow riskier, so that the higher returns of a quality portfolio are a compensation …

2016-01-18abs ↗pdf ↗

Do-AIQ framework evaluates AI algorithms' quality using DOE.

problem Quality evaluation of AI mislabel detection algorithms.
method Design-of-experiment approach with high-dimensional constraint space design and surrogate modeling.
result Established framework for evaluating AI algorithm quality robustly.

Framework evaluates quality of synthetic data generated with differential privacy.

problem Ensuring synthetic data retains statistical quality after applying differential privacy.
method Developed a framework to evaluate synthetic data quality from a practical researcher's viewpoint.
result Synthetic data can be evaluated against training data or underlying populations, and for specific tasks like inference or prediction.

MRI image quality affects statistical and predictive analysis of brain morphology.

problem Impact of MRI image quality on statistical and predictive analysis of brain morphology.
method Systematic testing of image quality on univariate statistics and machine learning classification using three large datasets.
result Low-quality MRI data significantly affects detecting significant sex/gender differences in smaller samples, but not in larger ones.