Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,982 papers · 148 categories

Trend · papers per month

1345 · Aug 201919922001200920172026
48 results for patent citations

The paper proposes a novel model to forecast patent citations using multi-attention recurrent networks.

problem Forecasting forward citations to patents to discover emerging technologies.
method The approach employs a sequence-to-sequence model with an attention-of-attention mechanism to capture dependencies in multiple time sequences.
result The proposed model outperforms state-of-the-art models in forward citation forecasting.

Improved AI patent classifier measures U.S. and China's AI patenting.

problem Measuring AI patents with high precision and generalization.
method Fine-tuning PatentSBERTa on manually labeled data from USPTO's AI Patent Dataset.
result Rapid growth in AI patenting in both countries, but different organizational patterns.

DREAM model improves computational efficiency for non-linear effects in relational event models.

problem Efficiently modeling non-linear effects in dynamic relational networks.
method Introduces Deep Relational Event Additive Model (DREAM) using Neural Additive Models.
result Demonstrates superior computational efficiency compared to traditional REM approaches.

Patent lawsuits are costly and time-consuming. An ability to forecast a patent litigation and time to litigation allows companies to better allocate budget and time in managing their patent portfolios. We develop predictive models for estimating the likelihood of litigation for patents and the expected time to litigati…

2016-03-23abs ↗pdf ↗

Estimates yearly improvement rates for nearly all technologies using US patent data.

problem Providing a comprehensive account of technological change rates.
method Mapping patents to technology domains, calculating average centrality, and estimating improvement rates.
result Variation in improvement rates from 1.9% to 228.8% per year, with software-based domains often the fastest.

Study predicts startup outcomes like funding, patenting, IPOs using machine learning.

problem Forecasting startup success metrics like funding, patenting, IPOs.
method Developed interpretable machine learning framework, used preprocessing, class imbalance handling, and compared multiple models.
result Achieved high AUROC values for patent, funding, and exit predictions.

A multi-task model tackles citation purpose classification with limited data.

problem Classifying citations based on their purpose is challenging due to limited labeled data and subjectivity.
method Combines linguistic features, TF-IDF, and an LSTM-with-attention model for multi-task learning.
result Improves classification accuracy compared to single-task models.

Measuring the impact of scientific articles is important for evaluating the research output of individual scientists, academic institutions and journals. While citations are raw data for constructing impact measures, there exist biases and potential issues if factors affecting citation patterns are not properly account…

2015-02-25abs ↗pdf ↗

Advances citation and subject label recommendation using multi-modal adversarial autoencoders.

problem Improving recommendation systems for citations and subject labels.
method Multi-modal adversarial autoencoders with adversarial regularization, sparsity, and input modality analysis.
result Adversarial regularization consistently improves recommendation performance.

Synthetic reference strings are as effective as real ones for training citation parsing models.

problem Lack of training data for citation parsing, especially with deep neural networks.
method Trained Grobid with human-labelled and synthetically created reference strings, and evaluated retraining and out-of-sample data impact.
result Synthetic and real reference strings are equally effective for training Grobid, with retraining improving performance.

Paper introduces ML for rare-event prediction in patent quality estimation.

problem Lack of predictive modeling in econ, management, tech forecasting.
method Introduces ML approach for optimizing predictive performance.
result Demonstrates synergy between ML and inferential statistics.

A combined model integrates latent factor and logistic regression for citation network analysis.

problem Insufficient representation by either latent factor or logistic regression alone.
method Proposes a combined model integrating latent factor and logistic regression, with parameter estimation through joint-likelihood and penalty terms.
result The proposed method captures both main technological trends and ad-hoc dependencies in citation networks.

Automatic measurement of semantic text similarity is an important task in natural language processing. In this paper, we evaluate the performance of different vector space models to perform this task. We address the real-world problem of modeling patent-to-patent similarity and compare TFIDF (and related extensions), t…

2018-09-24abs ↗pdf ↗

This paper asks, "Do classics exist in megaproject management?" We identify three types of classic texts: conventional, Kuhnian, and citation classics. We find that the answer to our question depends on the definition of "classic" employed. First, "citation classics" do exist in megaproject management, and they perform…

2017-09-06abs ↗pdf ↗

We describe a number of devices for pulling candy, called taffy pullers,that are related to pseudo-Anosov maps of punctured spheres. Though the mathematical connection has long been known for the two most common taffy puller models, we unearth a rich variety of early designs from the patent literature, and introduce a …

2016-07-30abs ↗pdf ↗

Bibliographic analysis considers author's research areas, the citation network and paper content among other things. In this paper, we combine these three in a topic model that produces a bibliographic model of authors, topics and documents using a non-parametric extension of a combination of the Poisson mixed-topic li…

2016-09-22abs ↗pdf ↗

Develops a deep learning framework to predict future tech directions for high-tech companies.

problem Difficult task in predicting future R&D trends for high-tech companies due to complexity and variety of factors.
method Deep Technology Forecasting (DTF) framework with three components: PCR, CTR, and DTT neural network.
result DTF framework precisely predicts future tech emphasis of companies using hybrid factors.

We study the relationship between firms' performance and their technological portfolios using tools borrowed from the complexity science. In particular, we ask whether the accumulation of knowledge and capabilities related to a coherent set of technologies leads firms to experience advantages in terms of productive eff…

2017-07-07abs ↗pdf ↗

Visuals in scientific papers are used to express complex ideas; this study uses them to identify knowledge domains.

problem Scientific figures are underutilized in literature analysis.
method Encoded scientific figures into visual signatures and used distances between signatures to compare communities of practice.
result Figures can differentiate knowledge domains as effectively as text or citation patterns.

New algorithm classifies and generates genomic sequences using RG-flow categorifier.

problem Classifying and generating genomic sequences for disease prediction.
method RG-flow based categorifier combining quantum field theory, holographic duality, and neural ODEs.
result RG categorifier can classify and generate new sequences from genomic data.

The paper analyzes tech specialization and diversification at various scales.

problem Trade-offs between specialization and diversification in economic development.
method Patent data and Economic Complexity framework.
result Technological Coherence positively impacts growth at metropolitan areas but negatively at larger scales.

We perform an optimal localization of asymptotically flat initial data sets and construct data that have positive ADM mass but are exactly trivial outside a cone of arbitrarily small aperture. The gluing scheme that we develop allows to produce a new class of NN-body solutions for the Einstein equation, which patently…

2014-07-17abs ↗pdf ↗

A novel model-selection method for dynamic networks using synthetic data.

problem Classifying and understanding the growth mechanisms of dynamic networks.
method Training a classifier on synthetic network data generated by nine random graph models, using dynamic features that count new links.
result Achieves near-perfect classification of synthetic networks, outperforming state-of-the-art methods.

Paper introduces a core-periphery model for identifying informative network structures.

problem Noise and bias in non-informative periphery structures obscure the informative core in complex networks.
method Spectral algorithms for core identification as a preprocessing step for network analysis.
result The proposed method outperforms traditional core-periphery methods in various downstream tasks.