The paper proposes a novel model to forecast patent citations using multi-attention recurrent networks.
problem Forecasting forward citations to patents to discover emerging technologies.
method The approach employs a sequence-to-sequence model with an attention-of-attention mechanism to capture dependencies in multiple time sequences.
result The proposed model outperforms state-of-the-art models in forward citation forecasting.
Improved AI patent classifier measures U.S. and China's AI patenting.
problem Measuring AI patents with high precision and generalization.
method Fine-tuning PatentSBERTa on manually labeled data from USPTO's AI Patent Dataset.
result Rapid growth in AI patenting in both countries, but different organizational patterns.
The paper models network formation using mixed logit models.
problem Modeling network formation in various fields.
method Mixed logit models, specifically the repeated-choice (RC) model.
result The RC model outperforms the multinomial logit (MNL) model in estimating network formation.
DREAM model improves computational efficiency for non-linear effects in relational event models.
problem Efficiently modeling non-linear effects in dynamic relational networks.
method Introduces Deep Relational Event Additive Model (DREAM) using Neural Additive Models.
result Demonstrates superior computational efficiency compared to traditional REM approaches.
Improved patent classification using fine-tuned BERT model.
problem Classifying large patent datasets efficiently and accurately.
method Fine-tuning a pre-trained BERT model on patent claims.
result Outperforms state-of-the-art methods by 20%.
This work fine-tunes GPT-2 for generating patent claims.
problem Generating coherent patent claims automatically.
method Fine-tuning OpenAI GPT-2 on patent claim language structure.
result Demonstrated the first machine-generated patent claims.
Surprise predicts breakthroughs in science and technology.
problem Predicting breakthroughs in science and technology.
method Using embeddings from high-dimensional stochastic block models, predicting combinations with AUC of 95%.
result Breakthroughs often occur when problems from one field are solved by researchers from another field.
BERT learns claim descriptions to identify patent novelty.
problem Identifying novel patent claims among existing documents.
method Training BERT on concatenated claims and descriptions, scoring BERT's output.
result BERT identifies relevant X documents for patent novelty.
Patent lawsuits are costly and time-consuming. An ability to forecast a patent litigation and time to litigation allows companies to better allocate budget and time in managing their patent portfolios. We develop predictive models for estimating the likelihood of litigation for patents and the expected time to litigati…
Estimates yearly improvement rates for nearly all technologies using US patent data.
problem Providing a comprehensive account of technological change rates.
method Mapping patents to technology domains, calculating average centrality, and estimating improvement rates.
result Variation in improvement rates from 1.9% to 228.8% per year, with software-based domains often the fastest.
Study predicts startup outcomes like funding, patenting, IPOs using machine learning.
problem Forecasting startup success metrics like funding, patenting, IPOs.
method Developed interpretable machine learning framework, used preprocessing, class imbalance handling, and compared multiple models.
result Achieved high AUROC values for patent, funding, and exit predictions.
Improved antibody humanness prediction using patent data.
problem Predicting the immunogenic response of antibody therapeutics.
method Multi-stage, multi-loss training process with contrastive learning and cross-entropy loss.
result The model outperforms baselines on five out of six inference tasks, improving humanness prediction.
This paper explores using NFTs for patents, offering a framework and addressing challenges.
problem Lack of research in applying NFT to intellectual property, especially patents.
method Developed a layered conceptual NFT-based patent framework.
result Promotes transparency and liquidity in patent markets.
Mass algorithm predicts M&A deals from patent data.
problem Hard to automatically predict M&A deals due to complex human skills.
method Machine learning-based similarity measure (MASS) applied to patent data.
result MASS outperforms LightGCN in forecasting M&A deals.
Bayesian model for discrete data with conditional transformations.
problem Handling discrete ordinal and count data with excess zeros.
method Bayesian framework with conditional transformation functions and modular MCMC algorithm.
result Flexible modeling of linear and nonlinear covariate effects for ordinal and count data.
Dolby has the best financial health, but competition for patents could create jobs.
problem Comparing stock valuation of companies using financial metrics.
method Analysis of financial statements over three years.
result Dolby has stable profit margins and generates billions in revenue.
SentiCite analyzes citations for sentiment and nature, improving on existing methods.
problem Identifying quality scientific work amidst many citations.
method Sentiment analysis of citations with motivation detection.
result SentiCite outperforms state-of-the-art methods with a F1-measure of 0.71.
A multi-task model tackles citation purpose classification with limited data.
problem Classifying citations based on their purpose is challenging due to limited labeled data and subjectivity.
method Combines linguistic features, TF-IDF, and an LSTM-with-attention model for multi-task learning.
result Improves classification accuracy compared to single-task models.
Measuring the impact of scientific articles is important for evaluating the research output of individual scientists, academic institutions and journals. While citations are raw data for constructing impact measures, there exist biases and potential issues if factors affecting citation patterns are not properly account…
Advances citation and subject label recommendation using multi-modal adversarial autoencoders.
problem Improving recommendation systems for citations and subject labels.
method Multi-modal adversarial autoencoders with adversarial regularization, sparsity, and input modality analysis.
result Adversarial regularization consistently improves recommendation performance.
Synthetic reference strings are as effective as real ones for training citation parsing models.
problem Lack of training data for citation parsing, especially with deep neural networks.
method Trained Grobid with human-labelled and synthetically created reference strings, and evaluated retraining and out-of-sample data impact.
result Synthetic and real reference strings are equally effective for training Grobid, with retraining improving performance.
Objectives: Text categorization has been used in biomedical informatics for identifying documents containing relevant topics of interest. We developed a simple method that uses a chi-square-based scoring function to determine the likelihood of MEDLINE citations containing genetic relevant topic. Methods: Our procedure …
Paper introduces ML for rare-event prediction in patent quality estimation.
problem Lack of predictive modeling in econ, management, tech forecasting.
method Introduces ML approach for optimizing predictive performance.
result Demonstrates synergy between ML and inferential statistics.
A combined model integrates latent factor and logistic regression for citation network analysis.
problem Insufficient representation by either latent factor or logistic regression alone.
method Proposes a combined model integrating latent factor and logistic regression, with parameter estimation through joint-likelihood and penalty terms.
result The proposed method captures both main technological trends and ad-hoc dependencies in citation networks.
Although information extraction and coreference resolution appear together in many applications, most current systems perform them as ndependent steps. This paper describes an approach to integrated inference for extraction and coreference based on conditionally-trained undirected graphical models. We discuss the advan…
Automatic measurement of semantic text similarity is an important task in natural language processing. In this paper, we evaluate the performance of different vector space models to perform this task. We address the real-world problem of modeling patent-to-patent similarity and compare TFIDF (and related extensions), t…
Corrects errors in Hans' pseudocovering spaces paper.
problem Mathematical errors and citation issues in Hans' pseudocovering spaces paper.
method Identifies and corrects errors in Hans' previous work.
result Addresses mathematical and citation errors in Hans' pseudocovering spaces paper.
Paper develops an attention mechanism for long-term scientific impact prediction.
problem Predicting the long-term impact of scientific papers based on citation records.
method Develops an attention mechanism to predict long-term scientific impact.
result Emphasizing the limited attention can better stand on the shoulders of giants.
This paper asks, "Do classics exist in megaproject management?" We identify three types of classic texts: conventional, Kuhnian, and citation classics. We find that the answer to our question depends on the definition of "classic" employed. First, "citation classics" do exist in megaproject management, and they perform…
We describe a number of devices for pulling candy, called taffy pullers,that are related to pseudo-Anosov maps of punctured spheres. Though the mathematical connection has long been known for the two most common taffy puller models, we unearth a rich variety of early designs from the patent literature, and introduce a …
Estimates citation impact to recommend best publication venue.
problem Choosing optimal publication venue for academic papers.
method Treatment effect estimation and bias correction method.
result Effective recommendation of publication venues based on citation potential.
Bibliographic analysis considers author's research areas, the citation network and paper content among other things. In this paper, we combine these three in a topic model that produces a bibliographic model of authors, topics and documents using a non-parametric extension of a combination of the Poisson mixed-topic li…
Bibliographic analysis considers the author's research areas, the citation network and the paper content among other things. In this paper, we combine these three in a topic model that produces a bibliographic model of authors, topics and documents, using a nonparametric extension of a combination of the Poisson mixed-…
Develops a deep learning framework to predict future tech directions for high-tech companies.
problem Difficult task in predicting future R&D trends for high-tech companies due to complexity and variety of factors.
method Deep Technology Forecasting (DTF) framework with three components: PCR, CTR, and DTT neural network.
result DTF framework precisely predicts future tech emphasis of companies using hybrid factors.
We correct a mistake on the citation of JSJ theory in \cite{Ni}. Some arguments in \cite{Ni} are also slightly modified accordingly.
Novel graph network learns hierarchical network structure.
problem Lack of information in hierarchical network topology.
method Hierarchical clustering for multiscale decomposition, graph convolutional layers.
result Competitive performance on citation network benchmark.
We study the relationship between firms' performance and their technological portfolios using tools borrowed from the complexity science. In particular, we ask whether the accumulation of knowledge and capabilities related to a coherent set of technologies leads firms to experience advantages in terms of productive eff…
Visuals in scientific papers are used to express complex ideas; this study uses them to identify knowledge domains.
problem Scientific figures are underutilized in literature analysis.
method Encoded scientific figures into visual signatures and used distances between signatures to compare communities of practice.
result Figures can differentiate knowledge domains as effectively as text or citation patterns.
New algorithm classifies and generates genomic sequences using RG-flow categorifier.
problem Classifying and generating genomic sequences for disease prediction.
method RG-flow based categorifier combining quantum field theory, holographic duality, and neural ODEs.
result RG categorifier can classify and generate new sequences from genomic data.
System identifies high impact research from academic papers.
problem Identifying high impact research amidst a growing volume of papers.
method Combines visual and text classifiers on a large dataset of PDFs and citation counts.
result Improved accuracy in predicting high impact research.
Public data boosts MICCAI papers' citations, but sharing remains low.
problem Slow adoption of public benchmarks in medical image computing.
method Analysis of MICCAI papers from 2014-2018.
result Public data boosts citations, but sharing is low.
The paper analyzes tech specialization and diversification at various scales.
problem Trade-offs between specialization and diversification in economic development.
method Patent data and Economic Complexity framework.
result Technological Coherence positively impacts growth at metropolitan areas but negatively at larger scales.
We perform an optimal localization of asymptotically flat initial data sets and construct data that have positive ADM mass but are exactly trivial outside a cone of arbitrarily small aperture. The gluing scheme that we develop allows to produce a new class of N-body solutions for the Einstein equation, which patently…
Higher-order optimization problems naturally appear when investigating the effects of a patent with finite length, as in the pioneering work of Futagami and Iwaisako (2007). In this paper, we establish the Euler equations and transversality conditions necessary for analyzing such higher-order optimization problems. We …
NoLBERT avoids lookback and lookahead biases for better econometric inference.
problem Information leakage in language models affects econometric inference.
method Pretrained on text from 1976-1995, avoiding lookback and lookahead biases.
result NoLBERT outperforms domain-specific baselines and predicts higher profit growth.
New machine learning model faster, more accurate, and can identify hard-to-classify samples.
problem Improving classification models in machine learning.
method Quadratic Multiform Separation approach.
result Produces comparable predictive accuracy, runs faster, and identifies hard-to-classify samples.
A novel model-selection method for dynamic networks using synthetic data.
problem Classifying and understanding the growth mechanisms of dynamic networks.
method Training a classifier on synthetic network data generated by nine random graph models, using dynamic features that count new links.
result Achieves near-perfect classification of synthetic networks, outperforming state-of-the-art methods.
Paper introduces a core-periphery model for identifying informative network structures.
problem Noise and bias in non-informative periphery structures obscure the informative core in complex networks.
method Spectral algorithms for core identification as a preprocessing step for network analysis.
result The proposed method outperforms traditional core-periphery methods in various downstream tasks.