Improved patent classification using fine-tuned BERT model.
problem Classifying large patent datasets efficiently and accurately.
method Fine-tuning a pre-trained BERT model on patent claims.
result Outperforms state-of-the-art methods by 20%.
Estimates yearly improvement rates for nearly all technologies using US patent data.
problem Providing a comprehensive account of technological change rates.
method Mapping patents to technology domains, calculating average centrality, and estimating improvement rates.
result Variation in improvement rates from 1.9% to 228.8% per year, with software-based domains often the fastest.
Improved AI patent classifier measures U.S. and China's AI patenting.
problem Measuring AI patents with high precision and generalization.
method Fine-tuning PatentSBERTa on manually labeled data from USPTO's AI Patent Dataset.
result Rapid growth in AI patenting in both countries, but different organizational patterns.
This work fine-tunes GPT-2 for generating patent claims.
problem Generating coherent patent claims automatically.
method Fine-tuning OpenAI GPT-2 on patent claim language structure.
result Demonstrated the first machine-generated patent claims.
The paper proposes a novel model to forecast patent citations using multi-attention recurrent networks.
problem Forecasting forward citations to patents to discover emerging technologies.
method The approach employs a sequence-to-sequence model with an attention-of-attention mechanism to capture dependencies in multiple time sequences.
result The proposed model outperforms state-of-the-art models in forward citation forecasting.
BERT learns claim descriptions to identify patent novelty.
problem Identifying novel patent claims among existing documents.
method Training BERT on concatenated claims and descriptions, scoring BERT's output.
result BERT identifies relevant X documents for patent novelty.
Patent lawsuits are costly and time-consuming. An ability to forecast a patent litigation and time to litigation allows companies to better allocate budget and time in managing their patent portfolios. We develop predictive models for estimating the likelihood of litigation for patents and the expected time to litigati…
Study predicts startup outcomes like funding, patenting, IPOs using machine learning.
problem Forecasting startup success metrics like funding, patenting, IPOs.
method Developed interpretable machine learning framework, used preprocessing, class imbalance handling, and compared multiple models.
result Achieved high AUROC values for patent, funding, and exit predictions.
Improved antibody humanness prediction using patent data.
problem Predicting the immunogenic response of antibody therapeutics.
method Multi-stage, multi-loss training process with contrastive learning and cross-entropy loss.
result The model outperforms baselines on five out of six inference tasks, improving humanness prediction.
Paper introduces ML for rare-event prediction in patent quality estimation.
problem Lack of predictive modeling in econ, management, tech forecasting.
method Introduces ML approach for optimizing predictive performance.
result Demonstrates synergy between ML and inferential statistics.
This paper explores using NFTs for patents, offering a framework and addressing challenges.
problem Lack of research in applying NFT to intellectual property, especially patents.
method Developed a layered conceptual NFT-based patent framework.
result Promotes transparency and liquidity in patent markets.
Mass algorithm predicts M&A deals from patent data.
problem Hard to automatically predict M&A deals due to complex human skills.
method Machine learning-based similarity measure (MASS) applied to patent data.
result MASS outperforms LightGCN in forecasting M&A deals.
Dolby has the best financial health, but competition for patents could create jobs.
problem Comparing stock valuation of companies using financial metrics.
method Analysis of financial statements over three years.
result Dolby has stable profit margins and generates billions in revenue.
New machine learning model faster, more accurate, and can identify hard-to-classify samples.
problem Improving classification models in machine learning.
method Quadratic Multiform Separation approach.
result Produces comparable predictive accuracy, runs faster, and identifies hard-to-classify samples.
The paper models network formation using mixed logit models.
problem Modeling network formation in various fields.
method Mixed logit models, specifically the repeated-choice (RC) model.
result The RC model outperforms the multinomial logit (MNL) model in estimating network formation.
New algorithm classifies and generates genomic sequences using RG-flow categorifier.
problem Classifying and generating genomic sequences for disease prediction.
method RG-flow based categorifier combining quantum field theory, holographic duality, and neural ODEs.
result RG categorifier can classify and generate new sequences from genomic data.
Automatic measurement of semantic text similarity is an important task in natural language processing. In this paper, we evaluate the performance of different vector space models to perform this task. We address the real-world problem of modeling patent-to-patent similarity and compare TFIDF (and related extensions), t…
We describe a number of devices for pulling candy, called taffy pullers,that are related to pseudo-Anosov maps of punctured spheres. Though the mathematical connection has long been known for the two most common taffy puller models, we unearth a rich variety of early designs from the patent literature, and introduce a …
Develops a deep learning framework to predict future tech directions for high-tech companies.
problem Difficult task in predicting future R&D trends for high-tech companies due to complexity and variety of factors.
method Deep Technology Forecasting (DTF) framework with three components: PCR, CTR, and DTT neural network.
result DTF framework precisely predicts future tech emphasis of companies using hybrid factors.
DREAM model improves computational efficiency for non-linear effects in relational event models.
problem Efficiently modeling non-linear effects in dynamic relational networks.
method Introduces Deep Relational Event Additive Model (DREAM) using Neural Additive Models.
result Demonstrates superior computational efficiency compared to traditional REM approaches.
We study the relationship between firms' performance and their technological portfolios using tools borrowed from the complexity science. In particular, we ask whether the accumulation of knowledge and capabilities related to a coherent set of technologies leads firms to experience advantages in terms of productive eff…
The paper analyzes tech specialization and diversification at various scales.
problem Trade-offs between specialization and diversification in economic development.
method Patent data and Economic Complexity framework.
result Technological Coherence positively impacts growth at metropolitan areas but negatively at larger scales.
We perform an optimal localization of asymptotically flat initial data sets and construct data that have positive ADM mass but are exactly trivial outside a cone of arbitrarily small aperture. The gluing scheme that we develop allows to produce a new class of N-body solutions for the Einstein equation, which patently…
Higher-order optimization problems naturally appear when investigating the effects of a patent with finite length, as in the pioneering work of Futagami and Iwaisako (2007). In this paper, we establish the Euler equations and transversality conditions necessary for analyzing such higher-order optimization problems. We …
NoLBERT avoids lookback and lookahead biases for better econometric inference.
problem Information leakage in language models affects econometric inference.
method Pretrained on text from 1976-1995, avoiding lookback and lookahead biases.
result NoLBERT outperforms domain-specific baselines and predicts higher profit growth.
GSR optimizes tasks in scientific workflows, improving performance across diverse applications.
problem Uncertainty in task selection and evaluation in scientific workflow optimization.
method Generate-Select-Refine (GSR) framework that alternates between task generation and optimization.
result GSR outperforms existing LLM-based optimizers in various scientific applications.
Surprise predicts breakthroughs in science and technology.
problem Predicting breakthroughs in science and technology.
method Using embeddings from high-dimensional stochastic block models, predicting combinations with AUC of 95%.
result Breakthroughs often occur when problems from one field are solved by researchers from another field.
We describe a fully data driven model that learns to perform a retrosynthetic reaction prediction task, which is treated as a sequence-to-sequence mapping problem. The end-to-end trained model has an encoder-decoder architecture that consists of two recurrent neural networks, which has previously shown great success in…
Model predicts diverse chemical reactions for target compounds.
problem Making generalizable and diverse retrosynthetic reaction predictions.
method Transformer architecture with novel pre-training methods and a latent variable model.
result Improves performance on USPTO-50k dataset, generating more diverse predictions.
FinTech negatively impacts Chinese banks' financial sustainability.
problem Impact of FinTech on financial sustainability of Chinese commercial banks.
method Three-stage network DEA-Malmquist model and two-way fixed effects model.
result FinTech primarily undermines financial sustainability by eroding loan efficiency and profitability.
Unlike other industries in which intellectual property is patentable, the financial industry relies on trade secrecy to protect its business processes and methods, which can obscure critical financial risk exposures from regulators and the public. We develop methods for sharing and aggregating such risk exposures that …
Tag2Vec learns tag representations in hybrid networks with semantic and hierarchical information.
problem Lack of semantic and hierarchical information in tag networks.
method Tag2Vec model that combines nodes and tags into hybrid networks, using parameterized random walks and hyperbolic Skip-gram model.
result Tag2Vec outperforms other models in learning rich semantic tag representations.
Understanding cities is central to addressing major global challenges from climate and health to economic resilience. Although increasingly perceived as fundamental socio-economic units, the detailed fabric of urban economic activities is only now accessible to comprehensive analyses with the availability of large data…
Framework quantifies financial NLP robustness under regime shifts.
problem Semantic and causal drift in financial news narratives.
method Four metrics: FCAS, PCS, TSV, NLICS.
result Transformer models are more affected by semantic drift.
Bayesian model for discrete data with conditional transformations.
problem Handling discrete ordinal and count data with excess zeros.
method Bayesian framework with conditional transformation functions and modular MCMC algorithm.
result Flexible modeling of linear and nonlinear covariate effects for ordinal and count data.
Startups is a popular phenomenon that has a significant impact on global economy growth, innovation and society development. However, there is still insufficient understanding about startups, particularly, how to start a new business in the relation to consequent performance. Toward this knowledge, we have performed an…
The availability of large idea repositories (e.g., the U.S. patent database) could significantly accelerate innovation and discovery by providing people with inspiration from solutions to analogous problems. However, finding useful analogies in these large, messy, real-world repositories remains a persistent challenge …
Proposes LVGP for multi-source data fusion in science and engineering.
problem Differences in quality and comprehensiveness of data sources.
method Latent Variable Gaussian Process (LVGP) framework.
result Improved predictions for sparse-data problems.
Industry evolution caused by various reasons, among which technology progress driving industry development has been approved, but with the new trend of industry convergence, inter-industry convergence also plays an increasing important role. This paper plans to probe the industry synergetic evolution mechanism based on…
Research explores how local communities and corporations interact in finance.
problem Impact of local government subsidies and corporate bankruptcy on bond yields.
method Difference-in-differences analysis, econometric models, deep-learning model.
result Corporate subsidies and bankruptcy filings affect bond yields significantly.
Paper presents a new framework for sequence classification.
problem Sequence classification in real-world applications.
method Reference-based sequence classification framework.
result New sequence classification algorithms achieve comparable accuracy.
Dual-stage sEMG classification improves gesture recognition accuracy.
problem Improving accuracy in hand gesture recognition from sEMG signals.
method Dual-stage classification approach: first stage groups similar activities, second stage classifies within groups.
result Dual-stage classification yields significantly higher accuracy than single-stage approach.
A novel method for classification with rejection using ensemble of cost-sensitive classifiers.
problem Avoid risky misclassification in error-critical applications.
method Learning an ensemble of cost-sensitive classifiers.
result Improved classification accuracy and flexibility in loss selection.
The number of possible methods of generalizing binary classification to multi-class classification increases exponentially with the number of class labels. Often, the best method of doing so will be highly problem dependent. Here we present classification software in which the partitioning of multi-class classification…
Few-shot image classification is improved by correcting CNNs' texture bias.
problem Few-shot image classification performance is hindered by CNNs' texture bias.
method Corrected CNNs' texture bias using a simpler method than state-of-the-art approaches.
result State-of-the-art performance on miniImageNet task achieved.
New NHCAs improve multi-category classification efficiency.
problem Efficient multi-category classification for real-world problems.
method Twin SVM (TWSVM), Generalized eigenvalue proximal SVM (GEPSVM), Regularized GEPSVM (RegGEPSVM), and Improved GEPSVM (IGEPSVM) with OAA, BT, and TDS approaches.
result TDS-TWSVM outperforms other methods in classification accuracy.
Paper compares XGB and BPNN for music style classification.
problem Efficient music style classification using different methods.
method Feature extraction for timbral texture, rhythmic content, and pitch content; comparative evaluation of XGB and BPNN.
result XGB outperforms BPNN for small datasets in music classification.
Classification outperforms regression in portfolio construction, yielding higher Sharpe ratios.
problem Determining which machine learning approach (classification vs. regression) is more effective for portfolio construction.
method Used stacking ensemble of gradient boosted tree, random forest, and neural network models.
result Classification yields higher Sharpe ratios and economically significant alphas compared to regression.