EconPapers    
Economics at your fingertips  
 

Harmonizing and Combining Large Datasets - An Application to Firm-Level Patent and Accounting Data

Grid Thoma (), Salvatore Torrisi (), Alfonso Gambardella, Dominique Guellec, Bronwyn Hall and Dietmar Harhoff ()

No 15851, NBER Working Papers from National Bureau of Economic Research, Inc

Abstract: This paper discusses methods for the harmonization and combination of large-scale patent and trademark datasets with each other and other sources of data. Dictionary- and rule-based approaches to the consolidation of applicant names in patent data are presented and shown to have both benefits and drawbacks in isolation. We combine the two methods and develop a set of rules and dictionaries to consolidate European, Patent Cooperation Treaty (PCT) and US patent data with firm accounting data. The resulting data encompass about 131,000 patent applicant names from 46 countries, covering 58.8 percent of EPO applications and 50.6 percent of PCT applications by business organizations during the time period from 1979 to 2008. For US data, the resulting dataset includes around 54,000 assignee names and 51.3 percent of US granted patents during approximately the same time period.

JEL-codes: C81 O34 (search for similar items in EconPapers)
Date: 2010-03
Note: PR
References: View references in EconPapers View complete reference list from CitEc
Citations: View citations in EconPapers (71)

Downloads: (external link)
http://www.nber.org/papers/w15851.pdf (application/pdf)

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:nbr:nberwo:15851

Ordering information: This working paper can be ordered from
http://www.nber.org/papers/w15851

Access Statistics for this paper

More papers in NBER Working Papers from National Bureau of Economic Research, Inc National Bureau of Economic Research, 1050 Massachusetts Avenue Cambridge, MA 02138, U.S.A.. Contact information at EDIRC.
Bibliographic data for series maintained by ().

 
Page updated 2025-03-31
Handle: RePEc:nbr:nberwo:15851