Managing Edit Checks and Database Cleaning with Stata
Jacqueline L. Buros ()
Additional contact information
Jacqueline L. Buros: Perfuse Core Laboratories and Data Coordinating Center
North American Stata Users' Group Meetings 2006 from Stata Users Group
Abstract:
We have developed a set of ado-files for use in data management, specifically designed to manage user-written edit checks and to complement the process of data cleaning. Collectively, these tools enable us to identify, distribute, and track edit-checks in several large multi-center clinical trials using Stata software. Our approach is successful because the coding is simple and the entire process is visible and familiar to most users. It does not depend on any particular database structure. The framework approximates an object-oriented environment, with the objects being (a) the database, open at the time a command is called, (b) an edit-check, consisting of a Stata do-file, a query message and a list of variables to be identified for review, and (c) the edit-check history, implemented as a Stata dataset. These objects can be manipulated directly or by using a command in Stata. Actions managed by command include creating or modifying an edit-check, generating a query-clean dataset, preparing and tracking a set of edit-check documents, and summarizing the edit-check history. Here, we present a brief overview of our process and describe the use of the commands in the context of clinical research.
Date: 2006-07-23
New Economics Papers: this item is included in nep-cmp
References: Add references at CitEc
Citations:
Downloads: (external link)
http://repec.org/nasug2006/Buros.NASUG2006.ppt.zip (application/zip)
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:boc:asug06:11
Access Statistics for this paper
More papers in North American Stata Users' Group Meetings 2006 from Stata Users Group Contact information at EDIRC.
Bibliographic data for series maintained by Christopher F Baum ().