Judith ZIDAR, coordinator, National Agricultural Text Digitizing Program
(NATDP), National Agricultural Library (NAL), illustrated the technical
details of NATDP, including her primary responsibility, scanning and
creating databases on a topic and putting them on CD-ROM.
(ZIDAR remarked a separate arena from the CD-ROM projects, although the
processing of the material is nearly identical, in which NATDP is also
scanning material and loading it on a Next microcomputer, which in turn
is linked to NAL's integrated library system. Thus, searches in NAL's
bibliographic database will enable people to pull up actual page images
and text for any documents that have been entered.)
In accordance with the session's topic, ZIDAR focused her illustrated
talk on image capture, offering a primer on the three main steps in the
process: 1) assemble the printed publications; 2) design the database
(database design occurs in the process of preparing the material for
scanning; this step entails reviewing and organizing the material,
defining the contents--what will constitute a record, what kinds of
fields will be captured in terms of author, title, etc.); 3) perform a
certain amount of markup on the paper publications. NAL performs this
task record by record, preparing work sheets or some other sort of
tracking material and designing descriptors and other enhancements to be
added to the data that will not be captured from the printed publication.
Part of this process also involves determining NATDP's file and directory
structure: NATDP attempts to avoid putting more than approximately 100
images in a directory, because placing more than that on a CD-ROM would
reduce the access speed.
This up-front process takes approximately two weeks for a
6,000-7,000-page database. The next step is to capture the page images.
How long this process takes is determined by the decision whether or not
to perform OCR. Not performing OCR speeds the process, whereas text
capture requires greater care because of the quality of the image: it
has to be straighter and allowance must be made for text on a page, not
just for the capture of photographs.
NATDP keys in tracking data, that is, a standard bibliographic record
including the title of the book and the title of the chapter, which will
later either become the access information or will be attached to the
front of a full-text record so that it is searchable.
Public-domain text, read in full here on John Shaqi.
Reviews
Reviews
No reviews yet
Be the first to share your thoughts on this work.
Elsewhere in the archive
Join the Discussion
Join the discussion
Sign in to leave a comment or review.
Sign InorCreate an account