Although AM has learned much from its experiences with various collections
and various service bureaus, ERWAY concluded pessimistically that no
breakthrough has been achieved. Incremental improvements have occurred
in some of the OCR technology, some of the processes, and some of the
standards acceptances, which, though they may lead to somewhat lower costs,
do not offer much encouragement to many people who are anxiously awaiting
the day that the entire contents of LC are available on-line.
******
+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
ZIDAR * Several answers to why one attempts to perform full-text
conversion * Per page cost of performing OCR * Typical problems
encountered during editing * Editing poor copy OCR vs. rekeying *
+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Judith ZIDAR, coordinator, National Agricultural Text Digitizing Program
(NATDP), National Agricultural Library (NAL), offered several answers to
the question of why one attempts to perform full-text conversion: 1)
Text in an image can be read by a human but not by a computer, so of
course it is not searchable and there is not much one can do with it. 2)
Some material simply requires word-level access. For instance, the legal
profession insists on full-text access to its material; with taxonomic or
geographic material, which entails numerous names, one virtually requires
word-level access. 3) Full text permits rapid browsing and searching,
something that cannot be achieved in an image with today's technology.
4) Text stored as ASCII and delivered in ASCII is standardized and highly
portable. 5) People just want full-text searching, even those who do not
know how to do it. NAL, for the most part, is performing OCR at an
actual cost per average-size page of approximately $7. NAL scans the
page to create the electronic image and passes it through the OCR device.
ZIDAR next rehearsed several typical problems encountered during editing.
Praising the celerity of her student workers, ZIDAR observed that editing
requires approximately five to ten minutes per page, assuming that there
are no large tables to audit. Confusion among the three characters I, 1,
and l, constitutes perhaps the most common problem encountered. Zeroes
and O's also are frequently confused. Double M's create a particular
problem, even on clean pages. They are so wide in most fonts that they
touch, and the system simply cannot tell where one letter ends and the
other begins. Complex page formats occasionally fail to columnate
properly, which entails rescanning as though one were working with a
single column, entering the ASCII, and decolumnating for better
searching. With proportionally spaced text, OCR can have difficulty
discerning what is a space and what are merely spaces between letters, as
opposed to spaces between words, and therefore will merge text or break
up words where it should not.
Public-domain text, read in full here on John Shaqi.
Reviews
Reviews
No reviews yet
Be the first to share your thoughts on this work.
Elsewhere in the archive
Join the Discussion
Join the discussion
Sign in to leave a comment or review.
Sign InorCreate an account