An SGML document consists of text that is marked up with descriptive tags
that specify the function of a given element within the document. As a
formal language construct, an SGML document can be parsed against a
document-type definition (DTD) that unambiguously defines what elements
are allowed and where in the document they can (or must) occur. This
formalized map of article structure allows the user interface design to
be uncoupled from the underlying database system, an important step
toward interoperability. Demonstration of this separability is a part of
the CORE project, wherein user interface designs born of very different
philosophies will access the same database.
NOTES:
(6) The CORE project is a collaboration among Cornell University's
Mann Library, Bell Communications Research (Bellcore), the American
Chemical Society (ACS), the Chemical Abstracts Service (CAS), and
OCLC.
Michael LESK The CORE Electronic Chemistry Library
A major on-line file of chemical journal literature complete with
graphics is being developed to test the usability of fully electronic
access to documents, as a joint project of Cornell University, the
American Chemical Society, the Chemical Abstracts Service, OCLC, and
Bellcore (with additional support from Sun Microsystems, Springer-Verlag,
DigitaI Equipment Corporation, Sony Corporation of America, and Apple
Computers). Our file contains the American Chemical Society's on-line
journals, supplemented with the graphics from the paper publication. The
indexing of the articles from Chemical Abstracts Documents is available
in both image and text format, and several different interfaces can be
used. Our goals are (1) to assess the effectiveness and acceptability of
electronic access to primary journals as compared with paper, and (2) to
identify the most desirable functions of the user interface to an
electronic system of journals, including in particular a comparison of
page-image display with ASCII display interfaces. Early experiments with
chemistry students on a variety of tasks suggest that searching tasks are
completed much faster with any electronic system than with paper, but
that for reading all versions of the articles are roughly equivalent.
Pamela ANDRE and Judith ZIDAR
Text conversion is far more expensive and time-consuming than image
capture alone. NAL's experience with optical character recognition (OCR)
will be related and compared with the experience of having text rekeyed.
What factors affect OCR accuracy? How accurate does full text have to be
in order to be useful? How do different users react to imperfect text?
These are questions that will be explored. For many, a service bureau
may be a better solution than performing the work inhouse; this will also
be discussed.
SESSION VI
Marybeth PETERS
Public-domain text, read in full here on John Shaqi.
Reviews
Reviews
No reviews yet
Be the first to share your thoughts on this work.
Elsewhere in the archive
Join the Discussion
Join the discussion
Sign in to leave a comment or review.
Sign InorCreate an account