Forty-Five Years of Digitizing Ebooks: Project Gutenberg's PracticesNewby, Gregory B.
History
Forty-Five Years of Digitizing Ebooks: Project Gutenberg's Practices
Newby, Gregory B.
Project Gutenberg
From 2002-2004 an important innovation was developed, in support of
the creation of new Project Gutenberg eBooks. This was Distributed
Proofreaders, an early example of what is now known as crowdsourcing.
Through Distributed Proofreaders, volunteers engage in a portion of
the eBook creation process — whether it is copyright clearances,
proofreading (a page at a time!), or formatting, checking, and
finalization before uploading. Those portions, when coordinated
together, lead to the creation of new eBooks from printed sources.
Distributed Proofreaders has become the single largest source for new
eBooks to the collection, accounting for approximately half of all
titles. Distributed Proofreaders has also innovated substantially in
the use of HTML+CSS (cascading style sheets) for very attractive
presentation of eBooks in Web browsers.
SCANNING
By the early 1990s, scanning and optical character recognition (OCR)
started to become widely available. Hart received a full scanning
station via a grant from a computer manufacturer, which was used to
produce several of the first 100 eBooks. The scanner was a flatbed
model, which required the user to hold the book open, scan a page (or
pair of pages) for ingest to the OCR software, then flip to the next
page.
The OCR software would then automatically recognize the characters from
the scan, and create an editable view of the text. Proofreading and
formatting would then occur in the same way as for a typed text.
A few years later, Project Gutenberg worked with Distributed
Proofreaders to acquire sheet-fed scanners. These scanners, which are
still in operation, are faster. They also tend to produce an image
that is properly aligned, versus the skewing that sometimes occurs
with flatbed scanners. An important difference is the printed books
are damaged: prior to scanning, the spines of the books are cut off, in
order for the individual pages to be ingested by the scanner.
[Illustration: 0006]
Figure 3: Image from the Doré illustrations of Dante’s Inferno
It has been Project Gutenberg’s intention to make all the original
images from the scanners available, alongside the finished eBook. This
is to have a more complete record of the eBook’s source(s), and also to
facilitate improvements by finding typos. Most eBook producers to date
have chosen to not provide the scans, however.
Scanners are used for images within printed books, which are typically
included as JPEG, GIF or PNG items within HTML and other formats. Inline
images may be at a lower resolution, and then clickable to obtain higher
resolution images. Color scanners are used, whenever possible, for color
images.
Project Gutenberg has no prohibition against using items scanned by
other parties. Several excellent sources of scans are freely available,
including Google Books, Gallica, and The Internet Archive. Scans, and
raw OCR output (if available), may then be transformed into Project
Gutenberg eBooks by volunteers.
Public-domain text, read in full here on John Shaqi.
Reviews
Reviews
No reviews yet
Be the first to share your thoughts on this work.
Join the Discussion
Join the discussion
Sign in to leave a comment or review.
Sign InorCreate an account