Forty-Five Years of Digitizing Ebooks: Project Gutenberg's Practices — John Shaqi
Forty-Five Years of Digitizing Ebooks: Project Gutenberg's PracticesNewby, Gregory B.
History
Forty-Five Years of Digitizing Ebooks: Project Gutenberg's Practices
Newby, Gregory B.
Project Gutenberg
Today, we know US copyright is based on the creative expression of ideas
through authorship. Markup and spelling changes do not qualify. As a
result, Project Gutenberg volunteers are able to “harvest” public domain
materials on the Internet, once they are determined to match public
domain print materials. This is not a frequent occurrence, however,
since most volunteers prefer to work on items that are not yet
digitized.
Similarly, Project Gutenberg claims no copyright on the “sweat of the
brow” labor which is applied to make eBooks from print sources. There
were a few earlier items where such copyright was claimed erroneously,
but this is no longer done.
EBOOKS, OR PICTURES OF BOOKS?
Project Gutenberg has over 50,000 eBooks in its collection. This is far
fewer than Google Books, or The Internet Archive, or other large-scale
digitization projects of historical items. An important distinction
is that Project Gutenberg engages in the proofreading, formatting,
markup/encoding, and other activities described above. Those other very
large projects are primarily devoted to scanning, and then provide raw
OCR output with a few automatically generated formats.
Such items are only partial eBooks — really, they are pictures (scans)
of books, with some additional automated features. These are valuable,
but do not provide the reading experience or quality of presentation
that Project Gutenberg strives for. Using current technology, it takes
human intellect and effort to convert a picture of a book to a true,
functional, eBook.
PAST INNOVATIONS AND FUTURE INITIATIVES
Project Gutenberg has evolved its practices over the years, and has
often been a leader in the creation and distribution of eBooks. Some
past innovations include the following, and all are still in active use
today:
Development of an open content trademark license (1991-
1993), which is intended to guarantee to readers that public
domain items remain free, while placing restrictions on the
trademarked name “Project Gutenberg” to protect against
abusive practices by those who would sell the public domain
items;
File/directory-based access to the collection, guaranteeing
ease of copying (by file, or subcollection, or the entire
collection), mirroring, and large-scale redistribution
(1994);
Anonymous access for all readers, requiring no logins or
authorization for any items (1994);
Web-based access to content, and development of procedures
to assure HTML is valid and well-formed (1996);
The Copyright How-To, including the Rule 6 How-To for non-renewed
items (2000 & 2008);
Support of Distributed Proofreaders (2002-2004), for
crowdsourced proofreading and other aspects of new eBook
creation;
Implementation of eBook reader formats, for free use on
mobile phones, tablets, and other devices (2009);
Free redistribution of metadata as a separate download (2007
& 2012);
Public-domain text, read in full here on John Shaqi.
Reviews
Reviews
No reviews yet
Be the first to share your thoughts on this work.
Join the Discussion
Join the discussion
Sign in to leave a comment or review.
Sign InorCreate an account