Forty-Five Years of Digitizing Ebooks: Project Gutenberg's PracticesNewby, Gregory B.
History
Forty-Five Years of Digitizing Ebooks: Project Gutenberg's Practices
Newby, Gregory B.
Project Gutenberg
Today, EPUB and MOBI (also known as Kindle) formats are the most
common. Free software for conversion, called ebookmaker (previously
called epubmaker) is used to create derivative formats. This helps to
assure compatibility for different reader devices.
UPLOADING A NEW EBOOK
Volunteers upload the master format for their completed eBook to the
Project Gutenberg server, where it undergoes automated and manual
checks before the new eBook is posted and announced online. Prior to the
upload, the copyright clearance must be completed.
Upon uploading, automated checks include:
HTML checks for validity of the HTML encoding (via the W3C
validator);
HTML checks for internal link structure;
Spelling checks (English, with limited support for other
languages);
Typo/scanno checks (seeking common scanner/OCR errors, such
as “he” for “be” and vice-versa);
Conversion checks.
The conversion check consists of using the ebookmaker application to
automatically generate derived formats. Ideally, resulting files will
include:
Plain text in UTF8 encoding;
Automatically generated HTML (if HTML is not the master
format).
EPUB and MOBI
For HTML, EPUB and MOBI, pairs of files are generated: one with images,
and one without. The set of files without images is intended to be
friendlier to readers with limited bandwidth, or without the necessary
storage space for any images included with the eBook.
After uploading, a team of human experts — known as the “whitewashers,”
after a scene in Mark Twain’s “The Adventures of Tom Sawyer” — does
final formatting, attaches the Project Gutenberg header and footer, and
uploads the new item to the server at www.gutenberg.org.
CATALOGING AND MIRRORING
The Project Gutenberg catalog database includes metadata from
within each eBook: the author, title, available file formats,
upload/publication date, language, etc. Human catalogers eventually add
additional metadata, including Library of Congress Subject Headings.
This catalog is available for free download in machine readable form
(XML/RDF or MARC).
Organizations that desire to redistribute Project Gutenberg’s content,
freely and without limitations, are invited to do so. The catalog may
be used for this purpose, and various mechanisms are available
to automatically maintain a copy of the collection itself (i.e.,
“mirroring”), including for generated content.
“NO SWEAT OF THE BROW COPYRIGHT”
An important innovation during the evolution of Project Gutenberg was
to clarify the notion of “authorship” and its critical role for
establishing copyright. In early days, it was common to think that
applying HTML markup, or reformatting, or spelling changes, qualified
an item for a new copyright. Historically, some print publishers even
claimed new copyrights simply for typesetting a new edition.
Public-domain text, read in full here on John Shaqi.
Reviews
Reviews
No reviews yet
Be the first to share your thoughts on this work.
Join the Discussion
Join the discussion
Sign in to leave a comment or review.
Sign InorCreate an account