Forty-Five Years of Digitizing Ebooks: Project Gutenberg's Practices — John Shaqi
Forty-Five Years of Digitizing Ebooks: Project Gutenberg's PracticesNewby, Gregory B.
History
Forty-Five Years of Digitizing Ebooks: Project Gutenberg's Practices
Newby, Gregory B.
Project Gutenberg
In practice, the majority of Project Gutenberg eBooks rely on a single
printed source. However, even those items might benefit from other
sources — such as when some pages are missing, or illustrations come
from a different version, or when typos/errata reports come from other
sources.
It is a principal of Project Gutenberg that the eBooks in the collection
are denoted as Project Gutenberg eBooks. Even if the publisher imprint
and frontispiece from a printed work is included, there is no assurance
that the content exactly matches that printed work. And, in fact,
it will not match: minimally, the headers/footers will be removed, and
paragraphs will flow together such that they span the pages of the
printed source. Many other adjustments are typically made, as mentioned
above.
For this reason, Project Gutenberg’s online catalog metadata does not
include a citation to the source(s) used to create an eBook. Instead,
Project Gutenberg should be cited as the publisher. For example, a
bibliographic citation might have a form such as this:
Carroll, Lewis. “Alice’s Adventures in Wonderland.” Urbana, Illinois:
Project Gutenberg. Available: www.gutenberg.org/ebooks/11
OTHER CONTENT TYPES
Project Gutenberg is, arguably, the oldest continuously operating online
content project in the world. From 1971 until the mid-1990s, there were
relatively few online resources for literary content. For this reason,
and also due to a general willingness to experiment and reach out to
broader audiences, Project Gutenberg has a great variety in the content
types offered.
Among the first 100 items, there are mathematical constants and a
musical performance. Government publications, notably the 1990 US Census
and the CIA World Factbook from 1990 onward, were also included. The
next few hundred items include movies, photographs of ancient cave
paintings, and the first non-English items (Virgil’s Aeneid, Cicero’s
Orations, and Caesar’s Commentaries, all in Latin).
Hundreds of audio eBooks are in the collection. Many were automatically
generated via text-to-speech software. There are also a number
of readings/performances by human readers, including from Project
Gutenberg’s partner, Librivox (www.librivox.org). Today, automated
text-to-speech is accessible by most people with a computer or
mobile phone, so there is less emphasis on that format. Human
readings/performances continue to be of interest, especially when the
performance, as well as the original Project Gutenberg source eBook, is
granted to the public domain.
LANGUAGES OTHER THAN ENGLISH
Non-English languages have some additional characteristics that were not
well-suited for the plain text ASCII of Project Gutenberg’s early days.
By the early 1990s, it was necessary to display accented characters, to
accommodate languages such as French and Spanish. Later, languages such
as Chinese would require entirely separate character sets.
Public-domain text, read in full here on John Shaqi.
Reviews
Reviews
No reviews yet
Be the first to share your thoughts on this work.
Join the Discussion
Join the discussion
Sign in to leave a comment or review.
Sign InorCreate an account