Project Gutenberg also publishes eBooks in well-known formats like HTML, XML or
RTF. There are Unicode files too. Any other format provided by volunteers (PDF,
LIT, TeX and many others) is usually accepted, as long as they also supply an
ASCII version where possible.
But a large scale conversion into other formats is handed over to other
organizations. For example Blackmask Online, which uses Project Gutenberg's
collections to offer thousands of free eBooks in eight different formats based
on the Open eBook (OeB) format. Or Manybooks.net, which converts Project
Gutenberg's eBooks into formats readable on PDAs. Or Bookshare.org, the main
digital library for the visual impaired community in the US, which converts
books from Project Gutenberg into Braille format and DAISY (Digital Audio
Information System) format.
What is entailed exactly, once copyright clearance is received? Digitization is
done by scanning the book page after page to get "image" files. Then volunteers
run an OCR (Optical Character Recognition) software to convert "image" files
into text files. Then each text file is proofread (i.e. re-read and corrected)
by comparing it to the "image" file or the original page of the print version.
There is an average of 10 mistakes per page for a good OCR package and... many
more mistakes if the quality of the scanner and the OCR package is not great.
The book is proofread twice on the computer screen by two different people, who
make any corrections necessary. When the original is in poor condition, as with
very old books, it is keyed in manually, word by word. Some volunteers
themselves prefer to type short texts, or works they particularly like. But most
books are scanned, "OCRized" and proofread.
Digitization in "text format" means a book can be copied, indexed, searched,
analyzed and compared with other books. It is possible to search the content of
the book with the "Find" button available in any browser and any software,
without a specific search engine. Project Gutenberg provides a "Nearly Full
Text" search (on the first 100 K of each file) using Google, with a database
updated approximately monthly. It also provides a search of book metadata
(author, title, brief description, keywords) as a participant in Yahoo!'s
Content Acquisition Program, with a database updated weekly. (Please see the
bottom of the Online Book Catalog.) In the Advanced Search, several fields can
be filled: author, title, subject, language, category (any, audio book, music,
pictures), LoCC (Library of Congress Catalog classification), filetype (text,
PDF, HTML, XML, JPEG, etc.), and eText/eBook No. A field "Full Text" was
recently added as an experimental feature.
Public-domain text, read in full here on John Shaqi.
Reviews
Reviews
No reviews yet
Be the first to share your thoughts on this work.
Join the Discussion
Join the discussion
Sign in to leave a comment or review.
Sign InorCreate an account