Project Gutenberg is convinced that proofreading by human beings is a very
important step, and that this step makes all the difference. The use of scanned
books as is --converted to text format by OCR software with no proofreading--
gives a much lower quality result. After running OCR software, the text is 99%
reliable, in the best of cases. After proofreading, the text becomes 99.95%
reliable (a high percentage which is also the standard at the Library of
Congress).
For this reason, Project Gutenberg's perspective is rather different from that
of the Internet Archive. In its Text Archive, books are scanned and "OCRized",
but they are not proofread. The main formats used are XML, TIF and DjVu. Books
are not proofread either in other main collections: Open Content Alliance (OCA),
Google Books Search or Microsoft Live Books Search.
Project Gutenberg provides a "Nearly Full Text" search (on the first 100 K of
each file) using Google, with a database updated approximately monthly. It also
provides a search of book metadata (author, title, brief description, keywords)
as a participant in Yahoo!'s Content Acquisition Program, with a database
updated weekly. Both are available in the Online Book Catalog (at the bottom of
the page). In the Advanced Search, several fields can be filled: author, title,
subject, language, category (any, audio book, music, pictures), LoCC (Library of
Congress Catalog classification), filetype (text, PDF, HTML, XML, JPEG, etc.),
and eText/eBook No. A field "Full Text" was also added as an experimental
feature.
On Project Gutenberg's website, a File Recode Service allows users to convert
books in one format (ASCII, ISO-8859, Unicode and others) into another, and vice
versa. A much more powerful conversion program may be launched in the future,
with a conversion into still more formats (XML, HTML, PDF, TeX, RTF), including
Braille and voice. It will then also be possible to choose the font and size of
characters and the background color. Another eagerly expected conversion is that
of a book from one language to another by machine translation software. This may
be possible in a few years, when machine translation is accurate to 99%. Still,
these books will certainly need some proofreading too by human translators.
4. SHARED PROOFREADING
The main "leap forward" of Project Gutenberg in the last few years is due to
Distributed Proofreaders. Distributed Proofreaders was launched in October 2000
by Charles Franks to help in the digitizing of public domain books. Originally
meant to assist Project Gutenberg in the handling of shared proofreading,
Distributed Proofreaders became the main source of Project Gutenberg books. In
2002, Distributed Proofreaders became an official Project Gutenberg site. In May
2006, Distributed Proofreaders became a separate entity and continues to
maintain a strong relationship with Project Gutenberg.
Public-domain text, read in full here on John Shaqi.
Reviews
Reviews
No reviews yet
Be the first to share your thoughts on this work.
Join the Discussion
Join the discussion
Sign in to leave a comment or review.
Sign InorCreate an account