IBM and the EU have expanded their research collaboration, which now includes more than two-dozen national libraries, research institutes, universities, and companies across Europe to provide new technology that will enable highly-accurate digitization of rare and culturally significant historical texts on a massive scale. Unlike past digitization projects where the result has been static, online libraries of texts, this unique widescale effort, called IMPACT (IMProving ACcess to Text), will offer new tools and best practices to institutions across Europe that will enable them to efficiently and accurately continue to produce quality digital replicas of historically significant texts and make them widely available, editable and searchable online.
[Snip]
"IMPACT is remarkable in that it not only allows these prominent centers of culture to ultimately bring people closer to perhaps never before seen historically significant texts of heritage -- but because it actually allows these people to become part of the preservation process," said Tal Drory, manager of the document processing group at IBM Research in Haifa. "IMPACT offers the first digitization system that combines the power of crowd computing with an adaptive optical character recognition (OCR) correction solution that can achieve excellent recognition rates across all kinds of documents – from the 15th century right up through the 19th century."
[Snip]
IMPACT technology streamlines, simplifies and accelerates the process of winnowing out questionable text scans, enabling reviewers to key in corrections to the text. Instead of displaying an entire scanned page, reviewers only see the actual letters or words in question. For example, the letter combination "r" and "n" ("rn") may appear indistinguishable from the letter "m." In those instances, the system collects many instances of the letter "m," and places these samples next to the letters in question, making it much easier to determine the letter's real identity.
In cases where an entire word is suspect, it is added to a collection of other questionable terms, which are then arranged in alphabetical order. Volunteer reviewers need only accept or reject suggested substitutes with one keystroke. In addition, the system uses adaptive dictionary enrichment, a method in which new words are added to a central dictionary based on cross-identification and correction by other users.
For example, a small book that normally takes four hours to key in manually, would take one hour using standard OCR technology with manual correction. Incorporating the new collaborative review technology cuts the process down to 30 minutes. IBM researchers explained that the new adaptive OCR system can further reduce the time, cutting it in half to 15 minutes.
The consortium partners include, among others: IBM Research – Haifa, Koninklijke Bibliotheek, The British Library, Osterreichische Nationalbibliothek, Universitat Innsbruck, Deutsche Nationalbibliothek, Bayerische Staatsbibliothek, Staats- und Universitatsbibliothek Gottingen, ABBYY Production, Instituut voor Nederlandse Lexicologie, National Centre for Scientific Research "Demokritos." Centrum fur Informations- und Sprachverarbeitung, University of Munich, University of Bath, University of Salford, Bibliotheque Nationale de France, Biblioteca Nacional de Espana and Poznan Supercomputing and Networking Center in Poland.
A family of resources to help information workers be more effective, raise the value of information in their organisations and contribute to success. Read more »
Recently I have found myself cooing over visualisation maps (and heat maps) of health and well being resources. The content rich data is overlayed with mapping technologies, and some interesting themes and patterns are emerging.
A lot of the talk around social media in the last year has been around information overload. Social media has provided us with new and exciting ways to create content. But it has also meant learning new ways to manage and engage with social media tools. Are we teetering on the edge of an information overload precipice?
Information overload is a figment of your imagination. Or a failure of your filter. Or a symptom of your technological submissiveness. Depends on who you ask.
What if you had to sort through 3.5 million articles and social media posts a day and try to pull out the most relevant items for your organisation? What if you then had to cobble it all together into something readable for your top groups and executives in your organisation?
Alacra Compliance saves time by aggregating information from both free and fee-based sources and enabling users to conduct an accurate federated search across these sources (coined “simultaneous search” by Alacra).