The Future of Information Extraction – Be Part of TUC 2024! ✨ Feb 15-16, In-Person and Online. Get your Ticket >>

+ Reading admiral de Ruyter’s journal – using existing transcripts to train Automated Text Recognition

Nicoline van der Sijs is part of a team of researchers working at the Meertens Institute in the Netherlands (one of the READ MOU partners).  The team has trained an Automated Text Recognition model to process the handwriting of Michiel de Ruyter, a Dutch admiral from the seventeenth century.

The model was trained with around 20,000 words of existing transcribed material from de Ruyter’s journals (see below for an example of his tricky handwriting!).  These transcriptions were matched automatically to corresponding digitised images of de Ruyter’s handwriting using Text2Img matching technology developed by the CITlab team at the University of Rostock (one of the READ project partners).

The resulting model is capable of recognising De Ruyter’s handwriting with a Character Error Rate (CER) of around 10%, which is an remarkable result for such a complex hand.

Image from the De Ruyter collection from the National Archives of the Netherlands, NL HaNA 1.10.72 20 0004

Professor van der Sijs and her colleagues are planning to use these transcriptions to compile an online corpus of de Ruyter’s writings for general access and scholarly linguistic analysis.

Researchers at the Meertens Institute are also interested in replicating these exciting results with other collections where existing transcriptions are already available, thanks to the hard work of volunteer transcribers.  The Stichting Vrijwilligersnet Nederlandse Taal (SVNT) is a network of about 100 volunteers who have been transcribing historic Bibles for more than ten years.  Other material transcribed by volunteers includes sailing letters from the seventeenth and eighteenth centuries and seventeenth-century printed newspapers.  The transcriptions that these volunteers have produced can be fed into our cutting-edge technology and used as training data for Automated Text Recognition.

  • Do you have existing transcriptions that you have produced or collected as part of a research project?
  • Send them to us and we can process them and train a model to recognise the writing in your documents!
  • To find out more about working with existing transcripts, consult our How to Guide or contact us.
SHARE THIS ARTICLE

Recent Posts

January 31, 2024
News
We’re pleased to announce the latest updates to our document editor, bringing you a more intuitive and cleaner interface. Our ...
January 17, 2024
News, Transkribus
Do I need to transcribe or translate handwritten text to be able to work with it? Well, that depends on ...
January 11, 2024
News, Transkribus
The process of managing and publishing historical documents has never been easier! Creating a website that presents your transcribed material ...