TrOCR_Fr - Digitization Pipeline for French Handwritten Archives (Github Repository)

An open-source pipeline for digitizing French handwritten archives, developed with Arnault Gombert. The pipeline is built around three core steps:

  • Layout parsing: Raw document images are first segmented into individual lines of text.
  • OCR: Each line image is then transcribed using a TrOCR model fine-tuned specifically on French handwritten text. The trained model is available for download on huggingface.
  • Named Entity Recognition: Finally, a NER module pulls out structured information from the transcribed content — names, places, dates, etc.