TrOCR_Fr - Digitization Pipeline for French Handwritten Archives (Github Repository)
An open-source pipeline for digitizing French handwritten archives, developed with Arnault Gombert. The pipeline is built around three core steps:
- Layout parsing: Raw document images are first segmented into individual lines of text.
- OCR: Each line image is then transcribed using a TrOCR model fine-tuned specifically on French handwritten text. The trained model is available for download on huggingface.
- Named Entity Recognition: Finally, a NER module pulls out structured information from the transcribed content — names, places, dates, etc.
