What are they good for?
Introduction
Recently, generative AI models, particularly Gemini 3.0, have shown impressive accuracy in transcribing images of historical handwriting, including the ability to use context clues and historical research to override the AI’s statistical training and arrive at more accurate results. After reviewing the literature on these AI developments, we designed a study to test the methods using local Carleton datasets of document images. To conclude this project, we will share our findings and the implications of using AI to transcribe historical documents to the wider Carleton community.
Literature Review
As of the academic year 2025-26, a limited number of peer-reviewed journals are available to observe what other scholars had attempted. In his article, Handwriting Recognition Roundup, Tom Schienfeldt provides several examples of historians using and testing Gemini 3.0’s performance in recognizing archival and historical scholarship. This shift demonstrates a critical step in realizing mass digitization.
In The Sugar Loaf Test: How an 18th-Century Ledger Reveals Gemini 3.0’s Emergent Reasoning, Mark Humphries modifies his previous prompting to trace the thought process and reasoning of Gemini 3.0. Besides simply transcribing and calculating the word error rate (WER), Gemini can also reformat the text to show its understanding in the output. Even when Gemini is given a fictitious currency to work with, it argues that this currency does not exist using its logical deductions from transcribing the document. More detailed experimentation can be found in Professor Humphries’s older posts: Has Google Quietly Solved Two of AI’s Oldest Problems? and Gemini 3 Solves Handwriting Recognition and it’s a Bitter Lesson.
In When the Machine Finally Learned to Read: Gemini 3 and the Question of “Good Enough”, Steve Little argues that AI transcription has gotten “good enough” such that measuring the WER has little significance because the AI model relies semantic context. This allows scholars to focus on using these transcripts in applications with keyword searching, indexing, and digital text analysis. Similarly, in Bottlenecks, Side Quests, and the Calculus of Historical Research, Cameron Blevins exclaims that historians no longer have to go on “side quests” to manually transcribe documents in order to do more exciting and relevant analysis for their research.
In Futzing with Newspaper OCR, Shawn Graham illustrates that by giving the AI model a specific role, such as a historian transcriber in the 20th century, the LLM performs better utilizing the field expertise to look for particularities. This method of crafting structured prompts provide more clues for the model to accurately interpret messy handwriting and page layouts.
Goals
Handwritten document transcription remains labor-intensive, but new AI systems may reduce that burden. However, it remains unclear how well these systems perform across different handwriting and whether prompt design improves results in meaningful ways. This project assesses how accurately and usefully current AI tools can transcribe handwritten documents across different genres, levels of legibility, and contextual density. Furthermore, we want to understand how and why the AI produces the transcription. Aside from comparing transcription quality across prompt strategies (same model, same pages), we aim to identify a stable default workflow for manual correction and build a reusable evaluation framework for a larger-scale study.
Data
To replicate the different AI prompting and testing its performance, we use
- (1) Susannah Ottaway’s historical primary sources at the Virtual Workhouse archive and
- (2) Jay Beck’s handwritten notes of cinematic reviews.
Methods / Tools
To set up a controlled experiment, we kept the model and page fixed and only vary the prompt conditions. By keeping the input smaller and more controlled, e.g., feeding one page at a time or splitting dense content, we get more reliable results. We tested with different versions of Gemini to analyze the accuracy of later versions. Based on the methods mentioned in the literature review articles, we crafted the prompts using zero-shot, one-shot, role-based, and context-enriched techniques.
- The zero-shot prompt is simply telling the AI to
- “Transcribe this handwritten page as accurately as possible. Preserve original spelling and line breaks.”
- The one-shot prompt is more descriptive in outlining what the AI should watch out for with an example
- “You are an archival handwriting transcription assistant. Transcribe only the visible handwritten text in the provided page image. Do not summarize, infer missing content, or modernize spelling. If text is unreadable, write [illegible]. If text is uncertain but partially readable, write [?text]. Example: Input image content (simulated): ‘Dear John, I received your letter yesterday. The farm is [smudged] this spring.’ Desired output: Dear John, I received your letter yesterday. The farm is [illegible] this spring. Now transcribe the provided page image. Return only the transcription text, no extra commentary.”
- The role-based prompt is similar to the one-shot prompt but without the example
- “You are an archival handwriting transcription assistant. Transcribe only the visible handwritten text in the provided page image. Preserve original line breaks as much as possible. Do not summarize, infer missing content, or modernize spelling. If text is unreadable, write [illegible]. If text is uncertain but partially readable, write [?text]. Keep punctuation/capitalization exactly as seen when legible. Return only the transcription text, no extra commentary.”
- The context-enriched prompt further provides the document details and backgrounds for the model
- For the 18th-century historical documents. “Gemini, you are transcribing an image of eighteenth-century handwritten English workhouse meeting minutes. Transcribe the text exactly as written. Preserve original spelling, punctuation, capitalization, and line breaks. Do not modernize or interpret the language. Do not add words that are not visible in the image. If a word is unclear, write [unclear]. If you are unsure of a word, write your best guess followed by [?]. Output only the transcription.”
- For the cinema handwritten notes. “Gemini or notebook lm, you are transcribing an image of handwritten film notes on Grossberg- Media Economy of Rock Culture notes from a rock and roll cinema course. Transcribe the text exactly as written. Preserve original spelling, punctuation, capitalization, arrows (→), and line breaks. Do not modernize or interpret the language. Do not add words that are not visible in the image. If a word is unclear, write [unclear]. If you are unsure of a word, write your best guess followed by [?]. Output only the transcription.”
Evaluations
The transcription below uses the one-shot method. The results show that the AI did not account for indents or special marks on the page; many symbols and columns are not aligned in the same form as the original document.

We refined the prompt to fix the page layout to closely match the actual document. This new prompt includes
- Fixed-Width Column Control: The original prompt lets the text freely wrap; the new master template explicitly forces monospace formatting and manual character-spacing to mimic the original ledger’s exact vertical column structure.
- Split-Screen Multi-Page Flow: The original prompt lumps all transcription text into a single output. The new version automatically detects two-page document spreads and creates a clean, repeated side-by-side view (Image Left / Text Right) across multiple pages.
- Advanced Structural Tracking: The perfected template adds instructions to use complex multi-line Unicode bracket symbols, preserving the physical layout of historical group entries (like lists of paupers) rather than flattening them.
- Automated Document Delivery: The new prompt moves beyond a simple text response by instructing the backend to compile and export a professionally scaled, presentation-ready A4 landscape PDF automatically.
You are an archival handwriting transcription assistant. Your task is to transcribe only the visible handwritten text in the provided image and compile it into a final, downloadable document for me. Follow these strict constraints:
- 1. Data Integrity: Do not summarize, infer missing content, or modernize spelling. If text is unreadable, write [illegible]. If text is uncertain but partially readable, write [?text].
- 2. Structure & Brackets: Keep all original grouping brackets (like large vertical curly braces). Maintain their exact height and structural scale using large multi-line Unicode bracket characters (e.g., ⎧, ⎪, ⎨, ⎩) so they span the correct number of rows.
- 3. Columnar Alignment: You must use fixed-width text formatting (monospace text blocks) with explicit spacing to keep numbers, tabulations, and columns in flawless vertical alignment with each other, matching the precise layout of the original account books.
- 4. Handling Multi-Page Spreads: If the provided image contains a two-page spread (a left page and a right page showing together), do not clump the entire transcription onto a single layout side. Instead, create a distinct two-page flow in your final document: – Page 1 of the output document will display the entire source image on the left, and ONLY the transcription of the left page on the right. – Page 2 of the output document will display the exact same source image on the left, and ONLY the transcription of the right page on the right.
- 5. Final Output Generation: Execute an internal Python rendering script to compile a downloadable, multi-page PDF in A4 landscape orientation.
- 6. Column Splits: On each output page, maintain a strict split: – LEFT SIDE (approx. 41% width): Place the document image scaled cleanly to fit the layout canvas without running off the borders. – RIGHT SIDE (approx. 59% width): Place the corresponding text transcription box utilizing a balanced font size (e.g., 8.5pt or 9pt monospace) so rows do not wrap incorrectly or spill onto unintended pages.
Return only the clean, monospaced transcribed plain text block in your chat response (using clear markers to separate the pages), followed immediately by the final downloadable project file. No extra commentary.
Our new transcriptions look much better.


Asides from transcribing historical documents that have circulated the Internet, we also attempted to transcribe recent, private handwritten notes, again using Gemini 3.0 with the one-shot prompt. The results are both surprising and unsurprising, although the LLM could transcribe even the most illegible handwriting, however, the output format always required further prompting and specifications.


A more refined prompt is required for exact formatting.
Scaling Up
Rather than transcribing page by page, we worked toward scaling the manuscript image transcription across multiple archive collections while preserving page layout and spatial relationships. We desire a reviewable output in a Markdown and HTML file following a scaling-up workflow.
Archival images + metadata
↓
JSONL export
↓
Collection-level split
↓
Batch-level split
↓
LLM transcription
↓
Reviewable output
This workflow is designed to scale manuscript transcription from the Virtual Workhouse archive. The process begins by exporting item metadata and image URLs from the Omeka API, splitting the exported data by collection, and then dividing each collection into smaller JSONL batches for Gemini transcription. Gemini returns transcription results as JSONL, and finally those results can then be converted into readable HTML files. The LLM prompt is shared in this open source repository. The sample transcription below exhibits the workflow’s capabilities to mass produce formatted transcription, however, we can still observe several discrepancies and uncertainties in the output.



Challenges
Several roadblocks we ran into are that we had to ask the LLM to enhance certain aspects of the presentation of the transcription repeatedly, such as creating a downloadable PDF output or effectively transcribing brackets or indentations. But this was solved with prompting the LLM to incorporate these instructions into its own prompt; these prompts’ format need to be more like a list with bullet points, rather than a paragraph so the LLM can easily follow.
Feedback
We have yet to receive professional feedback from either source of the data used for our results. We will update this section soon, hopefully.
Implications
The lessons learned from this AI transcription project are that using LLM can help improve the research process. By incorporating AI transcription, rather than manually plowing through handwritten documents, the LLM can spit out a not-so-perfect transcription but a PDF good enough that researchers can focus on pattern recognition by searching for keywords or applying more advanced research processes that demand a text file. Given more time, we would love to experiment on the produced PDF files for its pattern-seeking ability and prompt the LLM to produce data visualizations based on the found patterns, such as names, significant data points, etc.