Hello,
Thank you for your kind words and for your question.
The project began when I gradually acquired a collection of old letters belonging to one family. After reading them, I realised that the story emerging from them should not remain with me alone. I therefore began thinking about how to make the letters accessible to others while, if possible, preserving them for the future.
At first, I mainly considered museums and similar institutions. Then I discovered Omeka S, and almost immediately felt that it was exactly the environment I had been looking for. Each document can have its own identity, images of the original, metadata, a transcription, a translation, and relationships to other documents. The original can remain with its current owner, while its digital form and content become publicly accessible.
The original collection of letters gradually developed into a much broader project. Today, Neruda Archives has two public parts. The main website presents the individual collections and the human stories behind the documents. Omeka S serves as the research catalogue and archive, where it is always possible to return to the original source.
Each published letter receives its own permanent NRA-L identifier and a stable public URL. Its record may contain images of the original, a transcription, an English translation, descriptive and postal metadata, information about the document’s provenance and rights, and a recommended citation. The images are also accessible through IIIF, while I am gradually making the public data available through the Omeka API, JSON-LD, and catalogue exports.
The workflow begins with scans of the physical document. I preserve the source files unchanged, and all adjustments are made only to derivative copies. Several image variants of each page are prepared for the initial reading of the handwriting. The first draft of the transcription is now most often produced with Gemini, but I do not treat it as a finished result. Codex compares it again with the original images and looks for errors or inconsistencies, while a separate check is carried out using ChatGPT.
People, companies, places, dates, and events are also verified using the content of the entire letter and the available sources. Context can sometimes reveal a misread name or word. One important rule always applies, however: illegible text must not be “corrected” merely because a particular solution would fit the historical context well. If a reading is not sufficiently certain, it must remain uncertain.
The results are used to prepare the metadata, document description, transcription, translation, and rights information. The record is initially created as non-public and is published only after the images, text, metadata, and technical outputs have been checked. After publication, its actual public presentation, API output, and IIIF manifest are verified once again.
The public catalogue is built on Omeka S and uses, among other things, Faceted Browse, IIIF Presentation, CSV Import, Custom Vocab, Custom Ontology, and Value Suggest. A small custom layer was also created to provide stable NRA-L URLs. Omeka’s visual design is intentionally simple and restrained because I did not want an elaborate interface to obscure the historical documents. Individual pages inherit their design and functionality from shared templates and modules, allowing the archive to grow gradually to contain thousands of documents while remaining consistent.
I am not a professional developer, and I certainly could not have written much of the code used here from scratch on my own. The technical implementation developed gradually with the help of AI: I described a problem or the intended result, tested the proposed solutions, looked for errors, and had them reworked. It involved many evenings and nights, dead ends, and corrections. I had very long discussions with the individual AI models, and occasionally I also swore at them quite a bit.
I used the same principle of linking the reader-facing presentation back to the original source in the book I. Lettres . Each letter in the book links back to its complete record in Omeka S.
If you or other members of the community are interested in any particular part of the workflow—image processing, metadata, IIIF, stable identifiers, or the technical implementation—I will be happy to describe it in more detail and share how it is handled in Neruda Archives.
Michal Neruda