Neruda Archives: historical letters, Omeka S, and a book linked back to every source

Hello,

I have been building Neruda Archives as an independent archive of historical letters that I own and preserve. The archive is still growing: new letters are added as they are scanned, described, transcribed and reviewed.

The project has two separate public parts. The main website introduces the collections and the human stories behind them. Omeka S holds the research catalogue and the evidence on which that work rests. Each published letter has a stable NRA-L identifier and a record with images of the original, metadata, a transcription, an English translation where available, rights information and IIIF access.

I wanted to keep those two functions separate. An editorial page can help a reader understand why a letter matters, while the Omeka record allows anyone to return to the document itself rather than simply trusting my reading of it.

The first book to come out of this work is I. Lettres. It contains sixty authentic letters and has been published as an ebook in Czech, English, French, German and Spanish. It is not a conventional book about historical correspondence. The letters are arranged without commentary and form a polyphonic human narrative that can be read almost like an existential novel, although nothing is invented. Every letter in the book links back to its complete archival record in Omeka S.

I have also used AI as a practical assistant when dealing with difficult handwriting, draft translations and contextual research. It can be useful, but it can also be confidently wrong. I therefore check public results against the original images and leave uncertain readings visibly uncertain.

The public catalogue currently uses Faceted Browse, IIIF Presentation and a small custom module for stable NRA-L routes. CSV Import, Custom Vocab, Custom Ontology and Value Suggest support the processing workflow.

Omeka S catalogue:

Ordinary Lives catalogue:

Example letter from the book:

Project website:

I. Lettres:

I am sharing the project here because Omeka has made it possible to keep the readable, human result connected to the original documents and their complete archival records.

Michal Neruda
Neruda Archives

3 Likes

Hi Michal,

Congratulations for this beautiful and meaningful work!

Can you please tell us more about what’s behind the scenes concerning your working method, the technical aspects, the theme design, and any other subject that you’d like to share?

Hello,

Thank you for your kind words and for your question.

The project began when I gradually acquired a collection of old letters belonging to one family. After reading them, I realised that the story emerging from them should not remain with me alone. I therefore began thinking about how to make the letters accessible to others while, if possible, preserving them for the future.

At first, I mainly considered museums and similar institutions. Then I discovered Omeka S, and almost immediately felt that it was exactly the environment I had been looking for. Each document can have its own identity, images of the original, metadata, a transcription, a translation, and relationships to other documents. The original can remain with its current owner, while its digital form and content become publicly accessible.

The original collection of letters gradually developed into a much broader project. Today, Neruda Archives has two public parts. The main website presents the individual collections and the human stories behind the documents. Omeka S serves as the research catalogue and archive, where it is always possible to return to the original source.

Each published letter receives its own permanent NRA-L identifier and a stable public URL. Its record may contain images of the original, a transcription, an English translation, descriptive and postal metadata, information about the document’s provenance and rights, and a recommended citation. The images are also accessible through IIIF, while I am gradually making the public data available through the Omeka API, JSON-LD, and catalogue exports.

The workflow begins with scans of the physical document. I preserve the source files unchanged, and all adjustments are made only to derivative copies. Several image variants of each page are prepared for the initial reading of the handwriting. The first draft of the transcription is now most often produced with Gemini, but I do not treat it as a finished result. Codex compares it again with the original images and looks for errors or inconsistencies, while a separate check is carried out using ChatGPT.

People, companies, places, dates, and events are also verified using the content of the entire letter and the available sources. Context can sometimes reveal a misread name or word. One important rule always applies, however: illegible text must not be “corrected” merely because a particular solution would fit the historical context well. If a reading is not sufficiently certain, it must remain uncertain.

The results are used to prepare the metadata, document description, transcription, translation, and rights information. The record is initially created as non-public and is published only after the images, text, metadata, and technical outputs have been checked. After publication, its actual public presentation, API output, and IIIF manifest are verified once again.

The public catalogue is built on Omeka S and uses, among other things, Faceted Browse, IIIF Presentation, CSV Import, Custom Vocab, Custom Ontology, and Value Suggest. A small custom layer was also created to provide stable NRA-L URLs. Omeka’s visual design is intentionally simple and restrained because I did not want an elaborate interface to obscure the historical documents. Individual pages inherit their design and functionality from shared templates and modules, allowing the archive to grow gradually to contain thousands of documents while remaining consistent.

I am not a professional developer, and I certainly could not have written much of the code used here from scratch on my own. The technical implementation developed gradually with the help of AI: I described a problem or the intended result, tested the proposed solutions, looked for errors, and had them reworked. It involved many evenings and nights, dead ends, and corrections. I had very long discussions with the individual AI models, and occasionally I also swore at them quite a bit.

I used the same principle of linking the reader-facing presentation back to the original source in the book I. Lettres . Each letter in the book links back to its complete record in Omeka S.

If you or other members of the community are interested in any particular part of the workflow—image processing, metadata, IIIF, stable identifiers, or the technical implementation—I will be happy to describe it in more detail and share how it is handled in Neruda Archives.

Michal Neruda

Thanks Michal for the detailed description of your working method. Which of course leads to some questions :slightly_smiling_face:

Today, Neruda Archives has two public parts. The main website presents the individual collections and the human stories behind the documents. Omeka S serves as the research catalogue and archive, where it is always possible to return to the original source.

Is there a technical/practical reason why you had to split the project in 2 different platforms (Framer and Omeka S) rather than publishing the editorial content in Omeka S? How do you manage the link between the two, for example to feed the collections page on the “main” site? As for the theme design, did you start from the default theme?

Each published letter receives its own permanent NRA-L identifier and a stable public URL.

Is the “NRA-L stable identifier” a persistent identifier stored in a global registry, like DOIs or Handle.net prefixes, or do you maintain it by yourself?

The images are also accessible through IIIF

I can’t find a IIIF manifest for the images, the Mirador viewer loads diretly the original JPEG files… Is this feature still a work in progress?

Kind regards,

Chris

Dear Chris,

Thank you for these very precise questions. They also highlighted two areas in which my public description of the project was not sufficiently exact.

The separation between Framer and Omeka S is a deliberate architectural decision rather than a consequence of a technical limitation in Omeka. I use Omeka S as the canonical research catalogue and archival record. I selected Framer as the editorial and presentation layer because it gives me greater long-term freedom to develop the visual and narrative presentation of the main website and to create future public-facing formats.

I do not maintain the two parts as independent catalogues. Each letter has an explicit primary collection assignment, and public catalogue data are regularly generated from the same publication system. The catalogue components in Framer consume these data but do not duplicate the complete records or their metadata. All letter-level links resolve to their canonical records in Omeka.

The public Omeka site did indeed begin with the default Omeka S theme. I have since substantially customised its shared templates, CSS, navigation and several functional elements for the project while retaining Omeka’s native catalogue functionality.

Thank you also for questioning the NRA-L identifiers. My original wording was not sufficiently precise and may have suggested that NRA-L was registered in a global identifier system. I apologise for the ambiguity.

NRA-L is a project-maintained persistent local identifier, not a DOI or Handle. It is independent of the internal Omeka item number, and each identifier has a stable public URL maintained by Neruda Archives. I have therefore corrected the wording to:

“Each published letter receives a project-maintained persistent NRA-L identifier and a stable public URL.”

Regarding IIIF, each public letter does have a IIIF Presentation 3 manifest. The manifests are now discoverable through the central IIIF Collection:

https://archive.neruda-archives.org/iiif-collection.json

For example, an individual letter manifest is available here:

https://archive.neruda-archives.org/iiif-presentation/3/item/744/manifest

You are correct, however, that the image bodies in the manifests currently reference the public JPEG access copies directly. I do not yet operate a tiled IIIF Image API service. My previous statement that the images were accessible through IIIF was therefore too broad. The more precise description is now:

“Each public letter has a IIIF Presentation 3 manifest. Images are currently delivered as public JPEG access copies rather than through a tiled IIIF Image API service.”

The JPEG files described by Omeka as “original” are the approved public access copies uploaded to Omeka, not the non-public TIFF preservation masters.

Thank you again. Your questions helped me improve both the project’s public documentation and the distinction between the current IIIF Presentation implementation and a possible future IIIF Image API service.

Kind regards,

Michal Neruda

Thanks for sharing this! It is interesting to see your approach in having multiple models for the transcription. It seems like this approach would be helpful in catching mistakes as you mentioned:

And I also appreciate how you have established a way for the user to verify the letters by going directly to the source material, which is definitely something that will carry much more importance as we have entered into an era where it is more and more difficult to determine authenticity of materials.

Great work!