somniabunt SL

← ./work / TIDE.archive

TIDE.archive

A searchable archive for thirty years of an illustrator's work

An illustrator has about 40,000 files spread across a NAS, two laptops and a phone, with names like scan_0042_final2.tif. TIDE is a study of an archive that sorts itself as files arrive and can answer a request like "the forest series, print quality, everything from 2019".

Type
Automation · AI search
Year
2025
Role
Architecture, pipelines, search
Status
Concept study

Concept A study of how I’d approach this kind of problem, not a client project.

$ cat ./targets — design targets, not measured results

40,000
files indexed
< 1 s
typical search
3
copies of every original
monthly
restore drill

Capture

Anything dropped into the inbox, synced from the phone album or migrated from the old NAS lands in S3-compatible object storage with versioning on, so an overwrite never loses the original. Each new object fires an event.

Enrichment

One queued job per file, fanned out to three workers. The first makes derivatives with libvips: a print master that keeps its ICC profile, and web sizes in AVIF and WebP. The second runs a vision model that writes a caption and suggests tags from the studio's own vocabulary. The third stores an embedding, so images can be found by what they show as well as by name.

Search

Metadata and vectors live together in Postgres with pgvector. The search API blends keyword and semantic results, so "rainy street, blue" works as well as a filename. The catalogue is a fast web app, and the chat assistant answers the same questions in a message with a signed download link.

Safety

Originals are never changed. Storage is copied to a second location every night, encrypted, and a scripted restore drill runs every month and posts its result. Tags stay suggestions until a person confirms them, and every edit is versioned.

What changes

Finding a file takes seconds instead of an afternoon, and sending a publisher a print-ready set is one message.

$ trace — how it works

capturestorageenrichindexserve Scanner600 dpi TIFFPhone uploadshared albumStudio NASold archiveObject storageS3 · versionedEvent queueone job per fileDerivativesweb · print · ICCVision modeltags · captionsEmbeddingsvector per imageImage CDNAVIF · WebPPostgrespgvector + tagsSearch APIkeyword + semanticCataloguebrowse · filterChat assistantask in plain wordsOff-site copyencrypted nightlyRestore drillmonthly, scripted
  1. 01Scanner · phone · NAS
  2. 02Object storage
  3. 03Event queue
  4. 04Derivatives · vision · embeddings
  5. 05Postgres + pgvector
  6. 06Search API
  7. 07Catalogue · chat

$ open ./result — what it would look like

archive.local/forest SeriesForestRiversFolkloreSketches Year202620252024 forest_01.jpg forest_02.jpg forest_03.jpg forest_04.jpg forest_05.jpg forest_06.jpg forest_07.jpg forest_08.jpg forest_09.jpg
The catalogue: every scan named, tagged and findable. (Illustrative.)