← ./work / TIDE.archive
TIDE.archive
A searchable archive for thirty years of an illustrator's work
An illustrator has about 40,000 files spread across a NAS, two laptops and a phone, with names like scan_0042_final2.tif. TIDE is a study of an archive that sorts itself as files arrive and can answer a request like "the forest series, print quality, everything from 2019".
- Type
- Automation · AI search
- Year
- 2025
- Role
- Architecture, pipelines, search
- Status
- Concept study
Concept A study of how I’d approach this kind of problem, not a client project.
$ cat ./targets — design targets, not measured results
- 40,000
- files indexed
- < 1 s
- typical search
- 3
- copies of every original
- monthly
- restore drill
Capture
Anything dropped into the inbox, synced from the phone album or migrated from the old NAS lands in S3-compatible object storage with versioning on, so an overwrite never loses the original. Each new object fires an event.
Enrichment
One queued job per file, fanned out to three workers. The first makes derivatives with libvips: a print master that keeps its ICC profile, and web sizes in AVIF and WebP. The second runs a vision model that writes a caption and suggests tags from the studio's own vocabulary. The third stores an embedding, so images can be found by what they show as well as by name.
Search
Metadata and vectors live together in Postgres with pgvector. The search API blends keyword and semantic results, so "rainy street, blue" works as well as a filename. The catalogue is a fast web app, and the chat assistant answers the same questions in a message with a signed download link.
Safety
Originals are never changed. Storage is copied to a second location every night, encrypted, and a scripted restore drill runs every month and posts its result. Tags stay suggestions until a person confirms them, and every edit is versioned.
What changes
Finding a file takes seconds instead of an afternoon, and sending a publisher a print-ready set is one message.
$ trace — how it works
- 01Scanner · phone · NAS
- 02Object storage
- 03Event queue
- 04Derivatives · vision · embeddings
- 05Postgres + pgvector
- 06Search API
- 07Catalogue · chat
$ open ./result — what it would look like