Skip to content
In this case study

MTL Archives

13,000+ historical photographs transformed into an AI-powered archive, search engine and daily game, with an automated editorial system reaching 2.66M social views.

My role

I designed and built the data pipelines, image enrichment, visual search, model evaluations, website, payments and editorial automation. The reach was measurable; whether it brought people back to the archive became the next question.

Role
Independent researcher, designer & engineer
Scope
Data pipelines · Search research · Brand · Product · Distribution
Period
2025–present
Download PDF

At a glance

  • 13,499

    Deduplicated serving records · June 2026 audit

  • 14,715

    Image embeddings in the January research projection

  • 2.66M

    Facebook + Instagram views · January–July 2026

Following an interest in the city

The archive was public, but finding a photograph still depended on knowing what to ask for. I built the ingestion, enrichment, search, website and distribution pipeline to make that collection easier to use. The social accounts reached about 2.66 million views in seven months; the website recorded 13,783 page views over that period. Learning why those numbers were so far apart became part of the project.

I don't remember exactly what led me to Montréal’s open data portal. I was interested in the city and followed that curiosity. I started browsing the datasets without a particular project in mind.

Among them, I came across the open photographic archive and aerial photothèque. There were photographs of streets, aerial views and records of the city at different points in its history. That was the collection I wanted to spend more time with. As I explored it, I started wondering how someone could find a photograph without already knowing its title or catalogue reference.

I started building MTL Archives in October 2025 as a way to explore that collection. It became a search engine, then a daily location game and a print-order flow. Along the way, it became a research project about the data itself: what a model notices in an old photograph, what makes a search useful, and whether attention on social media leads people back to the archive.

My work spans the ingestion scripts, model experiments, website, visual identity and daily editorial pipeline. The public product now presents 13,000+ records. The larger research datasets include earlier versions and duplicates, so I keep their counts separate.

Loading photograph…
An aerial photograph without the municipal document border. The January analysis placed many of these images together. Open archive record ↗

The first problem was the description

An early audit found that about 97% of the working records had synthetic descriptions. That needs a distinction: my cleaning pipeline had filled missing text with a title, date and archive reference. Those fallback sentences made the rows look complete without adding much meaning. They were not rich descriptions supplied by the city.

That matters for semantic search, which looks for meaning rather than an exact word match. If thousands of records say little more than “photograph, reference, date,” a better search model still has very little to work with. I needed to improve the evidence behind each result before improving its presentation.

Building the collection in layers

The extract, transform and load (ETL) pipeline downloads the city catalogues in batches, normalizes their fields, links records to image files and produces a manifest: a structured inventory of the collection. The source archive reference, or cote, stays attached to the record. Missing files, duplicate records and uncertain locations are data problems to track, not details to hide in the interface.

Some useful words are printed inside the image. I used Tesseract, an optical character recognition (OCR) engine, with French and English support to extract them. A vision-language model, which can interpret an image and produce text, supplies a separate description. Original metadata, extracted text and generated descriptions remain distinct.

From an archive export to a public product

Start with the city’s records.

Download the photographic archive and aerial survey catalogues in batches. Keep source identifiers and file references so each image can be traced back to its record.

Source records → downloaded files
The preparation jobs run separately from the public website. A visitor’s search does not recaption the collection.

The preparation jobs use Python and TypeScript. Cloudflare R2 stores images; D1 stores records; Vectorize stores the numerical representations used for similarity search. The public Worker answers requests, while the Next.js website presents the results.

The ingestion code streams records in batches and writes checkpoints so an interruption does not require starting over. Failures have their own logs. The serving database is also distinct from the research corpus: a June audit recorded 13,499 deduplicated production records and 14,822 development records. A larger file or vector count does not mean the website has that many distinct photographs.

Captioning was a job to measure

The first recorded large captioning run used LLaVA, the Large Language and Vision Assistant 1.5 7B on a rented graphics processing unit (GPU). In December, the run report recorded 12,304 captioned images, 23 errors, 10.86 hours and $14 in compute. It was a partial run, with roughly 2,500 images still remaining.

By May, I had moved to structured outputs: descriptions, visual categories and fields that later tools could use. The full-run report contains 14,822 output rows, including 14,706 captions, 79 captions that failed the required structure and 116 image or model errors. A valid structure tells me the program can read an answer; it does not prove the description is historically correct.

That run also had to recover from interrupted work. The first 12,100 rows were reconstructed from saved chunks, then the remaining 2,722 were resumed. The report estimates 24.47 GPU hours, but the recovered portion does not have the same complete timing record as the tail. I keep it as an estimate rather than presenting it as a measured invoice.

What the image model was noticing

For visual search I used CLIP, Contrastive Language–Image Pre-training, which represents images and text as numerical vectors. Similar vectors can help match a phrase to a picture. I also built an explorer to inspect the collection rather than judge the system only through a few search queries.

In January, the projection of 14,715 image embeddings showed a striking split: many plain aerial photographs grouped separately from survey documents with municipal headers and borders. Index cards formed their own tight group. The formatting was part of the signal, even when the underlying subject was similar.

The projection used UMAP, Uniform Manifold Approximation and Projection, to turn high-dimensional vectors into a view I could inspect. It suggested what to investigate; distances on that map were not proof of semantic similarity or a controlled explanation of the model.

One image, two different jobs
Loading photograph…Original archive image
Dimensions 1–64 of 512 · signed values

CLIP turns the image into 512 numbers. These values describe learned visual features. Search compares vectors for similarity; it does not treat any single bar as a named concept like “street” or “building”. The chart shows the first 64 actual values, using a fixed scale across images.

Actual saved outputs from the research explorer. Changing the image reveals its recorded vector or position; no model runs in your browser. UMAP summarizes a collection of vectors, rather than transforming one photograph in isolation.

That led to a practical question: should borders and document framing be removed before indexing? A later experiment tried deterministic cropping and tone adjustment on 12 flagged images. Two changed their predicted category. That was enough to justify reviewing individual cases, not enough to justify automatically cropping the archive. Borders can contain evidence worth preserving.

Stepping back to see the whole collection

A search result shows what the system found. I built the MTL Archives Explorer to ask a different question: how had it organized the collection? Each of its 14,715 points represents one photograph. I could move through the projection, open a record and compare a group of images instead of guessing from a handful of thumbnails.

The shape of 14,715 photographs

Select a region to highlight it. These colours follow the explorer’s eight annotated regions; they stay attached to the same photographs as you rotate the view.

Drag to rotate, or use the controls. The flat view preserves the saved UMAP layout. Date depth uses the explorer’s recorded-date mapping and seeded spacing, not a third UMAP dimension. Colour assignment follows the nearest of eight saved reference centres in the flat projection. The names are research annotations, not labels produced by UMAP or verified categories for every photograph. Directional names do not mark areas of Montréal. Open the full explorer to inspect individual photographs ↗

The flat view uses the saved UMAP coordinates. The 3D view keeps that layout and adds depth from the recorded date, with a little spacing to avoid stacked points. Its shape is a way to navigate the evidence, not the geography of Montréal or a third dimension discovered by UMAP. Colour views offer other ways to inspect it, including dates, photographer labels and annotated visual groups.

That made the split between aerial photographs, framed survey documents and index cards easier to investigate. I could move from an unusual group back to the images that formed it. Labels and anomaly highlights were prompts for review, not new archival facts. The useful outcome was a more specific question about the influence of document formatting, which led to the small cropping experiment.

The explorer is built with Three.js, a browser graphics library. It loads the saved positions and record identifiers first, then fetches the larger vector file when needed. Similar-image lookup compares the original CLIP vectors, not distances on the flattened map. Selected records can be collected and exported for follow-up. I kept this research workspace separate from the main site's simpler search, game and print flows.

Trying a replacement before changing the index

Should I replace CLIP, the image-and-text model behind search? On a diagnostic test of 500 images and 13 queries, CLIP’s mean reciprocal rank was 0.8269, compared with 0.4484 for SigLIP. I kept the existing index. A new model was not an improvement just because it was newer.

Recorded measureCLIP ViT-B/32SigLIP base
Mean reciprocal rank0.82690.4484
Queries with a rule match in top 513 / 138 / 13
May 26, 2026 · 500 images, 13 queries. Expected matches came from generated category/theme rules, not independent human relevance judgments.
Search the archive
Try a colour, a scene or a Montréal landmark

6 results for “tramway” · saved starting selection

The same smart search used by MTL Archives, combining text and visual retrieval. Landmark names can match catalogue descriptions; colours and scenes invite broader visual associations. Open a photograph for its full record. Images: Archives de la Ville de Montréal. Continue on MTL Archives ↗

Mean reciprocal rank measures how early the first expected match appears. The report’s “P@5” field actually measures whether any expected match appears in the first five results, not precision. Expected matches came from category rules rather than independent human judgments, so this test supported a local decision, not a general model ranking.

Broad ranking boosts from generated categories also underperformed the existing search on the recorded query set. I kept those signals available for inspection rather than letting them change scores. Better relevance labels were the next useful investment.

Designing a way into the archive

I worked through the identity and product states in Paper. The board covers the logo, typography, search, photo detail, game, print ordering, empty states and emails. That let me consider the same photograph as a search result, an archival record and a daily invitation to explore.

mtl archives
The product’s dotted rosette, rendered as a vector. The Paper exploration tested dense, medium, minimal and monochrome versions; this is the compact mark used by the site.

The mark draws on Montréal’s civic emblem. I tested it at icon size and kept the rest of the interface quiet so photographs and source references could carry the page.

The interface is bilingual, with familiar places and subjects as entry points. A person arriving from a phone should be able to browse before learning how the catalogue works. On the record, the source reference and location confidence remain available. The design needs to make discovery easier without making uncertain metadata look authoritative.

A reason to return, and a way to collect

The daily location game asks people to place a photograph on a map. It reuses the archive rather than requiring a separate content library. The code keeps challenges and guesses in the backend, while the map and photo controls belong to the interface. A daily game makes a different invitation from search: you can start with curiosity instead of a query.

Print ordering uses Stripe Checkout. The app validates the shipping details and quote, then a signed payment notification triggers confirmation and fulfilment emails. The print work itself remains manual. A successful browser redirect is not treated as proof of payment.

The newsletter is another return path. Signing up is explicit; playing the game does not subscribe someone automatically. Subscription state and delivery history live in the database. The daily scheduler checks Montréal’s local time, including daylight saving changes, rather than assuming the same server hour means morning all year.

These are working product surfaces, but their existence is not evidence of strong conversion. The saved business notes describe revenue as weak relative to attention. That is the next product problem, not a result I can claim to have solved.

Taking the archive to the feed

A searchable website still needs people to find it. I began using Instagram and Facebook as editorial experiments: exact places and dates, street transformations, lost landmarks and questions about what used to be there. Each post had to be worth looking at even if the viewer never ordered a print.

The daily pipeline selects an archive record, assembles its source material, drafts bilingual copy and produces a carousel or reel package. An image interpretation or a web search can suggest context, but unsupported exact locations should not quietly become facts. The code tracks location confidence and can reject copy that reintroduces a place name the evidence does not support.

Generating a package and publishing it are separate events. The pipeline records attempts, successful post identifiers and permalinks, which helps prevent duplicate delivery and makes later analysis possible. The home server holds operational state; an optional Obsidian mirror holds the editorial notes. A local fallback can prepare a package when the server is unavailable.

One archive, two editorial experiments

I tested two editorial directions: dramatic openings about lost or transformed places, and a more documentary voice built around a place, date and visible detail. The January pipeline assembled researched context and captions; later revisions separated reels from carousels and filtered generic mystery language.

The March 31 snapshot covered 135 first-quarter posts. Facebook reels with loss or erasure language averaged 100,260 views, against 35,974 without it. On Instagram, place-and-date openings averaged 4,866 views, against 2,315 without them. These were observational comparisons: subjects, formats and publication dates changed together.

February made the difference clear. Its 18 Facebook reels averaged 62,519 views in the saved cohort. Instagram’s five carousels averaged 9,468, compared with 1,533 for its 17 reels. One platform’s strongest format was not a recipe for the other.

I moved toward the documentary voice and made that choice repeatable in the caption pipeline. In March, Instagram shifted to 12 carousels and eight reels; monthly account views rose from 48,996 to 68,263. Facebook reel averages fell. That does not isolate the effect of a caption change, but it changed what I wanted to keep testing: specific local context and repeat archive use, not just the largest feed number.

Image viewer
The Miron quarry reel, published January 26. Captured September 9: 200,234 Facebook views and 3,739 Instagram views. These are cumulative post results, not February-only views. The same reel travelled very differently on the two platforms.

The reach was real. It did not stay there.

The saved January–July reports total about 2.66 million Facebook and Instagram account-level views. February was the peak: 1,363,500 Facebook views and 48,996 Instagram views. By July, Facebook was down to 6,560 while Instagram recorded 25,787. Showing only the peak would miss most of what the experiment taught me.

The spike, and what came after

February: 1,363,500 Facebook views, 48,996 Instagram views and 4,747 website page views from 875 reported visitors.

January–July 2026 saved reports. Each measure has its own scale. March/April use saved monthly summaries; May has partial social coverage. Website January covers January 1–30; later months include dashboard captures. These are separate measures, not an attributed conversion funnel.
Image viewer
Facebook, February 1–28, 2026, captured September 9. Meta rounds the views total to 1.4M and reports 447.1K viewers. The saved report supplies the more precise 1,363,500 views used above.
Image viewer
Instagram, February 1–28, 2026, captured September 9. The headline 197.6K includes 148,600 Facebook views and 48,996 Instagram views. I use the Instagram breakdown, not the combined headline, in the monthly chart.

By September 9, Meta Business Suite showed approximately 8.7K Facebook followers and 3.7K Instagram followers. Those are account totals at the time of review, rather than followers gained during the February experiment.

The saved March 19 account export recorded 3,340 followers and 425 posts on Instagram (@mtlarchives). That is a dated account snapshot, separate from the view totals above. You can also explore the published work on Facebook.

A small number of reels accounted for much of the observed Facebook attention. In the saved August post snapshot, the top five unique January reels represented 82.4% of that month’s reel cohort’s cumulative views. These are lifetime post counts, not views accrued during January, so I do not add them to the monthly account totals.

Concrete places and a reason to be curious appeared repeatedly among the stronger posts. On Instagram, examples included Parc Marquette, 1969 and Avenue du Mont-Royal at Saint-Denis, 1928. The saved analysis found different patterns for documentary carousels and curiosity-led reels. These were observational comparisons; timing, format and platform distribution changed together, so they do not establish a causal recipe.

The data needed its own cleanup. Facebook’s published-post export contained 196 reel rows for 98 unique reels. Counting the rows as separate pieces of content would double the denominator. Some monthly reports also lacked preserved daily exports. Those limitations remain in the supporting notes instead of being smoothed into an uninterrupted growth story.

What happened on the website

February brought 875 reported website visitors and 4,747 page views. March had fewer visitors, 714, but more page views, 5,296. The seven-month reports contain 13,783 page views in total. I do not sum monthly visitor counts and call that a unique audience: the same person can return in several months.

The gap between feed attention and website activity is the important result. The available exports do not reliably connect an individual post to a visit, game session or order. I can describe when activity rose and fell, but I cannot turn the social totals into an attributed conversion rate.

The later website captures show the smaller scale clearly: July recorded 130 visitors and 442 page views; August 1–10 recorded 33 and 122. The partial August period is kept out of the full-month chart. Better campaign links, preserved month-end exports and product events are needed to tell which forms of discovery lead to repeat archive use.

What changed how I work

The first mistake was making missing descriptions look complete. I now keep source metadata, extracted text and generated captions separate, and preserve the reference that lets someone check each one.

The experiments stopped me from rebuilding an index without evidence. Next time, I would create a small set of independently reviewed search tasks earlier. A category match is not the same as a useful answer. The explorer made a related point: a compelling cluster may reflect a document border rather than its subject.

Distribution taught me to measure what happens after attention. I would add campaign links and product events from the start, preserve month-end exports, and review the actual published captions. A successful pipeline run says nothing about whether someone returned, played, or ordered a print.

The code and source notes for this case study show the implementation and limits behind the story. If you are opening up a difficult collection, evaluating search, or building a product around specialist data, I’d be happy to talk.