AI platform for archives, museums, media, and research

Ask your collection.
Get cited answers.

arcafind turns digitized collections — documents, handwriting, images, audio, and video — into searchable, citable knowledge: structured, curatable, exportable, and publishable.

Media types
Documents · Image · Audio · Video
Search
Full text + semantic
Evidence
Evidence anchor per statement
Hosting
EU · Frankfurt
Fig. 01

The problem

Digitized is not automatically accessible.

Scans stay silent

Image files alone leave content, people, and relationships invisible.

Manual cataloging does not scale

Tens of thousands of pages cannot be transcribed by hand economically.

Keyword search falls short

Historical terms, spelling variants, and visual clues need hybrid search.

Fig. 02

Workflow

From digitized collection to research portal.

  1. 01

    Ingest collections

    Scans, documents, and media files — organized into collections and projects, with flexible metadata.

  2. 02

    Apply analysis profiles

    Your domain knowledge becomes a versioned analysis profile: the AI transcribes, detects entities, and structures metadata — with evidence and confidence per field.

  3. 03

    Review curatorially

    Your team stays in control: manual corrections take precedence over authority data and AI — traceable per field.

  4. 04

    Search and publish

    Hybrid search, publications, and standard exports (EAD, CIDOC CRM, CSV, JSON) open up the collection — internally, for clients, or publicly.

Fig. 03

Use cases

One platform, many collections.

arcafind is content-agnostic: your organization describes its knowledge, and the platform compiles the technology. Access can be tiered — internal for your team, protected for clients, open to everyone.

Museums & collections

Provenance research, acquisition records, object and image archives — opened up to the individual page, compatible with authority data and GLAM standards.

City & state archives

Files, finding aids, and estates become searchable — with evidence anchors that trace every statement back to the scan.

Public authorities & administration

Make historical and current records researchable — EU-hosted, GDPR-compliant, with tiered access rights.

Universities & research

Open special collections and research corpora for scholarly work — citable, with a source reference for every answer.

Publishers & media companies

Text, image, and AV archives become a product foundation: research tools for newsrooms, new digital offerings for audiences.

Content creators & creative industries

Structure and search your own media data — documents, images, audio, video — and offer it as new access points or products.

Discovery

Research starts with questions, not just matches.

Combine full-text, semantic search, image similarity, and source evidence to verify context instead of working through result lists alone.

Live demo on the public reference collection of Museum Ulm.

Why arcafind

More than an answer machine for archives.

arcafind is built for institutions: AI helps with enrichment, while your organization keeps control over sources, metadata, publication, and standards.

Institutional work at the core

Organizations, projects, collections, and publications reflect how institutions actually work — multi-tenant and multilingual.

Review before publication

Transcriptions, entities, and canonical versions remain inspectable and correctable.

Configurable analysis profiles

Prompt, schema, examples, and test cases are generated from your knowledge — versioned, testable, activatable per project.

Standards and preservation logic

Originals, derivatives, structured metadata, and exports to EAD, CIDOC CRM, CSV, and JSON are core platform concerns.

Trust

Built for sensitive data.

EU data residency

Architecture for Supabase Frankfurt and Vercel EU regions.

No training use

Designed for paid Gemini APIs without training on client data.

Preservation and derivatives

Originals stay preserved while web and AI derivatives are generated traceably.

Open standards

Export paths for EAD, CIDOC CRM, CSV, JSON, and XML.

Access & products

Collections become products.

Publish curated publications, research rooms, and search surfaces — internal for your team, protected for clients and partners, or public for researchers and the interested public.

Fig. 04

Reference project

Museum Ulm — Directorate files 1933–1945

The museum directorate's correspondence, digitized and searchable: around 30,000 pages in Sütterlin and Fraktur, opened up for provenance research — funded by the German Lost Art Foundation.

30,000+ pagesSütterlin & FrakturEvidence anchors & GND authority data

Ready to make your collection findable?

Start with a project, a collection, or a concrete research question.

Request access