Read + write surface with the current classification model.
Status: This is the active and only public API surface. (The legacy versioned API has been retired.) Owner: Branko · Last reviewed: 2026-06-23 Source of truth:
app/api/v2/**route handlers + this file. Taxonomy tables auto-generated from the database viatsx scripts/generate-taxonomy-docs.ts.
Machine-readable spec. An OpenAPI 3.1 description of every endpoint (incl. webhooks) is published as
/openapi-v2.yamland/docs/openapi.json. Import either into Postman or Insomnia, or feed it to an OpenAPI client generator (e.g.openapi-generator) to scaffold a typed client. The spec is hand-maintained, Redocly-linted in CI, and guarded against endpoint drift by a contract test (e2e/contracts/openapi-spec.spec.ts) that fails if a route and the spec disagree. This document remains the authoritative reference for semantics, vocabularies and examples.
extracted_metadata — what is inside itGet from zero to a delivered webhook in five steps.
A platform SuperAdmin creates your organization in the iDMS dashboard and assigns you as Org Admin. Then go to Settings → API Keys and create a key. The plaintext is shown once — copy it immediately. Keys live in idms_… format and are scoped to your organization.
Send the key on every request as a Bearer token:
curl -H "Authorization: Bearer idms_…" \
https://idms.mory.ai/api/v2/documents
Use external_id as a stable client-side identifier so retries don't create duplicates:
curl -X POST https://idms.mory.ai/api/v2/documents/upload \
-H "Authorization: Bearer idms_…" \
-F "file=@invoice-001.pdf" \
-F "external_id=client-system-2024-1041"
Second call with the same external_id returns the existing document with meta.idempotent_replay: true — no 409, no duplicate document.
In Settings → Webhooks, register your endpoint and check the events you want. Webhooks are POSTed with Content-Type: application/json, an X-IDMS-Signature header (HMAC-SHA256, hex), and an X-IDMS-Event header.
// Express — verify over the RAW body, constant-time. See Webhooks → Verifying the signature.
import crypto from "node:crypto";
app.post("/idms-webhook", express.raw({ type: "*/*" }), (req, res) => {
const expected = crypto
.createHmac("sha256", process.env.IDMS_WEBHOOK_SECRET)
.update(req.body) // req.body is the raw Buffer, not parsed JSON
.digest("hex");
const sig = req.headers["x-idms-signature"];
const ok = sig && crypto.timingSafeEqual(Buffer.from(sig, "hex"), Buffer.from(expected, "hex"));
if (!ok) return res.status(401).end();
const event = JSON.parse(req.body.toString()); // parse only AFTER verifying
// ... handle event.event
res.status(200).end();
});
document.processed eventWhen the AI pipeline finishes, your endpoint receives:
{
"event": "document.processed",
"timestamp": "2026-06-07T08:00:00.000Z",
"document_id": "550e8400-…",
"external_id": "client-system-2024-1041",
"filename": "invoice-001.pdf",
"doc_type": "invoice",
"doc_subtype": "service_invoice",
"tags": ["plumbing"],
"extracted_metadata": { "total_amount": "256.74", "currency": "CHF", "iban": "CH93 …" }
}
That's it. Pull the document state from GET /api/v2/documents/{document_id} if you need more than the webhook carries.
This is a read + write + upload API with the current classification model: endpoints for reading documents with the full classification surface, PATCH endpoints for manual corrections, idempotent upload, per-client output profiles, and six webhook event types. All routes live under /api/v2/*.
Highlights:
absender / empfaenger / cc).PATCH for classification / metadata / contacts / tags + bulk PATCH + tag add/remove — every write goes into an append-only audit log.external_id or repeated content hash returns the existing document; no 409 on duplicate filename (designed for high-volume bulk ingestion of 200 000+ files).GET /api/v2/documents?updated_since=<ts> plus the meta.server_time cursor lets you poll only what changed.abgeschlossen / finalized / deleted reject writes with 400 DOCUMENT_LOCKED. Reads always succeed.document.failed, classification.changed, tag.suggested, usage.threshold.reached) with an HMAC-SHA256 envelope.One key pool, one set of limits — the API key from Settings → API Keys works on every /api/v2/* endpoint with no extra setup.
Authorization: Bearer idms_<api_key>.404, never 200 with foreign data.429):
X-RateLimit-Limit: 100X-RateLimit-Remaining: <int>X-RateLimit-Reset: <unix seconds>Field-name note. Some JSON field names carry a
_v3suffix —document_classification_v3,intent_v3,v1_projection. Treat that suffix purely as a stable contract identifier, not an API version: a field named with_v3belongs to the current classification model, andv1_projectionmaps each value back to the legacy 16-type taxonomy that the API still surfaces unchanged. The suffix never changes, so you can hard-code these field names safely. The rest of this document just calls it "the classification model."
The classification model is opt-in per organization. By default an org runs the legacy extractor — the legacy model is the safe default so existing integrations never get a surprise upgrade. Opting in flips the org to the current model and adds the richer fields to every new document.
What your integration sees depends on your org's flag:
| Field on the read response | Org on the legacy default | Org opted in (current model) |
|---|---|---|
Top-level doc_type, doc_subtype, doc_intent | legacy vocabulary, populated (invoice, service_invoice, …) | legacy-vocabulary projection of the current-model value (same vocabulary, same semantics — backward-compat fallback) |
Nested document_classification_v3 | null — pipeline never ran the current extractor | populated with doc_type, doc_bereich, doc_subtype, intent_v3, status_lifecycle, zahlungsstatus (German wire values) |
| Tag fields | legacy vocabulary | legacy vocabulary kept, plus the richer per-document state in document_classification_v3 |
If your client code only reads the top-level fields, you'll see the legacy vocabulary regardless of opt-in — that's the backward-compat surface. To consume the richer model, read from the nested document_classification_v3 object. The mapping back to the legacy vocabulary for every current-model value is the v1_projection column in the taxonomy tables further down.
How to opt in. Contact your account manager. We flip a config flag and run a one-off reclassifier over your existing documents so the nested document_classification_v3 rows appear retroactively. No code change on your side — the surface stays the same, the nested object just becomes non-null.
Already opted in. A number of production organizations are already running the current model; the rest stay on the legacy default until they opt in as described above.
The model defines four orthogonal dimensions every document is classified across, plus a top-level Bereich for human-readable grouping in dashboards and reports.
| Dimension | What it answers | Cardinality |
|---|---|---|
Dokumententyp (doc_type) | What kind of document is this? | one |
Untertyp (doc_subtype) | Which variant within the type? | one (optional, depends on doc_type) |
| Tags | What is it about (substantively)? | many (1–5) |
Intent (intent_v3) | What does the recipient need to do? | one |
Plus structured extracted fields common to every document: status_lifecycle, zahlungsstatus, total_amount, currency, date_issued, booking_date, vat_rate, vat_amount, vat_amount_source, due_date, iban, signature, and linked entities (Bezüge) — property, unit, equipment.
A Bereich (area / domain) is purely a navigational grouping — the AI picks exactly one doc_type, and that type belongs to exactly one bereich. The 8 Bereiche cover the real-estate management domain (Finanzen / Finance, Beschaffung / Procurement, Verträge / Contracts, Immobilie / Property, Korrespondenz / Correspondence, Personal / HR, Vermarktung / Marketing, Auffang / Catch-all).
Several enum values stay German on the wire because they originated from the Swiss property-management domain and are stable contract identifiers — renaming them would break every integrator. Labels and surrounding prose are translated; the values themselves are not. The table below is the canonical mapping between the German wire values and their English meaning.
| Field | German wire value | English meaning |
|---|---|---|
zahlungsstatus | offen | open / unpaid |
zahlungsstatus | bezahlt | paid |
zahlungsstatus | teilbezahlt | partially paid |
zahlungsstatus | ueberfaellig | overdue |
zahlungsstatus | sonstiges | other / fallback |
status_lifecycle | neu | new — just received, not yet triaged |
status_lifecycle | in_bearbeitung | in progress — being worked on |
status_lifecycle | abgeschlossen | completed / locked — terminal write-state |
direction (routing) | eingehend | inbound — addressed to the property manager |
direction (routing) | ausgehend | outbound — issued by the property manager |
direction (routing) | neutral | neither inbound nor outbound |
contact role | absender | sender (from-party) |
contact role | empfaenger | receiver (to-party) |
contact role | cc | carbon copy |
These values are German because they reflect the Swiss real-estate domain vocabulary that originated this dataset; they are stable contract identifiers and will not be renamed even as labels and prose get English translations.
<!-- BEGIN AUTO-TAXONOMY -->Auto-generated from the taxonomy reference tables — DO NOT EDIT BY HAND. Run
tsx scripts/generate-taxonomy-docs.tsto refresh after a taxonomy migration. Current sizes: 8 bereiche, 31 doc_types, 121 doc_subtypes, 13 intents, 10 tag_groups, 54 tags.
finanzen)| doc_type | Label (DE) | Label (EN) | Subtypes | Legacy projection |
|---|---|---|---|---|
rechnung | Rechnung | Invoice | schlussrechnung, teilrechnung, akontorechnung, sammelrechnung, qr_rechnung | invoice |
gutschrift | Gutschrift | Credit Note | — | credit_note |
mahnung | Mahnung | Payment Reminder | — | reminder |
betreibung | Betreibung | Debt Collection | zahlungsbefehl, rechtsvorschlag, fortsetzungsbegehren, verlustschein | debt_collection |
quittung | Quittung | Receipt | schluesselquittung | receipt |
finanzbericht | Finanzbericht | Financial Statement | bilanz, erfolgsrechnung, jahresabschluss, quartalsabschluss, budget, kontoblatt, offene_posten_liste, liegenschafts_abrechnung, nebenkostenabrechnung_hknk, stwe_abrechnung | financial_statement |
bankdokument | Bankdokument | Bank Document | kontoauszug, zahlungsauftrag, dauerauftrag, zinsabrechnung, einzahlungsschein_wir | statement |
steuerdokument | Steuerdokument | Tax Document | steuererklaerung, steuerrechnung, steuerveranlagung, steuerwertschaetzung | statement |
finanzierungsdokument | Finanzierungsdokument | Financing Document | hypothekarvertrag, hypothek_offerte, schuldbrief, darlehensvertrag, kyc_unterlagen | contract |
beschaffung)| doc_type | Label (DE) | Label (EN) | Subtypes | Legacy projection |
|---|---|---|---|---|
offerte | Offerte | Quotation / Offer | — | quote |
bestellung_auftragsbestaetigung | Bestellung/Auftragsbestätigung | Purchase Order / Confirmation | bestellung, auftragsbestaetigung, submission_ausschreibung | quote |
lieferschein | Lieferschein | Delivery Note | — | delivery_note |
vertraege)| doc_type | Label (DE) | Label (EN) | Subtypes | Legacy projection |
|---|---|---|---|---|
vertrag | Vertrag | Contract | bewirtschaftungsauftrag, servicevertrag, wartungsvertrag, hauswartungsvertrag, mietvertrag, dienstbarkeitsvertrag, werkvertrag, nachtrag, aufhebung, kaufvertrag, arbeitsvertrag, rahmenvertrag, verwaltungsvertrag | contract |
reglement | Reglement | Regulation / Bylaws | nutzungsordnung, verwaltungsordnung, begruendungsurkunde, hausordnung | contract |
behoerdenentscheid | Behördenentscheid | Official Decision | verfuegung, bewilligung, urteil_entscheid, einsprache | other |
grundbuchauszug | Grundbuchauszug | Land Registry Extract | — | other |
versicherung | Versicherung | Insurance Document | gebaeudeversicherungsausweis, einzelauszug_sach, einzelauszug_haftpflicht, police, schadensmeldung, versicherungsfall | insurance |
immobilie)| doc_type | Label (DE) | Label (EN) | Subtypes | Legacy projection |
|---|---|---|---|---|
plan | Plan | Plan / Drawing | katasterplan, grundrissplan, gis_auszug, schliessplan | report |
zertifikat | Zertifikat | Certificate | sn_niederspannung, sn_brennerkontrolle, sn_anlageninbetriebnahme, sicherheitsschein, garantieschein, energieausweis_geak | report |
bericht | Bericht | Report | fact_sheet, zustandsanalyse, inspektionsbericht, wartungsbericht, schadensbericht, pflichtenheft, controlling_bericht, mieterspiegel, leerstandsliste, projektdatenblatt, terminplan_bauprogramm, baubeschrieb | report |
protokoll | Protokoll | Minutes / Protocol | bauabnahme, uebergabe, begehung, generalversammlung, ausschusssitzung, sitzungsprotokoll, schluesselprotokoll | minutes |
foto | Foto | Photo | — | other |
korrespondenz)| doc_type | Label (DE) | Label (EN) | Subtypes | Legacy projection |
|---|---|---|---|---|
brief | Brief | Letter | eigentuemerkorrespondenz, behoerdenkorrespondenz, mieterkorrespondenz, revisionsweisung, kuendigung, einladung_traktandenliste | letter |
e_mail | — | letter | ||
formular | Formular | Form | amtl_mietzinsaenderung, amtl_kuendigungsformular, mietbewerbung_anmeldeformular | form |
visitenkarte | Visitenkarte | Business Card | — | form |
personal)| doc_type | Label (DE) | Label (EN) | Subtypes | Legacy projection |
|---|---|---|---|---|
personaldokument | Personaldokument | HR Document | lohnabrechnung, lohndeklaration, arbeitszeugnis, bewerbung_lebenslauf, spesenabrechnung, unfallmeldung, stellenbeschrieb | other |
firmenunterlagen | Firmenunterlagen | Company Document | handelsregisterauszug, betreibungsauszug, statuten_gruendungsakte, aktienzertifikat, vollmacht | other |
ausweisdokument | Ausweisdokument | Identity Document | id_pass, aufenthaltsbewilligung, strafregisterauszug | other |
vermarktung)| doc_type | Label (DE) | Label (EN) | Subtypes | Legacy projection |
|---|---|---|---|---|
marketing_inserat | Marketing/Inserat | Marketing / Listing | inserat, expose_teaser, broschuere_prospekt, kampagne | other |
auffang)| doc_type | Label (DE) | Label (EN) | Subtypes | Legacy projection |
|---|---|---|---|---|
sonstiges | Sonstiges | Other | — | other |
| intent (key) | Label (DE) | Label (EN) | Legacy projection |
|---|---|---|---|
sonstiges | Sonstiges | Other | information |
ablegen | Ablegen | File / Archive | record_keeping |
antworten | Antworten | Reply / Respond | action_required |
bezahlen | Bezahlen | Pay | payment_required |
geltend_machen | Geltend machen | Assert claim | action_required |
genehmigen | Genehmigen | Approve | approval_needed |
kontakt_erfassen | Kontakt erfassen | Capture contact | record_keeping |
kuendigen | Kündigen | Cancel / Terminate | action_required |
pruefen | Prüfen | Review / Verify | action_required |
unterschreiben | Unterschreiben | Sign | action_required |
verbuchen | Verbuchen | Post (accounting) | record_keeping |
verlaengern | Verlängern | Extend / Renew | action_required |
weiterleiten | Weiterleiten | Forward | action_required |
anlagen_gebaeudeteile)| tag (key) | Label (DE) | Label (EN) |
|---|---|---|
heizung | Heizung | Heating |
dach | Dach | Roof |
fassade | Fassade | Facade |
sanitaer | Sanitär | Sanitary |
schliessanlage | Schliessanlage | Locking system |
aufzug | Aufzug | Elevator |
bodenbelag | Bodenbelag | Floor covering |
elektro | Elektro | Electrical |
fenster_tueren | Fenster & Türen | Windows & Doors |
lueftung_klima | Lüftung & Klima | Ventilation & Climate |
hausdienste)| tag (key) | Label (DE) | Label (EN) |
|---|---|---|
reinigung | Reinigung | Cleaning |
abfallentsorgung | Abfallentsorgung | Waste disposal |
hauswartung | Hauswartung | Caretaking |
schaedlingsbekaempfung | Schädlingsbekämpfung | Pest control |
schluesseldienst | Schlüsseldienst | Locksmith service |
umgebungspflege | Umgebungspflege | Grounds maintenance |
winterdienst | Winterdienst | Winter service |
sicherheit)| tag (key) | Label (DE) | Label (EN) |
|---|---|---|
brandschutz | Brandschutz | Fire safety |
ueberwachung | Überwachung | Surveillance / Monitoring |
energie_versorgung)| tag (key) | Label (DE) | Label (EN) |
|---|---|---|
gas | Gas | Gas |
solar | Solar | Solar |
strom | Strom | Electricity |
wasser | Wasser | Water |
zaehlerablesung | Zählerablesung | Meter reading |
finanzen)| tag (key) | Label (DE) | Label (EN) |
|---|---|---|
hypothek | Hypothek | Mortgage |
bank | Bank | Bank |
akontozahlung | Akontozahlung | Down payment |
buchhaltung | Buchhaltung | Bookkeeping |
budget | Budget | Budget |
finanzierung | Finanzierung | Financing |
steuern | Steuern | Taxes |
recht_verwaltung)| tag (key) | Label (DE) | Label (EN) |
|---|---|---|
bewilligung | Bewilligung | Permit / Authorisation |
eigentuemer | Eigentümer | Owner |
garantie | Garantie | Warranty / Guarantee |
streitfall | Streitfall | Dispute / Litigation |
mieter_mietverhaeltnis)| tag (key) | Label (DE) | Label (EN) |
|---|---|---|
mietzinsanpassung | Mietzinsanpassung | Rent change |
einzug | Einzug | Move-in |
auszug | Auszug | Move-out |
kaution | Kaution | Deposit / Security |
mieteranliegen | Mieteranliegen | Tenant request |
mieterwechsel | Mieterwechsel | Tenant change |
untermiete | Untermiete | Subletting |
bau_renovation)| tag (key) | Label (DE) | Label (EN) |
|---|---|---|
kueche | Küche | Kitchen |
badezimmer | Badezimmer | Bathroom |
balkon_terrasse | Balkon/Terrasse | Balcony / Terrace |
bauarbeiten | Bauarbeiten | Construction work |
malerarbeiten | Malerarbeiten | Painting work |
renovation | Renovation | Renovation |
personal_firma)| tag (key) | Label (DE) | Label (EN) |
|---|---|---|
firmenorganisation | Firmenorganisation | Company organisation |
lohn | Lohn | Salary / Wages |
personal | Personal | HR / Staff |
verkauf_einkauf)| tag (key) | Label (DE) | Label (EN) |
|---|---|---|
einkauf | Einkauf | Procurement |
verkauf | Verkauf | Sale |
vermietung_vermarktung | Vermietung/Vermarktung | Letting / Marketing |
Two enum-valued fields ride alongside the classification axes. Both are writable through the PATCH endpoints and seeded by the AI pipeline.
status_lifecycle| Value | Meaning |
|---|---|
neu | Newly ingested, not reviewed |
in_bearbeitung | A human is actively working on it |
abgeschlossen | Reviewed and accepted — locks the document for writes |
sonstiges | Other / unclassified state |
finalized | External system marked complete — locks |
deleted | Soft-deleted — locks |
zahlungsstatus is deprecated, and what replaces it (MOR-1426)zahlungsstatus is what the extraction model answered when asked, for an invoice,
to pick one of five German words. Nothing printed on an invoice says whether it was
paid, so the answer is a guess — and the corpus shows what that costs.
Measured read-only over all 33'161 completed production documents, 2026-09-18:
| n | |
|---|---|
zahlungsstatus = offen on rechnung | 7'483 — essentially every invoice |
| of those, already 90+ days past due on the day they were uploaded | 4'637 |
It is superseded by payment_status, derived from figures the document prints
and the moment it arrived: open · assumed_paid_by_age · paid · cancelled ·
not_applicable, with payment_status_source (due_date · no_due_date ·
age · not_payable · manual) recording how the answer was reached, and
payable saying whether money is owed at all.
Where the two disagree, the derived one is the one computed from printed dates.
paid is never derived — it needs a person or a bank statement.
Nothing is being removed. zahlungsstatus is still written, still returned and
still filterable; payment_status is not on this payload yet, and putting it there
is a deliberate decision rather than a diff, because assumed_paid_by_age is an
inference sitting beside figures your system treats as fact.
zahlungsstatus (only meaningful for Finanzen Bereich docs)| Value | Meaning |
|---|---|
offen | Unpaid |
bezahlt | Paid in full |
teilbezahlt | Partial payment |
ueberfaellig | Overdue |
sonstiges | Other / NA |
Every doc_type and intent row in the reference tables carries a v1_projection column so legacy-vocabulary receivers can interpret a classification from the current model. Read responses also surface the original legacy columns (doc_type, doc_subtype, doc_intent) alongside the current-model fields — see Documents — read.
All read endpoints share the standard envelope { data, meta, error } and the rate-limit headers above.
/api/v2/documentsList documents in the caller's organization, with the current classification fields nested alongside the legacy envelope.
Query parameters (all optional):
| Parameter | Type | Vocabulary | Description |
|---|---|---|---|
page | int | — | Page number (default 1, max 10000). |
per_page | int | — | Results per page (default 25, max 100). |
doc_type | string | legacy | Exact match on documents.doc_type (legacy type, e.g. invoice). |
subtype | string | current | Exact match on the current-model subtype (German wire value from the nested object, e.g. qr_rechnung). |
intent | string | current | Exact match on the current intent (e.g. payment_required). |
status | string | legacy | Exact match on documents.status (legacy pipeline status, e.g. completed). |
zahlungsstatus | string | deprecated | One of offen, bezahlt, teilbezahlt, ueberfaellig, sonstiges. A model guess, asked only for doc_type=rechnung and clamped to sonstiges on anything unexpected. Superseded by the derived payment_status; see the note below. Still written, still filterable, not removed without notice. |
bereich | string | current | Exact match on the current Bereich (e.g. finanzen). |
tag | string | current | Tag key from the controlled taxonomy. Repeatable for OR semantics: ?tag=plumbing&tag=heating. |
betrag_min | number | current | Minimum amount. The parameter is German-named; the field it filters is extracted_metadata.total_amount. Combine with betrag_max for a range. |
betrag_max | number | current | Maximum amount. Filters extracted_metadata.total_amount. |
faellig_after | date | current | Due date on/after this date (ISO YYYY-MM-DD). Filters on the extracted due date. |
faellig_before | date | current | Due date on/before this date (ISO YYYY-MM-DD). Filters on the extracted due date. |
external_id | string | — | Exact match on the caller-supplied external_id — look a document up by your own reference. |
updated_since | timestamp | — | Returns documents with updated_at strictly after this ISO-8601 timestamp, ordered updated_at ascending — delta sync. Pair with the meta.server_time cursor (below). A malformed value returns 400. |
profile | string | — | Apply an output profile. Profile key is provisioned by your account manager. |
The amount and due-date filters are all optional and combinable with every other filter. Parsing is server-side; a document whose stored due date is calendar-invalid is treated as a non-match rather than causing an error.
Blank and unparseable filter values on betrag_min, betrag_max, faellig_after and faellig_before:
| you send | effect |
|---|---|
| the parameter omitted entirely | no filter |
?betrag_min= (empty), or whitespace only | no filter — an empty value is absent, not zero |
?betrag_min=abc (unparseable) | no filter, silently — no 400, no warning |
?betrag_min=0 | a real filter at zero, applied as asked |
?faellig_after=01.01.2024 (not YYYY-MM-DD) | no filter, silently |
Two consequences worth reading twice, because both are easy to get wrong from the caller's side:
0 and empty are different. ?betrag_min=0 filters at zero and removes every document with no extractable total (just over half the corpus — multi-tenant and informational documents legitimately have no single total). ?betrag_min= applies no filter at all. Before 2026-08-09 these two were the same request, which is what the CHANGELOG entry for that date describes.200 with an unfiltered result set and no indication that your filter was dropped. This is deliberately not how updated_since and the usage endpoint's period behave — both return 400 on a malformed value. If your integration depends on a filter having been applied, validate the value before sending it rather than inferring it from the response.Example:
curl -H "Authorization: Bearer idms_..." \
"https://idms.mory.ai/api/v2/documents?doc_type=invoice&zahlungsstatus=offen&tag=plumbing"
Split-container wrapper rows are excluded from this list and from delta sync. When a multi-document PDF is split, the original file stays behind as a data-less container row (doc_type: "other", no classification, no webhook) and the real content lives in its children, which do appear here. Filtering the wrappers out keeps a poller's feed matching the webhook flow and the dashboard list. One exception: an exact ?external_id= lookup returns containers too — the container carries the external_id you supplied on upload while its children carry <external_id>:i/N, so excluding it would make your own lookup return nothing.
Returns paginated data: [...] with meta: { page, per_page, total, total_pages, server_time }. server_time is the server's query time (ISO-8601 UTC) — for delta sync, save it and pass it as the next updated_since so no change is missed between polls. Each row has the legacy envelope fields plus document_classification_v3, document_tags, document_tags_v3.
/api/v2/documents/:idRetrieve the full document with the legacy envelope, the full classification record, structured metadata, related contacts, and linked entities — property / unit / equipment (Bezüge).
Returns the same row shape as the list endpoint, plus a contacts: [...] array of normalized Person/Firma entities each with their role (absender / empfaenger / cc) — full field reference under GET /api/v2/documents/:id/contacts.
extraction_state (top-level on every document row, since 2026-08-18):
Whether the file's text was actually read. It exists because doc_type cannot say so. Every document has a doc_type — the column is NOT NULL and defaults to other — so a document nobody could read and a document the classifier was unsure about look identical in that field. They are not the same thing, and only one of them is our problem to fix.
| value | meaning |
|---|---|
extracted | the file was read; the classification describes the document |
unsupported_format | no extractor exists for this file type. Nothing has read the file. doc_type is other with a null confidence and carries no information about the content |
ocr_failed | an image extractor ran and errored |
ocr_unavailable | vision OCR is not configured |
null | the document was processed before this field existed. Not a claim that extraction succeeded |
If you branch on doc_type, treat extraction_state: "unsupported_format" as "unclassified, and not because the model was uncertain" — retrying, re-uploading the same file, or asking for a manual review will not change it. Supporting the format is work on our side.
The field is additive: nothing was removed or retyped, and a consumer that ignores it sees exactly what it saw before.
Extracted entity fields (top-level on every document row; each is null / [] when the document doesn't reference one):
| Field | Shape | Description |
|---|---|---|
sender | { name, address?, reference_id? } | The entity that ISSUED the document (the "From") |
receiver | { name, address?, reference_id? } | The entity ADDRESSED in the document (the "To") |
related_contacts | [{ name, role?, reference_id? }] | Contacts mentioned but neither sender nor receiver (the "CC") |
property | { name, address?, reference_id? } | Real-estate property referenced in the document |
unit | { name, type?, property_reference? } | Sub-unit within a property (apartment, parking, storage) |
equipment | { name, type?, reference_id? } | Specific asset referenced (heating, elevator, …) |
signature | { signed, signed_by?, signed_at?, confidence?, raw? } | Contract signing status (see the Signature changelog entry) |
sender / receiver / related_contacts are lightweight extraction snapshots embedded in the document row. The org-level deduplicated contact records (with emails, phones, registry numbers, …) live on the contacts array / endpoint below.
When called with ?profile=…, the response is rewritten per that profile — see Output Profiles.
/api/v2/documents/:id/contactsStandalone contacts list for one document. Identical data to the embedded contacts array on the detail endpoint — separated so clients can refresh just the contacts without re-pulling the full body.
Contact object fields:
| Field | Type | Description |
|---|---|---|
id | uuid | Org-level contact id (see stability note below) |
kind | string | person or firma (legal entity) |
role | string | Role on THIS document: absender / empfaenger / cc |
name | string | Display name ("Mark Müller" / "Müller AG") |
first_name, last_name, salutation | string | null | Person only |
date_of_birth | string | null | Person only — ISO YYYY-MM-DD |
function | string | null | Person only — role/job title as stated in documents (e.g. "Geschäftsführerin") |
email, phone | string | null | Primary email / phone |
emails, phones | string[] | All known emails / phones |
address | string | null | Full address as one string |
website, vat_number, hr_number | string | null | Firma only — website, UID/VAT, commercial-register number |
external_id | string | null | Caller-supplied id when the contact was imported |
{
"data": [
{
"id": "f7a34862-122b-4d45-b89c-836a9d715f28",
"kind": "person",
"role": "empfaenger",
"name": "Christian Käppen",
"first_name": "Christian",
"last_name": "Käppen",
"salutation": "Herr",
"date_of_birth": "1981-08-16",
"email": "c.kaeppen@example.ch",
"emails": ["c.kaeppen@example.ch"],
"phone": null,
"phones": [],
"address": "Ledergasse 11, 6004 Luzern",
"website": null,
"vat_number": null,
"hr_number": null,
"external_id": null
}
]
}
Contact id stability: contacts are deduplicated org-wide; when duplicates are merged, the surviving contact keeps its id and the merged duplicates' ids disappear. Treat document_id as your stable key and re-fetch contacts rather than persisting contact_id long-term. Document ids never change.
/api/v2/documents/:id/downloadReturns a short-lived signed URL that points at the document's original file bytes on Supabase Storage. The bytes never travel through this API — the client follows the URL directly to Storage (CDN-fronted, no extra auth).
Why a signed URL instead of streaming the bytes: avoids the serverless function duration ceiling on large downloads, keeps the function cheap, and lets the client hand the URL to any download tool unchanged.
Query params
| Name | Type | Default | Notes |
|---|---|---|---|
ttl | int (seconds) | 3600 | URL validity. Clamped silently to [60, 86400] — out-of-range values are not rejected. |
Response (200)
{
"data": {
"url": "https://<project>.supabase.co/storage/v1/object/sign/documents/...",
"expires_at": "2026-06-08T15:30:00.000Z",
"filename": "rechnung-q3.pdf",
"content_type": "application/pdf",
"file_size": 1842
},
"meta": { "ttl_seconds": 3600 },
"error": null
}
Client flow
url=$(curl -sS -H "Authorization: Bearer $IDMS_API_KEY" \
"$IDMS_BASE_URL/api/v2/documents/$DOC_ID/download" | jq -r .data.url)
curl -o "$DOC_ID.pdf" "$url"
Errors
| Status | Body | When |
|---|---|---|
| 401 | Invalid or missing API key | Missing or unknown Bearer token |
| 404 | Document not found | Document does not exist OR belongs to another org (no leak) |
| 429 | Rate limit exceeded. | 100 req/min/key cap |
| 500 | signed URL generation failed: … | Supabase Storage signing call returned an error |
Cross-org isolation: the lookup is scoped by both id AND organization_id. A key from one org asking for a document that belongs to another org receives 404, not 403 — existence is never disclosed.
Sammel-PDF caveats (relevant only if the document went through A2 split):
is_container = true) point at the original multi-document PDF — caller receives the whole bundle.source_document_id != null) share the parent's storage_path — caller receives the same bundle. The per-segment text lives in extracted_content on the child row and is already exposed through GET /api/v2/documents/:id; the original file is not re-encoded per child.extracted_metadata — what is inside itUntil 2026-08-26 this object was documented as an opaque bag. The OpenAPI spec
declared it { type: [object, "null"], additionalProperties: true } and named
nothing inside, so a client integrating against the spec could not learn that
total_amount exists, never mind vat_amount.
Nothing about the API changed. Every field below has been in the response
since the day it was first written — the read endpoints emit
extracted_metadata wholesale. What changed is that the spec now names them.
Read the counts as ages, not as reliability. A field near the bottom is
usually recent, not broken: vat_amount sits at 15 documents because the
historical backfill was deliberately not run, which is a decision recorded
in MOR-1001, not a failure rate.
Absent, null and empty are three different things. A key that is missing
means "not read". A key holding JSON null means "looked for, not found". Never
read either as zero — just over half the corpus legitimately has no single
total, which is why betrag_min=0 and betrag_min= are different requests.
Keys beginning with _ are internal and are stripped before serialisation.
They will never reach you. They exist and are numerous — _desc_enriched alone
is on 9'575 documents — so if you are reading a database export rather than the
API, expect them there and ignore them.
All counts measured 2026-08-26 across 33'065 production documents.
| field | type | documents | notes |
|---|---|---|---|
summary | string | null | 32'873 | short machine summary |
description | string | null | 32'865 | longer machine description |
suggested_tags | array | 31'842 | empty on all 31'842 — zero non-empty, ever (MOR-1010) |
total_amount | string | number | null | 31'829 | string on 15'762, null on 16'065, number on exactly 2. Parse defensively. Bare number, no currency symbol |
amount_direction | credit | null | new | MOR-1438 (2026-09-16): whose favour total_amount is in, when the page SAYS so — credit when the document's own total is a printed credit (c54c909a, an invoice, ends Total zu Ihren Gunsten -1'167.70; doc_type stays invoice, per MOR-1429). total_amount already carries the sign; this names why. Present only when the page's own credit label licensed the sign (not merely a negative total corroborated by the positions or by subtotal+VAT — see amount_direction_source). Absent on every other document, positive or negative. Not yet measured against production. |
amount_direction_source | printed_credit_label | null | new | how amount_direction was determined; the only value today, present exactly where amount_direction is credit (MOR-1001: a value and its provenance travel together) |
subtotal_amount | string | null | 0 | The printed Zwischentotal before VAT, bare number. Never derived from total − VAT and never summed from the positions. New: written from 2026-08-27, so an absent key means "not read yet", while a stored null means the document prints none |
adjustments | array | null | new | MOR-1289. What the document prints between its positions total and its final total, in its own order: Rabatt, Skonto, Anzahlung, Gutschrift, Rundung, Versandkosten. Each row is { label, amount } — the document's own wording, and the amount signed as printed (a deduction is negative, a surcharge positive). These are neither positions nor VAT. Printed only, never computed: where the positions miss the subtotal and the document names no reason, this is null rather than a difference we derived. New: written from 2026-09-09, so an absent key means "not read yet" and a stored null means the document prints none. Measured before it existed: 44 of 244 production documents carrying positions, a total and a stated VAT amount had a deduction on the page and nothing in the data |
currency | string | null | 31'829 | value on 24'697, null on 7'132 |
due_date | string | null | 31'828 | ISO date. MOR-1402: where the model read none and the page states a term in days and an issue date, assembled for an invoice or a reminder — see due_date_source. A printed due date is never replaced. |
due_date_source | string | null | new | derived_from_terms when due_date was assembled from payment_terms + date_issued (issue date + N calendar days; the net term where a Skonto term names two; refused for sofort and for a Skonto deadline alone); null otherwise. Invoice and reminder only.. MOR-1414: printed_label when the model read no due date and the page prints one under a due label |
skonto | object | null | new | MOR-1411 (2026-09-15): an early-payment discount the page OFFERS — information, never a position or a deduction. { pct, amount, until, days, amount_with_skonto, corroborated }: amount and amount_with_skonto only as PRINTED and corroborated against total_amount (±0.05), null when only the rate is printed (never computed); until the ISO date the offer names, days the term in days where the page states one instead. Null when the page prints no Skonto, says it is bereits abgezogen, prints two rates, or a printed figure does not close. total_amount stays the invoice amount; every deduction reader refuses a figure equal to this Skonto. Book on payment, not on receipt. |
payment_terms | string | null | 31'827 | free text as printed |
iban | string | null | 31'827 | |
validity_duration | string | null | 31'827 | free text, e.g. "12 Monate" |
qr_verified | boolean | 31'827 | true on 3'911. Present-and-false is the common case and is not the same as absent |
reference_number | string | null | 31'827 | changed 2026-09-04 (MOR-1108): the PAYMENT reference and nothing else — a QR-Referenz / ESR reference (27 digits, MOD10 check) or an ISO 11649 RF… creditor reference. Null where the document prints none. Until that date it also carried whatever the extractor answered, which measured as the invoice number on two of three sampled documents (5195152, a Komm.-Nr.; GC-26-2077, a Rechnung Nr.); those go to invoice_number now. No backfill was run — a document processed before 2026-09-04 may still hold an invoice number here |
line_items | array | null | 31'436 | array on 8'832, null on 22'604 |
section_block | object | null | — | MOR-1415: line_items came from a section block (Abonnemente 399.68 + the lines that sum to it); closes_on = total / total_minus_rounding / gross; replaced_model_rows when the model's rows (sum printed nowhere) gave way |
deduction_floor_waived | object | null | — | MOR-1435: a deduction row under the 100.00 floor (MOR-1141), taken because the page prints the whole equation — subtotal and VAT printed, their sum = the printed pre-deduction figure, deduction printed and > 0.05, subtotal + VAT − deduction = total to the rappen; {deduction, pre_deduction, label}. Also since MOR-1435: amounts_include_vat is true when the positive rows add up to subtotal + VAT and not to the subtotal |
printed_line_amounts | object | null | — | MOR-1434: rows corrected against their own printed line — rows_repaired[] (wrong_quantity_column: k × unit price is the printed amount, the stored product printed nowhere; printed_discount: <gross> <pct>% <net> → total = net, discount_pct), levies_appended[] (VOC/LSVA … Abgabe lines as rows); only when the result equals the stored subtotal to the rappen |
qr_vat_payload | object | null | — | MOR-1436: the Swiss QR bill's Swico /32/ corrected the VAT — rate:net pairs closing on the total (exactly or through 5-rappen rounding) write vat_amount (= total − Σ net), subtotal_amount (subtotal_source: qr_payload), vat_rate, vat_breakdown; a bare rate writes a missing vat_rate; {closes, written[], from}. /40/ net terms give due_date (due_date_source: qr_payload) only where none was read |
service_vat_table | object | null | — | MOR-1444: the page's per-service summary table (Leistung … Betrag MWST % MWST CHF Total CHF, closed by a Betrag line) corrected the VAT — every line closes on itself, the closing line on the column sums, its gross on the total (exactly or through 5-rappen rounding); writes vat_amount, subtotal_amount (subtotal_source: service_table), vat_breakdown (one bucket per rate, only with ≥ 2 rates) and a missing vat_rate where they differ; {closes, written[], from} |
customs_import_vat | object | null | — | MOR-1459: a customs import-VAT invoice (BAZG / EZV MWST - Rechnung whose Gesamtbetrag MWST [CHF] equals the total) — the whole amount is the import VAT, nothing on top: vat_amount 0.00 (derived), no rate, no breakdown, net = total, rows without VAT split; {import_vat, evidence} keeps the import VAT. MOR-1466: positions_read is added when the answer carried no rows and the page's own position table (closed by Total zu unseren Gunsten, rows summing to it and to the total) supplied them |
service_table_rows | object | null | — | MOR-1448: the rows ARE the per-service table's lines (same proof as service_vat_table) — total = gross, vat_net = net (Betrag, or Netto with Akonto columns), vat_rate; quantity/unit null; amounts_include_vat: true; written only where the model's rows differed; {rows, replaced_model_rows: {count, sum}}; a 5-rappen table-vs-total difference becomes a Rundung adjustment |
section_totals | object | null | — | MOR-1418: line_items are one row per section of a bill of sections (Total Mobile × n → printed sum); rounding_from_printed_totals = invoice amount − printed sum (≤ 5 Rp), also an adjustment; replaced_model_rows when the model's rows (sum printed nowhere, or equal to ONE section's total — MOR-1433) gave way; closes_on = printed_line | net | gross (MOR-1433: anchors Total <label>, Zwischentotal, Total je …, …summe <no.> Gesamt, amount also on the next line; a run with no sum line closes on the net or the total to the rappen); MOR-1440: when the model's table had a row without an amount, replaced_model_rows.sum is null and rows_without_amount counts those rows |
valid_until | string | null | 31'278 | ISO date |
betreff | string | null | 25'066 | the document's own subject line. In the scayla_v1 profile it is additionally surfaced as Subject. MOR-1422 (F44): where the classifier answered nothing or a generic label, the file name without the import's UUID and the extension stands in |
filename_hints | string[] | null | — | MOR-1422 C4: status words (def, prov, entwurf, retour, unterzeichnet, kopie) and office abbreviations (MV, SR, WAP, BR, KB, BA) parsed from the file name; present only when there are any; consumed by nothing before wave 5 |
date_issued | string | null | 24'996 | the date the DOCUMENT states — not the upload date |
date_issued_source | string | null | — | qr_slip_page (MOR-1166) or letterhead (MOR-1422: the model returned no date, or one the document prints nowhere, and the letterhead line was read deterministically). Absent when the model's date stood. MOR-1414: labelled_over_due when the model answered a printed DUE date and the labelled issue date replaced it (date_issued_was_due_date carries the dropped value) |
date_issued_not_in_text | object | null | — | MOR-1422 (F17): the model's date was printed in NO format in the document and was dropped; model is the dropped value. Where no letterhead date replaced it, date_issued is null and the document carries review reason RT18 |
received_date | string | null | — | MOR-1422: the receipt stamp (EINGEGANGEN 21. AUG. 2024, Eingang: 12.03.2024, Reçu le …), DD.MM.YYYY. Never the issue date. Forward only from the wave-2 promotion |
date_issued_was_due_date | string | null | — | MOR-1414: the model's date when it was a printed due date and was replaced |
vat_zero_printed | string | null | — | MOR-1414: the Mwst.: 0.00 line the page prints when vat_amount is 0.00 by that line (no other VAT line prints an amount, gross = net) |
pageCount | number | null | 21'128 | camelCase for historical reasons; not renamed, because clients read it |
booking_date | string | null | 12'460 | ISO date |
period_from / period_to | string | null | 5'382 / 5'301 | ISO dates |
vat_rate | string | null | 4'455 | always a string where present: value on 1'512, null on 2'943. A percentage |
valid_from | string | null | 3'586 | ISO date |
imageType | string | null | 1'710 | camelCase, as above |
headings | array of string | 700 | |
sheets | array of string | 337 | spreadsheet sheet names |
empty_content | boolean | 84 | the file carried no readable text |
alternatives | array | 34 | runner-up classifications |
vat_amount | string | null | 15 | 11 hold a value, 4 hold null. Never present without vat_amount_source |
vat_amount_source | extracted | derived | null | 15 | how vat_amount was obtained |
slides | array of string | 8 | presentation slide titles |
ai_skipped / ai_skip_reason | boolean / string | 2 | extraction was not attempted, and why |
vat_breakdown | array | null | 2 | array on 1, null on 1 |
qr_reference | string | null | 1 | Swiss QR bill reference |
qr_reference_type | QRR | SCOR | null | 1 | which standard qr_reference follows |
invoice_number | string | null | new | the number the supplier calls the document, read off the page by its label. NOT reference_number, which prefers the QR reference — on a QR document those differ, and this is the one an accountant keys in (MOR-1080) |
contract_number | string | null | new | the contract/policy number, read off the page by its own label (Vertrag Nr. / Vertrags-Nr. / Vertragsnummer / Police Nr. / Contrat n° / Contratto n.). A third field, not reference_number and not invoice_number: 163 production documents print a Vertrag Nr. and none of them also carry an invoice_number. Never derived from the QR reference even where the digits overlap (MOR-1258) |
contract_number_source | labelled | null | new | how contract_number was read; labelled is the only tier today (MOR-1258) |
vat_number | string | null | new | the supplier's Swiss UID, CHE-123.456.789. Validated by its mod-11 check digit. Null when the page prints several and whose is whose cannot be told — a wrong one is worse than none (MOR-1081) |
vat_number_source | sole | sender-adjacent | recipient-excluded | null | new | sole = the only UID on the page (7'228 of 8'050); sender-adjacent = one of several, resolved by whose name it sits beside; recipient-excluded (MOR-1460) = one of several, the others labelled on the page as the recipient's (Ihre MWSt Nr., customs Sped-Nr./TIN/UID) |
credit_balance | string | null | new | the settlement balance printed beside the total — a Nebenkostenabrechnung nets its total against what was already paid, and total_amount is still the total. Read only where an amount sits directly beside the label; 1'019 of 1'600 documents carrying the phrase answer null, which is the reader declining rather than guessing (MOR-1256) |
credit_balance_direction | customer | issuer | null | new | whose favour it is in, and it reverses on one word: Saldo zu IHREN Gunsten (the reader is owed — nothing payable) against zu UNSEREN Gunsten (the reader owes it — still due). 595 production documents print both phrases, so a direction is recorded only when an amount is attached to it. Measured: 234 customer, 347 issuer (MOR-1256) |
credit_balance_source | adjacent | null | new | how credit_balance was read. adjacent — the amount hanging off the right of the label — is the only reading that survived measurement; a positional one answered the Total on one layout and the Akonto on its mirror and was deleted (MOR-1256) |
qr_amount_present | boolean | new | whether the Swiss QR slip printed an amount. false is the variable-amount slip — the issuer saying no fixed sum is due. Absent, not false, where there is no slip: false is a statement about a slip that was read (MOR-1256) |
total_amount_repaired_from | string | null | new | what total_amount was before the page corrected it. 8,475.60 printed, 8.47 stored — three documents on three orgs. Fires only when the stored figure is absent from the text and exactly one printed number could have produced it (MOR-1038) |
subtotal_amount_repaired_from | string | null | new | what subtotal_amount was before the document's own VAT recapitulation corrected it. On a multi-rate Sammelrechnung the extraction can read one bucket's net as the invoice's: 22'128.57 stored against an invoice net of 23'662.66. Fires only when sum(net) + sum(VAT) = total_amount to the centime and the stored subtotal equals exactly one bucket — 1 of 33'116 documents (MOR-1131) MOR-1384 (2026-09-13): also set when subtotal_source is paired_with_invoice_number — the derived net the paired total disproved, whether the subtotal was then rewritten or cleared. |
subtotal_source | string | null | new | where subtotal_amount came from when it did not come from the model, or null. column_table (MOR-1187): read from the positions' own net column. paired_with_invoice_number (MOR-1384): on a split bundle child whose total was corrected by the invoice-number pairing (total_amount_source), the stored net was the bundle total minus VAT — a figure printed nowhere — and was replaced by the paired total minus VAT, only because the positions or the page print that figure. ledger (MOR-1400): the positions were read off the page's running totals, and the net is the running total that equals the sum of the positive rows — written when the model returned no subtotal or one printed nowhere. MOR-1416 (2026-09-15): printed_net when the model gave no subtotal and the page printed the net beside the VAT row that was removed (MWST Basis); see vat_row_removed_from. MOR-1432 (2026-09-16): management_fee_line — a Nebenkostenabrechnung's management fee is billed WITH its own VAT (3.5% Verwaltungshonorar zzgl. 8% MWST 108.25), which the model reads as the document's VAT line and the total minus that figure as the net; unlike every value above, this REPLACES a subtotal the model did return, because the fee line's own arithmetic proves the stored one wrong. See vat_source. |
vat_amount_repaired_from | string | null | new | what vat_amount was before a printed management-fee line corrected it, or null. Present only alongside vat_source: management_fee_line (MOR-1432) |
vat_source | management_fee_line | null | new | where vat_amount came from when it did not come from the model, or null. The only value today is management_fee_line: the fee row itself is untouched (a different repair, MOR-1416's vat_row_removed_from, never fires on the same document) — only vat_amount and subtotal_amount move, to the VAT on the fee alone (MOR-1432) |
management_fee_repair | object | null | new | MOR-1432 (2026-09-17). Present only alongside vat_source: management_fee_line. Which proof accepted the fee figure, and the numbers it turned on. proof is fee_arithmetic when the printed percentage times a base the page prints as rows closes on the fee to ±0.05, or labelled_row_in_complete_table when it does not and the figure is instead a row of the stored table carrying the fee's own label, inside a table that sums to the document total. The second proof exists because the printed percentage is not reliable: measured across the 98 production documents carrying this defect, the rate implied by the fee figure runs 0.35 %–15.66 % against a printed 3–3.5 %, and 003c3def prints 3.5 % while charging exactly 3.0 %. base is the net the arithmetic closed on, or null under the second proof, which consults no base. row_index is the fee row in line_items, or -1 where the model returned no row for the fee at all (32 of the 98) |
scale_repaired_from | object | null | new | what the three summary amounts held before the ×1000 repair, keyed by field — {"total_amount":"43.05748","subtotal_amount":"39.83116","vat_amount":"3.22632"}. A thousands separator read as a decimal point divides the figure by a thousand while keeping every digit, and the positions can be perfectly correct beside it, so nothing else objects. Fires only where the page PRINTS the ×1000 result and does NOT print the stored one — five documents in production (MOR-1164) |
invoice_number_source | labelled | positional | model | null | new | how it was read. positional is an alignment across a column header: sound, weaker, and worth showing a reviewer |
invoice_number_rejected_from | string | null | new | MOR-1471. What invoice_number would have held, when it was refused because it is the number printed under an Abrech.-Nr / Abrechnungs-Nr label — a social insurer's account number (contract_number, MOR-1468), not the document's own number. Absent where nothing was refused |
total_amount_source | paired_with_invoice_number | null | new | present only when total_amount did not come from the model: the page is a bundle overview (Summe Zahlbetrag) and the stored total was the bundle sum, so the amount printed beside this document's own invoice number was taken. 409.43 stored, 341.09 paired on the Swiss half of a split BMW Charging bundle (MOR-1376) |
vat_country | DE | AT | CH | FR | IT | LI | null | new | the country printed beside the VAT rate (19% MwSt. DE). Read off the page, never from currency or sender. Absent on an invoice that names none — every Swiss one. R9 judges the rate against that country's law (MOR-1377) |
reference_number_type | QRR | SCOR | null | new | which standard reference_number follows. Not interchangeable at a bank: a QRR sent against a plain IBAN is rejected, and so is a SCOR against a QR-IBAN (MOR-1108) |
reference_number_source | qr | page | model | null | new | how it was obtained. qr = machine-read off the slip; page = recovered from the document's own text and verified by the reference's own check digit; model = the extractor's answer, having passed that same check (MOR-1108) |
order_reference | string | null | new | the quote or order the document refers back to — Referenz: Offerte OF-26-1188 yields OF-26-1188. Neither the payment reference nor the invoice number. Forward only from 2026-09-04 (MOR-1108) |
order_reference_text | string | null | new | the referenced line as printed — Offerte OF-26-1188, Umgebung Neubepflanzung. The identifier is what a system matches on; this is what tells a person what it was for |
order_reference_kind | offerte | bestellung | auftrag | unspecified | null | new | which kind of preceding document it points at. unspecified = a bare Referenz: line with no order word beside it |
has_qr_bill | boolean | new | whether the page carries a Swiss/Liechtenstein QR-bill payment slip heading (Empfangsschein/Zahlteil/Payment part/Récépissé/Section paiement, in whichever of the four languages the standard prints it in). A flag and a filter, never a reclassification — doc_type/classification are untouched by this value; a reminder or a debt_collection notice carries a slip too, for reasons that are not "this is an invoice" (MOR-1255) |
line_items[]Shape verified against a production row, not against the extractor's prompt.
| key | type |
|---|---|
position | number — 1-based, as printed |
description | string |
quantity | number |
unit | string, as printed: Stk, h, Pl, m2 |
unit_price | number |
total | number |
vat_rate | number | null — new, MOR-1000 slice 2b |
vat_net | number | null |
vat_amount | number | null |
Per-position VAT is present only where the document attaches VAT to that
position — a MwSt column in the positions table, or a rate printed on the
line. null means the document did not print a rate there; it never means the
rate is zero. 0.0 is a real rate (reverse charge, exempt supply) and is
not the same as null.
vat_net and vat_amount are each null unless the document prints them.
Neither is computed from the other — a computed figure is not an extracted one,
and this API does not present one as the other.
Every position written before 2026-08-26 carries null on all three. No
archive backfill was run: 31'513 positions existed at that date and not one
carried a rate. Read the nulls as age, not as failure.
vat_breakdown[]The document's own per-rate recapitulation, in the order the document prints
it. Not a computed summary — where a document prints no recapitulation, this is
absent and vat_rate carries the single rate instead.
| key | type |
|---|---|
rate | number — percentage, e.g. 8.1 |
amount | number — VAT for this rate |
net | number | null — the net this rate was applied to. The key is net, not base |
vat_amount never travels without vat_amount_source. extracted means the
document stated the figure; derived means we computed it from a rate and a
net. A client posting VAT to a ledger needs to know which, and the field pair
exists so it never has to guess.
All PATCH endpoints:
400 DOCUMENT_LOCKED when the document's status_lifecycle is abgeschlossen / finalized / deleted. Reads still succeed.document_audit_log (actor, IP, prev/next snapshots, optional reason).Only the top-level PATCH /api/v2/documents/:id fires classification.changed webhook events — one per touched axis (see Webhooks). The sub-resource (/classification, /metadata, /tags) and bulk endpoints update the record and write the audit row but do not emit the event.
/api/v2/documents/:idAtomic top-level correction. Update classification axes and tags in one call.
{
"doc_type": "rechnung",
"doc_subtype": "service_invoice",
"doc_bereich": "finanzen",
"intent": "payment_required",
"status_lifecycle": "in_bearbeitung",
"zahlungsstatus": "bezahlt",
"tags": ["plumbing", "heating"],
"reason": "Manual correction after operator review"
}
All fields optional. Tags are replaced (final array = desired state). Unknown taxonomy keys return 422 — see Error codes.
/api/v2/documents/:id/classificationFine-grained classification + tag UPSERT. Same body shape as the top-level PATCH but scoped to the classification axes. Useful when you want a separate audit row for classification-only edits.
/api/v2/documents/:id/metadataMERGE update of document_classification_v3.extracted_metadata (JSONB). Specified keys overwrite, keys with null are deleted, unspecified keys remain. The legacy documents.extracted_metadata is not touched.
{
"reference_number": "RG-2026-001-CORRECTED",
"total_amount": "215.50",
"due_date": "2026-07-10",
"qr_verified": null,
"reason": "Corrected per supplier clarification"
}
/api/v2/documents/:id/contactsREPLACE the document's contact link rows. Payload is the desired final state — links not in the payload are deleted, new links are inserted, identical links unchanged.
{
"contacts": [
{ "contact_id": "contact-uuid-1", "role": "absender" },
{ "contact_id": "contact-uuid-2", "role": "absender" }
],
"reason": "Added secondary sender contact"
}
Anti-spoof: every contact_id is verified to belong to the caller's organization — out-of-org → 400 (same shape as "unknown contact" so no RLS leak).
/api/v2/documents/:id/tagsAdd one or more tags from the controlled taxonomy to the document. Body shape: { "tags": ["plumbing", "heating"] }. Idempotent UPSERT with confidence = 1 (manual sentinel). Unknown tag → 422. Optional X-Audit-Reason header attaches a reason to the audit row.
/api/v2/documents/:id/tagsRemove one or more tags. Body shape: { "tags": ["plumbing"] }. Silent no-op when the tag isn't present (idempotent).
/api/v2/documents (bulk)Up to 100 items per request. Each item is processed independently — one bad item never aborts the rest.
{
"items": [
{ "id": "<doc-1>", "doc_type": "rechnung", "tags": ["plumbing"], "reason": "..." },
{ "id": "<doc-2>", "zahlungsstatus": "bezahlt" }
]
}
Returns per-item result with status: "ok" | "error" and meta: { total, ok, error }.
/api/v2/documents/uploadUpload a single file with content-hash + external_id idempotency. Designed for high-volume bulk ingestion (200 000+ files): repeated filenames are NOT a duplicate signal — only external_id and content hash are.
Multipart fields:
| Field | Required | Description |
|---|---|---|
file | yes | The document file (PDF, image, Office) |
external_id | recommended | Caller's stable id — primary idempotency key |
metadata | no | JSON string with caller-side context, stored on the document |
force | no | "true" bypasses content-hash idempotency and processes the file as a NEW document (testing aid — see below) |
Size limit: 4 MB per file. Our serverless runtime caps multipart request bodies at 4.5 MB before they reach this handler — that limit is platform-level and not application-configurable. We pin the application-level limit to 4 MB so the descriptive 413 from this endpoint surfaces before the generic infrastructure 413. Pre-validate file size client-side before posting.
Files larger than 4 MB must use the presigned upload flow described below — that path streams directly to storage and lifts the per-file limit to 250 MB.
Idempotency / dedup decision:
upload request
│
▼
external_id provided?
│ yes │ no
▼ │
SELECT (org, ext_id) │
│ │
found? │
│ yes │
▼ │
return existing │
meta.idempotent_replay │
= true, │
dedup_by = "external_id"│
│ no │
└────────┬─────────┘
▼
compute sha256(content)
│
▼
SELECT (org, content_hash)
│
found?
│ yes
▼
return existing
meta.idempotent_replay = true,
dedup_by = "content_hash"
│ no
▼
── CREATE PATH ──
- upload to storage with 8-byte random suffix
- INSERT document (content_hash, external_id, original filename)
- audit log: action="upload_created"
- trigger pipeline in background
- return 200, meta.idempotent_replay = false
Response shape:
{
"data": {
"id": "550e8400-…",
"filename": "invoice-001.pdf",
"status": "pending",
"file_type": "pdf",
"file_size": 184320,
"external_id": "client-system-2024-1041",
"content_hash": "a1b2…",
"created_at": "2026-06-07T10:30:00.000Z"
},
"meta": { "idempotent_replay": false },
"error": null
}
On replay, data is the existing document and meta = { idempotent_replay: true, dedup_by: "external_id" | "content_hash" }.
Repeated filenames are accepted — deduplication is by external_id / content_hash, never by filename.
force=true (testing aid): bypasses the content-hash check and creates + processes a fresh document even for identical bytes. Two consequences, by design: (1) the forced row is stored without a content hash, so it never participates in future content dedup — the original document keeps winning; (2) force cannot reuse an existing external_id (400) — send a new one or omit it. The response carries meta: { idempotent_replay: false, forced: true }. Also accepted as a JSON field on /upload/finalize. Don't use it in production ingestion — it defeats the duplicate protection.
POST /api/v2/documents/upload also accepts a single .zip container through the same multipart file field (detected by .zip extension or a ZIP MIME type). Each entry inside the archive becomes its own document with its own idempotency.
Container rules:
.zip files are rejected per-entry (nested_zip_not_supported), never unpacked.external_id on the request, each entry gets "{external_id}:{entry_filename}" as its own idempotency key.Per-entry response — the endpoint returns one item per archive entry instead of one document:
{
"data": {
"container_kind": "zip",
"container_filename": "belege_2026_06.zip",
"items": [
{ "kind": "accepted", "id": "…", "filename": "rechnung_01.pdf", "status": "pending", "rejected_reason": null },
{ "kind": "idempotent_replay", "id": "…", "filename": "rechnung_02.pdf", "status": "completed", "rejected_reason": null },
{ "kind": "rejected", "id": null, "filename": "~$rechnung_03.docx", "rejected_reason": "office_lock_file" }
]
},
"meta": { "container_kind": "zip", "total_entries": 3, "accepted": 1, "idempotent_replay": 1, "rejected": 1 }
}
Per-entry rejected_reason values:
| Reason | Meaning |
|---|---|
directory_entry | Folder entry, nothing to ingest |
nested_zip_not_supported | .zip inside the archive — flatten before uploading |
entry_too_large | Entry exceeds the per-file size limit |
unsupported_type | Extension outside the supported document types |
empty_entry | Zero-byte or unreadable entry |
office_lock_file | Microsoft Office temporary/owner file (~$…) — a stub Word/Excel writes while a document is open, never a real document |
Container-level failures (whole archive, persisted as ONE failed document row so a subscribed document.failed webhook fires once): container not readable (corrupt/not a ZIP) and too many entries.
Note: standalone (non-container) uploads of Office lock files are rejected up front with
400on both multipart/uploadand presigned/upload/init.
For files larger than the multipart /upload cap of 4 MB, use the two-step presigned flow below — the file content streams directly to Supabase Storage and never touches our API function. (Multipart /upload is capped at 4 MB because our serverless runtime rejects request bodies over 4.5 MB before the handler runs — platform limit, not application-configurable.)
The presigned flow works for any file size up to 250 MB (matches the storage bucket's file_size_limit). Files over 250 MB are not supported — contact us if you have a use case.
/api/v2/documents/upload/initRequest a signed PUT URL for the upload. Accepts JSON; no file bytes here.
Request:
{
"filename": "invoice-001.pdf",
"file_size": 12345678,
"mime_type": "application/pdf",
"external_id": "client-system-2024-1041",
"metadata": { "source": "scan-batch-12" }
}
| Field | Required | Description |
|---|---|---|
filename | yes | Original filename, preserved on the eventual documents.filename |
file_size | yes | Bytes — must be ≤ 250 MB |
mime_type | yes | Must be in the allowed list (same set as /upload) |
external_id | recommended | Stable caller id — primary idempotency key |
metadata | no | Ignored on init; passed in finalize |
Response (new):
{
"data": {
"upload_url": "https://<project>.supabase.co/storage/v1/object/upload/sign/documents/<path>?token=…",
"upload_token": "eyJ…",
"storage_path": "<org_id>/<safe>-<random>.pdf",
"expires_at": "2026-06-10T09:30:00.000Z"
},
"meta": { "idempotent_replay": false },
"error": null
}
Response (idempotent on external_id):
{
"data": { "id": "550e8400-…", "filename": "invoice-001.pdf", "status": "completed", … },
"meta": { "idempotent_replay": true, "dedup_by": "external_id" },
"error": null
}
On idempotent hit there is no upload URL — the caller must not upload again. The existing document is returned verbatim.
curl -X PUT "<upload_url>" \
--data-binary "@invoice-001.pdf" \
-H "Content-Type: application/pdf"
The signed URL is valid for 2 hours. If it expires, call /upload/init again with the same external_id and we'll issue a fresh URL (or return the existing doc if a parallel finalize already completed).
Alternatively, with the @supabase/supabase-js SDK:
await supabase.storage.from("documents").uploadToSignedUrl(storage_path, upload_token, file);
/api/v2/documents/upload/finalizeOnce the PUT succeeds, register the document.
Request:
{
"storage_path": "<from init response>",
"filename": "invoice-001.pdf",
"mime_type": "application/pdf",
"file_size": 12345678,
"external_id": "client-system-2024-1041",
"content_hash": "a1b2…",
"metadata": { "source": "scan-batch-12" }
}
| Field | Required | Description |
|---|---|---|
storage_path | yes | Verbatim from the init response |
filename | yes | Original filename to persist |
mime_type | yes | Must be in the allowed list |
file_size | yes | Bytes — must be ≤ 250 MB |
external_id | recommended | Idempotency key |
content_hash | no | SHA-256 hex of the file content — enables content-based dedup. If omitted (and force is not set), the server computes it from the uploaded object, so content dedup applies either way |
metadata | no | Stored verbatim on documents.extracted_metadata |
Response: identical shape to /upload ({ data: documents_row, meta: { idempotent_replay, dedup_by? } }).
Errors:
| Status | When |
|---|---|
400 | missing/invalid field |
403 | storage_path not under your org folder |
413 | file_size over 250 MB |
422 | storage object not found at storage_path (upload incomplete or wrong path) |
429 | rate-limit (100 req/min, shared with /upload) |
Idempotency order: external_id → content_hash → insert. Either hit returns the existing document and does not touch the just-uploaded storage object. (Orphaned storage objects from never-finalized init calls are swept by a separate cleanup job; do not rely on this for correctness — finalize what you init.)
/api/v2/documents/upload/batchUp to 50 files per request. Each item runs the same idempotency rules independently.
Multipart fields:
| Field | Description |
|---|---|
files[] | Repeatable file field (max 50) |
external_ids[] | Optional, paired index-wise with files[] |
metadata[] | Optional, JSON string per file |
Response shape:
{
"data": [
{ "index": 0, "status": "created", "doc": { … } },
{ "index": 1, "status": "idempotent_replay", "doc": { … }, "dedup_by": "external_id" },
{ "index": 2, "status": "failed", "error": "Unsupported file type: image/heic" }
],
"meta": { "total": 3, "created": 1, "idempotent_replay": 1, "failed": 1 },
"error": null
}
One bad item never aborts the rest. Each successful item (created or replay) writes one audit row.
/api/v2/taxonomyReturns the full taxonomy (bereiche, doc_types with v1_projection, doc_subtypes with parent type, intents with v1_projection, tag_groups, tags). German labels included. The taxonomy is global today; per-org overrides will fold in transparently when shipped.
Useful for rendering dropdowns and filter pickers without hardcoding values that may grow.
/api/v2/usageReturns a usage snapshot for your organization (the org the API key belongs to). Read-only. Authenticated with your API key as a Bearer token, like every other endpoint — the org is always taken from the key, never from the request.
The period defaults to the current calendar month, month-to-date (UTC). Pass ?period=YYYY-MM to read a past month (e.g. ?period=2026-03); period_start is echoed back so the window is unambiguous. A malformed period returns 400.
api_calls.by_route for the current month comes from request telemetry that has 90-day retention. Past months are served from a durable monthly rollup, so per-route history for older months stays available beyond the 90-day window. The source field reports which: "live" (current month) or "rollup" (a past month).
# current month
curl -H "Authorization: Bearer idms_..." \
https://idms.mory.ai/api/v2/usage
# a specific past month
curl -H "Authorization: Bearer idms_..." \
"https://idms.mory.ai/api/v2/usage?period=2026-03"
Response:
{
"data": {
"period_start": "2026-06-01T00:00:00.000Z",
"source": "live",
"api_calls": {
"this_period": 1432,
"by_route": [
{ "route": "/api/v2/documents", "request_count": 980, "last_seen": "2026-06-19T17:42:11Z" },
{ "route": "/api/v2/documents/:id", "request_count": 452, "last_seen": "2026-06-19T17:40:02Z" }
]
},
"documents": { "total_stored": 21043, "ingested_this_period": 318 },
"search": { "this_period": 74 },
"storage": { "bytes_used": 48210334720 },
"remaining": { "rate_limit": 97, "rate_limit_reset": 1750352580 },
"limits": { "documents_max": null, "api_calls_max": null, "search_queries_max": null }
},
"meta": null,
"error": null
}
source is "live" for the current month (raw telemetry, 90-day retention) or "rollup" for a past month (durable monthly aggregate that outlives the 90-day window).documents.total_stored and storage.bytes_used are point-in-time (current totals), not reconstructed for a historical period. documents.ingested_this_period and api_calls.* are scoped to the requested month.remaining.rate_limit / rate_limit_reset mirror the X-RateLimit-* response headers (requests left in the current minute window, and the reset time as a Unix timestamp).limits.* are the quotas configured for your organization (documents_max, api_calls_max, search_queries_max). null means unlimited — that is the default, not a placeholder. They are reported here whether or not enforcement is switched on for your org; when it is, exceeding one returns 429 (see Error codes).search.this_period is the count of POST /api/v2/search calls in the requested month — pulled out of api_calls.by_route so it can be compared directly against limits.search_queries_max.{ "data": null, "error": { … } }): 401 for a missing or invalid API key, 429 when the per-key rate limit is exceeded — see Error codes.POST /api/v2/searchSearch your organization's documents by meaning and by keyword in a single call. The query is matched two ways — a semantic leg that understands paraphrase and synonyms (so "Mietzinserhöhung" also surfaces "Anpassung des Mietzinses"), and a keyword leg for exact strings like invoice numbers, names or references. The two result sets are fused (reciprocal-rank fusion) into one ranked list, and matched_via tells you which leg(s) surfaced each hit. Read-only; the org is always taken from your API key.
Availability is per organization. If search is not enabled for your org the endpoint returns 403 — contact us to enable it.
Request (JSON):
| Field | Type | Required | Description |
|---|---|---|---|
query | string | yes | The natural-language query. |
mode | string | no | hybrid (default) — semantic + keyword fused. semantic — meaning only. keyword — exact-match only. |
limit | integer | no | Max results, 1–100. Default 25. |
doc_type | string | no | Restrict to a document type (same vocabulary as GET /documents). |
bereich | string | no | Restrict to a Bereich. |
intent | string | no | Restrict to an intent. |
tags | string[] | no | Restrict to documents carrying all of these tag keys. |
created_after | string (YYYY-MM-DD) | no | Only documents created on or after this date. |
created_before | string (YYYY-MM-DD) | no | Only documents created before this date. |
sender | string | no | Absender — substring match on the document's sender name. |
receiver | string | no | Empfänger — substring match on the document's receiver name. |
property | string | no | Liegenschaft — substring match on the document's property name or address. |
betrag_min | number | no | Minimum amount. The parameter is German-named; the field it filters is extracted_metadata.total_amount. Combine with betrag_max for a range. |
betrag_max | number | no | Maximum amount. Filters extracted_metadata.total_amount. |
faellig_after | date | no | Due date on/after this date (ISO YYYY-MM-DD). Filters on the extracted due date. |
faellig_before | date | no | Due date on/before this date (ISO YYYY-MM-DD). Filters on the extracted due date. |
betrag_min, betrag_max, faellig_after and faellig_before parse exactly as they do on GET /api/v2/documents: an empty string, a whitespace-only string, null and an unparseable value all mean no filter, while 0 is a real filter at zero. Both endpoints share one parser, so the two surfaces cannot drift apart. An empty <input type="number"> submits "", which is why this distinction matters more here than it looks.
Response 200:
{
"data": [
{
"document_id": "6e27ff62-790a-44e6-9ed5-7d689cbc7602",
"filename": "Mietzinsanpassung_2024.pdf",
"doc_type": "letter",
"score": 0.032266,
"snippet": "…Anpassung des Mietzinses per 1. April 2024 gemäss Referenzzinssatz…",
"chunk_index": 2,
"matched_via": ["semantic", "keyword"]
}
],
"meta": { "query_id": "…", "mode": "hybrid", "took_ms": 214, "total": 1 },
"error": null
}
score is the fused relevance score (higher is better); use it for ordering, not as an absolute threshold.snippet is the passage centered on the first query-term match (≤ 280 chars, … marks truncation), with matched query terms wrapped in <mark>…</mark>. The text is HTML-escaped, so the field is safe HTML; strip the tags if you want plain text.matched_via is an array containing "semantic", "keyword", or both.meta.total is the number of results in data. Its meaning is unchanged.meta.keyword_match_count is how many documents in your organization match the query lexically, across the whole corpus, before limit — useful for "showing 25 of 412". It is null when the keyword component did not run (mode: "semantic").
hybrid mode it can be smaller than meta.total, and that is not an error: the semantic component contributes documents that were never counted lexically. There is no combined figure, and this is a property of the search rather than a limitation — the semantic component ranks by similarity with no match/no-match threshold, so "how many documents match semantically" is not a defined quantity. Treat keyword_match_count as a lower bound on how much more there is to find, never as a total.meta.mode reports the mode that actually ran — a hybrid request stays available even if the semantic component is briefly unavailable, in which case it transparently returns keyword-only results and meta.mode is "keyword".Errors: 400 missing/invalid query; 401 invalid/missing API key; 403 search not enabled for your organization; 429 rate limit (100/min, shared with all API calls); 503 for an explicit mode:"semantic" request while the semantic component is temporarily unavailable (a hybrid request degrades to keyword instead).
Example:
# Hybrid (default) — meaning + keyword, fused:
curl -X POST https://idms.mory.ai/api/v2/search \
-H "Authorization: Bearer idms_..." \
-H "Content-Type: application/json" \
-d '{ "query": "Mietzinserhöhung 2024", "limit": 5 }'
# Restrict to a type and a date window:
curl -X POST https://idms.mory.ai/api/v2/search \
-H "Authorization: Bearer idms_..." \
-H "Content-Type: application/json" \
-d '{ "query": "Heizkostenabrechnung", "doc_type": "invoice", "created_after": "2024-01-01" }'
POST /api/v2/search/answerOptional (search phase 3). Runs the same hybrid search as POST /api/v2/search, then asks an LLM to synthesize a natural-language answer over the top matching excerpts, with every factual claim cited inline ([1], [2], …). sources lists only the excerpts the model actually cited — a citation number the model invents never produces a fabricated entry, and excerpts it never cites are simply omitted. Read-only; the org is always taken from your API key.
Availability is per organization, and separate from plain search — enabling POST /api/v2/search does not enable this endpoint. If answer synthesis is not enabled for your org, it returns 403 — contact us to enable it.
The LLM call runs on the same EU/CH-compliant (Switzerland North) provider the rest of the AI pipeline uses — no document text leaves the same residency boundary as classification.
Request (JSON) — same filters as POST /api/v2/search (doc_type, bereich, intent, tags, created_after, created_before, sender, receiver, property), plus:
| Field | Type | Required | Description |
|---|---|---|---|
query | string | yes | The natural-language question. |
mode | string | no | hybrid (default) / semantic / keyword — same semantics as POST /api/v2/search. |
limit | integer | no | Number of context chunks fed to the LLM, 1–10. Default 5. |
Response 200:
{
"data": {
"answer": "The rent is CHF 1800 [1] and additional costs are CHF 150 [2].",
"sources": [
{ "marker": "[1]", "document_id": "6e27ff62-790a-44e6-9ed5-7d689cbc7602", "chunk_index": 0, "filename": "Mietvertrag.pdf" },
{ "marker": "[2]", "document_id": "8f31aa11-2222-4a9c-9c1e-1234567890ab", "chunk_index": 1, "filename": "Nebenkosten.pdf" }
]
},
"meta": { "query_id": "…", "mode": "hybrid", "took_ms": 640, "chunks_used": 2 },
"error": null
}
answer cites every claim inline; if the matching documents don't contain enough information, the answer says so instead of guessing.sources mirrors only the citation markers present in answer — never a fabricated reference.meta.chunks_used is 0 when no documents matched the query; in that case answer is a fixed "no information" message and no LLM call is made.Errors: 400 missing/invalid query; 401 invalid/missing API key; 403 answer synthesis not enabled for your organization; 429 rate limit (100/min, shared with all API calls); 503 the LLM provider is temporarily unavailable (or an explicit mode:"semantic" request while the semantic component is down).
Example:
curl -X POST https://idms.mory.ai/api/v2/search/answer \
-H "Authorization: Bearer idms_..." \
-H "Content-Type: application/json" \
-d '{ "query": "Wie hoch ist der Mietzins?", "limit": 3 }'
A profile is a per-org renderer applied to read responses. Pass ?profile=<name> to GET /api/v2/documents or GET /api/v2/documents/:id and the server rewrites the response shape.
Two profile types ship today:
Generic profile — JSONB config in output_profiles.fields:
{
"renamed": { "doc_type": "Belegart", "intent": "Absicht" },
"excluded": ["storage_path"],
"tag_groups_included": ["maintenance", "utility", "legal"],
"required_fields": ["Belegart", "Daten"],
"fallback_label": "Sonstiges"
}
Built-in 4-level routing profiles — for integrations whose target system uses a fixed taxonomy of folders + templates + sub-templates, the API ships routing engines that map our classification to verbatim names in the target system. Example response shape (profile keys like acme_v1 are provisioned per client by your account manager):
GET /api/v2/documents/<id>?profile=<profile_key>
{
"data": {
"id": "550e8400-…",
"modul": "Belege",
"ordner": "EK Einkauf",
"template": "EK Rechnung",
"belegart": "Default",
"direction": "eingehend",
"buchungsrelevant": false,
"rule_id": "rechnung_eingehend",
"used_fallback": false,
"ou": { "short_name": "TENANT.001", "full_name": "Tenant 001 (Property Management)" },
"status_set": "default_documents",
"status": "neu",
"required_fields": ["Subject", "Buchungsdatum", "Rechnungsdatum", "Periode"],
"fields": { "betrag": "256.74", "iban": "CH93 …" },
"contacts": [ … ]
},
"meta": {
"profile": { "id": "<uuid>", "name": "<profile_key>" },
"routing_rule": "rechnung_eingehend",
"missing_required_fields": ["Buchungsdatum"],
"missing_client_fields": ["Organisation", "Belegart"]
},
"error": null
}
Caller-supplied overrides on a routing profile: ?direction=eingehend|ausgehend|neutral and ?buchungsrelevant=true|false let you force the routing decision without changing the document content.
Profiles are configured server-side per org via the output_profiles table (is_default controls the default when ?profile= is omitted). Without a configured profile, the API returns the raw classification row shape. Ask your account manager for the profile key provisioned for your tenant.
A mandant is an attribute of a document, not a storage partition. One organisation holds all of a client's mandants: organization_id never changes, there is no organisation per mandant, and no document moves between organisations because of one.
In mory a mandant is a company or a person, not an entity of its own. Its identity is the pair {type, id} exactly as mory's document-mandants feed returns it. iDMS stores that pair verbatim and keys on it.
This is the substrate. The routing engine, the property register and the candidate filter are not built. Nothing decides a mandant today: routing_enabled is per mandant and defaults to false, and a document with no mandant supplied reports source: "unresolved" with all three signals not_present, which is true because nothing looked.
/api/v2/mandatesRequired: mory_mandant, name. type defaults to mandate, routing_enabled to false.
curl -X POST https://idms.mory.ai/api/v2/mandates \
-H "Authorization: Bearer idms_..." \
-H "Content-Type: application/json" \
-d '{
"mory_mandant": { "type": "company", "id": "0f2c7a44-9e11-4b70-8c2d-71a3f5e8d016" },
"name": "Liegenschaft Gasshof 11 AG",
"code": "M36",
"uid": "CHE-123.456.789",
"seat_address": { "street": "Ledergasse", "house_number": "11", "postal_code": "6003", "city": "Luzern" }
}'
Idempotent on the mory pair. A repeat with the same pair returns 200 with the existing record and meta.idempotent_replay: true. It does not create a second row and it does not return 409, the same way /upload behaves on external_id.
| Status | When |
|---|---|
| 201 | Created. meta.idempotent_replay: false |
| 200 | A replay of a pair this organisation already has |
| 401 | Missing or invalid key |
| 404 | parent_id unknown, or in another tenant |
| 409 | code already used by a live mandant here. The existing id is in error.details.existing_mandate_id |
| 422 | Validation |
| 429 | Rate limited or quota reached |
The error shape on these routes. Every other v2 route puts a string in
error. The mandant routes put an object,{ message, details }, because the contract requires the conflicting id to travel with the 409. The envelope keys are unchanged:data,meta,error.
/api/v2/mandatesFilters: updated_since, mory_mandant_type with mory_mandant_id, code, include_ended (default false), limit (default 100, max 1000), cursor.
meta.next_cursor is opaque and is null on the last page, so a client can stop without a second request.
/api/v2/mandates/:id/api/v2/mandates/:idPartial update of name, code, uid, parent_id, seat_address, routing_enabled, ended_at. Any other field is refused with 422 rather than ignored, mory_mandant included: the pair is the row's identity.
Setting ended_at retires a mandant and touches no document. A document from 2023 belonged to that mandant in 2023 and stays there. A retired mandant leaves the list unless include_ended=true, and releases its code for reuse.
/api/v2/documents/:id/mandate{
"mandate_id": "0d4a9b17-22e6-4f81-8c05-1b7d3e9a6f40",
"verdict": "filing_wrong",
"expected_updated_at": "2026-09-18T09:20:44.002Z"
}
mandate_id is required. verdict is required only when the document's current mandate_source is routed — omitting it there returns 422. For a document whose source is provided or manual, or that has none, verdict may be omitted and is stored as null. An explicit null is refused: omitting it says you did not answer, and null would claim there is no answer.
| Verdict | Means |
|---|---|
new_property | The address is real, the register did not have it |
filing_wrong | The register is right, the filing was not |
extraction_wrong | iDMS misread the document |
The verdict is recorded, not acted on. Nothing learns from it yet. It is asked for rather than inferred because the three are different repairs, and collapsing them is the quietest way to teach the system the wrong thing later.
It is required only for routed because only then did a machine choose. A verdict on a mandant the caller supplied, or one a person already set by hand, is a verdict about nobody's mistake — demanding one there does not collect a signal, it collects a guess.
expected_updated_at is optional. On a mismatch the call returns 409 and writes nothing, which is how two people resolving the same document at once is caught.
The call writes exactly one mandate_history row and sets mandate_source to manual. History is append only: one row on every change, including the first assignment, because once anything is reported per mandant for a period a silent reassignment changes last month's numbers.
Three routes take an optional mandant. There is no POST /api/v2/upload — that path never existed, and a contract written against it will 404.
| route | body | mandate_id | mory_mandant |
|---|---|---|---|
POST /api/v2/documents/upload | multipart | yes | yes, as a JSON string |
POST /api/v2/documents/upload/batch | multipart | yes | yes, as a JSON string |
POST /api/v2/documents/upload/finalize | JSON | yes | yes, as an object |
POST /api/v2/documents/upload/init | — | no | no |
init has neither on purpose: it only signs a URL and there is no row yet to carry a mandant. On the presigned path the mandant is named on finalize.
On the multipart routes mandate_id is a plain form field, exactly like external_id. The alternative pair travels in the same form as a JSON string:
mandate_id=8f14e45f-ceea-467a-9f8a-1b2c3d4e5f60
mory_mandant={"type":"company","id":"0f2c7a44-9e11-4b70-8c2d-71a3f5e8d016"}
type is company or person; anything else is 422. On finalize the same pair is a real JSON object in the body rather than a string.
If both are sent, mandate_id wins and the pair is not read.
The mandant must resolve inside the caller's organisation or the upload is refused with 422: one that does not resolve would otherwise be a silent misfiling. On batch the check runs before the loop and refuses the whole request — a batch half filed under a mandant and half not is worse than one that is refused.
An empty value is treated as absent. mandate_id= (or whitespace only) is not an error: the upload proceeds and the document simply has no mandant.
When a mandant is supplied, mandate_source is provided.
Existing idempotency is unchanged: external_id first, then content_hash, with meta.idempotent_replay exactly as before.
Webhooks are POSTed to the URL you register in Settings → Webhooks, with:
Content-Type: application/jsonX-IDMS-Event: <event-name> (e.g. document.processed)X-IDMS-Signature: <hex> — HMAC-SHA256 of the request body, hex-encoded, using the secret shown when you create the webhook[0, 60s, 300s]Node.js
import crypto from "node:crypto";
function verify(rawBody, signatureHeader, secret) {
const expected = crypto
.createHmac("sha256", secret)
.update(rawBody)
.digest("hex");
return crypto.timingSafeEqual(
Buffer.from(signatureHeader, "hex"),
Buffer.from(expected, "hex")
);
}
Python
import hmac
import hashlib
def verify(raw_body: bytes, signature_header: str, secret: str) -> bool:
expected = hmac.new(secret.encode(), raw_body, hashlib.sha256).hexdigest()
# constant-time compare — never use `==`
return hmac.compare_digest(signature_header, expected)
# Flask: pass request.get_data() (the RAW bytes), not request.json —
# verify(request.get_data(), request.headers.get("X-IDMS-Signature", ""), SECRET)
Important: hash the raw request body bytes. Re-serializing the parsed JSON changes whitespace / key order and breaks the comparison.
The API ships 6 event types, all in the clean root format — the delivered fields sit at the top level of the payload. (The two original events, document.processed and batch.completed, previously also carried a deprecated nested data mirror; it was removed on 2026-07-04.) New webhook subscriptions opt in to events — existing subscriptions don't auto-subscribe to the newer ones.
document.processed{
"event": "document.processed",
"timestamp": "2026-06-07T08:00:00.000Z",
"document_id": "550e8400-…",
"external_id": "client-system-2024-1041",
"filename": "invoice-001.pdf",
"doc_type": "invoice",
"doc_subtype": "service_invoice",
"tags": ["plumbing"],
"extracted_metadata": { "total_amount": "256.74", "currency": "CHF" }
}
When the org has opted into the current classification model, the payload additionally carries classification_v3: { doc_type, doc_subtype, doc_bereich, intent_v3, status_lifecycle, zahlungsstatus, tags }.
For those same opted-in orgs the payload also carries a top-level extracted_contacts array (sibling of related_contacts; omitted when the document has no contacts) — the org-level deduplicated contact records for this document. The same array, same shape, same gating is returned by the document detail read (GET /api/v2/documents/{id}, with or without a profile), so polling integrators and webhook consumers see identical contact data. Each entry is the documented contact object plus three fields specific to this array:
| Field | Type | Description |
|---|---|---|
confidence | number | null | Extraction confidence for THIS contact on THIS document (0..1). A property of the extraction event — deliberately not on the stored contact / contacts endpoint. null only for links created before 2026-06-12 or via manual PATCH. |
function | string | null | Person's role/job title as stated in the document (e.g. "Geschäftsführerin") |
relationships | [{ contact_id, relationship_type }] | Explicit pairing signal. contact_id references a SIBLING entry in the same array; relationship_type ∈ employee_of / represents / related_to. Emitted only when the document states the link — never inferred from roles or array order. |
external_id carries the caller-supplied id when the matched contact was created/updated through the API with one; contacts created purely by extraction have external_id: null.
batch.completed{
"event": "batch.completed",
"timestamp": "2026-06-07T08:00:00.000Z",
"batch_id": "b_001",
"total_files": 10,
"processed_files": 10,
"status": "completed",
"documents": [ { "id": "...", "status": "completed" }, … ]
}
document.processed on a duplicate upload (opt-in)By default, re-uploading a file whose content we already hold returns 200 with meta.idempotent_replay: true and emits nothing — the event fired when that content first arrived.
That is a poor fit if your model is "one inbound message = one record": you send a new external_id, receive a document carrying a different one, and no callback ever arrives for your reference.
For those integrations we can enable, per account, re-emission of document.processed when a content-hash replay carries an external_id we have not seen for that document.
document_id | the existing document — no new document is created |
external_id | the id you sent on that replay call |
dedup_replay | true — distinguishes this from fresh processing |
stored_external_id | the id the stored document carries, so you can reconcile both of your records against one document |
Everything else is a normal document.processed, so no branching is needed in your handler.
Guarantees:
document_id + external_id pair — repeating the same id never fires twice;200 returns, so you may see the response before the event — your handler must be idempotent;Off unless explicitly enabled for your account. Ask us if you want it.
document.failed (new — clean root)Fires after the pipeline exhausts its retry budget and parks the document at status="failed".
{
"event": "document.failed",
"timestamp": "2026-06-07T08:00:00.000Z",
"document_id": "550e8400-…",
"filename": "broken.pdf",
"external_id": "client-system-2024-1041",
"error_stage": "ai",
"error_message": "AI processing timed out after 60s"
}
error_stage ∈ download / extract / ai / persist / unknown. error_message is sanitized: filesystem paths → [path], URLs → [url], stack frames stripped, 500-char cap.
Stable error_message wordings (safe to match on exactly — they will not change without a changelog entry):
| Condition | Exact string |
|---|---|
| Corrupt / invalid Office file | File is corrupt or not a valid Office document (.docx/.xlsx/.pptx). |
| Password-protected file | File is password-protected and cannot be processed. Remove the password protection and upload it again. |
| Content declined by AI processing | The AI service declined to process this document's content. |
Note: two different non-failures also emit document.processed. Neither is an error and neither will be retried, but they are separate claims and they carry separate markers:
| what happened | marker on the payload | doc_type |
|---|---|---|
| the file was read and had no recognizable text (a pure image / scan) | extracted_metadata.empty_content = true | the deterministic call — photo for an image file |
| no extractor exists for the file's format, so nothing has read it | extraction_state: "unsupported_format", mirrored at extracted_metadata.extraction_state | other, with confidence: null — no information about the content |
The distinction matters on your side: the first is a document that genuinely has nothing to extract, the second is a gap in what we support. Re-uploading the same file will not change the second one. See extraction_state.
In both cases no AI analysis ran, so the analysis fields are empty by design: tags is [], doc_subtype / doc_intent / language are null, and no classification_v3 block is present. The envelope fields (document_id, external_id, filename, file_type) are always present, so the event is safe to correlate exactly like any other.
classification.changed (new — clean root)Fires when a PATCH endpoint changes a classification axis. One event per axis touched in a single PATCH (cleaner downstream parsing).
{
"event": "classification.changed",
"timestamp": "2026-06-07T08:01:00.000Z",
"document_id": "550e8400-…",
"field": "doc_type",
"old": "invoice",
"new": "rechnung",
"actor": { "kind": "api_key", "id": "k_xyz" }
}
field ∈ doc_type / doc_subtype / doc_bereich / intent_v3 / status_lifecycle / zahlungsstatus / tags.
tag.suggested (new — clean root)Fires when the AI proposes a tag that isn't in the controlled taxonomy.
{
"event": "tag.suggested",
"timestamp": "2026-06-07T08:02:00.000Z",
"document_id": "550e8400-…",
"suggested_tag": "solar_panel_maintenance",
"confidence": 0.84
}
Dedup is per (document_id, suggested_tag) — re-processing the same doc with the same suggestion never re-fires.
usage.threshold.reached (new — clean root)Fires when your organization's API-call count for the current calendar month first crosses the alert threshold configured in Settings. Fires at most once per calendar month — the first crossing in a month sends one event; further calls that month do not re-fire. The counter resets at the start of each month.
{
"event": "usage.threshold.reached",
"timestamp": "2026-06-19T17:45:00.000Z",
"organization_id": "550e8400-…",
"period_start": "2026-06-01T00:00:00.000Z",
"threshold": 10000,
"api_calls": 10034
}
threshold is the configured value; api_calls is the month-to-date count at the moment it crossed (always >= threshold). Use it to get ahead of quota / billing surprises rather than discovering them after the fact.
Existing webhook subscriptions list specific event types in their events column. The newer events (document.failed, classification.changed, tag.suggested, usage.threshold.reached) are not added to existing subscriptions automatically — pick them in Settings → Webhooks when you want them. This guarantees existing integrations never receive surprise events without explicit opt-in.
mandate block on document.processedOff by default, per organisation. The block is sent only when org_config.mandate_block_enabled is true for your organisation. With the flag off the payload is byte identical to the one you receive today: the key is not sent as null, not sent as an empty object, not sent at all.
No new event type was added. Conflict and uncertainty ride on document.processed by design.
When the mandant is known:
"mandate": {
"id": "7c1f0e42-3b8a-4c9d-9f21-5a6e0d8b4c33",
"mory_mandant": { "type": "company", "id": "0f2c7a44-..." },
"source": "provided",
"confidence": null,
"decided_at": "2026-09-18T09:20:44.002Z",
"signals": [],
"candidates": []
}
When it is not:
"mandate": {
"id": null,
"mory_mandant": null,
"source": "unresolved",
"state": "uncertain",
"confidence": null,
"decided_at": null,
"signals": [
{ "name": "property_register", "spoke": false, "reason": "not_present" },
{ "name": "addressee", "spoke": false, "reason": "not_present" },
{ "name": "mandate_code", "spoke": false, "reason": "not_present" }
],
"candidates": []
}
source is provided (mory sent it), routed (iDMS decided), manual (a person resolved it) or unresolved (no mandant yet).
state appears only when source is unresolved: uncertain when no signal spoke, conflict when two spoke and disagreed. They are different jobs for a person and must not be collapsed.
A signal is {name, spoke, evidence?, mandate_id?, reason?}. reason is an enum, never prose, so it can be translated: not_present, address_has_no_street_number, address_not_in_register, no_mandate_entity_named, entity_seat_excluded, register_ambiguous, low_text_content. Values may be added later; existing ones never change.
Since the routing engine is not built, a document with no mandant supplied reports all three signals as not_present. That is correct for this pass: nothing looked.
POST /api/v2/feedbackReport classification feedback from your side: a field your users corrected, or an AI draft your users deleted. Feedback is intake only — it never changes the document or its classification. We use it to measure and improve extraction quality.
Request (JSON):
| Field | Type | Required | Description |
|---|---|---|---|
document_id | uuid | yes | The iDMS document id (from document.processed or the read API). Must belong to your organization. |
event_type | string | yes | classification.changed — your user corrected a field. ai_rejected — your user deleted the AI draft entirely (negative signal, no field detail needed). |
field_path | string ≤ 200 | for classification.changed | Which field was corrected, e.g. doc_type, extracted_metadata.total_amount. Free-form path — use your own canonical naming consistently. |
old_value | string ≤ 2000, nullable | no | The value before the correction (as you received it). |
new_value | string ≤ 2000, nullable | no | The corrected value. |
Response 200:
{ "data": { "id": "<feedback-id>", "received": true }, "meta": null, "error": null }
Errors: 400 validation details in error; 401 invalid/missing API key; 404 document not found in your organization; 429 rate limit (100/min, shared with all API calls).
Example:
curl -X POST https://idms.mory.ai/api/v2/feedback \
-H "Authorization: Bearer idms_..." \
-H "Content-Type: application/json" \
-d '{
"document_id": "6e27ff62-790a-44e6-9ed5-7d689cbc7602",
"event_type": "classification.changed",
"field_path": "doc_type",
"old_value": "invoice",
"new_value": "contract"
}'
# Your user deleted the AI draft — pure negative signal:
curl -X POST https://idms.mory.ai/api/v2/feedback \
-H "Authorization: Bearer idms_..." \
-H "Content-Type: application/json" \
-d '{ "document_id": "6e27ff62-790a-44e6-9ed5-7d689cbc7602", "event_type": "ai_rejected" }'
All errors follow the standard envelope:
{ "data": null, "meta": null, "error": "Human-readable message" }
For locked documents, the body also carries a machine-readable code:
{ "data": null, "meta": null, "error": "Document is locked (status_lifecycle=abgeschlossen)", "code": "DOCUMENT_LOCKED" }
| Status | Meaning | Common causes |
|---|---|---|
200 | Success | — |
400 | Bad request | Invalid JSON; missing required field; bad enum value; DOCUMENT_LOCKED; out-of-org contact_id; bulk over 100 items; Office lock file (~$…) sent as a standalone upload |
401 | Unauthorized | Missing or invalid API key |
403 | Forbidden | A capability not enabled for your organization: POST /api/v2/search (semantic search) or POST /api/v2/search/answer (answer synthesis). Also POST /api/v2/documents/upload/finalize when storage_path is not under your org's folder |
404 | Not found | Document not in caller's organization, or no document with that id. A cross-org document lookup is a 404, never a 403 — existence is never disclosed |
413 | Payload too large | Single file over 4 MB (multipart /upload) or over 250 MB (presigned /upload/init). Multipart hard cap is the 4.5 MB serverless runtime body limit; use the presigned flow for larger files. |
422 | Unprocessable entity | Unknown taxonomy key (doc_type, doc_subtype, intent, tag) |
429 | Rate limited | > 100 requests/minute on this key — wait until X-RateLimit-Reset |
429 | Usage quota exceeded | A configured monthly quota for your organization is reached — api_calls_max (any v2 route), documents_max (upload) or search_queries_max (POST /api/v2/search). The message names the quota and the counts, e.g. Usage quota exceeded: api_calls_max reached (10000/10000 this period). X-RateLimit-Reset does not apply here — the window is the calendar month, so retrying at that timestamp fails again. Read limits.* from GET /api/v2/usage and contact us to raise the quota |
500 | Server error | Database error or unexpected internal failure. Retry; contact support if persistent. |
403 is about capability, never about which org a document belongs to. The API returns it in exactly three places: POST /api/v2/search and POST /api/v2/search/answer when the feature is not enabled for your organization, and POST /api/v2/documents/upload/finalize when the submitted storage_path lies outside your org's storage folder. Nothing about another tenant's data is disclosed by any of them.
Cross-org references still never produce a 403: they return 400 (POST/DELETE tag, contact link) or 404 (document lookup), so the response stays indistinguishable from a missing resource.
Most likely your key belongs to a different organization than the document. The API returns 404 instead of 403 for cross-org access so an attacker can't probe for the existence of foreign documents. Verify by listing /api/v2/documents — if your document isn't in the list, it's in a different org.
400 DOCUMENT_LOCKEDThe document's status_lifecycle is one of abgeschlossen / finalized / deleted. These are terminal write-states. To edit, either:
in_bearbeitung (also a PATCH), orexternal_id=<new-id>.abgeschlossen is meant to be set when a human has reviewed and signed off — toggling it back means you're explicitly re-opening the review.
idempotent_replay: true but I sent a brand-new file"You hit one of the two idempotency keys. Check:
dedup_by: "external_id" — you sent the same external_id twice. The API considers that intentional; the caller's id is the trust boundary.dedup_by: "content_hash" — the SHA-256 of the file content matches an existing document in your org. If you uploaded the file before (possibly to a different filename), the API returns the existing record. To force a new document, modify any byte of the file or send a fresh external_id.meta.missing_required_fields is non-empty on a routing-profile responseRouting profiles look up per-template required fields (Pflichtfelder) from a per-client configuration. When one is not satisfied, the response succeeds (200) and names it. Advisory only — fill it via PATCH /api/v2/documents/:id/metadata and repeat the GET to confirm.
Two lists, since 2026-08-26, and the difference is what you can do about them:
| field | meaning |
|---|---|
missing_required_fields | extraction could have supplied this and did not — a value we failed to read, or one the document does not state. Actionable. |
missing_client_fields | extraction cannot supply this, by definition — it comes from your own system or from the routing decision. Organisation, Belegart, Project Name, Account No., the cost-centre and OU fields. |
Why the split. Until that date the first list carried both, and of the 22 distinct Pflichtfeld names across 272 template blocks exactly one could ever be satisfied. A list that always contains the same 21 names is a list nobody reads — and when a genuinely missing field appeared in it, nothing distinguished it from the permanent noise.
Name matching now goes through an alias table. A Pflichtfeld is satisfied by its own name or by the metadata key that means the same thing: Rechnungsdatum ← date_issued, Periode ← period_from / period_to, Subject ← betreff. An alias only ever adds a way to satisfy a field, so nothing that resolved before stops resolving.
Buchungsdatum is deliberately not aliased. booking_date holds a value on 20 of 33'065 documents; an alias would change the message from "missing" to "missing" while making the gap look handled. That is an extraction question and has its own ticket.
Additive. missing_required_fields keeps its name, type and position; it simply stops carrying what you could never act on.
The most common cause is hashing the re-serialized JSON body instead of the raw bytes. Example anti-pattern in Express:
// WRONG — req.body is the parsed object, not the raw bytes
const sig = hmac(JSON.stringify(req.body));
// RIGHT — preserve the raw bytes
app.post("/idms-webhook", express.raw({ type: "*/*" }), (req, res) => {
const sig = hmac(req.body); // req.body is a Buffer here
});
For frameworks where you don't control the raw bytes, capture them in a middleware before JSON parsing.
tag.suggested subscription but no events arrive"Two checks:
Documentation only. No endpoint, field, status code or behaviour changed — every item below was already true of the running API; this page and the OpenAPI spec had not caught up. Listed worst-first, by what acting on the old text would have cost you.
Four claims were wrong and are corrected in place, because each one told you to do something that does not work. The wording that was there until today is quoted so you can tell whether you built against it:
| was documented | actually the case |
|---|---|
"limits.* are null today, meaning unlimited — placeholders for future per-org quotas" | Quotas are live and configurable per organization (documents_max, api_calls_max, search_queries_max). null is the default, not a placeholder. GET /api/v2/usage reports them whether or not enforcement is switched on for your org |
429 had one cause — the per-minute rate limit — and one remedy, "wait until X-RateLimit-Reset" | A 429 is either the rate limit or a monthly usage quota. On the quota one, X-RateLimit-Reset is the wrong clock: that window is the calendar month, so retrying at the header's timestamp fails again. The error message names which quota and the counts |
"The API never returns 403" | It returns 403 in three places, all about capability, none about tenancy: POST /api/v2/search and POST /api/v2/search/answer when the feature is not enabled for your org, and POST /api/v2/documents/upload/finalize on a storage_path outside your org's folder. Cross-org references are unchanged — still 404 (document lookup) or 400 (tag, contact link), never 403 |
| "five webhook event types" in API Scope | Six, and had been since usage.threshold.reached shipped on 2026-06-30. The list two lines further down already said six |
Four things were simply missing and are now written down:
GET /api/v2/usage in the OpenAPI spec — the ?period=YYYY-MM parameter, its 400 on a malformed value, and the source, search.this_period and limits.search_queries_max response fields. The prose here already described the first two; the machine-readable spec did not, so a generated client dropped them.usage.threshold.reached in the spec's webhooks block — documented in this page since 2026-06-30 but absent from the spec, so it did not reach generated clients or the rendered reference.GET /api/v2/documents (browse and delta sync), with the exact ?external_id= lookup exempted. In effect since 2026-08; never stated. See GET /api/v2/documents.meta.profile and meta.missing_required_fields_per_id in the spec's PageMeta — both present on a profile-mapped list response, both absent otherwise.If you want to check your own integration against this: the quota 429 and the three 403s are the two that change what your error handling should do.
missing_client_fields: the required list stops carrying what you cannot act onAdditive. No existing field changed name, type or position.
Routing-profile responses gain meta.missing_client_fields — Pflichtfelder that extraction cannot supply by definition, because they come from your own system or from the routing decision: Organisation, Belegart, Project Name, Organization Name, Account No., Chart of accounts, Reporting Structure, the cost-centre fields and the OU fields.
missing_required_fields keeps only what extraction could have supplied and did not.
Why. Of the 22 distinct Pflichtfeld names across 272 template blocks, exactly one could ever be satisfied — Subject, and only because betreff had been aliased to it by hand. The other 21 were reported missing on every document forever, which made the list noise rather than a signal.
Two of those 21 were only ever a naming difference and are now closed by an alias table:
| Pflichtfeld | satisfied by | documents |
|---|---|---|
Rechnungsdatum | date_issued | 24'996 |
Periode | period_from / period_to | 5'382 / 5'301 |
An alias only ever adds a way to satisfy a field, so nothing that resolved before stops resolving.
Buchungsdatum is deliberately not aliased. booking_date holds a value on 20 of 33'065 documents — that is an extraction gap, not a naming one, and an alias would make it look handled.
If you currently ignore missing_required_fields because it is always full: it is worth reading again.
.xls, .docm and .gif are now readThree formats that previously had no extractor at all are now processed normally, so documents in them get a classification derived from their actual content instead of other at no confidence.
.xls (legacy binary Excel) — read via SheetJS. Note this is a genuinely different format from .xlsx, not an older spelling of it. 150 documents on production were affected..docm (macro-enabled Word) — read exactly like .docx; the macro part is ignored..gif — now goes through vision OCR like every other image type. It was previously classified as a photo without anything having looked at it..heic / .heif still cannot be read and now say so via extraction_state: "unsupported_format" rather than receiving a type from nowhere. No vision provider we use accepts them; converting them is separate work.
Additive: no field changed shape, and documents already stored keep their current classification until they are reprocessed.
extraction_state, and a doc_type that stops guessingextraction_state on every document row (GET /api/v2/documents, GET /api/v2/documents/:id) and on the document.processed webhook payload. It reports whether the file's text was actually read: extracted, unsupported_format, ocr_failed, ocr_unavailable, or null for documents processed before the field existed. See extraction_state.doc_type. Until now, a file in a format we have no extractor for was still sent to the classifier — with a placeholder string standing in for its text — and came back with a real-looking doc_type at a low confidence. That classification described the placeholder, not your document. Such files now return doc_type: "other" with classification_confidence: null and extraction_state: "unsupported_format".other, which is the correct answer arriving in place of a wrong one. Existing documents are unaffected until they are reprocessed. Nothing was removed or retyped, so an integration that ignores the new field keeps working.POST /api/v2/search gains betrag_min / betrag_max (filter on the extracted total) and faellig_after / faellig_before (filter on the extracted due date) — the same semantics as the equivalent GET /api/v2/documents filters.
meta gains keyword_match_count: how many documents match the query lexically across the whole corpus, before limit. meta.total is unchanged and still means the number of rows in data. In hybrid mode keyword_match_count may be smaller than total; see Search for why that is expected rather than a defect.
All additive and optional; existing requests are unaffected.
Fixed: on GET /api/v2/documents, an empty betrag_min= / betrag_max= was parsed as 0 rather than as "no filter", which silently removed every document with no extractable total from the response. Sending the parameter with an empty value now correctly applies no filter. Sending 0 still filters at zero, as before.
POST /api/v2/search/answer (RAG synthesized answer, optional)New endpoint (search phase 3). Synthesizes a natural-language answer over the top hybrid-search matches, with every claim cited inline ([1], [2], …); sources lists only excerpts the model actually cited. Opt-in per organization, separate from plain search's opt-in — enabling POST /api/v2/search does not enable this endpoint. LLM calls run on the same EU/CH-compliant (Switzerland North) provider the rest of the AI pipeline uses. See Search.
data mirror removeddocument.processed and batch.completed previously duplicated their fields in a nested data object for backward compatibility. That mirror is removed — all six webhook events now ship the clean root format only (fields at the top level). Receivers must read payload.<field>, not payload.data.<field>.
POST /api/v2/search gains three optional filters — sender (Absender), receiver (Empfänger), property (Liegenschaft) — substring-matched on the document's entity names. Snippets are now centered on the first query-term match and wrap matched terms in <mark>…</mark> (HTML-escaped). See Search.
POST /api/v2/search (hybrid semantic + keyword)New endpoint. Meaning-based and keyword search fused into one ranked list, with matched_via per hit and the same doc_type / bereich / intent / tags / date filters as the documents list. Available per organization (403 until enabled). See Search.
usage.threshold.reached webhookGET /api/v2/usage now accepts ?period=YYYY-MM to read a past month. Current-month data is live (90-day-retention telemetry); past months are served from a durable monthly rollup that outlives the 90-day window, so older per-route history stays queryable. A new source field reports "live" or "rollup". Default behavior (no period) is unchanged — still the current month, month-to-date. A malformed period returns 400.usage.threshold.reached — fires once per calendar month when your org's API-call count crosses a configured alert threshold. Clean root payload { event, timestamp, organization_id, period_start, threshold, api_calls }. Like the other opt-in events, existing subscriptions are not auto-subscribed — enable it in Settings → Webhooks./api/v2/usage snapshot endpointGET /api/v2/usage returns a month-to-date usage snapshot for your organization: api_calls (total this_period plus a per-route by_route breakdown), documents (total_stored, ingested_this_period), storage.bytes_used, the current rate-limit remaining / rate_limit_reset (mirroring the X-RateLimit-* headers), and limits.* (null = unlimited until per-org quotas land). The organization is taken from the API key, so there are no path or query parameters.api_calls.by_route is sourced from request telemetry with 90-day retention, so the per-route history reaches back at most 90 days. Additive — nothing to change in existing integrations.extracted_metadata describing the timeframe a document covers:
period_from / period_to — the reporting period a document looks back on (tax documents, financial reports, bank statements, HR/payroll).valid_from — the start of a validity window; pairs with the existing valid_until (insurance, certificates, financing documents, contracts).DD.MM.YYYY strings. Populated only for the relevant document type and only when the document states the dates explicitly (a bare fiscal year like "2024" expands to 01.01.2024–31.12.2024); null/absent otherwise. A document carries at most one of the two pairs.GET /api/v2/documents, GET /api/v2/documents/:id) inside extracted_metadata, just like the other metadata fields. Additive only — no client change required; existing integrations are unaffected./api/v1/* API has been removed; v2 is now the only public API surface. v2 is a field-superset of v1 (every v1 field is still present, plus the current-model fields), so existing integrations migrate by swapping the base path /api/v1 → /api/v2 and reading document_tags[].tag_key (the redundant tag_value is dropped — it always equalled tag_key). The former v1-only batches and sync/documents endpoints are retired — use GET /api/v2/documents?updated_since=<ts> + the meta.server_time cursor for delta sync.GET /api/v1/batches, GET /api/v1/batches/:id, and GET /api/v1/sync/documents so this page is the single complete API reference. No behavior change — these endpoints already existed on the /api/v1/ path; they are now described here (with a note that the /api/v1/ path is what's live today). (Superseded: the entire /api/v1/* surface was removed — see the latest changelog entry. Use GET /api/v2/documents?updated_since= + meta.server_time for delta sync.)GET /api/v2/documents now supports betrag_min / betrag_max (number — filter on extracted_metadata.total_amount) and faellig_after / faellig_before (ISO date YYYY-MM-DD — filter on extracted_metadata.due_date). All optional, combinable with the existing filters; amount/date parsing is server-side and calendar-invalid stored dates are treated as no-match (never error). Previously documented as not-implemented.doc_type can now also be plan, land_register_extract, or authority_decision — non-image documents (building/floor plans and drawings, land-register/cadastral extracts, authority decisions and permits) that were previously dumped into other. Additive values in the existing open vocabulary; doc_type is a string, not a fixed enum — v1 consumers that whitelist types should add these if they want to surface them, otherwise they behave like any other type.doc_subtype is now strict: it is always either a known subtype for its doc_type or null. Previously the model could emit free-text or German variants that leaked into the field; those now collapse to null. No action needed — already-valid subtypes are unchanged; only invented/foreign values stop appearing.photo doc_typedoc_type can now be photo — image files (jpg/png/…) the classifier cannot assign to a real document type are typed photo instead of other. Filterable via ?doc_type=photo. Additive value in the existing vocabulary; v1 consumers that whitelist types should add photo if they want to surface it (otherwise it behaves like any other type).extracted_metadata._ai_possible_duplicate: true is now set on documents whose filename carries a copy marker ("- Kopie", "(1)", "- Copy"). Non-destructive — the document is still processed normally; the flag only surfaces likely duplicates for review. Absent when not a copy. Additive.extracted_metadata.vat_amount — the VAT/MwSt amount in money as a number string without a currency code (e.g. "34.96"), nullable. Appears in the read API and the document.processed webhook alongside vat_rate.extracted_metadata.vat_amount_source — "extracted" or "derived", nullable. It is always present when vat_amount is, and absent when it is not.
"extracted" — the document printed the amount and we read it. Never computed: a document that states a rate and a subtotal but not the VAT amount itself yields null, not the product of the two."derived" — we computed it as total_amount minus the sum of line_items[].total, and only where that arithmetic is about VAT: the document states a rate, every line carries a total, and the lines reconcile with the total at that rate. Documents whose lines simply equal their total carry no VAT amount at all, because the difference there is zero and 0.00 would assert that the document charges no VAT.vat_amount_source is the field that tells you whether the number came off the page.vat_rate keeps its meaning exactly.extracted_metadata.vat_rate — the VAT/MwSt rate as a percentage number without the % sign (e.g. "8.1", "2.6", "3.8"), nullable. Extracted ONLY when the document explicitly states the rate; never inferred. Appears in the read API and the document.processed webhook alongside the other amounts.extracted_contacts on the detail read (v3 orgs)GET /api/v2/documents/{id} now returns the same top-level extracted_contacts array the document.processed webhook delivers (same builder, same v3-org gating, same one-entry-per-contact + relationships semantics). Applies to the raw shape and to every output profile.external_id + updated_since filtersGET /api/v2/documents?external_id=<id> — exact lookup by your external_id.GET /api/v2/documents?updated_since=<ISO8601> — documents with updated_at after the
timestamp, ordered updated_at ascending (delta sync); malformed timestamp → 400.external_id for idempotent
lookups and updated_since for delta-sync polling.meta now also carries server_time (the server's query-time timestamp) —
save it and pass it as the next updated_since for incremental polling._-prefixed markers are no longer exposedextracted_metadata._ai_possible_duplicate (and any other internal _-prefixed bookkeeping key) is no longer included in API responses or webhook payloads. It remains an internal review marker only. This reverses the 2026-06-13 note above that surfaced _ai_possible_duplicate to clients.extracted_metadata.booking_date (format DD.MM.YYYY, nullable) — extracted ONLY when the document explicitly labels a booking date ("Buchungsdatum", "gebucht am"); never inferred from the issue or due date. Appears in the read API and the document.processed webhook payload alongside date_issued.error_message: File is password-protected and cannot be processed. Remove the password protection and upload it again. Previously an encrypted PDF could complete silently with empty content.document.failed section now lists all stable error_message wordings integrators may match on exactly; they will not change without a changelog entry.POST /api/v2/feedback — report a field your users corrected (event_type: "classification.changed" with field_path, old_value, new_value) or an AI draft your users deleted (event_type: "ai_rejected", no field detail needed).404 for documents outside your organization.Add-only. Newest entries on top. Older entries are preserved verbatim.
force=true testing flagPOST /upload (multipart field) and POST /upload/finalize (JSON field) accept force: true to bypass content-hash idempotency and re-process identical bytes as a new document. Forced rows store no content hash (excluded from future dedup); reusing an existing external_id with force returns 400. Response meta gains forced: true. Built for integration test loops — production ingestion should not use it.extracted_contacts on document.processed (v3 orgs)document.processed for v3-opt-in organizations now carries a top-level extracted_contacts array — the document's deduplicated contact records with per-contact confidence (0..1), the new function field (person's role/job title) and an explicit relationships pairing signal referencing sibling entries. See Webhooks.function (e.g. "Geschäftsführerin"); also returned on the contacts endpoint./openapi-v2.yaml. Generated from the route handlers; validated with Redocly. The markdown reference stays authoritative for semantics.contact_id values of merged duplicates no longer resolve. Treat document_id as the stable key and re-fetch contacts instead of caching contact_id long-term (now documented on the contacts endpoint). Document ids, payload structure and webhooks are unchanged.~$… stubs) are now rejected up front: 400 on multipart /upload and presigned /upload/init, and per-entry rejected_reason: "office_lock_file" inside ZIP containers.error_message "File is corrupt or not a valid Office document (.docx/.xlsx/.pptx)." instead of a raw parser error — visible on the document row and in the document.failed webhook payload.POST /upload/init + POST /upload/finalize flow — file bytes stream directly to storage, lifting the per-file limit from the multipart 4 MB cap to 250 MB. Multipart /upload now returns a descriptive 413 pointing at the presigned flow.date_of_birth), the document-level entity fields (sender / receiver / related_contacts / property / unit / equipment), and the ZIP container upload contract (entry rules, per-entry rejected_reason values) are now fully documented. No behavior change._v3 field-name suffix_v3 suffix on intent_v3, document_classification_v3, and v1_projection is a stable contract identifier, not an API version. A field tagged _v3 belongs to the current classification model; a field tagged v1 belongs to the legacy taxonomy the v1 API surfaces unchanged.document_classification_v3 object is null on v1 orgs and populated when opted in.document_classification_v3 was null. They assumed it was a bug, not a per-org config. The docs now state this up-front so integrators don't lose time chasing a non-bug.POST /api/v2/documents/upload/init returns a Supabase Storage signed PUT URL + the bucket-scoped storage_path + a 2-hour expiry. JSON body, no file bytes.POST /api/v2/documents/upload/finalize registers the document after the caller has PUT the file to the signed URL: idempotency (external_id → content_hash) → storage object existence check → documents row + batch + audit log + pipeline trigger./upload endpoint unusable for real-world property-management documents. The presigned flow streams the file content directly to Supabase Storage and never touches our API function, lifting the per-file cap to the storage bucket's file_size_limit (now 250 MB per migration 056).Authorization: Bearer idms_… key as the rest of v2, same 100 req/min rate-limit shared with /upload, same { data, meta, error } envelope.upload_url is a plain signed PUT — curl -X PUT --data-binary @file … works. Callers using @supabase/supabase-js can also use uploadToSignedUrl(storage_path, upload_token, file).GET /api/v2/documents/:id/download returns a short-lived Supabase Storage signed URL for the document's original bytes. Client follows the URL directly — bytes do not travel through the API function. Default TTL 3600 s, clamped to [60, 86400] via ?ttl=.IDMS_API_KEY, 100 req/min/key, { data, meta, error }).id and organization_id; a cross-org hit returns 404 (not 403) so existence is never disclosed.GET /api/v2/documents/:id.doc_bereiche, doc_types, doc_subtypes, intents, tag_groups, tags) now carries label_en alongside the existing label_de. Backfilled for all 237 taxonomy entities (8 + 31 + 121 + 13 + 10 + 54). Schema change applied via migration 052; data backfill via migration 053. Strictly additive: a v2 consumer that only reads label_de is unaffected.GET /api/v2/taxonomy now returns label_en on every row in bereiche, doc_types, doc_subtypes, intents, tag_groups, and tags. Null in the response means the row pre-dates the backfill — currently nothing is null in prod, but the field is nullable on the wire so a future taxonomy migration can introduce a new row without immediately needing the translation.Label (EN) column. The new German enum glossary subsection there documents which enum values stay German on the wire (offen, bezahlt, neu, abgeschlossen, eingehend, ausgehend, absender, empfaenger, …) and what each one means. Those values are contract identifiers and will not be renamed.bezüge, Pflichtfelder, Empfänger-Regel, Modul/Ordner/Vorlage/Belegart) now lead with the English term and keep the German in parentheses at first mention.org_config.classification_version ∈ (v1, v3), default v1 for every existing org (the value v3 here is a model-version literal — the API itself stays at v2). Flip via INSERT INTO org_config … ON CONFLICT … DO UPDATE SET classification_version = 'v3'.document.processed so the payload can carry classification_v3 = { doc_type, doc_subtype, doc_bereich, intent_v3, status_lifecycle, zahlungsstatus, tags[] } ADDITIVELY (v1 fields untouched).UPDATE org_config SET classification_version = 'v1' WHERE … reverts behavior on the next processed document. Persisted rows in document_classification_v3 are preserved.tsx scripts/reclassify-org.ts --org <uuid> --dry-run (default) or --apply (writes). Rate-limited (default 1.5 s between docs). Refuses to run when the target org is still on v1.document.failed — fires after the pipeline exhausts retries. Payload (clean root, no deprecated data wrapper): { event, timestamp, document_id, filename, external_id, error_stage, error_message }. error_message is sanitized (paths, URLs, stack frames stripped, 500-char cap).classification.changed — one event per axis touched in a PATCH /api/v2/documents/:id. Payload: { event, timestamp, document_id, field, old, new, actor: { kind, id } }.tag.suggested — fires when the AI proposes a tag outside the controlled taxonomy. Dedup per (document_id, tag) via extracted_metadata._tag_suggested_fired. Payload: { event, timestamp, document_id, suggested_tag, confidence }.document.processed and batch.completed — ship the clean root format; the deprecated nested data wrapper on the two original events was removed 2026-07-04.POST /api/v2/documents/upload — single file, multipart. Idempotency via external_id first, then SHA-256 content_hash. Same content uploaded twice in the same org → meta.idempotent_replay = true + the existing document; no second storage object, no second pipeline trigger.POST /api/v2/documents/upload/batch — up to 50 files per request. Per-item result { index, status: "created" | "idempotent_replay" | "failed", doc?, error?, dedup_by? }. One bad item never aborts the rest.20260603105516.pdf) repeat by the thousands; v2 stores each accepted upload at ${orgId}/v2/${stem}-${random8hex}${ext} so storage collisions are impossible. The caller's original filename is preserved in documents.filename.documents.content_hash (nullable) + partial UNIQUE on (org, content_hash).GET /api/v2/documents/:id?profile=<profile_key> — runs the configured 4-level routing engine for that client: Modul (module) → Ordner (folder) → Vorlage (template) → Belegart (document kind) with direction (inbound/outbound resolved via a receiver rule — Empfänger-Regel) and buchungsrelevanz (posting-relevance) axes, OU resolution from the per-tenant org tree, per-template status set + required-field (Pflichtfeld) validation. Returns { modul, ordner, template, belegart, direction, buchungsrelevant, ou, status_set, status, required_fields, missing_required_fields, ... } with meta.routing_rule + meta.missing_required_fields. Profile keys are provisioned per client by your account manager.taxonomy_maps (40 routing rows for a typical client). Migration 049 extends taxonomy_maps with direction, buchungsrelevant, and per-client routing columns (target Modul / Ordner / Template / Sub-template).output_profiles.fields JSONB carries { renamed, excluded, included, tag_groups_included, required_fields, fallback_label }.PATCH /api/v2/documents/:id — atomic correction of doc_type / doc_subtype / doc_bereich / intent / status_lifecycle / zahlungsstatus + tags + optional reason.POST /api/v2/documents/:id/tags and DELETE /api/v2/documents/:id/tags — idempotent tag UPSERT/DELETE with confidence = 1 sentinel marking manual overrides.PATCH /api/v2/documents — bulk (max 100 items, per-item result, one bad item does not abort the rest).document_audit_log row with prev/next snapshots and a diff (fields_touched or added/removed).document_tags_v3.confidence = 1 + extracted_metadata._manually_corrected_fields[] + _manually_corrected_at so a future re-classifier can detect and preserve manual edits.GET /api/v2/documents — list with current-model filtering (doc_type, subtype, intent, status, zahlungsstatus, bereich, repeatable tag).GET /api/v2/documents/:id — full document with the current classification record, structured metadata, related contacts, and linked entities (Bezüge — property/unit/equipment).GET /api/v2/documents/:id/contacts — extracted Person/Firma contacts with role (absender/empfaenger/cc), normalized across all docs in the org.PATCH /api/v2/documents/:id/classification — correct doc_type, doc_subtype, intent, status_lifecycle, zahlungsstatus, and tag set.PATCH /api/v2/documents/:id/metadata — MERGE update of the current-model extracted_metadata (null deletes a key).PATCH /api/v2/documents/:id/contacts — REPLACE the document's contact links (anti-spoof org check on every contact).GET /api/v2/taxonomy — full taxonomy (8 bereiche, 31 doc_types, 121 subtypes, 13 intents, 10 tag_groups, 54 tags) with German labels and v1 projection.document_classification_v3 (doc_bereich, doc_subtype, intent_v3, status_lifecycle, zahlungsstatus) plus document_tags_v3.status_lifecycle ∈ {abgeschlossen, finalized, deleted} return 400 DOCUMENT_LOCKED. Reads unaffected.v1_projection so clients can translate current-model values back to v1 vocabulary for legacy integrations./api/v1/*.Dashboard file upload no longer silently overwrites files with the same name in a single drag-drop batch. Duplicate names are auto-suffixed (scan.pdf → scan (2).pdf) and surfaced in the upload toast.
New field signature on documents (list, detail, sync, and the document.processed webhook): { signed, signed_by?, signed_at?, confidence?, raw? }. Auto-detected from text (mainly for contracts) and overridable in the UI. Backward-compatible additive change.
Added JavaScript examples, rate limit headers, batch.completed webhook payload, error response bodies, pagination guide, API scope notes.
New fields: doc_type, doc_subtype, doc_intent, sender, receiver, property, unit, equipment. Structured tag taxonomy (59 tags, 8 groups — v1 vocabulary). Old classification field kept for backward compatibility.
Added POST /api/v1/documents/upload and GET /api/v1/sync/documents endpoints. Client webhook support via environment variables.
Documents, batches, tags, API key auth, rate limiting, per-org webhooks.