AI & Sub-processors

Version 1.0.0 · Effective 2026-08-12

LoremiAI is currently in Beta. This page explains exactly where your data goes, who can see it, and what is not yet finished. We would rather tell you plainly than let you find out later.

Last updated: 2026-08-12


The short version

  • Your documents are stored in Germany, and the search index is built on our own servers — your document text is never sent to a third party to make it searchable.
  • Generating answers and reading scanned documents (OCR) use external AI providers. Depending on which provider your workspace uses, this can mean your document content is processed outside the EU.
  • Our development team includes personnel located outside the EU who can access production systems for support and operations.
  • No AI provider we use trains its models on your content.

What happens to a document you upload

StepWhere it runsLeaves the EU?
Upload & storageOur servers in GermanyNo
Triage (deciding how to process the file)Our serversNo
OCR (reading scanned pages and images)External AI providerPossibly — see below
Parsing (extracting text from Office/PDF files)Our serversNo
Search indexing (embeddings)Our servers — local modelNo
Knowledge-graph extraction (Deep Reasoning tier only)External AI providerPossibly — see below
Answering your questionsExternal AI providerPossibly — see below

Two points deserve emphasis:

OCR routing. One category always goes to Google Gemini, regardless of your settings:

  • Arabic-language documents, in any file format — Gemini is materially more accurate on Arabic script, and we would rather give you a correct result than a worse one.

Everything else — scanned PDFs and image files, in English, German or any other non-Arabic language — uses the OCR provider configured for your workspace. That can be Mistral, which processes in the EU.

So: if you need OCR to stay in the EU, set Mistral as your OCR provider. Only Arabic documents will still be sent to Google.

Deep Reasoning tier. If your workspace uses the Deep Reasoning (knowledge-graph) package, your entire document is sent to an AI provider at indexing time to extract entities and relationships. This is a considerably larger disclosure than the Smart Search package, where only your question and the matching excerpts are sent. If this matters to you, use Smart Search.


Sub-processors

Sub-processorEntity & locationWhat it doesData it receivesTransfer safeguard
Hetzner Online GmbHGermany 🇩🇪Hosting, object storageAll stored dataNone required (EU)
Mistral AI SASFrance 🇫🇷OCR only (not answer generation)Document contentNone required (EU)
OpenAI Ireland LtdIreland 🇮🇪 — processing in the USAAnswer generationDocument excerpts, queriesData Processing Agreement + EU Standard Contractual Clauses / Data Privacy Framework
Google (Gemini API)USA 🇺🇸OCR for images + Arabic; answer generationDocument contentData Processing Agreement + EU Standard Contractual Clauses / Data Privacy Framework
Anthropic (Claude)USA 🇺🇸Answer generation, if selected for your workspaceDocument excerpts, queriesData Processing Agreement + EU Standard Contractual Clauses
Alphametic Services & Technologies s.a.r.l (AST)Lebanon 🇱🇧Development, technical support, operationsPotential access to production dataEU Standard Contractual Clauses (Module 3) + Transfer Impact Assessment

A Data Processing Agreement (DPA, German: Auftragsverarbeitungsvertrag) is the contract Art. 28(3) GDPR requires with anyone who processes personal data on our behalf.

None of these providers use your content to train their models.

If you configure your own endpoint — an Azure OpenAI resource in your Azure subscription, or a self-hosted model — that provider is not our sub-processor. Your content goes directly to infrastructure you control, under your own agreement with that provider.

About AST

AST is our development partner. Named AST personnel can access production systems to operate and support the service. Because Lebanon has no EU adequacy decision, this access is a third-country transfer, governed by Standard Contractual Clauses and subject to a Transfer Impact Assessment. Access is limited to named individuals, routed through an audited connection with session logging, and bound by individual confidentiality undertakings.

Under Art. 28(2) GDPR you may object to any sub-processor listed here. Contact us and we will discuss options — which, depending on the objection, may include moving your workspace to EU-only providers or ending the contract without penalty.


Which providers your workspace uses

AI providers are configured per workspace by an administrator under Admin → LLM. There are two separate roles, and they use different providers:

RoleWhat it doesProviders available
OCRReads scanned documentsMistral (EU) or Google Gemini
GenerationAnswers your questions, extracts the knowledge graphOpenAI, Azure OpenAI, Google Gemini, Anthropic Claude, or any OpenAI-compatible endpoint

Mistral is available for OCR only — not for generating answers.

Keeping everything in your own region

Because generation always runs through an external model, the way to keep that data in a specific region is to bring your own endpoint:

  • Azure OpenAI — create the resource in your own Azure subscription in the region you need (for example West Europe, Germany West Central, or UAE North), and enter that key. Your content then goes to your Azure tenant in your region, not to us or to OpenAI.
  • A self-hosted model — any OpenAI-compatible endpoint (Ollama, vLLM) on infrastructure you control. Nothing leaves your network at all.

Combined with our local search indexing, either option keeps document content within your chosen region for the entire pipeline. For OCR, pair it with Mistral and PDF inputs as described above.

We are working toward making this a guided setting rather than manual configuration.


PII masking

LoremiAI can detect and mask personal data (names, emails, phone numbers, IBANs and similar) before content is sent to an AI provider. Workspace administrators control this under Settings.

Please read this honestly: masking meaningfully reduces exposure, but it is automated detection, not a guarantee. It can miss things, particularly in unusual formats or languages. Masked content is still personal data in the legal sense. Do not treat masking as a licence to upload material you could not otherwise justify sending to a third-party processor.


What is not finished yet

We are in Beta, and these are known gaps:

  1. Google Gemini runs on the Developer API rather than an EU-resident enterprise deployment. Migration to a configuration with EU data residency is planned.
  2. Arabic documents always go to Google for OCR, whatever your provider setting.
  3. We have no EU-hosted generation model of our own. Keeping answer generation in a specific region currently requires you to supply your own Azure OpenAI or self-hosted endpoint, as described above.
  4. Region configuration is manual. It works, but an administrator has to set the providers rather than flipping a single switch.
  5. No formal external audit. We have no ISO 27001 or SOC 2 certification.
  6. No uptime guarantee. Beta carries no SLA.

Given the above: do not upload special categories of personal data (Art. 9 GDPR — health, biometrics, religious or political views, trade union membership, sexual orientation) or material subject to specific professional secrecy obligations during the Beta.


Changes

We will notify workspace administrators by email at least 30 days before adding or replacing a sub-processor, so you have time to object under Art. 28(2) GDPR. Material changes to this page will also require re-acceptance at next sign-in.

Questions: ⁦info@loremi.ai