Sovereign AI · Local Language Models · ADVISORI

A private LLM for your enterprise,hosted in Europe.Your data stays with you.

A private LLM, instantly available via API: performance on par with the top US models, at a fraction of the cost.

Operated in our own data centres in Berlin, Frankfurt, Vienna, Zurich and Bolzano. No CLOUD Act, no kill switch, no data leaving your control.

0.0×
cheaper than frontier API pricing
5
locations: Berlin · Frankfurt · Vienna · Zurich · Bolzano · soon: Hamburg — or on your premises
Instant
ready via API — no onboarding project
0
data shared with US model providers
Sovereignty ticker
16 Jul Kimi K3 debuts at #3 worldwide — right behind Fable 5 and GPT-5.6; the open weights are out since 27 Jul and the model runs in our catalogue12 Jun US government switches off Claude Fable 5 worldwide — 19 days offline, for all customers30 Jun US lifts the Fable 5 ban — with permanent strings attached: future model releases only in coordination with the US government8 Jul Fable 5 now credits-only: $10 / $50 per million tokens — the most expensive Anthropic model ever26 Jun GPT-5.6 launches for only ~20 government-vetted partners — Washington decides who gets frontier AI29 Jun US Supreme Court ruling shakes the legal basis of EU-US data transfers — noyb calls for revoking the adequacy decision28 Jun Austria asks Brussels to bring Anthropic to the EU — a consequence of the Fable 5 shutdownfrom 2 Aug EU AI Act: GPAI enforcement begins — fines of up to 3% of global revenue for model providers
Fundamentals

What a private LLM means for your enterprise

A private LLM is a language model running on your own hardware or in a dedicated European data centre instead of at a US provider. No business data leaves the controlled perimeter, no outside vendor can switch off your access, and the operator decides which model runs in which version.

Distinction

Private does not necessarily mean your own server room

What matters is control, not where the metal sits. A model in a dedicated data centre under European law serves the same purpose as one in your basement, without you having to buy hardware and hire people to run it. Both count as a private LLM; the difference is effort and capital outlay.

Distinction

Why an EU endpoint alone is not a private LLM

Many providers advertise a European access point. If the company behind it is subject to US law, the CLOUD Act still applies, and cutting off access remains a decision outside your influence. The sovereignty section below sets out the four levels and shows exactly where that line runs.

Use

What enterprises actually run on a private LLM

Document analysis on confidential contracts, knowledge search across internal repositories, minutes and summaries from meetings, code assistance on your own source, triage of customer enquiries. The common denominator is always the same: data a company will not, or may not, hand to someone else’s cloud.

The demo

Performance in Action

A short look at sovereign operation: from prompt to answer — entirely on our infrastructure across five European sites.

Models & pricing

Open frontier-class models — available now, transparent pricing

All models are already running and usable via API immediately — billed by usage, per million tokens, no base fee, no minimum spend; GLM 5.3 and the new DeepSeek V4 Pro follow on 1 September. On request we also run the same models dedicated or on-premise at a fixed price.

Now in the catalogue

Kimi K3 is live — sovereignly operated by us

Ranked #3 worldwide — right behind Claude Fable 5 and GPT-5.6, ahead of Opus 4.8 — and the leading open model. Sovereignly operated by us: 1M-token context, €3.00 input / €15.00 output per 1M tokens. And the catalogue keeps growing: on 1 September, GLM 5.3 — currently the strongest open coding model — and the new DeepSeek V4 Pro follow.

Secure capacity now
ModelClassContextLocationsInput€ / 1M tokensOutput€ / 1M tokens
GLM 5.3from 1 SepAll-round1Mfrom 1 Septbatba
DeepSeek V4 Pro 0813from 1 SepAll-round1Mfrom 1 Septbatba
Kimi K3Best sellerAll-round1MBerlin · Zurich3.0015.00
DeepSeek V4 ProAll-round1MAll 5 locations1.753.50
Qwen3.5 397BAll-round256KAll 5 locations0.603.60
Kimi K2.6All-round256KBerlin · Zurich0.954.00
GLM 5.2All-round1MAll 5 locations1.404.40
GLM 5.1All-round200KFrankfurt · Vienna1.404.40
MiniMax M3All-round1MVienna · Zurich · Bolzano0.301.20
Mistral LargeAll-round256KBerlin · Frankfurt · Zurich0.501.50
DeepSeek V4 FlashEfficiency1MAll 5 locations0.290.35
gpt-oss-120bEfficiency128KBerlin · Vienna0.150.60
Qwen3.6 27BEfficiency256KAll 5 locations0.402.70
Qwen3 Coder 480BCode256KAll 5 locations0.401.60
Kimi K2.7 CodeCode256KBerlin · Zurich · Bolzano1.445.18
Devstral SmallCode128KVienna · Bolzano0.100.30

Prices net, as of August 2026; billed by actual usage. More models on request — the catalogue grows continuously.

Dedicated or on-premise operation: fixed price based on sizing instead of token billing — talk to us.

Fugi · The orchestration

Many models. One conductor.

Fugi is our orchestration layer on top of the model pool: a small coordinator model receives every request and decides for itself what to do — answer directly, or assemble a team of specialists, hand out sub-tasks, cross-check results and condense them into one answer.

FUGI · the conductorClaude CodeCodexOpenCodeout of the boxYour requestFUGIlearned, not hard-wiredmodel pool · swappableGLM 5.2ThinkerQwen3 CoderWorkerDeepSeek V4VerifierKimi K3MiniMax M3Fugi itselfyour rule: DACH locations only 🔒One verified answer ✓simple → answers directlyThe team beats the star — composition instead of one giant model.
The difference

Learned, not hard-wired

Classic routers follow hand-maintained if-then rules. Fugi’s coordinator is itself a trained model: its strategies — who takes what, when to verify, when one answer suffices — emerge from experience and improve with every training run.

The effect

The team beats the star

Complex tasks are decomposed and assigned to the strongest specialists — with a verification step before the answer. A federation of open models reaches frontier-level results without needing one giant model.

The sovereignty

The pool is yours

New models dock on without retraining, deprecated ones drop out, and your rules apply at system level — such as “only models at DACH locations”. The whole federation runs on our infrastructure across five European sites.

Fugi orchestrates the model pool across all five sites — access comes via the waitlist.

Price calculator

Pick your model — and see the difference

Choose a model, set your monthly volume: the chart plots your monthly output cost against Claude Fable 5 at market API pricing — and how the savings grow with every token.

Kimi K3 is, per output token, 3.3× cheaper

Kimi K3 · local at ADVISORIClaude Fable 5
010,00020,00030,00040,00050,00002505007501000€ / monthM output tokens / monthYour savings: €10,500 / monthClaude Fable 5Kimi K3
Kimi K3: €4,500/moClaude Fable 5: €15,000/mosavings / month: €10,500

Comparison basis: Claude Fable 5 at the verified market API price ($10 input / $50 output per 1M tokens, July 2026 — the most expensive Anthropic model ever listed). Calculated over output tokens, because that is where the cost lives: Fable 5 always thinks, and thinking tokens are billed as output. User equivalence: around 0.3M output tokens per employee and month — a typical mix of assistant requests, knowledge-base queries (RAG) and agentic software development. Prices net.

The performance proof

On par. Proven.

Open frontier-class models are no longer a compromise. The head-to-head: GLM-5.2 (open, in our catalogue) versus Claude Opus 4.8 (US frontier) — measured, not claimed. And the gap keeps closing: Kimi K3, ranked #3 worldwide, is now openly available and running in our catalogue.

GLM-5.2 · openOpus 4.8 · US
Terminal-Bench 2.1agentic82.778.9
AIME 2026mathematics99.295.7
Composite scoreoverall9193
Cost / 1M outputAPI price$4.40$25.00
Licence & accessMIT · openproprietary · API
Operable in the EU

5.7× cheaper per output token. On agentic terminal tasks and mathematics it already beats Opus 4.8 today — both with 1M-token context. The narrow composite-score gap costs one fifth of the price.

benchlm.ai · llm-stats.com · as of June 2026

GLM 5.2 is in the price table above — available via API right away →

The integration

Swap the base URL. Done.

Our endpoint is OpenAI-compatible: existing code keeps running unchanged — you only change the base URL and the model. No SDK switch, no migration project.

from openai import OpenAI

client = OpenAI(
    base_url="https://api.advisori.de/v1",
    api_key="IHR_ADVISORI_KEY",
)

antwort = client.chat.completions.create(
    model="glm-5.2",
    messages=[{"role": "user", "content": "Hallo!"}],
)
print(antwort.choices[0].message.content)
Compatible out of the box:Claude CodeOpenCodeCodex

You receive your API key with your capacity release from the waitlist.

Your models live here.ATCHyour European jurisdiction ✓BerlinNVIDIA B200Heiligenhaus · inactiveFrankfurtNVIDIA B200Hamburg · soonZurichBolzanoViennaYour companyshort distances · your jurisdictionUS cloud~ Atlantic ~CLOUD Act ✗No ocean between you and your AI — 5 sites live · 1 in build-out · 100% Europe
The engine room

Your models live in Berlin, Frankfurt, Vienna, Zurich and Bolzano — Hamburg is next.

Our inference servers are located in Berlin, Frankfurt, Vienna, Zurich and Bolzano — running on NVIDIA B200, the current GPU generation. Hamburg is in build-out. High performance, high quality, short distances. And entirely in Europe: operations, data and jurisdiction stay where you are.

5
sites live: Berlin · Frankfurt · Vienna · Zurich · Bolzano — Hamburg in build-out
B200
current-generation NVIDIA GPUs
100 %
Europe — operations, data, jurisdiction
Security & compliance

Built by security people — not bolted on

ADVISORI is a security company. Model operations follow the same standards we apply when advising banks and insurers.

Data handling

Zero data retention

Prompts and responses are not stored — they are processed and discarded. Your data feeds no training and reaches no third party.

Encryption

Encrypted, end to end

TLS 1.3 in transit, AES-256 for everything at rest. Hardened and monitored by the ADVISORI FTC security team.

Contract

DPA included

A data processing agreement under Art. 28 GDPR is part of the contract — with processing exclusively at our European sites (DE, AT, CH, IT). EU AI Act evidence included.

Governance

Control for your IT

Dedicated API keys per team, budget limits and an audit trail across all requests — your IT keeps track of who uses what.

Multimodal guardrails — your policies, enforced at the endpoint

GUARDRAILS · the checkpointmultimodalTextImageDocument“…IBAN DE89 3704…”PII maskInjection filterYour policy 🔒“…IBAN ▮▮▮▮ ▮▮▮▮…” · masked ✓Your modelsGLM · Qwen · DeepSeek …the answer is checked too ✓Prompt injectionblocked ✗Audit logchecked ✓masked ✓blocked ✗Nothing reaches the model unchecked — and nothing leaves it unchecked.
Inputs

PII protection, multimodal

Personal data is detected and masked in text, images and documents before it reaches the model — and, on request, in the responses as well.

Defence

Prompt-injection defence

Attacks via manipulated inputs, documents or web content are detected and blocked at the endpoint — before they can steer the model.

Your rules

Company-own policies

Your guidelines as enforceable guardrails: topic blocklists, compliance requirements, department-specific approvals — centrally maintained, effective per team and API key.

Evidence

Every decision in the audit log

What was checked, masked or blocked is documented traceably — per request, per rule, audit-proof for your compliance.

The guardrails apply before and after every model call — regardless of which model Fugi selects. Also available as a standalone service in front of any model: Fugi Guardrails →

Certified security & quality of ADVISORI FTC
ISO 27001ISO 9001ISO/IEC 42001SOC 2 Typ IIEU AI ActDSGVO

The same certifications ADVISORI holds when advising regulated companies — no “in preparation”, no “expected”.

The operation

What sovereign operation looks like

ADVISORI · inference consoleModels: DeepSeek · GLM · Qwen · Kimi 🔒100 % Europa ✓Request · live> analyse contract_2026.pdfanswer in 1.2 s ✓route: Berlin · B200leaves the building: neverUtilisationBerlin100Frankfurt100Vienna100Cost ticker1M tokens: €4.40 instead of $25.005.7× cheaper ✓Privacystored: 0 byteszero retention ✓audit trail: onAPI · models · budgets · audit · statusNVIDIA B200 · 5 European sitesgreen powerYour own AI power plant. Already running — secure a slot.
Why now

When Washington pulls the plug

What was long a hypothetical risk has been documented reality since June 2026 — shutdown, access lists, overnight price changes.

Precedent

The off switch has been used

No longer hypothetical: in June 2026, Claude Fable 5 was switched off worldwide for 19 days by US government directive — for all customers. No contract, no SLA protected against it.

Access

Access is granted politically

GPT-5.6 launched exclusively for around 20 partners individually approved by Washington. Who gets to use frontier AI is increasingly decided by the US government — European companies are not at the table.

Cost

Price changes overnight

The same model that was just included in the subscription costs $10 / $50 per million tokens from July 8 — no grandfathering. Local models flip the equation: fixed infrastructure, plannable costs. The price calculator above shows the difference.

Law

CLOUD Act & structural dependency

US authorities can access data held by US providers — regardless of which data centre it sits in. And according to the first UN AI report, 75% of the world’s AI computing power is located in the US.

Regulation

GDPR & EU AI Act

Documentation duties, data minimisation, transparency. Easiest to satisfy when the model runs where your data already lives: with you.

Chronicle: the last 30 days
  1. 12 Jun 2026US government orders the shutdown of Claude Fable 5 & Mythos 5 — worldwide, for all users. Source →
  2. 26 Jun 2026OpenAI launches GPT-5.6 for only ~20 government-vetted partners — access approved customer by customer. Source →
  3. 01 Jul 2026First UN AI report: 75% of global AI computing power is located in the US, 15% in China. Source →
  4. 08 Jul 2026Fable 5 back online after 19 days — throttled, and from July 8 only via usage credits ($10 / $50 per million tokens). Source →
  5. 16 Jul 2026The open world catches up: Moonshot unveils Kimi K3 — ranked #3 worldwide, ahead of Opus 4.8. The weights were released on 27 Jul. Source →
Digital sovereignty

Digital sovereignty: what the term actually demands

Digital sovereignty is usually debated as a stance and rarely written down as a checklist. For running a private LLM, digital sovereignty can be pinned down precisely: four questions, each with a verifiable answer. Who can switch off your access? Whose law governs your data? Who decides on model changes and pricing? And could you take over operations yourself if you had to?

EU data residency

The data physically sits in Europe, but the provider remains subject to US law. That satisfies storage-location requirements and nothing else, and has nothing to do with digital sovereignty yet: the CLOUD Act still applies, and shutdown remains someone else’s decision. Most offerings advertising European data centres stop here.

Location settled, control not

European operator

Operations and contract sit with a European company, while the model still comes from a US provider on that provider’s terms. A clear step up for digital sovereignty; technically the dependency remains. If the model is switched off or repriced, it hits you just the same.

Law settled, model not

Open weights in a European data centre

The model is open source or available with open weights and runs at a European operator. Now nobody can switch it off: you hold the weights, and the version stays frozen until you change it. This is where a private LLM genuinely begins, where most enterprises should land, and what ADVISORI delivers by default.

Shutdown risk eliminated

Full self-operation

Hardware, model and operations entirely in house. Maximum digital sovereignty, and the only route for classified environments or strict network separation. The price is capital expenditure on hardware, your own specialists, and responsibility for availability and updates.

Fully sovereign, at a cost

The decisive line runs between levels two and three: only with open weights does dependency on someone else’s release decision end. Everything below that is data protection, not digital sovereignty. Why a purely European access point does not cross that line is covered in the next section.

The decisive question

An “EU endpoint” is not sovereignty

Many providers advertise “EU data residency” — meaning a gateway in Europe that receives your requests and forwards them to US model providers. What matters is not where the request arrives, but where the model computes.

US cloud APIOpenAI, Anthropic & co. directlyAI gateway“EU data residency” / EU endpointLocal modelsADVISORI
Where does the model compute?In the US cloudStill at the US model providerOn your premises or in our data centres (DE/AT/CH/IT)
Who sees your prompts?The model providerGateway and model providerOnly you
CLOUD Act access possible?YesYes — at the model operatorNo
Cost modelAPI price per tokenAPI price + markupFixed infrastructure, unlimited usage
Can third parties switch it off?YesYesNo — you run it
Data protection

Private LLM and the GDPR: what self-operation changes

Running a private LLM in your own legal jurisdiction does not turn the GDPR into a formality, but it moves the question from a hard-to-evidence level to an easily evidenced one. Three things change concretely.

Third country

The transfer disappears instead of being justified

Calling a US model means justifying a third-country transfer under the GDPR and safeguarding it continuously. If the model runs in a European data centre, the transfer simply does not happen. That is the cleaner route: processing that does not take place needs no justification.

Evidence

Article 32 asks for effectiveness, not intent

Article 32 GDPR requires technical and organisational measures reflecting the state of the art that ensure confidentiality, integrity, availability and resilience, plus a process for regularly testing their effectiveness. Under self-operation the logs, access records and configuration are yours, and so is the evidence. With an external service you hold the provider’s certificate and little else.

Roles

The allocation of roles becomes unambiguous

With us, ADVISORI remains a processor under Article 28 GDPR, with a data processing agreement, a named data centre and no training on your data. Under full self-operation the processor disappears entirely. Either way the chain is short and documented in writing, rather than running through several sub-processors across shifting jurisdictions.

How to set up and document such an operation in detail, from role allocation to deletion policy, is covered in our guide on GDPR-compliant AI. It does not replace an assessment of your specific use case, which is exactly what we have advised on for years.

The offering

Open frontier-class models — three operating paths, one accountable partner

We select the right open model for your use cases, build the operation and take responsibility for it — including hardening, updates and monitoring.

Option A

On-premise with you

Operation on your own hardware, in your own building. Maximum control — the right choice for classified material, R&D and strictly regulated data.

Option B

European data centre

Dedicated operation on NVIDIA B200 infrastructure at one of our five European sites — no hardware investment of your own, under the law of the respective location (DE/AT/CH/IT), with a clear SLA.

Option C

Hybrid with failover

Local models as the default, defined fallback paths for peak load — orchestrated via Synthara AI Studio, with rules you define.

Industry models · fine-tuning

Your model speaks your industry

Beyond the catalogue, we fine-tune open frontier-class models on your industry and your documents: domain language, forms, processes, regulation. Training runs on our B200 clusters across five European sites — the fine-tuned model belongs exclusively to you and runs as its own endpoint.

INDUSTRY MODEL · the fine-tuningOpen base modelGLM · Qwen · DeepSeek …+Your documentsdomain language · forms · regulationFINE-TUNINGYour documents stay at our European sites 🔒Your industry modelexclusively yours · own endpoint 🔒“Your domain language?”“I speak it.”Base model + your documents = a model that speaks your language.

Clicking your industry adds it straight to your waitlist request — we follow up with a concrete proposal.

Optional extension

The API is complete. Synthara takes it to your business teams.

Your endpoint works on its own — your developers integrate directly via the API, no extra software needed. If you also want to bring AI to teams without developers, you can optionally put Synthara AI Studio on top of the same sovereign models: AI assistant, no-code studio, workflows and reporting — with automatic failover between your models.

Discover Synthara AI Studio →
How fast it goes

Three steps to your sovereign endpoint

Sign up

Join the waitlist — a business email is all it takes. We release capacity in waves and reach out in the order of sign-ups.

today · 2 minutes

Receive access & DPA

You get your API key, your site assignment and the data processing agreement — everything your IT and privacy teams need for approval.

after 3–5 days

Swap the base URL, go

The endpoint is OpenAI-compatible: one line of configuration and your applications, Claude Code or OpenCode run on sovereignly hosted models. You pay only for what you use — industry models and an on-premise path are always available as a next step.

day 1
Frequent questions

What decision-makers ask us most often

What is local AI?
Local AI means language models (LLMs) run entirely on your own infrastructure — on-premise in your company or in a dedicated European data centre — instead of a US provider’s cloud. Inputs and documents remain under your control: no data outflow, no CLOUD Act access, GDPR and EU AI Act compliant, with performance on par with open frontier-class models.
Are open models really good enough for enterprise use?
Yes. Open frontier-class models sit on par with the frontier providers in independent benchmarks — for the vast majority of enterprise tasks (documents, assistance, extraction, code) the difference is no longer measurable in practice. In the proof of concept we demonstrate this on your own tasks before you decide.
What sets you apart from providers advertising “EU data residency”?
With a gateway offering an EU endpoint, a server in Europe receives your request — but the computing still happens at the US model provider, where the CLOUD Act applies. With us, the model itself runs on your infrastructure or in our data centres in Berlin, Frankfurt, Vienna, Zurich and Bolzano. Your prompts never reach a US provider.
Do we need our own GPU hardware?
Not necessarily. The data-centre option (Berlin, Frankfurt, Vienna, Zurich, Bolzano) requires no investment of your own. For on-premise we advise on the right sizing — often far less hardware is needed than customers expect, because modern open models have become very efficient.
What does it cost compared to our current API usage?
Token prices are published transparently on this page — depending on the model, a fraction of frontier API costs. Use the price calculator to compare your preferred model directly with Claude Fable 5. Dedicated or on-premise operation is a fixed price based on sizing: the more you use, the bigger the advantage.
How quickly are we productive?
Via the API endpoint: immediately. The models are already running — once your access is activated, you integrate within minutes. Capacity is currently allocated; the waitlist secures you the next available slot. Dedicated or on-premise operation typically takes 3–4 weeks from analysis to go-live.
Who operates and maintains the models long term?
Both paths are open: we enable your team to take over operations — or ADVISORI runs them as a managed service with SLA, updates and monitoring. Many customers start managed and take over step by step later.
Is a data processing agreement (DPA) available?
Yes. A DPA under Art. 28 GDPR is part of the contract, complemented by your individual technical and organisational measures on request. Processing happens exclusively at our European sites (DE, AT, CH, IT); prompts and responses are not stored and feed no training.
Can we enforce our own policies (guardrails)?
Yes. Multimodal guardrails run at the endpoint: PII detection and masking in text, images and documents, prompt-injection defence and content filters. On top of that we implement your company-own policies — topic blocklists, compliance requirements, approvals per team or API key. Every decision lands in the audit log, and the guardrails apply regardless of which model handles the request.
What availability do you commit to?
The endpoint runs redundantly across our five sites in Berlin, Frankfurt, Vienna, Zurich and Bolzano. For dedicated instances and on-premise operation we agree an individual SLA with defined response times — the specific commitments are set jointly in the contract.
Next step

Secure the next available slot

Current capacity is allocated — we are expanding continuously. Join the waitlist: we reach out in the order of sign-ups as soon as your access is ready.

Join the waitlist

Questions first? Write to us: info@advisori.de · ADVISORI FTC GmbH