Practical Reference
AI-Operable IT-Services — a feasibility study with a runnable demonstrator
How must IT services be designed so that AI agent systems can steer large parts of their lifecycle — operation, optimization, evolution — autonomously and under control? The study answers this not only conceptually, but with a real, governance-led demonstrator.
What it is about
Value does not arise from the single, ever more powerful agent, but from the interplay of four building blocks. AI-Operability thereby becomes an independent quality dimension — on a par with security, maintainability and scalability.
machine-understandable, observable and steerable through clearly defined interfaces (API-First, Self-Describing, Semantic Observability).
separates what an agent technically can do from what it is allowed to do — as executable code in the decision path, not as a PDF.
specialized roles (Operations, Security, Cost, Evolution, Auditor …) on a shared, generalizable framework.
persistent, model-independent memory for knowledge, patterns and experience — it outlives models and providers.
What actually runs in the demonstrator
Between 10 and 12 June 2026, a runnable, governance-led demonstrator was built in a real cloud sandbox: an AI-operable reference application, a populated AI Service Brain, an agent system on a shared framework and a live-enforced governance. The following evidence is taken from the whitepaper and is marked there as preliminary:
- Governance executable and faithful to its implementation: the live-enforced policy (Open Policy Agent) is, across 156 checks, 100 % congruent with its specification — 0 deviations (F-011).
- Autonomy ladder demonstrated on real data: observe → approval → autonomous; the level is set externally by the policy, not by the agent code. One approval let the agent — not the control center — scale for real, fully audited and reversible (F-002).
- The AI Service Brain is the decisive quality lever: in a blind A/B comparison, “with Brain” clearly wins on domain-specific questions (quality ≈ 3.5 vs. 0.5 out of 5); blind grounding harms results on off-topic questions → relevance gating (F-006).
- Closed learning loop: an agent recommendation was implemented and its effect measured — the model-provider switch measurably reduced latency (~3–10×), the scaling lever did not (F-005).
- Model choice as a governed, reversible trade-off: Claude Opus ↔ EU-resident Gemini switchable at runtime, without redeploy — coupled to trust and residency levels (F-004).
Relation to the seven missions
The study is the real-world counterpart to the architecture principles of this site: each of the seven missions finds a concrete, partly measured equivalent in the demonstrator. It thereby complements the didactic Service-Agent example with real evidence.
Limits & context
All findings are explicitly preliminary: one cloud sandbox, two services, a short observation period, a partly synthetic corpus and heuristically set confidence values. The study deliberately separates concept (target form) from implementation status — it is a research program, not a product proof.
This page summarizes the study in curated form and remains independent of the hr/ARD context in which it was created. The screenshots come from the demonstrator (as of June 2026).
Whitepaper reading sample
The study is available as a whitepaper (81 pages, v3.7, June 2026, in German). A reading sample — cover, table of contents and chapters 1–3 (executive summary, starting point, vision and target picture) — can be read directly. The full version is available on request by e-mail.
The concepts behind this reference?
Nova-7 explains governance, AI readiness and agent architecture from the seven missions.