AI.BAY Health Cluster  --  Freistaat Bayern

AI-BAY-HEALTH

A sovereign, multi-modal foundation model for medicine -- unifying imaging, omics, biosignals, clinical text, and tabular data into a single shared representation of human health.

Imaging Pathology Omics Signals Clinical Text EHR / Tabular
Scroll
The Concept

What is a Foundation Model?

A foundation model is a large neural network pre-trained on vast, diverse data. Rather than training a narrow model for each task, a single powerful model learns general representations that any downstream application can build on -- by prompting, fine-tuning, or retrieval.

Pre-trained on broad data

Billions of parameters learn the structure of language, images, and biology from large, heterogeneous corpora -- developing a deep prior over the natural world.

A universal representation

Any input -- an image, a sentence, a DNA sequence -- maps to a semantically meaningful vector. Related concepts are close in this space, regardless of their original format.

The foundation for many tasks

With lightweight adaptation, the same model powers diagnosis, prognosis, retrieval, generation, and clinical reasoning -- no full retraining required.

What we integrate

All of medicine, in one space

Every patient generates data across many modalities that clinicians integrate intuitively but that AI systems treat as separate silos. We train specialised encoders for each domain and project them all into a unified 8192-dimensional shared embedding.

🧤

Imaging

CT, MRI, X-ray, ultrasound. Volumetric and 2D encoders (MedSigLIP, CT-CLIP, Rad-DINO) learn anatomy and pathology from pixels and voxels.

Subproject A  -- FAU / TUM
🧰

Pathology

Gigapixel whole-slide images from surgical biopsies. Patch-level encoders (UNI, TITAN) map tissue morphology into the shared space at cellular resolution.

BRIDGE -- Radiology-Pathology
🧬

Omics

Single-cell RNA, DNA sequences, proteomics. Foundation models (scGPT, Nicheformer) encode the molecular state of cells and tissues at nucleotide resolution.

FORTE -- Omics Integration
📈

Biosignals

ECG, EEG, wearable sensor streams. Time-series transformers capture longitudinal physiological dynamics -- heartbeat to brainwave -- in compact vectors.

Subproject C -- LMU
📋

Clinical Text

Radiology reports, discharge summaries, clinical notes in German and English. A large VLM backbone acts as the universal semantic pivot across all modalities.

Subproject A -- FAU / TUM
📊

EHR / Tabular

Lab values, ICD codes, medications, demographics, FHIR records. TabPFN and EHR-adapted transformers encode the full structured clinical history.

Subproject C -- LMU
The Architecture

Converging on a shared space

Each encoder produces a TokenBundle -- a structured pointer to modality-specific features. The fusion layer projects all bundles into a single SharedEmbedding of exactly 8192 dimensions, regardless of which modalities are available for a given patient.

Text as the Universal Pivot

Almost every medical dataset pairs data with text. We align imaging to text, and omics to text independently. Because Image ~ Text and Omics ~ Text, Image ~ Omics transitively -- without ever needing paired image-omics patients.

Sovereign by Design

Patient data never leaves the hospital. FAU, TUM, and LMU each maintain a local data lake and train locally. Only privacy-preserving gradients and embeddings are exchanged -- never raw clinical records.

Missing Modality Robustness

A patient with only a chest X-ray and a blood panel produces a valid SharedEmbedding identical in format to one with the full data. The fusion layer is explicitly trained to handle any combination of present and absent modalities.

MCP as the Integration Layer

Each modality service exposes its encoder as an MCP tool. The fusion orchestrator discovers and calls these tools at runtime -- the same protocol that governs modern LLM agent systems.

The Roadmap

From raw data to deployed model

AI-BAY-HEALTH executes in four phases across 36 months, with 1.5 million GPU-hours at NHR@FAU (H200) and 3 million GPU-hours at JUPITER (GH200 via CRESCENDO), funded by the Bavarian State Ministry of Science and the Arts.

Phase 0  -- M1 to M3

Sovereign Foundations

Establish decentralised data lakes at FAU, TUM, LMU. Implement data readers for DICOM, NIfTI, h5ad, EDF, FHIR. Ingest public datasets. Set up the Globus data fabric.

Milestone M3: readers operational
Phase 1a -- M4 to M9

Domain Expert Encoders

Train specialised unimodal models: MedSigLIP for imaging, scGPT and Nicheformer for omics, Signal Transformer for ECG/EEG, TabPFN for EHR. Lock the 8192-dim interface.

Milestone M9: all encoders trained
Phase 1b -- M10 to M12

Fusion and Alignment

Project all encoders into the shared space via contrastive text-pivot alignment. Macro-micro alignment between radiology and pathology. Knowledge graph regularisation via UMLS/SNOMED.

Milestone M12: cross-modal retrieval live
Phase 2 -- M13 to M15

VLM and Clinical Pilot

Connect the 8192-dim embeddings to a large VLM backbone via soft-prompt injection. Deploy the AI-BAY-Box at UKER, TUM Klinikum, and LMU Klinikum. Submit the Nature paper draft.

Milestone M15: pilot at 3 sites
What we hope to discover

Emergent properties of unification

When all clinical modalities converge on a shared embedding space, capabilities arise that no single-modality model can achieve. These are scientific hypotheses -- not guarantees -- and testing them is at the core of why this project exists.

🔍

Cross-Modal Retrieval

Given an MRI scan, retrieve the most semantically similar pathology slides -- across institutions, without exchanging raw images. The embedding space dissolves modality boundaries.

Benchmark: Recall@K across radiology-pathology pairs

Zero-Shot Rare Disease Diagnosis

A disease with only three known cases in the training data. Because the VLM has seen its textual description, it can identify the imaging pattern in a new patient -- never having seen that exact combination before.

Enabled by: text-pivot alignment and VLM reasoning
🤔

Missing Modality Imputation

When a patient cannot undergo MRI, the model predicts what the MRI embedding would have looked like from their available CT and lab data. A quantifiable bound on the prediction uncertainty is included.

Technique: masked autoencoder training across modalities
📊

Molecular Prediction from Images

Inferring a tumour's EGFR mutation status or immune infiltration directly from a radiological scan -- without genomic sequencing. A consequence of aligning imaging and omics in a shared space.

Requires: paired imaging-omics cohorts (BRIDGE / FORTE)
🏛

Federated Multi-Site Generalisation

Models trained at UKER, TUM Klinikum, and LMU Klinikum on disjoint patient populations generalise across all three sites, via shared embedding alignment -- without ever centralising data.

Privacy-preserving by construction; GDPR-compliant
🌟

Conversational Clinical Intelligence

A clinician asks: "Summarise the oncological risk for this patient given their last CT, three ECGs, and recent labs." The VLM backbone reasons across all modalities simultaneously, in natural language.

Target: AI-BAY-Box deployment at Milestone M15
Proudly supported by

Funded and powered by Bavaria

AI-BAY-HEALTH exists because of the visionary commitment of the Bavarian state government and the world-class computing infrastructure it has built at Bavarian universities. We are deeply grateful for this support.

Bayerisches Staatsministerium fuer Wissenschaft und Kunst

Funded by the Free State of Bavaria

AI-BAY-HEALTH is funded by the Bayerisches Staatsministerium fuer Wissenschaft und Kunst (StMWK) under the Hightech-Agenda Bayern -- the Free State's flagship programme for science and technology investment. The Ministry oversees higher education, research, and innovation across all Bavarian universities, and has made sovereign AI one of its strategic priorities, including the launch of a Europe-wide unique Bavarian AI base model initiative. This support unites eleven universities and research centres to build shared foundation model infrastructure for medicine that will serve patients and clinicians across Bavaria and beyond.

www.stmwk.bayern.de ↗ Press release: Bayern startet KI-Basismodell ↗
NHR@FAU -- National High Performance Computing at FAU Erlangen-Nurnberg

Computing infrastructure: NHR@FAU

All large-scale training runs for AI-BAY-HEALTH are executed on the Helma H200 GPU cluster at NHR@FAU (Friedrich-Alexander-Universitat Erlangen-Nurnberg), part of Germany's National High Performance Computing network. The project is allocated 1.5 million GPU-hours on H200 hardware plus an additional 3 million GPU-hours on the JUPITER GH200 system via CRESCENDO -- the compute backbone that makes training models at the scale of medicine possible.

1.5M GPUh  -- NHR@FAU Helma H200 3M GPUh  -- JUPITER GH200
hpc.fau.de ↗
baiosphere -- Bavarian AI initiative baiosphere

Coordinating framework: baiosphere

AI-BAY-HEALTH operates within Bavaria's overarching baiosphere initiative, jointly steered by three Bavarian state ministries (Science and the Arts, Digital Affairs, and Economic Affairs). baiosphere is not a funder; it is the organising framework that connects the sovereign Bavarian foundation model efforts across health, robotics and industry, links research with infrastructure and the regional innovation ecosystem, and acts as the public face of the initiative. The medical foundation model announced in April 2026 is one of baiosphere's flagship efforts.

baiosphere.org ↗ Press release: Bavarian AI foundation model ↗
Consortium -- 11 Bavarian Universities & Research Centres