Planning a data platform, analytics system, or AI solution? Our team can help design scalable architectures and deliver production-ready solutions tailored to your business.
Client context
A university clinical hospital, one of the largest in its region, with 32 clinical departments, operating under strict data protection and regulatory requirements. Patient and operational data are highly sensitive: nothing can be processed outside the client's own infrastructure, and no production process can depend on external cloud services.
At the same time, the hospital generates years of structured historical data: medical records, medication usage, stock levels, bed occupancy, that remained largely unused for decision support.
This is a pattern we see across regulated organizations, from hospitals to pharmaceutical and industrial environments. Valuable operational data exists, but compliance constraints rule out standard cloud-based AI, and internal teams lack the infrastructure to run AI in production on their own terms.
The challenge
The organization needed production-grade AI, not a pilot, inside its own walls.
The requirements were unambiguous:
- all data processing and model training had to run entirely on-premise, in the client's server room
- clinical and operational staff needed AI support in daily workflows: preparing structured medical documentation, selecting materials for medication administration, and ordering medications and supplies
- the solution had to integrate with the hospital's mission-critical information system (HIS)
- production operation could not depend on the availability of an LLM, external services, or any generative process running in real time
- after handover, the client's own IT team had to be able to maintain and extend the system independently, with no vendor lock-in
Most AI initiatives in regulated environments fail at exactly these points. They remain proof-of-concepts, they quietly route data through external APIs, or they create permanent dependency on the vendor.
What it took to deliver results
We delivered a complete on-premise AI platform covering three functional modules on shared infrastructure.
1. AI-assisted medical documentation. The platform analyzed the hospital's historical medical records and generated structured documentation templates, covering care plans, medical orders, referrals, discharge recommendations, and final coding (ICD-10 / ICD-9), for the most frequent diagnoses and procedures in each of the 32 clinical departments. The contracted scope covered a minimum of 256 department-specific documentation schemas.
2. AI-based recommendations for medication administration. Based on historical usage data, the system recommends the disposable materials and equipment needed for administering specific medications, adjusted to local practices of each department. Recommendation quality is verified against a contractual threshold of precision@k ≥ 0.70, validated against expert-reviewed test sets.
3. Demand forecasting for medications and supplies. Dedicated predictive models forecast daily demand across 3-day and 7-day horizons, accounting for annual seasonality, weekly variability, and holiday effects, generating order proposals from wards to the hospital pharmacy and from the pharmacy to external suppliers.
The architecture principle that made it production-grade: the LLM is used exclusively in the content-preparation layer. All generated outputs are validated, versioned, and published to a local reference database, and it is this database, not the model, that production systems query. Live operation has zero dependency on model availability.
The platform is built on:
- on-premise GPU server infrastructure for local LLM inference
- a local LLM, so no data leaves the client's infrastructure at any stage
- MLOps tooling for training, validation, versioning, monitoring, and rollback
- a governed content lifecycle: every AI-generated artifact moves through draft, validation, approved, published, and retired states, with full metadata and audit trail
- JSON-based data pipelines with encrypted transfer and a dedicated intermediate database
- HIS integration via API and database interfaces
The solution
The platform turns years of dormant historical data into a structured foundation for daily decision support, while keeping every byte of sensitive data inside the hospital's own infrastructure. Physicians gain access to pre-structured documentation templates matched to the diagnostic and treatment process. Nurses gain material recommendations consistent with local department practice. Pharmacy and logistics gain daily, seasonality-aware order proposals in place of manual estimates.
Because the production path is decoupled from the generative layer, the system was built to deliver this support with the stability and predictability required of mission-critical hospital infrastructure, ready for the client's team to bring into daily clinical and logistics workflows.
How it works
To meet these constraints, the solution required capabilities that rarely exist in one team:
- designing and delivering the full on-premise infrastructure: an AI server with GPU acceleration, deployed and configured in the client's server room
- standing up a complete local environment for LLM inference, with MLOps tooling for the full model lifecycle
- building secure data pipelines from the client's source systems: encrypted transfer, JSON-based exchange, dedicated intermediate database
- integrating with a mission-critical hospital system (HIS) without disrupting its operation
- engineering an architecture that keeps generative AI out of the production-critical path
- structuring the entire delivery around knowledge transfer, so the client's engineers can operate and extend the platform independently
Impact on operations
Historical and current data from the hospital's systems flow through secure pipelines into the local AI environment. Models are trained on this data, entirely within the client's infrastructure.
Generated outputs (documentation schemas, material recommendations, demand forecasts) are validated against quality thresholds, versioned, and published to a local reference database. The hospital's HIS reads from this database in production, a deliberate separation of the generative layer from the operational layer.
Every schema and recommendation carries full metadata: identifier, scope of application, linked diagnoses and procedures, version, status, and publication dates. The client's team can modify dictionaries, classification rules, and mappings, and generate new content for additional diagnoses, medications, and products, without involving the vendor.
Forecasting models run in a daily cycle, delivering order proposals before a fixed cut-off time each morning.
Business impact
The delivery demonstrates a capability that most organizations in regulated environments are still searching for:
- Full data sovereignty. AI training and inference with zero data leaving the client's infrastructure.
- Production deployment, not a pilot. Contractual quality gates (precision@k ≥ 0.70), end-to-end testing, and formal acceptance testing before handover.
- No runtime dependency on generative AI. Offline-first architecture built for environments where downtime is not an option.
- Governed AI content lifecycle. Versioning, validation states, and audit trail designed for regulated settings.
- No vendor lock-in. Complete documentation, source code, and knowledge transfer enabling the client's team to independently extend the system with new diagnoses, medications, and forecasting models.
- Compliance-first model selection. Both the LLM and the supporting models were selected against EU AI Act alignment and provenance from trusted vendors. The platform runs on Gemma, Google's open model family.
- Measured results at handover:
- Documentation module: 320 schemas generated across all 32 departments for the top 10 most frequent diagnoses per department, each with the full 14-section structure, PII protection built into the generation pipeline, and version control on every generation cycle.
- Recommendation module: precision@k of 0.80 achieved against a contractual requirement of 0.70 (0.85 on the validation sample, 0.94 micro-averaged), covering 100 medications across 32 departments and 1,456 recommendation packages. Across the full quality distribution, the median score reached 0.875, with 469 of 693 samples at or above the 0.70 threshold.
- Forecasting module: MASE below 1 at every level of the hierarchy, outperforming a naive seasonal baseline, with hospital-wide WAPE in the 15 to 17 percent range.
For organizations in pharma, life sciences, and other regulated industries facing the same constraint (our data cannot leave the building), this deployment is proof that production-grade, locally hosted AI is achievable today.
We’ll review your goals, technical constraints, and opportunities to design a solution that fits your organization.





