Savana: Turning Clinical Text into Regulatory-Grade Evidence
Deep Dive | Edition 23
Welcome back to the deep dive, where we break down the AI tools and data reshaping how new drugs are discovered. In each edition, we speak directly with the teams behind these tools to explain what they solve, how they work and where they are going next.
The most valuable dataset in healthcare is the one no one can query.
For: Real-world evidence teams, pharma clinical operations, health data scientists, and anyone building regulatory submissions from observational data.
Savana transforms unstructured clinical text into validated, structured databases using clinical NLP, operating across 300+ sites in 14 countries.
Regulators will not accept outputs from large language models. Savana’s discriminative AI provides the reproducibility that submissions demand.
We spoke with founder Ignacio Medrano about why free-text clinical data remains the most underleveraged layer in healthcare.
Kiin Pioneer Programme
We built a platform that helps researchers speed up their entire science, from literature review and biomarker discovery to bioinformatics and computational chemistry. If your workflow involves pulling findings from five different places before you can actually act on any of them, this is for that.
The Pioneer Programme gives academic labs and non-profits one year of free access, plus support from our science team. No cost, no data transfer, all IP stays with your institution. Applications close August, cohort starts September.
This week we spoke with Ignacio Medrano, neurologist-turned-CEO and founder of Savana.
The problem: hospitals produce millions of words every day that never become data
Every hospital generates enormous volumes of clinical text. Radiology reports, discharge summaries, histopathology findings, physician notes and inside those documents sits rich information about what happens to patients and what doctors are thinking. Yet none of it is structured, queryable, or feeding into models.
Real-world evidence has always mattered for regulatory submissions. The bottleneck was never demand, it was collection: patient by patient, variable by variable, building registries manually for years. A single retrospective study might take 18 months of chart review before producing its first result.
“There’s no reason for humans to collect data manually anymore,” Ignacio Medrano told us. “We created a new way of doing clinical research where humans don’t have to collect data. They can dedicate their time to more interesting things, like thinking or analysing.”
The timing matters here. The European Health Data Space regulation is pushing hospitals toward structured data sharing. Simultaneously, AI scribes mean even more unstructured text will be generated in the coming years. The raw material is growing faster than anyone’s ability to manually process it.
The approach: validated NLP across 300 sites and six languages
Savana operates through its Smart Health Alliance, a network of over 300 sites across 14 countries. The platform has three layers: a hospital-facing suite for structuring local data, an interoperability layer for cross-border sharing and governance, and Next Generation Registries where pharma companies access continuously updated, AI-generated databases across multiple sites and countries.

The registries are the sharpest departure from traditional study design. Unlike static observational studies, they stream. Researchers can add new variables a year into a study and the system retrieves them retrospectively across five years of raw data, then continues capturing them prospectively. You no longer need to define every variable at protocol stage.
Amazingly, the entire system operates multilingually: Spanish, English, German, French, Portuguese, and Italian.
Why it’s different: validation that regulators will actually accept
Many NLP tools can extract information from text. The question is whether a regulator will trust the output.
“Most linguistic approaches don’t offer scientific robustness,” Ignacio Medrano explained. “We have a validation methodology that guarantees that when we transform a fragment of text into a variable, it is accurate and reliable. If we do it several times, we get the same result.”
This distinction matters more now than it did two years ago. The FDA and EMA do not accept outputs from large language models as evidence. LLMs are stochastic: the same input can produce different outputs. Savana’s discriminative AI approach provides reproducibility and auditability. So principal investigators at participating sites also review extractions through a guided interface, keeping quality anchored to clinical expertise.
For pharma teams evaluating real-world evidence platforms, this is the core differentiator. Plenty of vendors can structure text. The question is whether that output can appear in a regulatory submission without manual validation layered on top.
Real-world impact
Three weeks after COVID patients began being treated, Savana built a complete database across a European region of 2 million people, with all variables analysed and a paper submitted. Traditional registries would take 12-18 months to reach that point. The analysis identified that patients with mild symptoms would become severe or die three weeks later, enabling primary care to intervene early.
In oncology, working with Pfizer, Savana built a thrombosis prediction model for anticoagulation decisions in patients with solid tumours. It was validated against the Spanish oncology society’s database and is now referenced in their clinical guidelines. That progression, from unstructured notes to clinical guideline, is the full value chain.
Published collaborations span BMS, Johnson & Johnson, and Gilead across use cases from market access to external control arms.
The future
The more interesting direction is multimodal integration. Clinical text captures what happens to patients, but genomics, proteomics, and radiomics each add layers that text alone cannot provide.
“Doctors are inherently multimodal,” Ignacio Medrano said. “When you get into the practice of a doctor, they work with your data across all these layers. AI needs to be multimodal too.”
As AI scribes become ubiquitous, Savana’s role grows rather than shrinks. Scribes produce unstructured output. Transforming that into reliable databases still requires validated, discriminative AI. The generative layer creates the text. The validated extraction layer makes it usable as evidence. Different problems, different architectures.
Kiin’s view
Savana occupies a pretty defensible position. The regulatory gap between “we extracted this with GPT-4” and “we extracted this with a validated, reproducible pipeline” is not closing any time soon. The moat is not the NLP itself, it’s the validation methodology, the 300-site network, and the multilingual coverage that took a decade to build. Competitors can build extractors but replicating the clinical validation infrastructure across 14 countries is a different proposition entirely.
Thanks for reading Kiin Bio Weekly!
💬 Get involved
We’re always looking to grow our community. If you’d like to get involved, contribute ideas or share something you’re building, fill out this form or reach out to me directly.
Subscribe now to stay at the forefront of AI in Life Science and keep up with this upcoming season of deep dives.
Connect With Us
Have questions on this or suggestions for our next deep dive? We’d love to hear from you!
📧 Email Us | 📲 Follow on LinkedIn | 🌐 Visit Our Website





