Dozens of formats
Data arrives as CSV, Excel, XML, EDC exports, CDISC, FHIR, or a database dump — and every new source means a new, brittle import.
AI clinical-data platform
MedCupola is a light core with plug-in modules: import any format, ask questions in plain English, and keep every installation validated — on your own servers.
The problem
Data arrives as CSV, Excel, XML, EDC exports, CDISC, FHIR, or a database dump — and every new source means a new, brittle import.
Most platforms are all-or-nothing. You buy the whole thing and live with the parts you don't need.
“AI” is often a chat box on the side, disconnected from the data, with no guardrails and no evidence.
Proving a system is correct becomes a painful, one-time project instead of something that simply stays true.
The idea
There is a central core that provides the essentials — power, life support, a docking ring. Over the years, modules are added or removed: a laboratory here, a solar array there. Each attaches to a standard interface, and the station keeps working throughout.
A small, domain-agnostic kernel: accounts, storage, a central record of every action, a model gateway, and the ability to host add-ons.
Each capability is a pip-installable package that plugs in, activates, and can be swapped — independently of everything else.
Migrate from anything
Point MedCupola at almost any source. The format is detected, the mapping is proposed, and the data is loaded, checked, and kept validated from then on.
CSV, Excel, XML, JSON, EDC exports, CDISC, FHIR, or a database dump. The format is detected automatically.
The AI proposes how each source field maps to the target schema. A person reviews and approves.
The data is loaded, checked, and reported — and every subsequent change is revalidated automatically.
AI-first
The agent suite ships in the core, not as a later add-on. Agents plan, use tools over a standard interface, and answer with evidence — under guardrails that keep them read-only unless a person approves.
The SQL agent turns a question into a safe, read-only query, runs it, and shows its work.
Document search grounds every response in the exact source, so nothing is invented.
Data-quality and mapping agents flag missing, invalid, duplicate, or inconsistent records with evidence.
Modules
Run and pay for only what you need, then plug in more. New formats are new adapters; new clinical needs are new modules.
Synthetic study schema, seed data, and vector search (pgvector).
IQ/OQ/PQ, perpetual revalidation, traceability, and gating.
Runtime, SQL, RAG, assistant, data-quality, mapping, guardrails, eval.
Blinded endpoint and event review: intake, dossier assembly, committee voting, consensus, and export.
EDC XML → validate → map → load: the first proof of pluggability.
Lab import with reference ranges; deterministic checks; human-approved schema mapping.
The catalog runs much deeper. As the platform grows, these plug in exactly the same way — a new format or a new clinical need is just a new module, with no change to the core.
Trust
A dedicated module confirms each installation is correct and re-checks the entire system whenever anything changes.
Inference runs locally. A small model is always on; extra GPU power is switched on only when needed.
The pilot
We are running a focused pilot on synthetic data — deploying the platform, migrating real-shaped datasets through the 1-2-3 pipeline, and proving the AI agents end to end. We are looking for a partner to shape it.
Request a pilot