Newsroom

Why we invested in Sci2sci

Published
September 8, 2026
|
Updated
September 8, 2026

Pharma R&D is the industry that would benefit most from intelligence over its own data, yet it is the most underserved in getting it. sci2sci is building the layer that changes that. Biopharma first, high-reg industries beyond it next.

That’s why we invested.

The problem

Data in highly regulated industries is not accessible for AI agents, because it's scattered across systems and sits behind compliance restrictions. Wherever AI is used, decisions need to be traceable, because regulators are demanding a clear link from every decision back to its source and generic AI tools simply aren’t built for that. Additionally, organizational knowledge and the context between projects and documents stays with the people that created it. Resulting in organizations spending time and money inefficiently.

The market

Biopharma data is growing 30 to 35% annually, and regulation requires much of it to be retained for up to 25 years. Two large software categories have formed around that: compliance-native life-sciences software, which produced Veeva and IQVIA, and horizontal data infrastructure and governance, which produced Databricks, Collibra and Glean. Both categories are now extending into AI governance, but neither solves the problem of making scattered R&D data usable by AI agents while keeping every output traceable, auditable and defensible to a regulator.

The solution

sci2sci is building the data intelligence and governance layer for biopharma and beyond.

Their VectorCat product suite connects the scattered data sources into one secure system without any migration, enriches every asset with metadata, structures it into one harmonized format and flags sensitive information, which lays the compliance basis for agents to access datasets.

Integrity Cortex builds on that and turns the governed data into a neurosymbolic memory layer made of executable code that deterministically verifies every LLM output against its source, catches hallucinations and errors in the underlying data and makes the data estate “FAIR": findable, accessible, interoperable, and reusable. Regulators require exactly this auditability, while generic AI tools can simply not deliver it.

Why we're convinced

Our conviction rests on the combination of team, product maturity, early validation and the commercial dynamics of the model.

The team:: Angelina Lesnikova and Valerii Kremnev come at this problem from both of its sides. Angelina spent her neuroscience PhD on the mechanisms of memory and years inside pharma, so she knows the science behind the architecture and the industry they sell into. Valerii built classification systems that had to stay correct at continental scale, which is the machinery and logic deterministic verification requires.

sci2sci's product is fully featured and already running in production today with a concrete feature pipeline behind it, which is rare at pre-Seed. They track their portfolio publicly, spanning seven commercial products, six open-source tools and five research papers (sci2sci.com/portfolio). sci2sci is already deployed inside a listed global biopharma champion and runs multiple pilots with relevant customers.

The economics are fit for a data-heavy future: sci2sci prices along data volume, the one asset class that grows in effectively every company and every sector, resulting in ACVs growing with the customer automatically. With increasing usage and data coverage, their products sit in a sticky and hard to replace position in the enterprise AI stack. Customers bring their own cloud and inference while sci2sci deliver the data governance layer underneath.

That’s why we invested.