CASE STUDIES

CORTEX IN PRACTICE

How CORTEX is being used across real scientific environments.

Biotech archive

1.8M files becomes 8 relevant records.

A large, heterogeneous archive transformed into a structured, queryable scientific record.

~73,000 searchable files · 8 relevant records · 9,125× reduction

“It’s like night and day.”

The problem

The archive held approximately 1.8 million files, about 5.6 TB, accumulated over years across different systems and teams. Office documents sat alongside SAS and XPT datasets, CSV exports, GraphPad Prism files, and other formats, 128 file extensions in total. Nothing tied a file to the study, compound, or question it belonged to.

Finding a specific result meant knowing which person, which system, which folder, and which version to look in. Keyword search returned tens of thousands of candidates and gave no way to tell canonical records from copies. Duplicate detection later showed that 660,647 files, 48.5% of the archive, were duplicates of other files.

The approach

CORTEX established structured representations of the archive through a sequence of operations: file-level metadata extraction, domain-specific organization, duplicate detection, classification, metadata enrichment, entity resolution, metadata-aware filtering, and a review and correction workflow for cases the system could not settle on its own.

Entity resolution produced 3,375 canonical entities with 3,750 aliases, so a compound, study, or assay referred to by different names in different files resolves to one record. Duplicate detection removed 660,647 redundant files from the search space. The result is an archive where retrieval is filtered by what a file is about rather than where it happens to sit.

Scientific example

“AMT-101 quality stability”

Across the structured archive, roughly 73,000 files were findable for this query by content alone. After classification, deduplication, and metadata-aware filtering, CORTEX returned 8 relevant records: 9,125 times fewer documents to inspect, a 99.989% reduction in the search space, for a question a scientist actually asked.

The result

~73,000 findable files reduced to 8 relevant records.

9,125× fewer documents to inspect and a 99.989% reduction in the search space for a real scientific query. Across the archive, 660,647 duplicate files (48.5%) were identified and 3,375 canonical entities were resolved from 3,750 aliases. Retrieval now operates on structured, deduplicated, entity-resolved records rather than on raw files.

From the customer

“This is 50× better than Box AI. It’s like night and day.”

Senior research lead, biotech · a qualitative statement, not a measured benchmark

Why it matters

Retrieval over the archive now operates on a structured record rather than on files. The same structure that narrows a search from 73,000 candidates to 8 is what lets an AI system work from the laboratory’s actual evidence instead of its folder tree.

Research institute

50 TB of science becomes one connected map.

Scientific data distributed across millions of files and multiple systems reconstructed into a connected research map.

~2,000 M. tuberculosis strains · 53,000 compounds screened · 94% name similarity resolved

The problem

The program’s record lived in a roughly 50 TB directory containing millions of files: screening data, genomics, analysis outputs, and instrument data accumulated across projects and people. Identifiers for strains, compounds, and genes were inconsistent between files, and the relationships between a screen, the compound tested, and the strain used were not recorded anywhere a machine could read.

In the scientist’s own description, the hard part was not storage but querying, and legacy material resisted being moved into any new system.

The approach

CORTEX treated the directory as a mapping problem rather than a search problem. It extracted the entities present in the files, including strains, compounds, genes, assays, and results, resolved the identifiers used for each across files, and recorded the relationships between them as a connected scientific representation.

Entity resolution was the central operation. Names that look alike but denote different things were kept apart, and names that differ but denote the same thing were merged, with the evidence for each decision retained for review.

Scientific example

BRD-8000 is not BRD-8000.1. EfpA is Rv2846c.

BRD-8000 and BRD-8000.1 have a name similarity of 0.94, yet they are distinct compounds. CORTEX kept them separate rather than collapsing them on string similarity. Conversely, EfpA and Rv2846c are two identifiers for the same M. tuberculosis gene, a common name and a locus tag. CORTEX resolved them to a single entity so that results recorded under either name connect to the same node.

The result

Scientific identities resolved across millions of files.

Roughly 2,000 strains and 53,000 screened compounds, with their relationships, are represented as one connected map that can be queried by entity rather than by file path. Near-identical names remain distinct; aliases resolve to one record. Searching now operates on the map, not the directory.

From the customer

“I can give you a 50 terabyte directory right now that is a dumpster fire of millions of files.”

Senior research scientist, before the deployment

Why it matters

The institute’s screening and genomics record is now a connected scientific representation in which identities are resolved and relationships are explicit. That is the state an AI system needs in order to reason across a program rather than across files.

Preclinical program

3 studies becomes one connected research record.

Three historical fibrosis studies reconstructed into a navigable research record with source-level provenance.

183 animal records · 5,418 observations · 35 endpoints · 1,897 relationships

The problem

The program had run several versions of the same studies over time. Results existed, but it was unclear which data set was reproducible, which was believable, and where any given number had come from. The record was spread across Excel workbooks, GraphPad Prism files, PowerPoint decks, Word documents, PDFs, and native XPT datasets, each holding part of the picture in its own structure and naming.

Three fibrosis studies, IPF-133, IPF-149, and IPF-178, were documented in 40 source files. Reconstructing what happened in a study meant reading across all of them by hand.

The approach

CORTEX reconstructed the three studies into a common scientific structure: Study → Treatment → Animal → Endpoint → Observation → Source. Each observation is linked to the animal it was measured on, the treatment that animal received, the endpoint it belongs to, and the file and location it was taken from.

Discrepancies between sources were preserved for review rather than silently normalized away. Conflicting doses, inconsistent dates, missing measurements, formula errors, uncertain units, and incomplete assay documentation were recorded as 48 study-level review issues attached to the affected records.

Scientific example

Records 3-1 and 3-5 are the same animal.

The two records appeared in different source files under different identifiers. String matching treats them as two animals. CORTEX identified them as the same animal through their shared cage designation and ear identifier, experimental context the files carried but no identifier captured, and merged their observations under one record while keeping both source references.

The result

5,418 observations connected across 183 animals.

183 animal records, 5,418 observations, and 35 endpoints from three studies now sit in one record with 294 entities, 1,897 explicit relationships, and source-level provenance. A reader can move from the program to a study, to a result, to the evidence behind it, and 48 review issues mark exactly where the sources disagree.

From the customer

“The team has run six different data sets. Which one is reproducible? Which one’s believable? Where did that data come from?”

Research and operations lead, before the deployment

Why it matters

The three studies now form one traceable record in which every result links to its source and every disagreement between sources is visible. Provenance of this kind is what makes a research program legible to scientists and to AI alike.