How CORTEX is being used across real scientific environments.
Biotech archive
A large, heterogeneous archive transformed into a structured, queryable scientific record.
~73,000 searchable files · 8 relevant records · 9,125× reduction
“It’s like night and day.”
The problem
The archive held approximately 1.8 million files, about 5.6 TB, accumulated over years across different systems and teams. Office documents sat alongside SAS and XPT datasets, CSV exports, GraphPad Prism files, and other formats, 128 file extensions in total. Nothing tied a file to the study, compound, or question it belonged to.
Finding a specific result meant knowing which person, which system, which folder, and which version to look in. Keyword search returned tens of thousands of candidates and gave no way to tell canonical records from copies. Duplicate detection later showed that 660,647 files, 48.5% of the archive, were duplicates of other files.
The approach
CORTEX established structured representations of the archive through a sequence of operations: file-level metadata extraction, domain-specific organization, duplicate detection, classification, metadata enrichment, entity resolution, metadata-aware filtering, and a review and correction workflow for cases the system could not settle on its own.
Entity resolution produced 3,375 canonical entities with 3,750 aliases, so a compound, study, or assay referred to by different names in different files resolves to one record. Duplicate detection removed 660,647 redundant files from the search space. The result is an archive where retrieval is filtered by what a file is about rather than where it happens to sit.
Scientific example
Across the structured archive, roughly 73,000 files were findable for this query by content alone. After classification, deduplication, and metadata-aware filtering, CORTEX returned 8 relevant records: 9,125 times fewer documents to inspect, a 99.989% reduction in the search space, for a question a scientist actually asked.
The result
9,125× fewer documents to inspect and a 99.989% reduction in the search space for a real scientific query. Across the archive, 660,647 duplicate files (48.5%) were identified and 3,375 canonical entities were resolved from 3,750 aliases. Retrieval now operates on structured, deduplicated, entity-resolved records rather than on raw files.
From the customer
Senior research lead, biotech · a qualitative statement, not a measured benchmark
Why it matters
Retrieval over the archive now operates on a structured record rather than on files. The same structure that narrows a search from 73,000 candidates to 8 is what lets an AI system work from the laboratory’s actual evidence instead of its folder tree.
Research institute
Scientific data distributed across millions of files and multiple systems reconstructed into a connected research map.
~2,000 M. tuberculosis strains · 53,000 compounds screened · 94% name similarity resolved
The problem
The program’s record lived in a roughly 50 TB directory containing millions of files: screening data, genomics, analysis outputs, and instrument data accumulated across projects and people. Identifiers for strains, compounds, and genes were inconsistent between files, and the relationships between a screen, the compound tested, and the strain used were not recorded anywhere a machine could read.
In the scientist’s own description, the hard part was not storage but querying, and legacy material resisted being moved into any new system.
The approach
CORTEX treated the directory as a mapping problem rather than a search problem. It extracted the entities present in the files, including strains, compounds, genes, assays, and results, resolved the identifiers used for each across files, and recorded the relationships between them as a connected scientific representation.
Entity resolution was the central operation. Names that look alike but denote different things were kept apart, and names that differ but denote the same thing were merged, with the evidence for each decision retained for review.
Scientific example
BRD-8000 and BRD-8000.1 have a name similarity of 0.94, yet they are distinct compounds. CORTEX kept them separate rather than collapsing them on string similarity. Conversely, EfpA and Rv2846c are two identifiers for the same M. tuberculosis gene, a common name and a locus tag. CORTEX resolved them to a single entity so that results recorded under either name connect to the same node.
The result
Roughly 2,000 strains and 53,000 screened compounds, with their relationships, are represented as one connected map that can be queried by entity rather than by file path. Near-identical names remain distinct; aliases resolve to one record. Searching now operates on the map, not the directory.
From the customer
Senior research scientist, before the deployment
Why it matters
The institute’s screening and genomics record is now a connected scientific representation in which identities are resolved and relationships are explicit. That is the state an AI system needs in order to reason across a program rather than across files.
Preclinical program
Three historical fibrosis studies reconstructed into a navigable research record with source-level provenance.
183 animal records · 5,418 observations · 35 endpoints · 1,897 relationships
The problem
The program had run several versions of the same studies over time. Results existed, but it was unclear which data set was reproducible, which was believable, and where any given number had come from. The record was spread across Excel workbooks, GraphPad Prism files, PowerPoint decks, Word documents, PDFs, and native XPT datasets, each holding part of the picture in its own structure and naming.
Three fibrosis studies, IPF-133, IPF-149, and IPF-178, were documented in 40 source files. Reconstructing what happened in a study meant reading across all of them by hand.
The approach
CORTEX reconstructed the three studies into a common scientific structure: Study → Treatment → Animal → Endpoint → Observation → Source. Each observation is linked to the animal it was measured on, the treatment that animal received, the endpoint it belongs to, and the file and location it was taken from.
Discrepancies between sources were preserved for review rather than silently normalized away. Conflicting doses, inconsistent dates, missing measurements, formula errors, uncertain units, and incomplete assay documentation were recorded as 48 study-level review issues attached to the affected records.
Scientific example
The two records appeared in different source files under different identifiers. String matching treats them as two animals. CORTEX identified them as the same animal through their shared cage designation and ear identifier, experimental context the files carried but no identifier captured, and merged their observations under one record while keeping both source references.
The result
183 animal records, 5,418 observations, and 35 endpoints from three studies now sit in one record with 294 entities, 1,897 explicit relationships, and source-level provenance. A reader can move from the program to a study, to a result, to the evidence behind it, and 48 review issues mark exactly where the sources disagree.
From the customer
Research and operations lead, before the deployment
Why it matters
The three studies now form one traceable record in which every result links to its source and every disagreement between sources is visible. Provenance of this kind is what makes a research program legible to scientists and to AI alike.