01
Convert
Historical reports must first become structured data, one row per patient, with each value linked to the passage it came from.
Case study
01.
Clinical studies often depend on outcomes that take years to appear. A prospective study can define a cohort in advance, collect the same variables under one protocol, and follow patients until recurrence and survival mature. Its strength is consistency; its cost is time.
Dr. Blake Gilks, a pathologist at Vancouver General Hospital, describes the cost of the prospective route:
“The trouble with doing things prospectively is it takes two years to collect enough cases and then another four years to get follow-up information: so about six years before you have any results.”
| Trial | Patients | Research question and time horizon |
|---|---|---|
| GOG-0258 | 707 treated | Chemotherapy versus chemoradiation in advanced and high-risk disease. A later molecular analysis drew on a median follow-up of 113 months, more than nine years. |
| PORTEC-3 | 660 eligible | Pelvic radiotherapy versus combined chemoradiotherapy in high-risk disease. Long-term analyses followed the cohort for about a decade. |
| GOG-249 | 601 | Pelvic radiotherapy versus vaginal brachytherapy with chemotherapy in high-intermediate and high-risk early-stage disease. Primary analysis at a median 53 months. |
Researchers can also begin with patients whose outcomes are already known. They return to pathology archives, stored tissue, molecular records, and follow-up notes to reconstruct a cohort around a new question. This route can remove years of waiting; the work moves into reconstruction.
| Study | Patients | What the researchers reconstructed |
|---|---|---|
| Thompson et al., 2022 | 1,357 | Samples, clinicopathologic data, and outcomes from 10 tertiary centres and 19 community hospitals, each tumour then retrospectively assigned a ProMisE molecular subtype. |
| Kommoss et al., 2018 | 452 | Archived diagnostic material combined with clinical and outcome data to validate the ProMisE classifier in a population-based cohort. |
| Jamieson et al., 2025 | 154 | Pathology archives and previously classified cohorts from 2000-2023, used to study co-existent endometrial and ovarian carcinomas and possible treatment de-escalation. |
Gilks describes what that reconstruction takes:
“We never historically had the data that is now becoming the standard. For retrospective studies, we go back and add it case by case, which is expensive, but it’s the only way to make the records comparable using modern classifications.”
Records drafted for treatment differ from the datasets research needs.
02.
Patient records drafted for treatment differ from the controlled datasets created for research. Reports use different labels, tests were ordered inconsistently, and the same variable may be spread across pathology reports, operative notes, molecular results, treatment histories, and follow-up records. Three papers show distinct forms of this problem.
| Evidence | Reconstruction issue | Research consequence |
|---|---|---|
| Thompson et al., 2022 | Only 42% of patients had mismatch-repair testing at diagnosis (centre coverage 3.5% to 95.4%); only 21.1% had p53 immunohistochemistry. | A missing value may reflect local practice, historical timing, unavailable material, or a clinical decision. Each explanation means something different for the study. |
| Clements et al., 2025 | GOG-0258 treated 707 patients; archived tumour material was available for 426, and 416 could be classified using both MMR and p53 staining. | A later question may depend on specimens and tests outside the original protocol, even when the parent cohort was carefully governed. |
| Jamieson et al., 2025 | Among 43 patients with molecular results from both the endometrial and ovarian tumours, four had discordant classifications between sites. | A patient-level value may need specimen-specific interpretation and a documented rule for resolving disagreement. |
The difficulty extends beyond extracting words from reports. Researchers must establish what was measured, which specimen produced the result, which standard governed the interpretation, and how conflicting evidence should enter the study dataset. Even basic abstraction is usually manual.
This is the central bottleneck of retrospective research. The evidence may already exist, but turning it into a consistent cohort takes repeated review, domain knowledge, and quality control.
03.
In 2023, the International Federation of Gynecology and Obstetrics rewrote how endometrial cancer is staged. Apply the new rules to a cohort staged under the old ones, out of a study of over 500, and 26.8% of patients are classified in a fully different cancer stage under the new schema (Lim et al., 2025). So for many fields with shifting classifications, archival data shifts under researchers' feet.
“The two classification systems are not comparable, and it is not possible to convert the earlier classification to the current one based on old inspection reports. So it effectively hits the reset button for researchers.”
A case recorded as stage IA under FIGO 2009 may need histologic review, lymphovascular space invasion assessment, p53 immunohistochemistry, mismatch-repair testing, or POLE sequencing before a current stage can be assigned. The historical cohort still helps: archives often preserve tissue blocks that support later testing, so higher-resolution evidence can be generated years after diagnosis. The practical work can be intensive: a team may need to retrieve reports, locate the corresponding tissue, order new assays, identify which standard applies to each historical interpretation, join the new results to the correct patient and specimen, and preserve why the final study value differs from the original clinical label.
The reconstruction problem therefore has three layers
01
Historical reports must first become structured data, one row per patient, with each value linked to the passage it came from.
02
Those observations are reconciled with definitions that changed over time, such as FIGO 2009 and FIGO 2023, keeping both interpretations in view.
03
Information current practice requires, such as molecular subtype, may need regenerating from preserved tissue because the original report never captured it.
04.
Provinans is designed around this reconstruction layer. Within it, Claude Sonnet 5 reads each record against four connected sources of context:
This lets the model parse a record within the standard that originally governed it and compare that interpretation with the standard the current study requires. For a case staged under FIGO 2009, it can identify the historical staging language, retrieve the FIGO 2023 criteria, and show which newer requirements the available record supports, distinguishing a direct mapping from a new inference and flagging cases where a molecular result could change the stage.
The feasible contribution is concentrated in work that is repetitive, document-heavy, and rule-dependent. The model can:
A study with 200 patients and ten records each already holds about 2,000 records. Across that volume, locating evidence, pre-populating joins, and separating routine mappings from difficult cases removes a large amount of repetitive review. The larger value is prioritization: a clear histologic label moves through quickly, while a discordant MMR result, an obsolete stage label, or an incomplete molecular profile is surfaced for deliberate review.
05.
Every proposed value in Provinans is attached to a ledger entry. The ledger records:
This gives the group a shared source of truth for how the cohort was assembled, and makes the most nuanced parts of the data visible: a conflict becomes something the team can inspect, discuss, and resolve with the evidence and rule in view. The ledger preserves the historical interpretation beside the current research value, so when a standard changes again the team can find the cases that depend on the affected rule and revisit them without rebuilding the cohort from the beginning.
06.
This week, in consultation with Dr. Blake Gilks, the work focused on the first portion of this workflow: establishing a shared source of truth between labels, encoding how standards align or conflict, beginning to populate Sonnet with judgment-aware mappings, and refining how annotators review legacy reports.
The system is being developed to classify reports with awareness of their governing standards, compare those standards with the research group’s target schema, and surface the cases where interpretation matters most. The immediate output is a populated, reviewable patient record. The broader goal is a cohort whose joins, classifications, and expert judgments stay transparent throughout the study.
Sources referenced