Skip to content
Avicen
EVALUATION PROGRAMME · 2026

Twenty-five real cases, before and with Avicen

The complete write-up of the 2026 evaluation programme behind every number on this site: the corpus, the method, the measured results, the capabilities involved, and the limits of what we can claim.

The question we set out to answer

Medical experts do not lack judgement. They lack time, and the time goes to a specific place: rebuilding the case before the analysis can even begin. Reading hundreds of pages, sorting scans, reconstructing a chronology, confronting assessments that disagree, finding the one document that supports each finding.

We wanted a number, not an impression: how much of that preparation can be absorbed, and does absorbing it change what the expert actually sees in the case?

The corpus: 25 real cases, selected for difficulty

The programme ran on 25 real, anonymized cases from practicing Swiss medical experts: insurance and judicial assignments from their own files, not demonstration material. Cases were selected for the properties that make expert review expensive: volume, fragmentation, and disagreement between authors.

Cases reviewed25
Largest case1,100 pages
Most documents in a single case24
Most authors in a single case8
Formats presentScanned PDFs, Word files, spreadsheets, handwritten notes, old faxes, DICOM imaging, audio dictations
LanguagesFrench, German, Italian and English, including mixed-language files

Method: the same case, twice

Each case was reviewed twice by the same expert. First with their usual method: the case as it arrived, the tools they normally use. Then with Avicen: the case uploaded exactly as it arrived, with no sorting, renaming or preparation.

Three measures were recorded for both passes:

  • Preparation time: everything before the expert can start forming an opinion, from first opening the file to holding a usable chronology.
  • Material inconsistencies found: discrepancies capable of affecting an assessment, such as conflicting dosages, capacity ratings, dates or attributions.
  • Verification time: the time from doubting a statement to reading its source passage.

Cases were anonymized by the experts before upload. Processing ran entirely on dedicated infrastructure, and no data was used to train any model.

What changed, measured

90%

less preparation time

more material inconsistencies identified

10×

faster source verification

Measured on 25 real anonymized cases, before and with Avicen, by practicing medical experts (2026 evaluation programme). Figures rounded for publication.

Reading the numbers

The preparation figure is the least surprising to anyone who has opened a 1,100-page case: when the chronology, the findings and the links to evidence are already assembled, preparation shrinks from evenings to minutes. What remains is the part that cannot be absorbed: the expert reading of what matters.

The inconsistency figure is the one that matters professionally. The experts did not read badly the first time: a case of that size simply exceeds what any reader can hold in memory. Five times more material inconsistencies were identified once the case was fully cross-referenced, and several were the kind that change conclusions.

The verification figure changes behaviour rather than workload: when the source of any statement is one click away, checking stops being a cost. Experts verified more, not less.

Three findings, in detail

Two dosages for the same treatment

1,100 pages · 14 documents · 5 authors

One report documented 60 mg per day. A later summary listed 30 mg for the same period, with no explanation of the change. The dosage appeared across 14 documents with no consolidated treatment history, and the discrepancy had gone unnoticed in the initial review. Avicen flagged the 60 mg and 30 mg statements side by side, with their dates and source passages.

A capacity change without a clinical event

1,100 pages · 24 documents · 8 authors

Between two opinions, the assessment moved from complete incapacity to 50% capacity. No new examination, diagnosis or documented clinical improvement appeared between them. With both opinions side by side and the interval empty of clinical events, the unsupported change was visible in minutes.

A treatment escalation with a split rationale

760 pages · 18 documents · 6 authors

The decision sat in a clinical note; the evidence sat in a lab table, an imaging report and a later summary. The dosage increase followed worsening laboratory values while imaging remained stable, yet a later report attributed the decision mainly to the imaging. The gap between the stated rationale and the record sat on one screen, ready to cite.

What did the work

The results above are not produced by one feature. They come from a chain, and every link in it was exercised by this corpus:

Ingestion, as it arrives

  • Eight format families and 21 file extensions, including DICOM imaging and audio dictation, processed without preparation.
  • Scanned pages read at 300 DPI with per-word confidence; dedicated detection for handwriting, checkboxes, signatures, stamps and barcodes.
  • Audio transcribed with word-level timestamps; the transcript becomes a document, and citations point to both the sentence and the exact timecode.
  • Working languages: French, English, German, Italian and Spanish, including mixed-language files.

Reconstruction

  • Composite files are split into their internal source units, each attributed an author, a specialty and a date: a 24-document case reads as 24 sources, not one block.
  • Chronology markers are extracted and validated against a real calendar; the timeline orders events across every document and author.
  • Coverage is enforced: the analysis layer must account for 100% of the case content, or processing fails rather than continuing silently.

Verification, fail-closed

  • One citation model across chat, analysis, summaries and reports, anchored to the sentence, the table cell, the image region or the audio timecode.
  • Every generated citation is re-verified against its source before display. When a citation cannot be verified at full precision, it says so instead of pretending.
  • Page quality propagates: an illegible page caps the confidence of everything derived from it.

Analysis and reports

  • Findings across five categories, including divergences between assessments, each carrying its evidence and a verification flag for the expert.
  • A corroboration engine surfaces every mention of the same fact across the whole case, normalizing units and tolerating notation differences.
  • Reports follow a customizable structure of 19 chapters and 63 sub-chapters; Word and PDF exports keep an inline marker at each claim and an evidence appendix with page, excerpt and corroborating sources.

Control and confidentiality

  • Shared workspaces with three roles and 47 permissions; comments with mentions; 42 audited action types.
  • All model inference runs on dedicated GPU infrastructure; connections to external model providers are blocked in code, not just by policy. Data stays on HDS-certified hosting in Switzerland or the EU and is never used for training.

Limits of this study

  • 25 cases is a meaningful corpus for a working method, not an epidemiological sample. We publish the corpus profile so the scope is clear.
  • The measurements were recorded by the participating experts themselves, on their own cases, not by independent observers.
  • All cases came from one practice context: Swiss expert assessment, predominantly French-language.
  • Figures are rounded for publication, and this page will be updated if a later edition of the programme changes them.
  • One boundary is deliberate. Avicen embeds a large multilingual terminology base (over 600,000 drug names from five national registries and the WHO INN list, 173,000 ICD-10 codes and 97,000 LOINC lab codes across its five working languages) to recognize medications, diagnoses and lab tests however the record writes them. What it does not apply is external clinical rules: a potential finding that would depend on one, such as a drug interaction, is suppressed rather than guessed. Everything the platform states must be verifiable in the case itself.
  • This is a vendor-run evaluation, not a peer-reviewed study. We publish it because the profession deserves numbers with their method attached, not numbers alone.

Citing this work

Avicen (2026). Twenty-five real cases, before and with Avicen: the 2026 evaluation programme. https://www.avicen.io/evaluation

Journalists and researchers: for methodology details beyond this page, write to contact@avicen.io. We answer within 24 hours.

The 26th case could be yours.

A 30-minute call to scope one of your cases, then a free, guided 14-day evaluation. No patient data is needed to start.

Start my free evaluation

Reply within 24 hours. No commitment.