Atlas OS · Analysis and delivery

The complete chain

Between raw data and a delivered result there used to be manual work. Since this month that stretch runs as a machine – with the run as an object in the database, a checksum for every file, and a release that a person still grants.

On 22 August 2026 at 11:06, the last open stage of a whole-genome run reported the state “finished”. Behind it lie four days, 103.6 gigabytes of result data and not one manual touch of the data. Everything that happened during those four days was in our database before a person had read it: which tool at which version, on which compute node, from when to when, with which checksum going in and which coming out.

9

stages in one run, each logged separately

4

days from raw data to the finished delivery bundle

103.6 GB

of result data delivered, with a checksum per file

6

delivery lines with a checksum, each released on its own

0

manual interventions in the data between start and delivery bundle

Nine stages · four days · no manual touch

01

The gap sat in the middle

The front of the chain has been in place for months and has been written up: an order becomes a quotation, the quotation becomes an order confirmation, the confirmation becomes sample logistics with shipment tracking, the sample becomes an arrival in the lab. The back is in place just as much: portal, documents, invoice, release. What lay in between is run as manual work across our industry to this day. Raw data becomes alignments, alignments become variants, variants become an annotated and filtered selection, that becomes a report, and that becomes a delivery. Every one of those handovers was a person carrying a file from one place to another – and having to explain afterwards, from memory, where it came from.

  1. Quote
  2. Confirmed
  3. In transit
  4. At the lab
  5. Data delivered
  6. Completed

Order phase: Data deliveredResults released and available to download.

That middle is closed. For the first time the chain runs from the order through to a downloadable result without anyone pushing it along by hand at some point.

02

The run is an object now, not a story

Until now a compute run was something that happened on a machine and existed afterwards only as a folder. Today it is a record: pipeline and version, compute node, period, key figures, plus every single stage with its own tool, its own version, its own time window and its own state. The run hangs off the order, and every result file hangs off the run.

That sounds like bookkeeping. It is the difference between a result you have to take on trust and a result that brings its own provenance with it. Anyone asking in two years' time how a particular variant ended up in a particular table gets a row as the answer, not a recollection.

QN-2026-9xxx_P2026-3xx

Completed
Pipeline
nf-core/sarek 3.8.1
Period
18 Aug 2026, 19:54 – 22 Aug 2026, 11:06
Compute node
atlas-p0
Stages
9

Run completed without intervention. All nine stages logged with tool, version, time window and checksum.

StageTool and versionPeriodState
Alignmentbwa-mem 0.7.1818 Aug 20:48 – 19 Aug 11:43Completed
Duplicate markinggatk4 MarkDuplicates 4.6.1.019 Aug 11:44 – 13:36Completed
Recalibrationgatk4 BaseRecalibrator/ApplyBQSR19 Aug 13:38 – 16:21Completed
Variant callinggatk4 HaplotypeCaller 4.6.1.019 Aug 16:21 – 22:52Completed
AnnotationEnsembl VEP 11520 Aug 11:22 – 13:30Completed
Candidate filterkandidaten.sh20 Aug 13:30 – 13:34Completed
Additional analysissamtools/mosdepth, Ensembl REST20 Aug 14:08 – 21 Aug 15:29Completed
Splice scoringPangolin 1.0.220 Aug 22:34 – 22 Aug 11:06Completed
Structural variantsManta 1.6.021 Aug 00:50 – 01:13Completed

03

The machine navigates, the person decides

Every stage is selected, parameterised, started, monitored and evaluated without anyone setting it off. More interesting than running through is stopping. In one of the runs, a coverage weakness in a difficult genomic region could have stayed invisible under averaged window values. The analysis noticed the discrepancy itself, repeated the calculation at interval level, put both results side by side and wrote the rule that followed from it into the log. In another case a scanned document could not be read; instead of filling in the missing details, the run reported what it had been unable to determine and took the route via image recognition.

That is the actual point. A machine that only runs through is a risk. A machine that knows where its statement ends is a tool.

At the end of every chain there is therefore no automaton but an order of steps. Bioinformatics checks the run against its key figures. The technical and scientific team checks the report against the data. Where the case calls for it, medical sign-off goes on top. What has disappeared are the manual steps. What has remained are the decisions – and those are meant to remain.

How much falls away is best shown by a run itself. The annotation before it writes one row per transcript and one per gene for every variant – 45,042,328 and 6,632,525 rows. The cascade below counts variants, not rows, and therefore starts again from the beginning.

  1. 4,981,959variants from variant calling

  2. 261,763after frequency threshold

  3. 933after consequence class

  4. 795candidates in the delivered table

Every stage is evidenced in the delivery bundle as a file of its own.

04

Nothing leaves the building that has not been released

Finished files are not sent, they are fed in. Every line carries a name, a size and a checksum, and it sits withheld at first. Only a switch on each line releases it to the customer, and that switch writes into the order history who released what and when. An internal working draft can therefore sit in the same order as the customer document without ever becoming visible.

For large data a rule applies that we took from providers who move billions of objects a day: the portal authorises, it does not transport. No shared folder password, no link that travels on. The address is created at the moment of the click, is valid for minutes or hours depending on file size, lets an interrupted transfer resume, and leaves a row behind on every retrieval. For the customer, six kilobytes and a hundred gigabytes therefore look the same.

  • Analysis report on the genome analysis, 11 pages

    QN-2026-9xxx_auswertungsbericht.pdf

    Size 233.5 KBChecksum 0be87d69…24afaa71

    Released
  • Candidate variants with frequency, consequence and database comparison

    QN-2026-9xxx_kandidatenvarianten.tsv.gz

    Size 72.5 KBChecksum 43424c0f…14328ec1

    Released
  • Quality report of the sequencing and analysis run

    multiqc_report.html

    Size 2.6 MBChecksum ed749303…d7b1fd13

    Released
  • Supporting files and individual analyses, 45 files

    QN-2026-9xxx_belege.zip

    Size 140.7 KBChecksum 4e0e22e9…b1f82328

    Released
  • Internal working draft for the lab team, not for distribution

    QN-2026-9xxx_arbeitsfassung.docx

    Size 91.6 KB

    Withheld

Raw data and analysis files of the delivery, ten files, 103.6 GB

The address is created on click and is valid for a limited time. Every retrieval is logged.

05

Why this holds for every question

The chain does not ask what biology stood at the beginning. Raw data comes in at the front, defined deliverables fall out at the back, and in between lies an interchangeable stretch. Whether a genome, an exome, a trio, a microbiome, an expression profile or a protein panel from the partner network stands at the start changes the tools in the middle. It changes nothing about what stands around them: run object, stages, manifest, checksums, withholding, release, log.

That is exactly why the effort is one-off and the benefit recurring. A new question costs a pipeline, not a new organisation.

06

The evidence

One order has run through this chain in full, with three further runs over the same stretch. The fourth case comes from an ongoing validation project and is not a genome but an exome – enriched rather than sequenced in full, and from a different type of sequencing instrument than the three before it. The same stretch, a different data source, a different method, without a line of special handling.

Four cases are no proof of throughput. They are something else, and at this moment that is worth more: evidence that the architecture holds. What runs through four times without special handling also runs four hundred times – the difference is then compute time, not organisation.

What this post leaves out, it leaves out on purpose. No customer is named, no finding appears, and where a collaboration is still under way, we wait for the people involved before we show anything. The people whose data we process decide when it is spoken about.

07

The rest of the way

Two things are still manual work today, and both are on the list for the coming weeks. Someone starts the run, instead of arriving raw data setting it off by itself. And someone looks at the finished document before it goes out, because a report is more than the text in it. The laboratory side has already gone the opposite way: the instruments have long been programmed rather than operated, and the handling of the samples will follow before long.

The ambition is immodest and worth writing down for exactly that reason. Nine hundred and ninety-nine runs in a thousand should pass through without intervention. The thousandth should stand out because a stage says so itself – not because someone happened to look.

Our measure is not how fast we produce files. It is the time from a biological question to a confirmed result, and the number of hands that have to reach in along the way.

Order, run and delivery lines sit in the portal at app.atlasbiolabs.com. How the front of the chain comes about – from the question to the binding quotation – is written up here.