Skip to content
All cases

Designing how scientists review AI analysis

My role
Product Designer & CPOSole designer in a team of five, responsible for product design and product direction.
Period
Aug 2023 → presentzero to live experiments in nine months, across three product generations
Scope
Research · product model · UX/UI · design system · ownership~40 screens over 3 user roles, on a web platform and on the instrument itself
Field
Drug discovery and developmental toxicologybuilt on research published in Nature Methods and featured in Method of the Year 2023
8 hours~30 min
Scientist's hands-on time per run
3–7%
Results flagged for detailed review
Zero9 months
To the first live experiments
~40 screens
Across 3 roles and 2 surfaces, on ~40 components

Workflow figures are based on internal measurement and are approximate. Every run requires a scientist’s sign-off.

Interfaces are reconstructed in English with synthetic experiment data.

The platform on a laptop in a lab: a zebrafish run graded well by well into lethal, sub-lethal and tolerated, beside dose-response and lethality analytics.

In one paragraph

Turning a research method into a usable product

EmbryoNet uses AI to analyse images of developing embryos and identify how compounds affect them.

The platform is based on research from the University of Konstanz published in Nature Methods.

I joined as the sole designer and CPO in a team of five. I designed the web platform and the interface for EmbryoScan, the portable microscope developed by the team. My work covered research, experiment setup, result review, correction and the design system.

The main challenge was helping scientists inspect and verify the model’s output within their existing workflow. Results needed supporting evidence, a record of changes and human sign-off. The first live experiments ran through the product nine months after we started.

Five stages of a run: biology, drug library, time-series imaging, AI analysis, and the interface where a person decides ONE RUN, END TO END 01 Organoids & embryos Grown from stem cells, or embryos on a plate 02 Drug library One compound per well, at the lab's own doses 03 Time-series imaging EmbryoScan images every well, for hours 04 −BMP +RA −Wnt −Nodal AI analysis Phenotype named, and the pathway behind it 05 The interface Where a person decides what to believe Compound in, mechanism out — with a human at step five, every time. 240,000+ images from a 24-hour time-lapse, analysed in ~80 minutes on a consumer graphics card.

The argument in three screens

A design that failed, what replaced it, and what a scientist does when the model is wrong. Each is shown in full further down; these link to where.

  1. Results — sorted by confidence96 wells
    Well
    Compound
    Phenotype
    Confidence
    A4
    CMPD-123
    Lethal
    0.94
    C7
    CMPD-204
    Lethal
    0.88
    B5
    CMPD-123
    Sub-lethal
    0.71
    G6
    CMPD-407
    Sub-lethal
    0.68
    D9
    CMPD-204
    Lethal
    0.61
    F3
    CMPD-311
    Tolerated
    0.55
    Showing 6 of 96Hide below0.60
    1The confidence score we shipped and pulled
  2. EmbryoNetAI Technologies
    5 creditsLog out
    Zebrafish toxicity screen · Q3
    Run P-07Danio rerio·96 wells·analysed 24 min ago
    Toxicity-v2Export FAIR package
    Results by well
    LethalSub-lethalToleratedNeeds review
    1
    2
    3
    4
    5
    6
    7
    8
    9
    10
    11
    12
    A
    A1
    A2
    A3
    A4
    A5
    A6
    A7
    A8
    A9
    A10
    A11
    A12
    B
    B1
    B2
    B3
    B4
    B5?
    B6
    B7
    B8
    B9
    B10
    B11
    B12
    C
    C1
    C2
    C3
    C4
    C5
    C6
    C7
    C8
    C9
    C10
    C11
    C12
    D
    D1
    D2
    D3
    D4
    D5
    D6
    D7
    D8
    D9?
    D10
    D11
    D12
    E
    E1
    E2
    E3
    E4
    E5
    E6
    E7
    E8
    E9
    E10
    E11
    E12
    F
    F1
    F2
    F3?
    F4
    F5
    F6
    F7
    F8
    F9
    F10
    F11
    F12
    G
    G1
    G2
    G3
    G4
    G5
    G6?
    G7
    G8
    G9
    G10
    G11
    G12
    H
    H1
    H2
    H3
    H4
    H5
    H6
    H7
    H8
    H9
    H10
    H11
    H12
    How this run split
    30lethal16sub-lethal46tolerated4needs review
    Review thresholdSet by this lab · Toxicology
    Adjust
    Needs review4 of 96
    B5CMPD-123 · 3.0 µM
    Model saysSub-lethal
    Two phenotypes present in the same well
    ConfirmOpen frames
    D9CMPD-204 · 0.3 µM
    Model saysLethal
    Diverges from both replicates
    ConfirmOpen frames
    F3CMPD-311 · 0.3 µM
    Model saysTolerated
    Imaging interrupted at 14 h
    ConfirmOpen frames
    G6CMPD-407 · 10 µM
    Model saysSub-lethal
    Phenotype not in the trained set
    ConfirmOpen frames
    The other 92 are settled. Open any of them from the plate.
    Not signed off4 results still open. Signing enters this run in the audit trail.
    Sign off run
    2A queue of what needs a person
  3. EmbryoNetAI Technologies
    5 creditsLog out
    Run P-07 · Needs review, 4 wells
    Well B5CMPD-123·3.0 µM·24 h · 360 frames
    PreviousNext in queue
    Frame 14:00
    Marked as the divergence210 / 360
    B5 · 14:00 · bright-field
    What the model readSub-lethalAssigned from 06:00 onward
    What you are recordingLethalFrom 14:00 onward
    Frames affected150 of 360
    Correct from this frame
    Phenotype
    Lethal
    Signalling pathway
    BMP — loss of signal
    The correction is applied at 14:00 and carried forward. Everything after it is re-analysed rather than relabelled.
    Correct and recompute
    This lab contributes de-identified corrections to model training.Change what is shared
    Time-lapse
    Dimmed frames will be recomputedEvery 2 h of 24 h
    00:00
    02:00
    04:00
    06:00
    08:00
    10:00
    12:00
    14:00
    16:00
    18:00
    20:00
    22:00
    3Correcting the model from a frame

The research

What research changed

I conducted around 40 interviews with principal investigators, postdocs and lab technicians across three product generations.

Four findings shaped the design:

  • Scientists needed to inspect the evidence behind a result.
  • They organised experiments by well position, so the interface needed to preserve the plate layout.
  • Tables were more useful for reviewing a whole experiment; images mattered when examining individual results.
  • Labs wanted the product to fit their existing data systems rather than replace them.

Setting up a run

Setting up an experiment

Labs already used spreadsheets to record compounds, concentrations and controls by well position.

I kept that spatial model in the interface: scientists could select a range of wells and assign values directly on the plate grid.

Setup was available on both the web and the instrument, so the person preparing the plate could record its contents where they worked.

EmbryoNetAI Technologies
5 creditsLog out
Zebrafish toxicity screen · Q3
Plate P-08Danio rerio·96 wells·not started
Import plate mapStart run
Plate map
SelectedControlEmpty
1
2
3
4
5
6
7
8
9
10
11
12
A
A1veh
A20.1
A30.3
A41
A53
A610
A730
A80.1
A90.3
A101
A113
A1210
B
B1veh
B20.1
B30.3
B41
B53
B610
B730
B80.1
B90.3
B101
B113
B1210
C
C1neg
C20.1
C30.3
C41
C53
C610
C730
C80.1
C90.3
C101
C113
C1210
D
D1neg
D20.1
D30.3
D41
D53
D610
D730
D80.1
D90.3
D101
D113
D1210
E
E1neg
E20.1
E30.3
E41
E53
E610
E730
E80.1
E90.3
E101
E113
E1210
F
F1neg
F20.1
F30.3
F41
F53
F610
F730
F80.1
F90.3
F101
F113
F1210
G
G1neg
G20.1
G30.3
G41
G53
G610
G730
G80.1
G90.3
G10
G11
G12
H
H1neg
H20.1
H30.3
H41
H53
H610
H730
H80.1
H90.3
H101
H113
H1210
Assigned93 / 96
Compounds4
Concentrations6
Controls8
Assign to selection
SelectionB2 — B76 wells, one row
Compound
CMPD-123
Concentration
0.1 — 30µM
Series
Half-log, 6 steps
Imaging protocol
Bright-field · 24 h · 4 min
Mark as vehicle controlMark as negative controlClear wells
Apply to 6 wellsOr fill the row by dragging
3 wells have no compoundG10, G11, G12 will be imaged and analysed with nothing recorded against them. Once the run starts this cannot be corrected.
A plate grid with compound, concentration and control assignment.

Triage, or designing for the part where the model is wrong

A queue instead of a spreadsheet

My first design displayed a confidence score beside each result.

In testing, scientists compared those scores as though they were experimental measurements. The interface was making model confidence look like scientific data.

Results — sorted by confidence96 wells
Well
Compound
Phenotype
Confidence
A4
CMPD-123
Lethal
0.94
C7
CMPD-204
Lethal
0.88
B5
CMPD-123
Sub-lethal
0.71
G6
CMPD-407
Sub-lethal
0.68
D9
CMPD-204
Lethal
0.61
F3
CMPD-311
Tolerated
0.55
Showing 6 of 96Hide below0.60
The first version, with a confidence score on every result.

I changed the design so confidence routed results into a review queue instead.

Model confidence routes results into settled and needs-review lanes; both lanes are signed off by a person Model output 96 results, each with a label and a confidence CONFIDENCE THRESHOLD Set by the lab. We ship defaults. Settled Classified, not yet accepted 93–97% Needs review A queue, ordered by where the scientist's attention is worth most The frames are one click away 3–7% Signed off by a person Nothing is auto-accepted AUDIT TRAIL 21 CFR Part 11 · GDPR

Around 3–7% of results were flagged for detailed inspection, with footage available alongside them. Scientists could still inspect any result, and the full run required human sign-off.

EmbryoNetAI Technologies
5 creditsLog out
Zebrafish toxicity screen · Q3
Run P-07Danio rerio·96 wells·analysed 24 min ago
Toxicity-v2Export FAIR package
Results by well
LethalSub-lethalToleratedNeeds review
1
2
3
4
5
6
7
8
9
10
11
12
A
A1
A2
A3
A4
A5
A6
A7
A8
A9
A10
A11
A12
B
B1
B2
B3
B4
B5?
B6
B7
B8
B9
B10
B11
B12
C
C1
C2
C3
C4
C5
C6
C7
C8
C9
C10
C11
C12
D
D1
D2
D3
D4
D5
D6
D7
D8
D9?
D10
D11
D12
E
E1
E2
E3
E4
E5
E6
E7
E8
E9
E10
E11
E12
F
F1
F2
F3?
F4
F5
F6
F7
F8
F9
F10
F11
F12
G
G1
G2
G3
G4
G5
G6?
G7
G8
G9
G10
G11
G12
H
H1
H2
H3
H4
H5
H6
H7
H8
H9
H10
H11
H12
How this run split
30lethal16sub-lethal46tolerated4needs review
Review thresholdSet by this lab · Toxicology
Adjust
Needs review4 of 96
B5CMPD-123 · 3.0 µM
Model saysSub-lethal
Two phenotypes present in the same well
ConfirmOpen frames
D9CMPD-204 · 0.3 µM
Model saysLethal
Diverges from both replicates
ConfirmOpen frames
F3CMPD-311 · 0.3 µM
Model saysTolerated
Imaging interrupted at 14 h
ConfirmOpen frames
G6CMPD-407 · 10 µM
Model saysSub-lethal
Phenotype not in the trained set
ConfirmOpen frames
The other 92 are settled. Open any of them from the plate.
Not signed off4 results still open. Signing enters this run in the audit trail.
Sign off run
Flagged results with evidence available for review.

Review controls

  • Every run requires a scientist's sign-off.
  • Labs can adjust the review threshold for their work.
  • Source images remain accessible from every result.

Correcting a frame, and sending the run back

Correcting a result

Scientists could mark the frame where a classification became incorrect, correct it and rerun the analysis from that point.

The interface showed which subsequent frames would be recalculated.

EmbryoNetAI Technologies
5 creditsLog out
Run P-07 · Needs review, 4 wells
Well B5CMPD-123·3.0 µM·24 h · 360 frames
PreviousNext in queue
Frame 14:00
Marked as the divergence210 / 360
B5 · 14:00 · bright-field
What the model readSub-lethalAssigned from 06:00 onward
What you are recordingLethalFrom 14:00 onward
Frames affected150 of 360
Correct from this frame
Phenotype
Lethal
Signalling pathway
BMP — loss of signal
The correction is applied at 14:00 and carried forward. Everything after it is re-analysed rather than relabelled.
Correct and recompute
This lab contributes de-identified corrections to model training.Change what is shared
Time-lapse
Dimmed frames will be recomputedEvery 2 h of 24 h
00:00
02:00
04:00
06:00
08:00
10:00
12:00
14:00
16:00
18:00
20:00
22:00
Frame-level correction and the portion of the sequence to be recalculated.

Sharing data for model training was a separate choice. Labs had to opt in before their de-identified experiment data and corrections could be used for that purpose.

A scientist marks the frame where the model diverged, corrects it, and the run is recomputed forward from that frame ONE RUN, AS A TIME SERIES the model diverges here t = 0 8 h recomputed forward from the corrected frame Mark the frame where it went off Correct it the scientist's call, recorded Recompute send the run back WHERE CORRECTIONS GO — ONLY IF THE LAB OPTS IN De-identified experiment data, and how the team worked it through Training set the next generation Opt-in first, training after.

What we decided not to build

Choosing what not to build

The original plan included storage and reporting alongside imaging and analysis.

Interviews showed that labs already had systems for that work. I argued for a narrower scope: keep imaging, analysis, review and correction, and export results for use in the lab’s existing systems.

This let our five-person team focus on the workflow the product was built to improve.

The product was cut back to imaging, analysis, triage and correction; storage, reporting and planning leave as a FAIR package for the lab's own LIMS THE ORIGINAL PLAN — THE WHOLE WORKFLOW, ALL OF IT OURS EmbryoNet what only we could do Imaging Analysis Triage Correction WHAT WE CUT Storage Reporting Planning Their protocols FAIR package findable, accessible, interoperable, reusable The lab's LIMS the home they already had Storage Reporting Planning Their protocols Results leave as a package the lab takes into the systems it already has.

Three roles, two surfaces

One run, three views of it

I designed separate views for running an experiment, comparing results across experiments, and administering access and sign-off.

Three roles with different horizons and primary views, across two surfaces: the microscope and the web platform THE SAME RUN, HELD BY THREE PEOPLE 01 Running this experiment HORIZON One plate, today PRIMARY VIEW The plate, then the queue Set it up, run it, adjudicate what the model flagged. 02 Across a group of experiments HORIZON A study, over weeks PRIMARY VIEW A table before a picture Comparison, dose–response, what replicated and what did not. 03 Administrator HORIZON The organisation PRIMARY VIEW Sign-offs and access The layer that makes the audit trail real rather than nominal. Across a group of experiments, a table carries more than a gallery of frames. TWO SURFACES, ONE SET OF OBJECTS UNDERNEATH On the microscope Experiment setup · run status and log · calibration On the web Everything after a run, and everything across more than one

The instrument interface focused on experiment setup, run status and calibration. Detailed review and comparison stayed on the web platform.

EmbryoScanBench 2 · Konstanz
Running
Experiment setup
Run status
Calibration
Elapsed
06:12of 24:00
Plate P-08
Log
06:12Cycle 93 captured · 96 wells
04:00Focus re-acquired on B7, C2
02:18Chamber returned to 28.5 °C
00:00Run started · plate P-08 · protocol BF-24-4
Cycles captured93of 360
Chamber28.5 °Ctarget 28.5 ± 0.5
Storage61%free on this instrument
Stop run
A stopped run keeps every cycle already captured.
Instrument controls for setup, run status and calibration.

The design systemForty components

I built a shared system of around 40 components, including plate grids and time-series views, to support the web and instrument interfaces.

Results

Outcomes and my contribution

The product reduced scientists' hands-on time per run from around eight hours to thirty minutes.

This was a combined result of the analysis technology and the workflow around it. My contribution was designing setup, targeted review, evidence inspection and correction.

We reached the first live experiments in nine months. Across three product generations, I designed around 40 screens covering three user roles and both the web platform and instrument interface.

What I would do differently

What I would improve

Three things I would change on a similar product.

  • Support multiple models and result types earlier in the product structure.
  • Design partial and interrupted runs as normal workflows from the start.
  • Test basic navigation, feedback and empty states as carefully as the specialist tools.
Next caseMortgage Retention DecisioningA mortgage retention product for a US fintech. I designed the workflow that helps advisors review prioritised opportunities, prepare offers and contact borrowers.Read the case studyThe retention dashboard on a laptop: recapture offers with the number of eligible borrowers and total equity, a choropleth of customers by state, a loan product breakdown and a payoff trend chart.

Open to Senior Product Designer roles

Based in Porto, Portugal, with authorisation to work here. Get in touch about a Senior Product Designer role or to discuss my work in more detail.