The Human in the Loop: How AI is Supporting Genomic Discovery
From cell segmentation to protein-structure analysis, artificial intelligence is helping researchers investigate increasingly complex biological data- but prediction remains the beginning of scientific inquiry, not its conclusion.
Advances in sequencing are allowing researchers to capture information at increasingly greater scale and resolution. Yet, generating more detailed data does not automatically produce a clearer biological answer.
As datasets become larger and more multidimensional, researchers face a growing challenge: how do we identify meaningful biological signals within increasingly large and complex datasets?
Artificial intelligence is beginning to help researchers address these questions.
Its value is not in replacing laboratory scientists, bioinformaticians, or domain experts, instead, its value lies in extending their ability to recognise patterns, prioritise findings and connect information across datasets that would be difficult to interpret manually.
The Boundary Problem
One of the clearest applications of AI in genomics can be found in spatial transcriptomics.
Spatial transcriptomics measures gene expression while preserving information about where transcripts occur within a tissue. This allows researchers to study not only which genes are active, but also how different cell types are organised and interact within their surrounding environment.
Before this information can be interpreted at the cellular level, transcripts must be assigned to individual cells. This depends on a process known as cell segmentation: identifying the boundaries of each cell within a tissue image.

Segmentation is therefore not simply an image-processing step; errors in boundary estimation can propagate and impact downstream biological interpretation. Studies including Petukhov et al. and Jin et al. have similarly identified segmentation and annotation as important analytical challenges in spatial transcriptomics.
At Genomics WA, we are exploring AI-assisted approaches to refine cell-boundary estimation using tissue imagery and morphological information to support more accurate cell annotations, retain more usable transcript information and provide researchers with a stronger foundation for downstream analysis.
When a Gene is Not the Whole Story
AI may also help researchers interpret the increasing complexity revealed by long-read RNA sequencing.
As a single gene can produce multiple RNA transcripts through alternative splicing, long-read sequencing can reveal previously unannotated or condition-specific isoforms.
But greater resolution introduces another challenge: determining which of those transcripts are likely to be biologically meaningful.
AI and machine-learning methods could assist by helping researchers:
- identify patterns of differential isoform usage
- predict whether a transcript contains a plausible protein-coding region
- assess whether a sequence change may alter RNA splicing
- prioritise isoforms associated with cell types or conditions
- integrate transcript findings with genomic variants and clinical or phenotypic information
Testing What an Isoform Could Change
For isoforms that are predicted to encode proteins, the analysis could progress one step further.
Changes in transcript structure can alter the amino-acid sequence of the resulting protein, potentially removing part of a functional domain, introducing a premature stop, or producing a protein with substantially different structural properties.
AI-based protein-structure prediction tools such as AlphaFold can generate a three-dimensional model from an amino-acid sequence, allowing researchers to model different protein isoforms and investigate whether changes in their sequences may affect predicted domains, conformations, or interaction regions.
This could help turn a transcriptomic observation into a more focused biological hypothesis. However, the distinction between prediction and evidence is critical.
A plausible protein structure does not prove that the transcript is translated, nor establish that the resulting protein is stable, or responsible for a disease phenotype. Structure-prediction systems generally represent a limited set of molecular states and may not account for protein dynamics, cellular conditions, post-translational modifications or interacting molecules.
The European Bioinformatics Institute similarly describes predicted structures as complementary to experimental evidence and highlights limitations involving molecular dynamics, binding partners, and non-protein components.
Prediction is the Beginning, not the Answer
The longer-term opportunity for AI lies in connecting information across multiple biological layers, potentially assisting researchers in identifying relationships across multiple streams of data, helping recognise variants and unusual combinations of molecular features which may not be immediately apparent through conventional analyses.
While this has potential applications across multiple areas including rare disease, cancer, pharmacogenetics, and population genomics, identifying an association is not the same as establishing a biological mechanism, diagnostic result, or treatment recommendation.
Models can inherit and amplify biases from training datasets, potentially underrepresenting populations, tissue types or disease states. Results can also be affected by sample quality, experimental design, annotation choices and differences between analytical pipelines.
Meaningful use of AI therefore depends on transparent methods, representative data, appropriate controls, and researchers who understand both the biological question and the limitations of the model.
The most credible role for AI in genomics is therefore collaborative: extending the capacity of researchers while scientific oversight, critical judgement and responsibility of the conclusions remain firmly human.
Greater Resolution Requires Greater Infrastructure
Increasing biological resolution also changes what is required behind the scenes, for both researchers and sequencing facilities.
Long-read sequencing, single-cell analysis and spatial transcriptomics can generate increasingly large and complex datasets. Training and applying machine-learning models requires substantial processing capacity, particularly when genomic information is being analysed along large multimodal datasets.
At Genomics WA, the development of an AI-assisted cell segmentation approach has involved training a model across a large collection of images using dedicated machine-learning infrastructure.
Storage is only one part of the equation. Facilities also require sufficient computing capacity, efficient data transfer and appropriate archival infrastructure, whilst sensitive human genomic data requires suitable governance, security, and access controls.
Importantly, infrastructure needs to support more than the initial analysis. Genomic datasets can retain value well beyond the experiment that generated them, and as analytical methods improve, existing sequencing, imaging and spatial data may be revisited with new questions or new computational approaches.
Building the Next Layer of Genomic Capability
Genomics WA is continuing to develop sequencing, bioinformatics, and analytical capabilities alongside advances in long-read, single-cell, and spatial genomics. The objective is not to apply AI for its own sake, but to investigate how emerging computational approaches can improve the interpretation of increasingly high-resolution datasets.
As our ability to generate high-resolution biological data increases, both researchers and sequencing facilities must be equipped to manage, analyse, and interpret the complexity that comes with it.
AI can help researchers navigate that complexity, but meaningful discovery still depends on the people who design the experiment, test the predictions, and place those findings into biological context.







