Approaches & positioning

Why this lab should exist
Global agriculture must produce substantially more food from increasingly constrained land, water and nutrient resources, under greater climatic uncertainty. Much of that gain must come from better crops and better use of those crops. Genomics has transformed our ability to generate and identify genetic variation, but breeding can only exploit that variation as quickly as useful phenotypes can be recognised. High-throughput sensing has greatly expanded how much phenotype we can measure, but the harder bottleneck is increasingly understanding what those measurements mean biologically, early enough to make a decision.
Optical remote sensing and crop phenology capture emergent effects of genotype and environment, with equifinality in their generative process. Yield and other endpoints of crop growth can likewise arise through very different physiological routes. This limits early-season selection for yield, stress diagnosis and crop modelling: a sensor may detect that a crop is different without revealing why, and a powerful model cannot recover the correct mechanism from inputs that do not distinguish competing explanations.
We integrate fundamental crop physiology with data science to break that equifinality. We do this by reconstructing organised physiological space from rich phenomics, locating crops within it, and using causal machine learning to explain the mechanisms producing their trajectories.
The research programme
Our programme investigates:
- which components of crop physiology remain sufficiently persistent through development and environmental plasticity to provide reliable, organised biological reference points;
- how rich phenomics can reconstruct that organised physiological space and position crop observations within it; and
- how causal, knowledge-guided machine learning can distinguish the mechanisms and transition pathways responsible for crop performance.
Our work shows that frozen phenomic representations can retain broad functional organisation across years even when exact coordinates change. Spectral features can retain interpretable biological information while following genotype-specific developmental trajectories, allowing physiological differences between cultivars to be assigned. Mechanistically constrained neural models can recover physiologically meaningful and identifiable latent processes.
These ideas form one integrated system. Rich physiological phenotyping supplies reference biology from which phenomic representations are produced. Individual crops are assigned coordinates and mobility constraints within those representations. Causal mechanistic and hybrid models then explain the processes producing their movements. Hyperspectral sensing, pulse-amplitude-modulated fluorescence and other scalable sensors become practical ways of locating crops within the resulting physiological system.
Our strategic advantage
Our advantage is the combination of crop physiology, deep phenotyping and causal computation with mechanistically structured machine learning. Rather than developing sensors without physiology or physiological models disconnected from observation, we work across the entire inference chain measurement → physiological state → causal mechanism → breeding or management decision.
Our ambition is to make physiological state-space a useful organising layer for crop improvement, as genotyping has become. We envisage breeding programmes that identify useful physiological strategies early enough to fast-track selection; crop models that learn population-specific mechanisms from sensor observations while retaining physiological meaning; and autonomous agricultural systems that recognise a crop’s state, diagnose why it is departing from a desirable trajectory, and support the intervention most likely to improve productivity, resilience and resource efficiency.
Our research infrastructure
Our capability is organised as a connected research stack rather than a collection of instruments or isolated analytical services. Agronomy and crop physiology define the questions, experiments and calibration measurements. Automated phenotyping turns field observations into analysis-ready trajectories. Governed data infrastructure preserves context and provenance. Causal machine learning then connects measurements to mechanisms, using compute that scales from routine local training to demanding cloud workloads.
Agronomy and physiology trials designed around biological contrasts, perturbations and independent calibration.
Repeated, sensor-aware acquisition converted into spatially aligned plot and plant phenotypes.
Studies, protocols and multimodal assets organised through shared metadata, provenance and access rules.
Physiological hypotheses compiled into trainable causal pathways with interpretable latent states.
Routine GPU training on premises, with AWS used when data volume, memory or computation demands expansion.
In our typical neural models, each layer is designed to preserve the scientific meaning established by the layer before it. A model output can therefore be traced back through its graph, dataset, processing history, sensor protocol and field experiment.
Experimental and measurement capability
The stack begins in a crop science laboratory and in the field. Our agronomy and physiology trial expertise allows us to create the contrasts needed to distinguish canopy development, light interception, photosynthetic capacity, resource-use efficiency, partitioning and stress response. Experimental perturbations and targeted reference measurements are planned alongside sensing, because latent physiological traits are only credible when the experiment contains enough information to identify and validate them.
Our sensor capability is selected around those questions. SPECIM hyperspectral cameras provide dense spectral information for biochemical and physiological inference. Multispectral and thermal drones extend repeated measurement across whole trials, while autonomous drone operations through DJI Dock are on the way to eventually supporting consistent, high-frequency acquisition. Fluorometry tools such as MultiSpeQ contribute direct physiological calibration, and emerging spectroscopy tools such as SpectroPod allow us to test lower-cost and more deployable measurement strategies. The value of this portfolio lies in how the modalities constrain one another such that field observations provide scale, laboratory and proximal tools provide physiological anchors, and repeated sensing reveals dynamics.
Automated field phenotyping
Our drone-phenomics software turns aerial acquisition into repeatable plot-level data. The workflow coordinates processing jobs, builds calibrated orthomosaics and digital elevation products, detects and uses ground-control information, and aligns repeated outputs to a common reference so plots remain spatially comparable through time.
Governed research data
The CBL data platform provides the organisational layer for this work. It treats tables, vectors, rasters, image collections and documents as registered scientific assets rather than loose files. Researchers describe an asset through a structured manifest; the platform validates the submission, generates its canonical metadata sidecar and records searchable information about the study, experimental unit, protocol, spatial and temporal coverage, ownership and provenance.
This metadata-first approach makes heterogeneous research outputs discoverable and joinable without forcing every project into one monolithic database. A versioned rulebook provides a common data contract across the user interface, backend and programmatic access tools. The platform therefore supports reproducibility at the point where it is most often lost: between data acquisition, processing and model assembly.
Causal machine learning with Causality
We train mechanism-aware models through our in-house Causality platform. A researcher expresses a scientific hypothesis as a directed graph in a human-readable specification. Nodes represent observed inputs, latent physiological states, processes and outcomes; edges define permitted information flow; and constraints encode knowledge such as plausible ranges, non-negativity, ordering, saturation or conservation.
Causality validates the graph and compiles it into a trainable model. This allows neural components to learn flexible relationships without discarding the physiological architecture. The platform supports tabular, longitudinal, image and multimodal inputs, with explicit bindings between dataset variables and graph nodes. Sparse measurements can supervise intermediate states, allowing large volumes of indirect sensor data to be anchored by smaller sets of destructive or instrument-intensive measurements.
Causal machine learning commitments
Across this stack, our modelling commitments are:
- Mechanistic grounding in physiology and biophysics.
- Experimental perturbation to separate contributors and support identification.
- Structured modelling that respects causal architecture.
- Uncertainty and alternatives when the available measurements cannot uniquely identify a mechanism.
- Reproducibility through versioned data, code and model specifications.
Phenogenic fields
A phenogenic field is a reusable geometry of realised plant function. Rather than rebuilding a model for every target, measured phenotypes locate genotypes or plots within a frozen reference space. Yield, quality, stress response and stability can then be added as annotations.
The programme asks which minimal coordinate phenome can preserve useful biological neighbourhoods across time, environments and sensors. Its flagship test is early phenogenic selection for increased realised gain per year, cycle or cost?
One coherent programme
Structure
Causal graphs make scientific claims explicit and connect governed observations to named mechanisms.
Geometry
Phenogenic fields use harmonised longitudinal measurements to capture reusable patterns of biological resemblance without treating every feature as causal.
Action
Automated measurement and scalable computation let both approaches be judged by intervention, transfer, selection and independent biological evidence.
