Current Work
My current work focuses on how generative and agentic AI systems behave in deployment, and how that behaviour can be evaluated, governed, and made accountable.
I am especially interested in systems configured through prompts, tools, memory, workflows, interfaces, users, institutional settings, and downstream consequences. These systems cannot be adequately assessed through isolated prompts and outputs. They require evaluation methods that trace behaviour over time and connect it to governance evidence.
This work extends my PhD research on evaluation as governance into practical methods for organisations building, deploying, regulating, or assuring AI systems that act across real contexts.
I am open to research, advisory, policy, evaluation, and governance collaborations with organisations working on advanced AI systems.
Research areas
- Agentic AI evaluation: Evaluating systems whose behaviour unfolds across steps, tools, memory, delegation, and changing context.
- Deployed configurations: Treating AI systems as configured sociotechnical arrangements, not as standalone models.
- Behavioural trajectories: Studying how AI behaviour develops over time, where it drifts, how errors propagate, and what consequences follow.
- Evaluation as governance evidence: Designing evaluation methods that support oversight, accountability, contestation, and institutional decision-making.
- Responsible AI capability: Assessing whether organisations have the expertise, processes, and judgement needed to govern AI responsibly.
- Sociotechnical AI assurance: Developing methods that connect technical behaviour to organisational settings, human judgement, and downstream effects.
Current projects
Evaluating Agentic AI Systems: From Models to Deployed Configurations
I am developing a paper on trajectory-based evaluation for agentic AI systems. The central argument is that behavioural trajectories become governance evidence when they connect deployed configuration to institutional consequence.
The paper shifts evaluation away from isolated model outputs and towards systems in use: prompts, tools, memory, workflows, users, organisational roles, and accountability structures.
Traversals, Not Tokens: Movement, Time, and Evaluation in Generative AI
I am developing a conceptual paper on generative AI evaluation that shifts the unit of analysis from outputs to movement. The paper argues that responses are traces of traversal through a structured probability landscape, shaped by prompt, model, context, and environment.
It introduces semantic friction to describe resistance, instability, and variation when systems move through less probable semantic terrain, and ripple propagation to describe how outputs condition later interaction, decisions, and downstream effects.
This paper sits alongside my work on agentic AI evaluation by asking a more foundational question: what exactly are we measuring when we evaluate generative systems?
Constellations of AI evaluation and governance
I am also developing a broader conceptual paper on the constellation of ideas behind my work: evaluation as governance, deployed configurations, sociotechnical recursion, semantic hyperspace, pluralist measurement, and institutional accountability.
The paper asks how these ideas fit together as a research programme for AI systems that are no longer adequately understood as models alone, but as configured sociotechnical systems acting through organisations, infrastructures, and publics.
Responsible AI capability and readiness
I am developing applied work on how organisations assess Responsible AI capability, identify governance gaps, and match expertise to deployment-specific risk.
This work asks a practical question: how can organisations tell whether they have the right forms of AI governance capability for the systems they are actually deploying?
Governance workshops and briefings
I design and deliver workshops, talks, and briefings on generative AI, agentic systems, evaluation, governance, and accountability.
Recent work has focused on Australia’s AI governance landscape, the limits of model-centred evaluation, and the practical infrastructure needed for accountable AI deployment.
Open to collaboration
I am open to research, advisory, policy, evaluation, and governance collaborations with organisations working on generative or agentic AI systems.
I am especially interested in collaborations involving:
- agentic AI evaluation
- AI governance infrastructure
- evaluation design and assurance
- sociotechnical risk assessment
- responsible AI capability and readiness
- law, policy, and institutional accountability
- workshops, briefings, and applied research partnerships
For collaboration, advisory, speaking, workshop, media, or role enquiries, please get in touch.
How this connects to my PhD
My PhD, Measuring the Machine: Evaluating Generative AI as Pluralist Sociotechnical Systems, argued that AI evaluation is not neutral measurement. Evaluation helps shape what AI systems appear to be, what organisations optimise, and whose values become visible in practice.
My current work carries that argument into the governance problem now emerging around agentic AI: systems that do not merely generate outputs, but act through configurations, workflows, tools, and institutional settings.
Read more: PhD Research
