Abstract
Fisheries modelling requires specialist training, but the availability of suitably skilled modellers is limiting delivery of science to inform management. Agentic AI can automate computing workflows and will increasingly be used in ecological modelling. However, the quality of AI generated results is of concern because of risks of logically flawed code and inconsistent answers. Here, we tested whether a general‐purpose AI agent (Roo Code) can write computer code and interpret results to complete three common ecological and fisheries modelling workflows: (1) fitting a von Bertalanffy growth curve to data, (2) writing code and interpreting results of a generalised linear model relating fish abundance to habitat and (3) completing a yield per recruit analysis. We used replicate prompts and a structured evaluation rubric to ask if the agents' results were accurate and if responses were consistent. The agent successfully completed all three tasks in some replicates; however, accuracy varied widely across replicates. Of the large language models we tested, Claude Sonnet 4.5 had the greatest accuracy and always produced accurate responses for the simplest modelling task, but was less consistently accurate for complex tasks. All models sometimes overlooked user instructions for which methods to use. Our findings show that agents can sometimes complete ecological and fisheries modelling, but that they are not yet consistent and accurate enough to rely on for management advice. We outline prompt designs and workflows that improve reliability and discuss the importance of ensuring human oversight for credibility of modelling that is used in management.